10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This work demonstrates an objective way to select parameter values for a quadratic integrate-and-fire model so that its bifurcation diagram matches a specific target diagram, generated from the Wang-Buzsaki model. The method is useful for the field and is presented with convincing evidence. The method is currently limited in its ability to be applied to data, but improves our mathematical tools to treat a rarely studied type of bifurcation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      From a big picture viewpoint, this work aims to provide a method to fit parameters of reduced models for neural dynamics so that the resulting tuned model has a bifurcation diagram that matches that of a more complex, computationally expensive model. The matching of bifurcation diagrams ensures that the model dynamics agree on a region of parameter space, rather than just at specially tuned values, and that the models share properties such as qualitative features of their phase response curves, as the authors demonstrate. A notable point is the inclusion of extracellular potassium concentration dynamics into the reduced model - here, the quadratic integrate-and-fire model; this is straightforward but nonetheless useful for studying certain phenomena.

      Strengths:

      The paper demonstrates the method specifically on the fitting of the quadratic integrate-and-fire model, with potassium concentration dynamics included, to the Wang-Buzsaki model extended to include the potassium component. The method works very well overall in this instance. The resulting model is thoroughly compared with the original, in terms of bifurcation diagrams, production of various activity patterns, phase response curves, and associated phase-locking and synchronization properties.

      Weaknesses:

      It is important to note that the proposed method requires that a target bifurcation diagram be known. In practical terms, this means that the method may be well suited to fitting a reduced model to another, more complicated model, but is not likely to be useful for fitting the model to data.

    3. Reviewer #2 (Public review):

      Summary:

      The authors derive an integrate-and-fire model to describe the dynamics of a more complex Wang-Buzsaki model and compare the two models. A detailed discussion of bifurcation schemes in both models is convincing and allows us to evaluate the simpler model.

      Strengths:

      The idea is interesting, and the mathematical approach appears to be convincing. In addition, differences between the simple and original models are also discussed.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      From a big picture viewpoint, this work aims to provide a method to fit parameters of reduced models for neural dynamics so that the resulting tuned model has a bifurcation diagram that matches that of a more complex, computationally expensive model. The matching of bifurcation diagrams ensures that the model dynamics agree on a region of parameter space, rather than just at specially tuned values, and that the models share properties such as qualitative features of their phase response curves, as the authors demonstrate. A notable point is the inclusion of extracellular potassium concentration dynamics into the reduced model - here, the quadratic integrate-and-fire model; this is straightforward but nonetheless useful for studying certain phenomena.

      Strengths:

      The paper demonstrates the method specifically on the fitting of the quadratic integrateand-fire model, with potassium concentration dynamics included, to the Wang-Buzsaki model extended to include the potassium component. The method works very well overall in this instance. The resulting model is thoroughly compared with the original, in terms of bifurcation diagrams, production of various activity patterns, phase response curves, and associated phase-locking and synchronization properties.

      Weaknesses:

      It is important to note that the proposed method requires that a target bifurcation diagram be known. In practical terms, this means that the method may be well suited to fitting a reduced model to another, more complicated model, but is not likely to be useful for fitting the model to data. Certainly, the authors did not illustrate any such application. Secondly, the authors do not provide any sort of general algorithm but rather give a demonstration of a single example of fitting one specific reduced model to one specific conductance-based model.

      We thank the reviewer for this critical assessment. It is true that we demonstrate our approach using as target the bifurcation diagram of a more realistic model. In principle, the method would be applicable to experimental systems if parameter space is sampled under appropriate experimental conditions, e.g., by recording at different extracellular potassium concentrations. However, this is challenging: generating reliable” experimental bifurcation diagrams” would require tightly controlled experimental repetitions under various conditions, particularly for reconstructing two-dimensional bifurcation diagrams. We hope that this work encourages the development of such methods. We have added a discussion of this point; see the paragraph starting at line 549.

      We have also included a general algorithm (Table I) and provide a second example illustrating the procedure (Supplementary Fig. S2).

      Finally, the main idea of the paper seems to me to be a natural descendant of the chain of reasoning, starting from Rinzel - continuing through Bertram; Golubitsky/Kaper/Josic; Izhikevich; and others - that a fundamental way to think about neuronal models, especially those involving bursting dynamics, is in terms of their bifurcation structure. According to this line of reasoning, two models are “the same” if they have the same bifurcation structure. Thus, it becomes natural to fit a reduced model to a more complicated model based on the bifurcation structure. The authors deserve credit for recognizing and implementing this step, and their work may be a useful example to the community. But the manuscript should have described and cited this chain of works to put the current study in the correct context.

      We have added a paragraph in the Discussion section (starting at line 517) to better situate the manuscript within the relevant literature and to explicitly acknowledge the chain of work.

      Reviewer #1 (Recommendations for the authors):

      Please see my public review. In line with my comments, I recommend that the authors either (a) provide a general algorithm for fitting at least a class of reduced models (i.e., those that satisfy some general assumptions) to a class of bifurcation diagrams, or (b) provide at least one more example of implementing their method. Step (b) would not need to be done to the same degree of thoroughness as the example they provided (e.g., the PRCs and synchrony need not be considered), but to me, this step would be very important if (a) is impractical. Otherwise, the paper should probably be rewritten to de-emphasize the message that this is a general method; instead, this should be a paper about specifically fitting the QIF (with potassium dynamics) to the Wang-Buzsaki model (with potassium dynamics).

      We provide a general algorithm for deriving a quadratic integrate-and-fire model with dependence on a biophysical parameter by fitting the bifurcation structure of a given class I conductance based neuron model near an SNL bifurcation induced by this parameter; see Table I.

      In addition, we provide a second example of the reduction procedure: motivated by Hesse et al. (Nature Communications, 10.1038/s41467-022-31195-6, 2022), we derive a QIF model that captures dependence on temperature instead of potassium concentration; see Supplementary Fig. S2.

      Not surprisingly, I also think it’s essential that the authors describe and cite the chain of works on thinking of neuronal models in equivalence classes based on bifurcation diagrams, and make clear that this paper builds on the ideas set forth in that chain.

      We thank the reviewer for this comment. As mentioned above, we have added a paragraph in the Discussion section, starting at line 517, to acknowledge this chain of work.

      Also, the authors should make clear that their method is not one for fitting a model directly to data, which will require rewriting at least the first paragraph of their Discussion section.

      We thank the reviewer for helping us make our manuscript clearer. To avoid confusion, we have clarified this point already in the Introduction (see lines 52-57) and have included a new paragraph in the Discussion (starting at line 549).

      Other specific corrections are:

      (1) Typos should be fixed, as the paper has several. The first line of the abstract has one (“concentrations” → “concentration”), for starters. “Original model” on pg. 3 is missing “be” in “can defined”. “ceases” → “cease” on pg. 13. “nerons” → “neurons” on pg. 19. “standart” → “standard” pg. 24.

      Done. Additional typos were also corrected.

      (2) The abstract mentions “consequences in networks” in its second sentence. This is misleading because studying network dynamics is not at all the main emphasis of the paper, but rather a corollary application of the main ideas, so some restructuring of the abstract is needed. Similarly, the final abstract sentence overstates somewhat what was done with studying synchronization and should be rewritten more precisely.

      We have restructured the abstract accordingly.

      (3) For readers who are interested in the ideas here but not familiar with the QIF model, it will be very difficult to follow the first paragraph of Results. Elementary aspects of QIF dynamics should be explained here (e.g., what is the saddle-node bifurcation), and a basic figure panel about this should be included in Figure 1.

      We added a supplementary figure (Figure S1) adapted from Izhikevich for readers who might not be familiar with the QIF model.

      (4) Bottom lines of page 3 should be reworded to make clear that the slow variables are averaged over each member of a family of fast subsystem limit cycles. Also, “one limit action potential cycle” is an awkward phrase.

      We rephrased this sentence (see paragraph starting at line 140).

      (5) Text under system (1) – why isn’t c mentioned? Also, references to Figure 3 should be to Figure 2 here. And authors should state what they mean by “target model” and be clear about whether it includes potassium dynamics and/or pump current.

      c scales the parabola corresponding to the branch of fixed points, given by c(I<sub>app</sub> − I<sub>SN,0</sub> − I<sub>pump</sub>) = −a(ν − ν<sub>SN</sub>)<sup>2</sup>. Consequently, it also affects the position of the homoclinic bifurcation: In the previous version of the manuscript, c was inadvertently omitted from the expression for the branch of fixed points; this has now been corrected. See paragraph starting at line 154.

      Figure references have been corrected.

      By target model, we mean the conductance-based model that includes potassium dynamics and a pump current, in our case System 4. We clarified this point at the beginning of the Results section (see paragraph starting at line 116). Throughout the manuscript, we now explicitly indicate when we refer only to its fast subsystem and whether the pump current is included. In particular, note that the parameter derivation shown in Figure 4 is performed on the fast subsystem of the target model, in the absence of the pump current. I<sub>pump</sub> can be considered as a potassium-dependent contribution to the applied current, and can be added a posteriori to the QIF model. See paragraph starting at line 162.

      (6) Next par: is the “saddle-node bifurcation” that with I<sub>app</sub> as bifurcation parameter? Please clarify.

      Yes, it is the saddle-node bifurcation with I<sub>app</sub> as bifurcation parameter. We have clarified this in the manuscript; see the paragraph starting at line 170.

      (7) Bottom pg. 5: does “beyond” mean above? below?

      We meant above (larger values of ). In the text, we have replaced “beyond” with “larger than”. See paragraph starting at line 178.

      (8) Formula for v<sub>r</sub> at top of page 6: Please specify what formula for I<sub>pump</sub> is being used here.

      The formula for I<sub>pump</sub> is given in Eq. 5d. We are using the same formula throughout the paper.

      Note that to clarify the reduction procedure, we derive QIF parameters to match the bifurcation diagram with respect to the applied current of the fast subsystem of the target model when I<sub>pump</sub> = 0. Reintroducing I<sub>pump</sub> produces the same horizontal shift in this bifurcation diagram for both the QIF and Wang–Buzsáki versions.

      We have restructured the paragraph starting at line 178 to clarify these aspects.

      (9) Three lines below this: I don’t understand what “matching...is appreciable” and “in the continuity of...” mean. Please revise and also explain why a closer matching of v<sub>r</sub> to the min voltages in Figure 4d was not used, and exactly how the v<sub>r</sub> that is shown was chosen.

      With “matching...is appreciable”, we meant that values assigned to v<sub>r</sub> should be close to the minimum voltage values reached during spiking. With ”in the continuity of...”, we meant that when .(SNIC case), we choose v<sub>r</sub> by extrapolating the linear fit performed on the values of v<sub>r</sub> assigned when , (homoclinic case).

      When , v<sub>r</sub> was chosen so that the homoclinic bifurcation occurs at the same value of applied current as in the fast subsystem of the target model. This is explained in the paragraph starting at line 178 (see Eq. 2). This criterion also allows the minimum voltage values reached during spiking to be captured reasonably well (compare the green dotted line and the purple solid line in panel e of Figure 4).

      We have rewritten the paragraph starting at line 186 to clarify these aspects.

      (10) Bottom page 6 - reference to Figure 3e should be 4e. Also, the text mentions the shrinkage of spike amplitude, but the figure shows that vth increases over most of the K+ range before decreasing, so a correction is needed.

      We corrected the figure reference.

      The maximal voltage of the limit cycles of the target model’s fast subsystem (upper purple curve in Figure 4e) increases slightly between and , by less than 1mV. It then decreases by about 27mV before the fold of limit cycles. The sigmoidal function vth () allows us to capture this substantial decrease in the QIF model. The small preceding increase is not captured. We reformulated the text to avoid confusion (see paragraph starting at line 209).

      (11) Figure 4d: Why is E<sub>K</sub> plotted here? It should be mentioned in the caption and text. More generally, the caption for Figure 4e should be expanded to mention what the purple curves are, what is the black curve for K < K<sub>SNL</sub>, and what the other structures shown are. Finally, the text describes that theSNIC/SNL/Hom is determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>, so it’s not clear what is I<sub>app,SNL</sub> in the caption - please clarify.

      In conductance-based models, higher weakens the potassium concentration gradient, thereby raising E<sub>K</sub>. The sodium and potassium reversal potentials typically bound voltage oscillations during spiking (see for example Chander and Chakravarthy, PLOS ONE, 10.1371/journal.pone.0048802, 2012), so the minimum voltage of spikes is expected to be higher when is larger. We plotted E<sub>K</sub> in Figure 4d to show that the increase of the reset voltage v<sub>r</sub> at larger reflects this effect in the QIF version of the model. We clarified this in the caption of Figure 4 and in the text (see paragraph starting at line 199).

      Purple curves show families of limit cycles, while black curves, including the one for , show families of fixed points. The green curve shows the linear fit of v<sub>r</sub> from panel d. All these have now been included in the legend.

      In the QIF model, the onset bifurcation (SNIC, SNL, or homoclinic) is indeed determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>. Figure 4e shows the bifurcation diagram of the fast subsystem of the conductance-based model (Wang-Buzsáki). This is a bifurcation diagram with respect to , for a fixed value of applied current. We chose to fix I<sub>app</sub> at its value at the SNL bifurcation, denoted I<sub>app,SNL</sub>. I<sub>app,SNL</sub> is near 0.22 (see Figure 3a).

      (12) Eqn. (2a): Shouldn’t v<sub>SN</sub> depend on potassium like I<sub>SN</sub> and I<sub>pump</sub> do? What is the formula for Ipump there? Why isn’t the RHS of (2b) dependent on v as in the original model? Please clarify.

      For simplicity, we did not include a potassium dependence for v<sub>SN</sub> in the QIF model. Instead, we set it to its value at the SNL bifurcation (Figure 4b). This is explained in the paragraph starting at line 199: “We notice that v<sub>SN</sub> and the normal form coefficient a are relatively conserved in this interval. We fix them to their value at [K<sup>+</sup>]<sub>o,SNL</sub>.”

      The formula for I<sub>pump</sub> is given in Eq. 5d. The potassium dynamics depends on the voltage via the reset rule in Eq. 3d. This allows us to capture the small increments in at each action potential in the original model (see for example Figure 6c,h). We have clarified these two points in the manuscript (see paragraph starting at line 223).

      (13) Figure 3: Please indicate the criticality of the Hopf bifurcations shown.

      The legend of Fig. 3 now indicates that the Hopf bifurcations are subcritical. The same clarification has been added to the following figures as well.

      (14) Bottom pg. 9: Why isn’t there an I<sub>pump</sub> term as in eqn. (6a)? Please clarify. 

      You are correct, the I<sub>pump</sub> term should be included in the equation for the averaged slow subsystem of the QIF model (see paragraph starting at line 243); it was accidentally omitted in the manuscript. Thank you for pointing this out.

      (15) Top pg. 11: It’s important to reference the slow averaged dynamics here, which allows K+ to increase. Also, this first paragraph should already explain that this averaged dynamics is only relevant along the family of FS periodic orbits, not during the recovery when the FS has a branch of stable equilibria.

      The averaged slow subsystem is indeed only relevant along families of limit cycles of the fast subsystem. Along families of equilibria, averaging is not necessary and the standard slow subsystem can be used. We now explicitly define this standard slow subsystem for both Wang-Buzsáki and the QIF model (see paragraphs starting at lines 241 and 628). In panels e and j of Figure 6, both systems are now represented.

      In the paragraph starting at line 273, we now refer to the averaged slow subsystem to explain the overall increase of during bursts (purple curves in Fig. 6e,j), and to the standard slow subsystem to explain the decrease of during quiescent phases (black curves).

      (16) Pg. 11, par 3: This is unnecessarily confusing. Please try to reword and clarify this paragraph.

      We have simplified this paragraph (starting at line 293). The key point is that the reduction to the averaged slow subsystem is not valid too close to the homoclinic bifurcation.

      (17) Pg. 13, end of Scenario 2: Is there any evidence this is a canard effect and not a noise effect? If so, please mention the evidence; otherwise, perhaps take this out.

      What happens there appears to be a noise-induced canard effect: in the beginning of the burst, the system follows a family of stable limit cycles of the fast subsystem, i.e. a stable object. However, at some point, noise induces a transition to a portion of trajectory where the system evolves near the saddle branch, i.e. a repelling object, for a substantial amount of time. Such phenomena have been thoroughly investigated in the literature, and can also be obtained in a deterministic way; see for example Marin et al. (Physical Review E 90, 042718, 2014). Bursting traces similar to the one in Figure 7c are observed experimentally (see, for example, Figure 4c of Marin et al.), which is why we considered it worth mentioning. We have revised the paragraph starting at line 332 to clarify this point.

      (18) I only see 4 curves in Figure 9a,c, but the legend has 5. Are two on top of each other? Please clarify.

      Yes, the curve for = 7.21mM lies beneath the curve for = 5.21mM. This has been clarified in the figure caption.

      (19) Text should note that the QIF iPRC does not develop a negative region at high K+ and high phase, as WB iPRC does.

      We have updated the paragraph starting at line 409 to mention this.

      (20) Pg. 16, line 4: “at the network scale” is cryptic - a more precise phrase would be preferable.

      We have reformulated the sentence to clarify its meaning (see paragraph starting at line 380).

      (21) Pg. 16, line 9: Reordering of words could make this clearer.

      Done (see paragraph starting at line 385).

      (22) Pg. 16: I am confused by line 14 because the big changes in the iPRC in Figure 9c do not align with the spike phase in Figure 9d. Please clarify what is meant here.

      In the QIF model, at a given phase, the iPRC is the inverse of the slope of the voltage trace as a function of phase. Flatter slopes in Figure 9d therefore correspond to larger iPRC values in Figure 9c. This is illustrated in Figure S5. We have revised the paragraph starting at line 393 to make this point clearer.

      (23) Pg. 16: Please clarify what is meant by a “delta synapse”.

      By “delta synapse,” we meant a configuration in which each spike induces an instantaneous voltage jump in the postsynaptic neuron, modeled using the Dirac delta distribution. We have replaced the term “delta synapse” with “pulse-coupled,” which is more commonly used in the literature, and have added a clarification at its first occurrence in the manuscript.

      (24) Discussion, line 2: Delete comma.

      Done.

      (25) Importantly, as noted above, the first par. needs to be rewritten since the presented method won’t work directly from data or from a target model for which most of the parameters, and hence the bifurcation diagram, are not known.

      As mentioned above, we have included a new paragraph in the Discussion, starting at line 549, to clarify this point.

      (26) Pg. 18: ”Originally” → ”Typically”, perhaps?

      Done.

      (27) Note the work of Marder et al. on temperature-related neural variability.

      We thank the reviewer for this comment. We have added two relevant references from the work of Marder and colleagues addressing temperature-dependent neural variability and ionic concentrations in our manuscript (see the sentence starting on line 537).

      (28) Pg. 20: Cut the ”Potassium dynamics and network models” subsection since it does not add anything substantive as written (or else expand it and include it in the subsection below).

      We have expanded this paragraph and incorporated it into the subsequent subsection, as suggested by the reviewer.

      (29) Finally, it’s a bit confusing that the authors refer to the potassium concentration as a slow variable yet have an instantaneous jump in this quantity at reset in their QIF model (i.e., instantaneous is VERY fast). Some explanation about this should be provided. Do they make the general assumption that ∆<sub>K</sub> is small, for example, such that this reset reflects the slow nature of K+ evolution (i.e., during the reset period, K+ would only change slowly, and hence by a small amount)?

      Yes, ∆<sub>K</sub> is chosen to be small, to capture the behavior of the original model (compare for example panels c and h of Figure 6). As a result, in the QIF model, in the same way as in the original model, despite the fact that the dynamics of includes a fast component, on average evolves slowly. By using the averaging method, we can determine whether overall increases or decreases.

      We clarified this in the manuscript, in the paragraph starting at line 243.

      Reviewer #2 (Public review):

      Summary:

      The authors derive an integrate-and-fire model to describe the dynamics of a more complex Wang-Buzsaki model and compare the two models. A detailed discussion of bifurcation schemes in both models is convincing and allows us to evaluate the simpler model.

      Strengths:

      The idea is interesting, and the mathematical approach appears to be convincing. In addition, differences between the simple and original models are also discussed.

      Weaknesses:

      A comparison to experimental data is necessary to support the theoretical work.

      As mentioned above in our answer to Reviewer 1, we demonstrate our method using as target the bifurcation diagram of a more realistic neuron model. Ideally, one would want to derive phenomenological models that capture bifurcation structures obtained from data. However, this is challenging and beyond the scope of the present study. We hope that this work encourages the development of such methods. We have revised the Introduction (see lines 52-57) and added a paragraph in the Discussion (see the paragraph starting at line 549) addressing this point.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is well-structured; however, it appears that it has been edited with less care. Please see comments below:

      (1) Page 2: “A third bifurcation, the saddle-homoclinic orbit (HOM) bifurcation,”: provide a reference for the bifurcation.

      We have added a reference to the book by Izhikevich (see paragraph starting at line 63).

      We have added additional references in the Introduction that we considered helpful.

      (2) Page 3: “while a larger concentrations it is mediated by...”: remove “it”.

      There was indeed a typo in this sentence. The intended phrasing is: “while at larger concentrations it is mediated by...”. We have corrected it accordingly (paragraph starting at line 134).

      (3) Figure 2, caption: “dashed lines for unstable branches”: this is a dotted line.

      Corrected to “dotted lines”. Thank you.

      (4) Page 4: “is smaller than vSN (Fig. 3c),”: this figure panel does not exist, as well as the Fig.3d referred to afterwards. Please correct.

      We intended to refer to Fig. 2. Figure references have been corrected. See paragraph starting at line 154.

      (5) Page 6: “Fig. 3a-d shows ISN,0, vSN and a for [K]+o between 4 and 16 mM”: Fig.3c+d do not exist, please correct. Similar comment to “the absence of pump (Fig. 3e).” on the same page.

      We intended to refer to Fig. 4. Figure references have been corrected (paragraphs starting at lines 199 and 209).

      (6) Page 8: “(panel A)” → ”panel (a)”.

      Done.

      (7) Figure 4e: What is the meaning of the green dotted curve?

      This curve represents the linear fit of the reset voltage v<sub>r</sub> from panel d of Fig. 4, to show that the minimal voltage values of the limit cycles are also well captured. We have added this curve to the legend and included a brief explanation in the figure caption.

    1. eLife Assessment

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants and the authors provide solid evidence for contrasting symbiont-associated phenotypes under controlled conditions: Rickettsiella increases plant damage while reducing wing production and dispersal, whereas Regiella reduces feeding damage and, at some time points, aphid population growth. The stable establishment of these novel symbiont-host associations and complementary experiments spanning individual aphids, whole plants, populations, and mesocosms are notable strengths and the revised manuscript better clarifies the experimental approaches and the context dependence of symbiont effects; however, the underlying mechanisms remain unresolved, horizontal transmission is inferred rather than directly demonstrated, and limited replication and temporal variability constrain some population-level conclusions. Potential applications to pest management therefore require further validation under field conditions. The study will interest researchers working on insect symbiosis, plant-insect interactions, and biologically based pest management.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      It is a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

    4. Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their limitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      [1] Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.

      [2] Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.

      [3] Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.

      [4] James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants. The authors provide solid evidence that the two symbionts can generate contrasting phenotypes: Rickettsiella increases plant damage while reducing wing formation and dispersal, whereas Regiella reduces aphid population growth and feeding damage. The successful establishment and stable transmission of these novel symbiont-host associations, combined with experiments spanning individual, whole-plant, population and mesocosm scales, are notable strengths of the work; however, the mechanisms underlying these effects remain unresolved, evidence for horizontal transmission is indirect, and some population-level conclusions rely on relatively small sample sizes or effects that are not consistently detected across time points, and therefore the potential application of these findings to pest management remains promising but speculative. The study will be of broad interest to researchers working on insect symbiosis, plant-insect interactions, and biologically based approaches to pest management.

      We have made revisions to the manuscript to cover issues raised around sample numbers and mechanisms. We appreciate that we have not been able to finalize the exact mechanism underlying plant damage effects. We note that while comparisons of population performance were limited by the number of populations we could feasibly maintain; sample sizes were substantial in some of the other experiments. We also do substantiate effects through a combination of experimental approaches that start with controlled conditions and then encompass more complex environments.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

      We thank the reviewer for recognizing the strengths of the work and for providing extensive comments.

      Weaknesses:

      There are some major aspects of this paper that I thought could be strengthened. My main concern is that the manuscript feels broader than it is conceptually focused. A wide range of outcomes is measured, which gives the study breadth, but it also makes the central question harder to identify. As written, the paper reads more strongly as a proof-of-principle demonstration of ecologically relevant phenotypes than as a tightly framed test of a specific biological idea.

      The broad range of tests in our work was intentional. Rather than relying on a single experimental approach or scale, we designed the study using multiple complementary approaches to independently evaluate the effects of endosymbionts. Thus, our conclusions are supported across multiple experimental contexts rather than by a single experiment or scale. For example, the effects on plant damage were consistent across different host plants, including wheat and barley (Figures 1 & S1), and across different experimental scales, ranging from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level, using single aphids maintained on individual plants under favorable conditions (Figures S5 & S6), and at the population level under crowding conditions (Figures 3 & 4).

      Despite the breadth of measurements, the study is focused on establishing robust evidence for the contrasting effects of the two endosymbionts on aphid dispersal and plant feeding damage. To help address this concern, we have moved the summary table from the Supplementary Materials to the main text (now Table 1), which provides an overview of the experimental approaches and main findings and should help readers more clearly see how the different experiments are connected.

      A second issue is that the biological basis of the reported phenotypes remains less developed than the phenotypic description itself. The authors make a genuine effort to address mechanism through JA, JA-Ile, SA, and metabolomic profiling, but these analyses only partially explain the main results. The negative result for the canonical defense markers is informative, yet it still leaves a substantial gap between the observed variation in plant damage and the processes responsible for it.

      Our analyses of JA, JA-Ile, SA, and the metabolomic profiles provide some initial insights, but we appreciate that they do not fully explain the differences in plant damage observed between treatments. The primary aim of this study was to evaluate the phenotypic effects of the endosymbionts and their potential for agricultural application, rather than to provide a comprehensive mechanistic explanation. We have accordingly avoided overinterpreting the mechanistic results and now explicitly state that elucidating the underlying biological mechanisms will be an important direction for future research in Discussion section.

      I also think some caution is needed in how the two symbionts are compared. The authors explain why some follow-up experiments were designed differently for Rickettsiella and Regiella, and that rationale is understandable. Still, because the downstream assays were not fully matched, the paper is strongest when each symbiont is interpreted on its own terms rather than as a strict comparison.

      Our initial plant-damage experiment was designed as a first comparison to test whether different endosymbionts can have diverse and contrasting effects on their aphid host population and plant damage. We then investigated the individual phenotypes of each endosymbiont in greater detail, particularly in relation to their potential agricultural applications. Specifically, our results suggest that Regiella may reduce plant damage, whereas Rickettsiella may reduce dispersal. Based on these early findings, some further experiments were conducted with slightly different experimental setups. Nevertheless, many of the experiments conducted for the two endosymbionts were broadly comparable. We designed the additional experiment carried out only with Rickettsiella to test whether reduced alate production observed in our earlier experiments also translated into reduced dispersal at the population level. An equivalent experiment with Regiella was not undertaken because we failed to detect an effect of this endosymbiont on alate frequency in our preceding experiments. We have clarified this rationale in the Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”).

      Overall, I would suggest softening the Significance Statement so that it more clearly reflects what is directly shown here, namely that introduced symbionts can alter plant damage and dispersal-related phenotypes under controlled conditions, rather than implying that the study directly tests management utility in agricultural settings.

      We have done this in the Significance Statement. We appreciate that the current experiments have been carried out under controlled conditions, rather than in agricultural settings. Pending permit approval, we are currently planning to extend this work to contained field settings to test whether effects on plant damage and dispersal ability are also observed under less controlled conditions.

      Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      One thing that I struggled a little with was the rapid spread of Regiella in the shared plant experiments. Possibly this could be attributed to an increased reproductive output (due to faster developmental time, and/or an increase in fecundity) or efficient horizontal transmission. However, the other experiments performed indicate a slight negative impact (Figure 4a) or no influence (Figure 4C, 4D, Figure S6) of Regiella infection on host fitness. Given these other results, it seems that Regiella must spread fairly efficiently between hosts, which comes as a surprise, and there are very few examples of horizontal transmission of Regiella like this in the literature. The manuscript would benefit from a clear and direct demonstration of horizontal transmission, rather than it being inferred indirectly. The similar spread observed in the Rickettsiella mixed cages is less surprising, because there are several examples where this has been demonstrated.

      We agree that the rapid spread of Regiella in the shared-plant experiments cannot be readily explained by host fitness alone, though we have noted fitness benefits of Regiella in a different transinfection in oat aphids (Yu et al., 2025). In a previous study with transinfected green peach aphids and despite a substantial fitness cost, we found that Rickettsiella can spread relatively rapidly in a population and show high stability (Gu et al., 2023) and perhaps transmission of Regiella follows a similar pathway. However, whereas Rickettsiella may spread through plant tissues, Regiella showed relatively low levels of horizontal transmission through this pathway.

      We certainly agree that more work is required to establish the mechanism and dynamics of horizontal transmission in this system. Rather than focusing on mechanism, our objective here was to examine endosymbiont spread where intact plants were available and where there was a mixed aphid population. Note that we also conducted an additional experiment in which Regiella-infected aphids were present at a frequency of only 10% of the initial population, and in this situation Regiella nevertheless still increased in frequency including to a low Cp value in most (8/10) replicates, highlighting its persistence and potential to increase in populations.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Yu et al., A persistent bacterial Regiella transinfection in the bird cherry-oat aphid Rhopalosiphum padi increasing host fitness and decreasing plant virus transmission. Pest Manag Sci 81, 2791-2799 (2025).

      It is also a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

      We conducted the final mesocosm-dispersal experiment specifically with Rickettsiella because our earlier individual-plant experiments had already shown a clear reduction in alate production in Rickettsiella-infected aphids, together with effects on plant damage (Figure S1I) and population growth (Figure 3C). In contrast, we did not detect a significant difference in alate frequency between Regiella-infected and wild type aphid strains in the similar set up experiments (Figure S1I and Figure 4C). We therefore designed the additional experiment to test whether the reduced alate production observed with Rickettsiella also translated into reduced dispersal at the population level. We did not conduct the same experiment with Regiella because there was no difference in alate frequency in our earlier experiments. We have also added this explanation in Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”). We do appreciate however that future experiments on dispersal of Regiella will be worthwhile resources permitting.

      Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      We acknowledge that more replication is always desirable, but we would also argue that significant effects were not marginal and replication was substantial in many cases (e. g. 9-10 replicate plants per damage treatment evaluation). We were also focused on using multiple experimental approaches and scales to independently and repeatedly evaluate the effects of endosymbionts on aphid fitness, wing development, plant damage, and aphid dispersal. Thus, conclusions are not based on a single experiment or experimental scale. For plant damage, for example, we observed consistent effects across different host plants, including wheat and barley (Figure 1 and Figure S1), as well as across different experimental scales, from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level using single aphids maintained on individual plants under favorable conditions (60 replicates per treatment) (Figures S5 & S6) and at the population level under crowding conditions (Figures 3 & 4). These complementary experimental designs allowed us to examine whether the observed phenotypes were consistent across different environmental and population contexts. We did face challenges in high levels of replication of independent aphid strains in population cage experiments but attempted to replicate as much as possible given the resources that were available.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their dlimitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      (1) Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.

      (2) Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.

      (3) Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.

      (4) James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      We agree that qPCR-based measurements of endosymbiont gene copy number have limitations even if they are the standard approach used in most studies. In the current set of experiments, qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples and the only one feasible given the number of monitoring events and samples required. Nevertheless, we acknowledge the limitations of this approach (and have pointed this out ourselves in a recent COIS paper – Hoffmann et al 2026). We now mention this under further work and provide a couple of references (see Discussion).

      Reference:

      Hoffmann, A. A., Yang, Q. and P. A. Ross. Aphid endosymbionts revisited: molecular detection, diversity, and population dynamics. Curr Op Insect Sci (in press). (2026)

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Overall, I found the study interesting and worthwhile, particularly because it shows that novel symbiont associations can generate contrasting phenotypes in an important pest species. My main suggestion would be to sharpen the framing of the paper and to bring the mechanistic and applied discussion into slightly closer alignment with the current evidence base. With those points addressed, I think the manuscript would read as a clearer and more balanced contribution.

      We have rephrased the Discussion around the mechanistic component of the work to sharpen this and link more directly to evidence. For instance:

      “Despite this, our metabolomic analyses suggest that endosymbionts in D. noxia may influence some other aspects of wheat metabolism in a spatially structured way. Within aphid-feeding areas, wheat exposed to Rickettsiella aphids showed increased valine and decreased 3-phenyllactic acid. Changes in valine have been reported in plants responding to herbivory and other forms of stress (43-45), while 3‑phenyllactic acid has been associated with antimicrobial and defense-related activity (46). In non-feeding areas, wheat exposed to Regiella aphids showed reduced urea and increased pantothenic acid. These changes may reflect differences in nitrogen metabolism and allocation (47), and in metabolic processes involving pantothenic acid (48) respectively. However, the present metabolomic data do not establish the functional consequences or causal mechanisms underlying these changes. They indicate that aphids carrying different endosymbionts are associated with some distinct metabolic responses in wheat, including responses that differ between aphid-feeding and non-feeding areas. Our findings are consistent with previous work showing that phloem‑feeding insects can induce changes in plant metabolites in response to herbivory (49, 50) and provide a basis for future studies to determine how endosymbionts influence aphid-induced plant responses and contribute to contrasting plant phenotypes.”

      We have not undertaken a complete reframing of the paper but further emphasized the focus on phenotypic contrasts in a few places including incorporating some changes to the comments below.

      (2) Lines 130-135: Please clarify more explicitly whether the main aim of the paper is to test a specific biological hypothesis about endosymbiont-mediated aphid-plant interactions or to provide a broader proof-of-principle survey of symbiont-associated phenotypes. As it stands, the framing moves between multitrophic biology and pest-management relevance, which makes the central conceptual contribution harder to identify.

      We have rephrased this sentence as “By integrating these factors, we investigate how endosymbiont infection influences aphid fitness and aphid–plant interactions, providing insights into the ecological consequences of novel microbial associations and their potential relevance to sustainable pest management.”

      (3) Line 205: For the metabolomic analysis, please consider adding a formal multivariate test for strain effects, especially for the comparisons shown in Figure 2C and 2D, or otherwise interpret the PCA more cautiously as descriptive rather than inferential.

      We have emphasized the descriptive component and rephrased this sentence as “However, within either area, there was no clear separation among aphid strains (Figures 2C & 2D), suggesting broadly similar metabolomic profiles among strains of the same aphid clone carrying different symbionts.”

      (4) Please clarify how the top and bottom feeding leaves were handled analytically in the analyses, and explain the rationale for collapsing them into a single "feeding area" category. If possible, it would be helpful to show whether leaf position itself influenced the plant-response patterns.

      We combined the upper and lower leaves together to provide a representative measure of the plant-level responses, rather than focusing on responses at a particular leaf position. This approach was consistent with the main objective of our study, which was to investigate whole plant responses to aphids hosting different endosymbionts, rather than differences in responses among different plant parts. In addition, combining the two portions provided sufficient plant material for the GC-MS analysis and helped ensure reliable metabolite measurements from the same material. Because the two leaf positions were combined prior to GC-MS analysis, we were unable to separately test the effect of leaf position on the metabolomic response in this dataset. We have clarified this point in the revised manuscript in Materials and methods section (“Plant defense responses”).

      (5) Line 228: The use of 19{degree sign}C and 25{degree sign}C is not unusual in aphid work, but it would still help the reader if the manuscript stated more explicitly why these two temperatures were chosen in this study.

      We selected 19°C and 25°C because they represent contrasting temperature conditions within the range suitable for Russian wheat aphid development, allowing us to assess whether temperature influences Rickettsiella transmission and population dynamics. In particular, our previous observations indicated differences in the rate of Rickettsiella spread between these temperature conditions (Gu et al., 2023).

      Reference:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      (6) Lines 390-392: It would help to discuss more explicitly how the relatively modest effects in the individual life-history assays relate to the clearer signals seen at the whole-plant and population level.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with endosymbiont infection. In contrast, we also examined aphid performance at the population level (Figures 3 & 4), where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress. Under these conditions, the effects of endosymbiont infection differed from those observed in the individual assays, with Rickettsiella-infected aphids showing greater population growth and Regiella-infected aphids showing reduced population growth.

      These results suggest that the effects of endosymbionts on aphid fitness are context-dependent and may become more pronounced as population density increases and density-dependent stress develops. Thus, relatively modest effects observed at the individual level where experiments are often undertaken may not translate to differences at the population level, and plant level effects may subsequently influence feeding pressure and plant damage. We have added this perspective in the Results and Discussion sections.

      (7) Line 678 & 688: In both whole-plant experiments, please explain how 3 and 4 replicate plants were selected.

      The replicate plants were randomly selected from the available plants for each treatment to minimize potential selection bias. We have clarified this procedure in the revised manuscript in Materials and methods section.

      (8) Lines 686-690 / Figure 4: In the Regiella whole-plant experiment, the Methods state that 16 plants were established per treatment and that 4 plants per treatment were removed at each time point (days 4, 8, 12, and 16). However, in Figure 4B-D, day 12 appears to include 5 data points. Please clarify this apparent mismatch between the described sampling scheme and the data shown in the figure.

      We thank the reviewer for pointing out this and we have corrected this mistake. We initially established 16 plants for the wild type and 17 plants for the Regiella treatment. Four plants per treatment were originally planned to be sampled at each time point (Days 4, 8, 12, and 16). However, because Day 12 was a key time point at which an obvious difference in plant damage was observed between the treatments, we selected one additional plant each treatment for measurement at Day 12, resulting in five data points for this treatment at that time point. The remaining plant was therefore measured at Day 16. We have clarified the sampling procedure in the revised Materials and methods section and figure legend.

      (9) Figure 2A and Figure 4A: These schematics are helpful overall, but the brown supporting sticks stand out quite strongly and may make the panels a little harder to interpret at first glance. I wonder whether they could be simplified, made less prominent, or replaced with photographs of the actual setup if those are available.

      We have revised Figures 2A and 4A to simplify the supporting structures and reduce their visual prominence.

      Reviewer #2 (Recommendations for the authors):

      Some of the statistics were not entirely clear to me, particularly the tests reported in the results which differ from what is described in the figure legends:

      (1) Lines 306-308: "Total nymph numbers were higher on Rickettsiella aphids from Day 21 to Day 25" and indicates that this is based on GLM testing, but the figure does not indicate statistical significance, and the figure legend states independent sample t-tests were used. Similar comment for the following paragraph and corresponding figure.

      The GLMs were used to test the overall patterns in nymph numbers across the relevant time periods, including Days 21–25, rather than testing each time point independently. We also conducted independent-sample t-tests to assess differences between treatments at individual time points. We have clarified this distinction in the Statistical section.

      (2) Line 43: Does not seem like the appropriate reference (reference is on plant virus transmission, not salivary toxins).

      We thank the reviewer for pointing this out. We have removed this reference and replaced it with reference 34 (Luna et al., 2018) that directly supports the statement regarding aphid salivary toxins.

      Reference:

      Luna et al., Bacteria associated with Russian wheat aphid (Diuraphis noxia) enhance aphid virulence to wheat. Phytobiomes J 2, 151-164 (2018).

      Reviewer #3 (Recommendations for the authors):

      Minor editorial comments:

      (1) Figure S10 - the figure legend needs improvement as the current version does not help the reader understand the figure. Please also include a key.

      We have changed the figure legend with reference to our aim, and also explained use of the Cp values. “Figure S10. Rickettsiella Cp values in (A) routine screening of laboratory Rickettsiella colonies and (B) the mixed cage experiment assessing endosymbiont frequency changes over time at 19 °C and 25 °C. The dark red area represents overlap between the 19°C and 25°C experiments. Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA. This figure shows the typical range of Cp values observed in our laboratory Rickettsiella -infected aphid colonies. We used this range as a reference for identifying aphids that acquired Rickettsiella through horizontal transmission, as horizontally infected aphids generally showed much higher Cp values than vertically infected aphids.”

      (2) Define Cp.

      We have explained it as “Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA.”

      (3) Supplemental Information: Line 113 - T is missing from Table.

      This has been added.

      (4) Move Table S2 to the main paper - this table provides a useful summary of the work.

      This has been moved.

      Main Manuscript:

      (1) Line 145: Serratia was not detected at G0 - was it detected later? It seems possible that titer could be very low to begin and increase in later generations; please clarify.

      We have previously monitored the aphid populations for the presence of Serratia transinfected from the same donor resource across subsequent generations, and Serratia was not detected at any later generation. We also did not detect Serratia in the other aphid species we examined, including green peach aphids (Gu et al., 2023 & 2025) and oat aphids (Yang et al., 2026). Therefore, we believe that the absence of Serratia at G0 was not due to a very low initial titer followed by an increase in later generations but instead that Serratia had been lost from the aphid population.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Gu et al., Transinfections of the endosymbiont Rickettsiella viridis in different Myzus persicae (Hemiptera: Aphididae) clones show consistent deleterious effects and stable transmission. J Econ Entomol 118, 1544-1552 (2025).

      Yang et al., A Rickettsiella transinfection in Rhopalosiphum padi reduces fitness and alate production but not plant virus transmission. Pest Man Sci 82, 3894-3906 (2026).

      (2) Line 206: "different aphid strain" - I learned from the manuscript that a single clone of D. noxia is found in Australia. Further, from my reading of the manuscript, I understand that one clonal isolate was propagated and then infected with symbionts. I think it is important to reword this sentence so that it is clear that the aphid genetic background is held constant, and that the only differences here are the presence or absence of the different symbionts. My reaction to this sentence was that you are working with the same aphid strain hosting different symbionts.

      We have clarified it by adding “the same aphid clone carrying different symbionts” after the different aphid strains.

      (3) Measuring symbiont "density" is a tricky thing to do; I explain this above in the public review. I suggest considering some revisions to the manuscript to be sure that you accurately reflect what has been measured and what can reasonably be inferred from using qPCR to measure gene copy number.

      We agree that qPCR-based measurements of symbiont gene copy number have limitations. In this experiment, we had a relatively large number of samples, and qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples. While we acknowledge the limitations of this approach, the relative differences in gene copy number can still provide an indication of variation in endosymbiont abundance and infection status among treatments. Unfortunately, other approaches remain challenging given resource and expertise limitations.

      (4) I think that you may be undervaluing the results of the mixed infection experiments; I find them to be compelling. To me, the data suggest that the symbionts increase aphid fitness.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single wheat plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with symbiont infection. However, we also examined aphid performance under population-level conditions, where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress (Figures 3 & 4). Under these conditions, the effects on fitness were different from those observed in the individual assays. We therefore agree that our results suggest that symbionts can increase aphid fitness under some conditions, but that this effect may be context-dependent and can differ under population-level conditions where density-dependent stress occurs. We have also added this information to our Results section to make this clear to readers.

      (5) Lines 288-291: This sentence doesn’t make sense to me. What "minor fitness costs" are being referred to? If the infected lines are increasing in representation relative to the uninfected lines, that suggests that there are not fitness costs, but fitness advantages under the experimental conditions.

      We have rephrased it to “minor fitness effects” which we refer to the fitness test under favourable conditions.

      (6) The section that begins on line 293 - I find this part to not be particularly robust and suggest dropping it from the paper.

      We appreciate the reviewer’s concern regarding the robustness of this section. We included this experiment to monitor changes in aphid population over time and to help explain the differences in plant feeding damage observed between aphids carrying different endosymbionts. Importantly, we conducted this experiment using whole wheat plants to evaluate whether the patterns observed in our other experiments were also evident. We believe that these results provide important complementary evidence for interpreting the differences in plant damage among the aphids hosting different endosymbionts and therefore are relevant to the overall conclusions of the study. For this reason, we prefer to retain this section in the manuscript.

      (7) Line 298: three plants per time point - I do not think this sample size is reflected in the methods of the paper.

      We have mentioned in the method part with “At day 14, three replicate plants were randomly selected from the available plants for both treatments and we counted the total number of nymphs, alate adults, and apterous adults. This was repeated again at Days 21 and 25”.

      (8) Line 304: "the frequency of alates decreased as plant damage increased" - this seems to be counterintuitive!

      As plant damage increased, the total aphid population also increased, resulting in an increase in the absolute number of alates. However, the frequency (proportion) of alates decreased because the increase in the total aphid population was greater than the increase in the number of alates.

      (9) Line 404: "horizontal transmission through plant tissue and/or transfer via aphid contact or honeydew" - what evidence is there that this happens? I have not kept on top of the literature with respect to transmission of secondary symbionts in aphids, but back when I was very familiar with that literature, the data did not support transmission by any of those routes. If there is now evidence supporting transmission by these routes, please cite it here.

      Previous studies have provided experimental evidence that horizontal transmission of aphid-associated endosymbionts can occur through plants. For example, plant-mediated transmission has been demonstrated for direct detection of secondary endosymbiont in the plant tissue including Hamiltonella defensa (Li et al., 2018), Rickettsia (Shi et al., 2024) and Serratia symbiotica (Pons, et al., 2019a). Our previous research also demonstrates that Rickettsiella endosymbionts were detected in uninfected aphids after feeding by infected aphids regardless of physical contact (Gu et al., 2023). Endosymbionts have also been detected in aphid honeydew (Darby and Douglas, 2003) and some primary transmission route appears to be horizontal, through honeydew (faeces) and host plant phloem (Pons, et al., 2019a & 2019b; Perreau et al., 2021). Appropriate references have been added to the revised manuscript in the Discussion section.

      References:

      Li et al., Plant-mediated horizontal transmission of Hamiltonella defensa in the wheat aphid Sitobion miscanthi. J Agric Food Chem 66, 13367-13377 (2018).

      Shi et al., Rickettsia transmission from whitefly to plants benefits herbivore insects but is detrimental to fungal and viral pathogens. mBio 15, e02448-23 (2024).

      Pons et al., Circulation of the cultivable symbiont Serratia symbiotica in aphids is mediated by plants. Front Microbiol 10, 764 (2019a).

      Darby and Douglas, Elucidation of the transmission patterns of an insect-borne bacterium. Appl Environ Microbiol 69, 4403-4407 (2003).

      Pons et al., New insights into the nature of symbiotic associations in aphids: infection process, biological effects, and transmission mode of cultivable Serratia symbiotica bacteria. Appl Environ Microbiol 85, e02445-18 (2019b).

      Perreau et al., Vertical transmission at the pathogen-symbiont interface: Serratia symbiotica and aphids. mBio 12, e00359-21 (2021).

      (10) Line 511: revise to "to measure their relative densities relative to a host gene".

      We have revised it.

      (11) Throughout the manuscript, please replace "five aged-matched" with an accurate description of the aphids used in the experiment. Please pay particular attention to the figure legends. Simply state e.g. "five 10-day-old apterous females" etc.

      This has been replicated in the main manuscript and figure legends.

      (9) Line 727 - lowercase t for Tests.

      This has been added.

      (10) Line 738 - what happens when you don’t exclude the early time points? Do your significant results go away? Also, please define what is meant by "early time points".

      When the early time points (Day 14 for Rickettsiella and Day 4 for Regiella) were included in the analysis, the GLM still showed a significant effect of strain on alate production for Rickettsiella (F<sub>1,12</sub> = 26.434, P < 0.001). We excluded these early time points from the analysis presented in the manuscript because, at these stages, aphid population sizes were still similar between treatments. We also conducted independent-sample t-tests at individual time points and observed differences at the later time points, when aphid population sizes began to diverge. We therefore considered the later time points to be more informative for assessing fitness effects and population sizes under increasing crowding conditions. We have now clarified this as “The earliest time points in the experiments for Rickettsiella (Day 14) and Regiella (Day 4) were excluded” in the manuscript.

      (11) In the legends of all figures, please be explicit about sample sizes.

      We have added the relevant information about replicate number or sample sizes in the main and supplementary figures.

      (12) Line 1010 - replace "each leave" with "each leaf" - there was at least one other place, I think in the supplemental information, that leave was used instead of "leaf".

      We replaced this.

    1. eLife Assessment

      This valuable study provides a practical computational framework for inferring latent neural states directly from calcium fluorescence recordings, bypassing the traditional need for a separate spike deconvolution step. The evidence supporting the method is convincing, featuring rigorous validation across multiple latent variable model families (including HMM, GPFA, and LFADS) using both simulated and experimental data. To further strengthen method's generality, it would require further application to a broader range of experimental datasets, such as recordings from different brain regions or using different calcium indicators.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors elegantly combined latent variable models (i.e., HMM, GPFA and dynamical system models) with a calcium imaging observation model (i.e., latent Poisson spiking and autoregressive calcium dynamics (AR)).

      Strengths:

      Integrating a calcium observation model into existing latent variable models improves significantly the inference of latent neural states compared to existing approaches such as spike deconvolution or Gaussian assumptions.

      The authors also provide an open-source access to their method for direct application to calcium imaging data analysis.

      Weaknesses:

      As acknowledged by the authors, their method is dependent on the quality of calcium traces extraction from fluorescence videos. It should be noted that this limitation applies to alternative strategies.

      While the contribution of this study should prove useful for researchers using calcium imaging, the novelty is limited, as it consists of an integration of the calcium imaging model from Ganmor et al. 2016 with existing LVM frameworks.

      Comments on revised version.

      The authors addressed my comments and I have no further concerns.

    3. Reviewer #2 (Public review):

      Summary:

      This compelling study proposes a framework to implement latent variable models using population level calcium imaging data. The study incorporates autoregressive dynamics and latent Poisson spiking to improve inference of latent states across different model classes including HMMs, Gaussian Process Factor Analysis and nonlinear dynamical systems models. This approach allows for a more seamless integration of existing methods typically used with spiking data to apply on calcium imaging data. The authors test the model on piriform cortex recordings as well as a biophysical simulator to validate their methods. This approach promises to have wide usability for neuroscientists using large population level calcium imaging.

      Strengths:

      The strength of this study is the flexibility in the choice of models and relatively easy adaptation to user-specific use cases.

      Weaknesses:

      The weakness of the study lies in its limited validation of biological calcium imaging data. Calcium dynamics in a task-specific context in a sensory brain region might be very different from slower dynamics in a region of integration.

    4. Reviewer #3 (Public review):

      Summary:

      S. Keeley & collaborators propose a computational approach to infer time-varying latent variables directly from calcium traces (e.g., obtained with 2p imaging) without the need for deconvolving the traces into spike trains in a preliminary, independent step. Their approach rests on 1 of 3 families of latent models: GPFA, HMM and dynamical systems - which they augment with an observation model that maps latent variables to fluorescence traces. They validate their approach on simulated data as well as a single real dataset, showing that the approach improves latent variable inference and model fitting, compared to more traditional approaches (although not directly compared with the 2-step one; see below). They provide a GitHub repository with code to fit their models (which I have not tested).

      Strengths:

      The approach is sound and well-motivated. The authors are specialists of latent variable models. The manuscript is succinct, well-written and the figures are clear. I particularly liked the diversity of latent models considered, in particular latent models with continuous (GPFA) vs. discrete (HMM) dynamics, which are useful for characterizing different types of neural computations. The validation on both simulated and real data is convincing.

      Weaknesses:

      The main weakness point that I see is that the approach is tested only on a single real dataset (odor response dataset). The other model fits are obtained from simulated data. While the results are convincing, it would be useful to see the approach tested on other datasets, for instance datasets with different brain areas, different behavioral conditions, or different calcium indicators. This would help assess the generality of the approach and its robustness to different experimental conditions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors elegantly combined latent variable models (i.e., HMM, GPFA and dynamical system models) with a calcium imaging observation model (i.e., latent Poisson spiking and autoregressive calcium dynamics (AR)).

      Strengths:

      Integrating a calcium observation model into existing latent variable models improves significantly the inference of latent neural states compared to existing approaches such as spike deconvolution or Gaussian assumptions.

      The authors also provide an open-source access to their method for direct application to calcium imaging data analysis.

      Weaknesses:

      As acknowledged by the authors, their method is dependent on the quality of calcium trace extraction from fluorescence videos. It should be noted that this limitation applies to alternative strategies.

      While the contribution of this study should prove useful for researchers using calcium imaging, the novelty is limited, as it consists of an integration of the calcium imaging model from Ganmor et al. 2016 with existing LVM frameworks.

      Reviewer #2 (Public review):

      Summary:

      This compelling study proposes a framework to implement latent variable models using population level calcium imaging data. The study incorporates autoregressive dynamics and latent Poisson spiking to improve inference of latent states across different model classes including HMMs, Gaussian Process Factor Analysis and nonlinear dynamical systems models. This approach allows for a more seamless integration of existing methods typically used with spiking data to apply on calcium imaging data. The authors test the model on piriform cortex recordings as well as a biophysical simulator to validate their methods. This approach promises to have wide usability for neuroscientists using large population level calcium imaging.

      Strengths:

      The strengths of this study are the flexibility in the choice of models and relatively easy adaptation to user-specific use cases.

      Weaknesses:

      The weakness of the study lies in its limited validation of biological calcium imaging data. Calcium dynamics in a task-specific context in a sensory brain region might be very different from slower dynamics in a region of integration. The biophysical properties of the data would also be dependent on the SNR of the imaging platform and the generation of calcium indicator being used.

      Reviewers 1 and 2 correctly point out that our method depends on the quality of the upstream calcium trace extraction. As they noted, these traces can vary based on the specific indicator used, the signal-to-noise ratio (SNR) of the recording, and the region of the brain being imaged (such as sensory vs. integrative areas). As Reviewer 1 rightly mentions, this is a universal challenge that applies to alternative strategies as well, rather than a limitation unique to our framework. To address these points, we have added a paragraph to our Discussion section.

      Reviewer #3 (Public review):

      Summary:

      S. Keeley & collaborators propose a computational approach to infer time-varying latent variables directly from calcium traces (for instance, obtained with 2p imaging) without the need for deconvolving the traces into spike trains in a preliminary, independent step. Their approach rests on 1 of 3 families of latent models: GPFA, HMM and dynamical systems - which they augment with an observation model that maps latent variables to fluorescence traces. They validate their approach on simulated and real data, showing that the approach improves latent variable inference and model fitting, compared to more traditional approaches (although not directly compared with the 2-step one; see below). They provide a GitHub repository with code to fit their models (which I have not tested).

      Strengths:

      The approach is sound and well-motivated. The authors are specialists in latent variable models. The manuscript is succinct, well-written, and the figures are clear. I particularly liked the diversity of latent models considered, in particular latent models with continuous (GPFA) vs.

      discrete (HMM) dynamics, which are useful for characterizing different types of neural computations. The validation on both simulated and real data is convincing.

      Weaknesses:

      One advantage … is that one can inspect the quality of the deconvolution step independently from the latent variable inference step. For instance, if the inferred latent variables are not interpretable, how can one determine whether this is due to a poor choice of latent model (e.g., HMM with too few states), or a poor fit of the observation model (e.g., wrong parameters for the calcium dynamics)?

      We agree with the reviewer that integrating the calcium likelihood introduces additional parameters that require careful diagnostics. However, for the vast majority of imaging datasets, there is no simultaneous electrophysiology to verify the deconvolution step. If the final latent states are not interpretable, it remains impossible to determine whether the error originated in the initial spike inference from deconvolution or the subsequent model fitting.

      Our framework addresses this by maintaining the raw fluorescence as the fixed observation. We suggest for those using this model to employ cross-validation using this data to select model parameters. We outline how to do this below, but because the data itself does not change with each model fit, you can compare $P(\text{data} \mid \lambda)$ across any model configuration. In contrast, different deconvolution methods change the data itself (the spiketimes) making comparison across models impossible.

      Could the authors comment on whether their approach allows for instance to compare different forms of latent models (e.g., HMM vs. GPFA) in terms of model evidence, cross-validated log-likelihood or other model comparison metrics?

      We thank the reviewer for highlighting this. In short: yes. Because our framework integrates the calcium observation likelihood with various latent variable models, we can assess held-out prediction $P(\text{data} \mid \lambda)$ irrespective of the specific latent structure.

      However, because fitting the LVM requires inferring the latent state $\mathbf{z}$ to determine the firing rate $\lambda$, proper cross-validation across models involves holding out both neurons and timepoints. A principled approach—which our framework supports—is as follows:

      (1) Train both the latent states $\mathbf{z}$ and the model parameters (e.g., the mapping from latent space to observations) on a training portion of the recording.

      (2) On a held-out test segment, withhold a subset of "test" neurons and infer the latent states using only the "held-in" neurons.

      (3) Calculate the likelihood of the observed fluorescence for the test neurons given the inferred rates.

      We clarify this procedure in the revised manuscript. While a comprehensive benchmarking across all possible LVM architectures is beyond the scope of this study, we provide the statistical infrastructure for users to perform such comparisons. Furthermore, we would like to emphasize that while predictive likelihood is a rigorous metric for model selection, the primary utility of these LVMs often lies in the interpretability of the latent states themselves, which can remain biologically informative even if cross-validated performance is not the sole optimization target.

      While it certainly makes sense that models accounting for the full transformation of latent => spikes => fluorescence data should outperform the two-step (1) deconvolution => (2) latent variance inference approach, the amount of improvement is not clear. A direct comparison … would be useful

      We thank the reviewer for this point. Figure 4 was designed to address this comparison directly. By using a biophysical simulator, we generated a pseudo-realistic spiking network with ground-truth latent trajectories governed by a Gaussian Process. This allowed us to explicitly compare our unified approach against the traditional deconvolution-then-Poisson-GPFA pipeline. While a first-order (AR1) calcium likelihood did not show improvement over the two-step deconvolution method in recovering the ground-truth latents, the second-order (AR2) process demonstrated an improvement. Because there are no ground-truth parameters in the model, we use the reconstruction error of the latent values as our primary metric for recovery. These results suggest that when the observation model sufficiently captures the underlying calcium kinetics, the unified approach offers a more accurate estimation of the neural state.

      It would be useful to discuss the possible extension of the approach to other types of data that … have different observation models.

      We thank the reviewer for this helpful comment. We agree that the general framing of the likelihood has potential use in a wider range of data modalities.

      Specifically, all sensors (aside from some voltage sensors) have a rise and decay time in line with our model. Thus the autoregressive (AR) nature of the calcium likelihood we utilize makes the current implementation particularly well-suited for a broad range of fluorescence-based sensors with similar temporal profiles. The specific use and extension would depend heavily on the biological target of the sensor. For example Glutamate, dopamine, and similar indicators can be thought of as having a similar underlying Poisson firing model, as the release of these products is tied to neural firing. Other sensors that might relate to other biological processes, such as hemodynamics (via imaging or ultrasound) or broader neuromodulation (Norepinephrine imaging with nLight) might be more continually varying and therefore would require changing the Poisson with an appropriate alternative, for example a Gaussian Process or similar.

      Voltage imaging is the one exception that may require more complex observation models. However, the challenge in voltage imaging is not the ability to identify individual spikes, but more that the speed and scope of imaging is inherently limited by the speed of the voltage process and signal-to-noise ratios induced by the low quantum efficiency and membrane-bound nature of these indicators. If imaged well, single spikes would be clearly visible and the two-stage likelihood would not be necessary—one could simply use the spike times in a Poisson model just as with electrophysiology. We have added a paragraph in the discussion highlighting these points.

    1. eLife Assessment

      This valuable study provides a systematically curated atlas of non-canonical open reading frames translated across normal human and mouse tissues by uniformly integrating nearly 400 ribosome-profiling datasets with independent mass-spectrometry evidence for ncORF-encoded peptides. The convincing analyses connect ncORF evolutionary age and coding constraint with translation level, tissue distribution, and co-translation with canonical coding sequences, offering a broad comparative view of how non-canonical translation changes over mammalian evolution. Although the functions of most ncORF-encoded peptides remain unknown, the scale of their detection across tissues, together with their proteomic and evolutionary signatures, suggests that a substantial subset may have biological functions.

    2. Reviewer #1 (Public review):

      [Editors' note: This revised version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the substantive comments raised in the previous round of review. The qualifications and limitations are now more clearly acknowledged in the revised manuscript.]

      Summary:

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Comments on previous version:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

    3. Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Weaknesses:

      Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

      (1) Bias and representations of data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

      (2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TErelated mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

      (3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated. Comparisons with other noncoding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

      (4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in drosophila, worms, mouse, and human. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied, for their functions and corss-species conservations. The authors should explicitly show what is new here in their analyses.

      (5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

      (6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

      (7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

      Comments on revisions:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We thank Reviewer #1 for the constructive comments and recognition of the value of our ncORF atlas. We have addressed the key concerns by strengthening the analyses and clarifying the framing and limitations of our conclusions. We appreciate the reviewer’s support for publication and believe these revisions have further improved the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

      (2) Some analytical methods and standards were not clearly presented in the manuscript.

      We thank Reviewer #2 for the positive assessment of our comprehensive dataset and standardized analytical framework. We have clarified the analytical methods and criteria throughout the manuscript and better defined the scope and limitations of our bioinformatics-based analyses. We appreciate the reviewer’s constructive suggestions, which have helped improve the clarity and rigor of the manuscript.

      Recommendations for the authors:

      Reviewing Editor:

      We have evaluated the revision together with the reviewers' second-round assessments and your responses. The reviewers agree that the manuscript has improved in clarity and that the standardized integration of large-scale Ribo-seq datasets provides a valuable resource for the field. However, several important concerns remain insufficiently resolved. In multiple cases, the revision relies primarily on acknowledgment or reframing of limitations rather than additional analyses or clearer methodological justification, leaving the evidential support for several conclusions incomplete.

      Because the study is entirely computational, all analytical procedures, criteria, and thresholds should be explicitly defined and adequately justified to meet the expected standard of technical rigor. In particular, key components of the analytical framework require clearer description, including the definitions and criteria used for co-translation and ncORF classification. The limitations of the dataset should also be discussed more explicitly, especially regarding the heterogeneity of public Ribo-seq datasets, technical factors influencing detection sensitivity, and the interpretation of variable detection frequencies across samples.

      In addition, several conclusions remain largely descriptive, and the distinction between novel findings and confirmation of previous observations should be clarified more carefully. Conclusions should be framed strictly within the limits of the presented data and positioned appropriately relative to prior work, with suitable citation to avoid overstating novelty.

      We therefore ask the authors to refine the technical descriptions, ensure that all methods and analytical criteria are presented unambiguously, and expand the Discussion to clearly articulate the limitations of the dataset and analysis. The conclusions should also be revised to reflect an appropriately cautious interpretation of the findings.

      The primary strength of this study lies in the scale and standardization of the dataset as a community resource. Given this substantial resource value, we believe the manuscript could become suitable for publication provided that the issues outlined above are addressed clearly and transparently. With these revisions, the work will provide a useful foundation for future studies in this area.

      We therefore invite you to submit a final revised version addressing the points described above.

      We thank the Editor for the careful assessment and constructive guidance. In the final revision, we have clarified all key methodological definitions and analytical criteria, expanded the Discussion of dataset and detection limitations, and revised the conclusions to avoid overstatement. We believe these changes improve the technical rigor, transparency, and overall value of the manuscript as a community resource.

      Reviewer #1 (Recommendations for the authors):

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We appreciate the reviewer’s support for publication.

      Reviewer #2 (Recommendations for the authors):

      While the authors have made commendable efforts to revise the manuscript and address previous concerns, the revised version still falls short of fully resolving several critical issues regarding data interpretation and methodological transparency. I recommend the following modifications before the manuscript can be considered for publication:

      (1) Although the authors have annotated the detection frequencies of individual sORFs in the revised

      Supplementary Table 3 and Fig. S1B, the biological and technical implications of these data require further clarification:

      (a) As the data demonstrates, even the most abundant sORFs were detected in no more than half of the samples. If the authors attribute this low detection rate to technical limitations (e.g., batch effects, sequencing depth, or threshold stringency), this must be explicitly discussed and annotated in the main text to prevent readers from misinterpreting this as low biological penetrance.

      We thank the reviewer for this constructive comment. We have added a paragraph to the Discussion explicitly addressing the potential technical factors underlying the variable detection frequencies to avoid overinterpreting these frequencies as biological penetrance.

      (b) The authors did not fully address my previous query regarding tissue specificity. Given the diverse and complex origins of the analyzed ribo-seq datasets, it is crucial to know whether any of these sORFs are tissue-specifically translated. The authors should analyze and state whether certain sORFs are exclusively detected in specific sample categories (tissues/organs), and whether this translation pattern aligns with the tissue-specific expression of their corresponding host transcripts.

      We thank the reviewer for raising this important point. Strict tissue-exclusive translation is difficult to establish from heterogeneous public Ribo-seq datasets, as gene expression is inherently stochastic and most genes have some probability of being expressed across tissues, although expression levels may vary substantially between tissues. Moreover, failure to detect an ncORF in a given tissue may reflect low expression or insufficient sequencing depth rather than true biological absence. We therefore avoid making definitive claims about tissue-specific translation and instead quantify variation in ncORF expression across tissues using a tissue specificity index.

      (2) Echoing the concerns raised by Reviewer #1, I remain concerned that the observed uORF-CDS cotranslation might be a computational artifact or false positive. The current Methods section lacks sufficient detail on how "co-translation" was strictly defined and quantified. I strongly recommend that the authors include a schematic diagram (e.g., in Figure 6 or supplementary figures) that explicitly details their analytical strategy, statistical thresholds, and evaluation criteria for defining co-translation. As was pointed out, it is well-established that uORFs typically exert an inhibitory effect on the translation of the main CDS. To validate the accuracy and robustness of their analytical pipeline, the authors should use their collected dataset to demonstrate the prevalence and nature of this canonical inhibitory phenomenon. Successfully capturing this expected repression would serve as a crucial positive control for their methodology. While the authors provided a theoretically acceptable mechanistic model in the text to reconcile co-translation with uORF-mediated repression, this hypothesis currently lacks literature support. The authors must cite relevant prior studies that support this specific regulatory dynamic to strengthen their argument.

      We thank the reviewer for this important comment. We have further clarified the definition, detection criteria, and statistical framework for co-translation in the Methods, and have made the complete analysis code publicly available to facilitate reproducibility and independent evaluation. We have also added relevant literature supporting the proposed regulatory interpretation clarifying the relationship between our observations and the established inhibitory effects of uORFs.

    1. eLife Assessment

      This study presents valuable findings regarding cardiac and autonomic effects of seizures and epilepsy, with relevance to sudden unexpected death in epilepsy (SUDEP). They present solid evidence that genetic deletion of the potassium-chloride co-transporter in hypothalamic corticotropin-releasing hormone (CRH) neurons exacerbates bradycardia and enhances autonomic disturbances in a mouse model of temporal lobe epilepsy. This work will be of interest to neuroscientists working on epilepsy, the HPA axis, and autonomic control.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript entitled "Autonomic reflex plasticity associates with time-dependent SUDEP susceptibility in a murine model with hyperactive stress circuits" by Dr. Saunders and colleagues combined a traditional mouse model of SUDEP, ventral intrahippocampal kainite (vIHKA) injection, with a genetic model of chronic hyperactivity of central corticotropin-releasing hormone (CRH) neurons (Kcc2/Crh) that further increases the risk of SUDEP in the weeks following seizure.

      Strengths:

      Their results show during spontaneous seizures Kcc2/Crh mice had more pronounced reflex-like ictal bradycardias compared to WT controls that notably occurred prior (~10 sec) to seizure termination and had greater autonomic disturbances compared to WT controls, including a pronounced serotonin-mediated Bezold Jarisch reflex. These results show chronic hyperactivity of central corticotropin-releasing hormone (CRH) neurons (Kcc2/Crh) increased autonomic disturbances and risk of SUDEP in a kainic acid model of epilepsy.

      Weaknesses:

      This study could be improved with a more thorough assessment of heart rate, blood pressure and breathing during and following the seizures, and in particular the fatal event. It is unclear if the bradycardias were spontaneous, or a result of preceding central or obstructive apneas, oxygen desaturations, hypercapnia, arrhythmias, or other possible triggers.

      Considerable prior work in the literature suggests SUDEP could be mediated, in some patients, by a burst of parasympathetic activity to the heart. Were the heart rate changes in these animals during seizures inhibited or blocked by atropine, or atenolol? The injection of the 5HT agonist phenylbiguanide into the right jugular is not a selective approach for activating the Bezold Jarisch Reflex (BJR) which is caused by increased activity of intracardiac sensory neurons (generally activated with ischemia or a combination of low preload with high contractility). The results should be interpreted more cautiously, as a response to systemic administration of phenylbiguanide only.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors set out to evaluate the role of hypothalamic pituitary axis hyperactivity on cardiac and autonomic changes during epileptogenesis and following seizures in a mouse model of temporal lobe epilepsy. Epilepsy is very common. It can frequently result in death from sudden unexpected death in epilepsy, or SUDEP. SUDEP is thought to be at least in part due to seizure related cardiac and autonomic instability. Increased stress states are well known to be comorbid with epilepsy. This comorbidity is thought to increase the risk of SUDEP. Here the authors hypothesized that a mouse model of heightened stress in which there is hyperactivity of the CRH neurons in the hypothalamus would demonstrate exaggerated cardiac and autonomic effects of seizures and epilepsy.

      Strengths:

      For the chronic stress model, they employed the Kcc2/Crh mice that have a genetic deletion of the potassium chloride cotransporter in CRH neurons. They treated these mice and their wild type littermates with intra hippocampal kainic acid or saline, as epileptic and sham-treated animals respectively. The assessed cardiac activity, blood pressure, baroreflex, and the Bezold-Jerisch reflex during epileptogenesis. This in general is an interesting study. They make some interesting and potentially important observations regarding heart rate and blood pressure in seizures and epilepsy.

      Weaknesses:

      While the revised manuscript is much improved, there are still some concerns that should be addressed.

      (1) The low-pressure baroreceptor responses they show in Figure 4 are somewhat confusing. Should they not be seeing a reflex increase in heart rate when blood pressure is lowered with sodium nitroprusside? It would be helpful if they could describe whether the control responses were as expected or not, and if not, why? This makes it difficult to assess changes seen in the different genotypes and conditions.

      (2) It does not seem appropriate to label the assessments associated with Figure 5 as the Bezold Jarisch Reflex. This reflex involves bradycardia, hypotension, vasoconstriction, and hypopnea. They seem to be only looking at the cardiac component, which is likely mediated though peripheral 5HT3 receptors. Did they measure blood pressure and breathing? Can they include these? If they can only comment on HR, then the discussion should reflect this.

      (3) In Figure 1B, it would be helpful to show some short (e.g., 0.5 sec) snippets of ECG traces that exemplify the changes in HR (spikes/second).

      (4) The day 21 examples given in Figure 1B, do not seem to be representative of the data depicted in Figure 1C.

      (5) From the top panel examples in Figure 2A it looks like there might be greater EEG suppression following seizures in the Kcc2/CRH mice. Was this consistent? It might be worth looking into.

      (6) Can the authors include scale bars for the top panels in Figure 2A?

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents valuable findings regarding cardiac and autonomic effects of seizures and epilepsy, with relevance to sudden unexpected death in epilepsy (SUDEP). They present solid evidence that genetic deletion of the potassium-chloride cotransporter in hypothalamic corticotropin-releasing hormone (CRH) neurons exacerbates bradycardia and enhances autonomic disturbances in a mouse model of temporal lobe epilepsy. However, the evidence that this deletion produces chronic hyperexcitability of the hypothalamic-pituitary-adrenal axis was incomplete, leaving a mechanistic gap. This work will be of interest to neuroscientists working on epilepsy, the HPA axis, and autonomic control.

      We thank the editors and reviewers for their feedback. Although the loss of Kcc2 from CRH neurons in Kcc2/Crh mice has been confirmed (Melon et al., 2018) and leads to HPA axis hyperexcitability in response to stress or seizures, it does not chronically drive HPA axis hyperactivity in unstressed conditions (Basu et al., 2024).

      The following details have been added regarding the Kcc2/Crh model. We now describe the Kcc2/Crh as “hyperreactive” rather than “hyperexcitable/hyperactive” throughout the manuscript.

      -In the Introduction: “This loss of Kcc2 in CRH neurons has been previously confirmed and shown to cause an exaggerated HPA axis response to stress that is absent in baseline conditions (Basu et al., 2024; Melon et al., 2018)”

      -In the Discussion: “This aligns well with lack of elevated plasma corticosterone at baseline in Kcc2/Crh mice, compared to WT, because elevated PVN<sup>CRH</sup> neuron activity should otherwise increase this signal (Basu et al., 2024).”

      -In the Discussion: “Most notable, our model utilizes a developmental strategy to knock out Kcc2 from CRH neurons, which has been confirmed previously (Melon et al., 2018).”

      Public Reviews:

      Reviewer #1 (Public review):

      This study could be improved with a more thorough assessment of heart rate, blood pressure and breathing during and following the seizures, and in particular the fatal event.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2. Overall, pronounced bradycardia that occurred near seizure termination was followed by recovery of HR to pre-ictal baseline in the early post-ictal period (30 sec). In Results: “Independent of genotype, HR recovered to baseline levels during the immediate post-ictal period (0-10 sec, Fig. 2G; 10-30 sec, Fig. 2K).

      In pilot work, we determined that HR during spontaneous seizures were fundamentally different than HR during status epilepticus (see Author response image 1). Therefore, it was critical for us to examine HR during spontaneous seizures. We previously published that Kcc2/Crh+KA mice have a rate of 1-2 seizures per day. To limit additional stressors and seizure provocation, we employed radio telemetry (over tethered systems) and opted out of carotid instrumentation for BP as well as restricted environments of plethysmography chambers for these assessments of HR. After determining that heart rate was different between our mouse lines, we tested whether this change was mediated by central circuits regulating HR (ie: baroreflex, Bezold-Jarisch reflex) (Fig. 3, 4, 5).

      Author response image 1.

      In Discussion: “Although the present study includes only non-fatal seizures, our report of HR during spontaneous seizure events supports work suggesting physiological events during non-fatal seizures predict SUDEP risk (Lamrani et al., 2023; Ryvlin, Nashef, & Tomson, 2013; Schuele et al., 2011). It remains to be determined if ictal events during fatal and non-fatal spontaneous seizures are different and future studies could help clarify any distinctions. Our examination focused on HR (and not respiration or BP). Whether exaggerated BJR-mediated HR response co-occurs with greater magnitude BJR-mediated apnea and hypotension remains to be determined.”

      It is unclear if the bradycardias were spontaneous or a result of preceding central or obstructive apneas, oxygen desaturations, hypercapnia, arrhythmias, or other possible triggers.

      Our work demonstrates the occurrence of ictal bradycardia whereas identifying precipitating factor(s) and interaction(s) of this phenomenon will require alternate approaches. Normal activation of hypoxic and hypercapnic ventilatory responses would be expected to increase HR. However, whether these chemoreflex circuits undergo remodeling in Kcc2/Crh mice remains unknown. Obstructive apnea via laryngospasm can cause reflex bradycardia, but we did not record airflow or respiratory EMG in these studies to determine the existence of obstructive apnea. More testing is merited.

      In Discussion, “…Additional seizure-related disturbances such as central or obstructive apneas may contribute to BJR activation. Hypoxia, which could result from ictal apnea, is known to increase excitatory neurotransmission to cardiac vagal motor neurons that cause vagal bradycardia (Griffioen et al., 2007) and induces platelet activation (Tyagi et al., 2014) which is considered the main source of circulating serotonin for the BJR. As such, hypoxia resulting from apnea may increase likelihood of exaggerated BJR during seizures. Consistent with this…”

      Considerable prior work in the literature suggests SUDEP could be mediated, in some patients, by a burst of parasympathetic activity to the heart. Were the heart rate changes in these animals during seizures inhibited or blocked by atropine or atenolol?

      With this study targeting spontaneous seizures we were unable to test acute pre-treatment with atropine or atenolol. We did observe reduced mortality in Kcc2/Crh+KA mice that underwent chronic parasympathetic blockade via osmotic minipump of methylscopolamine (Figure 5).

      In Discussion:

      “Chronic inhibition of vagal parasympathetic motor output (the driver of BJR reflex bradycardia) improved mortality by 10% in Kcc2/Crh mice. Although this improvement provides some hope for patients at high risk for SUDEP with no treatment options, additional avenues of investigation are needed to more directly link seizure-related bradycardias to vagal parasympathetic motor output.”

      The injection of the 5HT agonist phenylbiguanide into the right jugular is not a selective approach for activating the Bezold Jarisch Reflex (BJR), which is caused by increased activity of intracardiac sensory neurons (generally activated with is chemia or a combination of low preload with high contractility). The results should be interpreted more cautiously, as a response to systemic administration of phenylbiguanide only.

      BJR can be experimentally triggered with intravenous infusion of various compounds including veratrum alkaloids (Cramer, 1915), 5HT (Fozard 1983), or 5HT3R agonists (Verberne & Guyenet, 1992). We now specifically refer to BJR in our study as that induced by PBG (a 5HT3R agonist), as others have done (Yamano et al., 1995; PMID: 8786638) and acknowledge endogenous BJR activation in the discussion.

      Added to the results: “As dysfunction of serotonergic signaling is implicated in the pathophysiology of SUDEP (Richerson & Buchanan, 2011), we investigated the Bezold Jarisch Reflex (BJR) (Fig. 5), a cardioinhibitory reflex that is reliably triggered experimentally by activation of cardiopulmonary vagal afferents containing serotonin type 3 receptors (5HT3R) (Fozard 1983; Yamano et al., 1995).”

      Added to Discussion: “Although our report is the first to link BJR to SUDEP, serum serotonin levels are elevated following generalized seizures (Murugesan et al., 2018), likely via release from activated platelets (Cloutier et al., 2018). This surge in serum serotonin could lead to endogenous activation of BJR, as bolus intravenous infusion of serotonin reliably triggers BJR experimentally (Fozard 1983; Whalen et al., 2000).”

      Reviewer #2 (Public review):

      Some of the conclusions may be a bit overstated as is and would benefit from more discussion and perhaps additional data.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2.

      The Discussion now includes more details regarding respiration, BP, obstructive apnea, properties of the BJR, and distinctions in seizure type based on whether evoked or lethal.

    1. eLife Assessment

      This important study provides new insights into the patterns of organelle inheritance in the protozoan parasite Toxoplasma gondii. The authors introduce an innovative dual-labeling approach to distinguish maternally inherited from de novo synthesized organelles, representing convincing evidence that different organelles follow distinct inheritance fates during parasite replication. Future studies will be needed to determine whether the residual body functions as a central recycling hub, as the current data are also consistent with alternative models.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting and metabolic compartments) are divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to better understanding the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites. In particular, this study shows that the residue body is a region of the cell syncytium that organelles can be actively transported from. Therefore, it is a space that can actively contribute to the segregation of the late segregating micronemes and rhoptries.

    3. Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of toxoplasmosis. Parasite invasion of host cells, intracellular replication, and subsequent egress, which results in destruction of the infected cell, are central to pathogenicity. This manuscript focuses on understanding how maternal resources, specifically cellular organelles, are shared between daughter parasites during cell division. Many organelles are present as a single copy, making their division and inheritance essential for successful replication. In T. gondii, our understanding of how organelles are divided during cell division remains limited, and this study helps address this important knowledge gap.

      Strengths:

      The major strength of this study is the use of a Halo-based pulse-chase assay to characterize patterns of organelle inheritance and to monitor protein synthesis, turnover, and movement. This approach will be of considerable interest to the field. Using this method, the authors identify three major modes of organelle inheritance:

      (1) Organelles present in multiple copies (such as micronemes and rhoptries) are partitioned between daughter parasites, with additional contributions from newly formed vesicles. Newly synthesized and pre-existing material remain as distinct populations within the cell.

      (2) Single-copy organelles, such as the Golgi and apicoplast, are expanded through the incorporation of newly synthesized material before division.

      (3) Cytoskeletal structures are synthesized de novo during each round of cell division.

      These findings provide a more refined understanding of organelle inheritance and demonstrate that secretory organelles are not generated entirely de novo during each round of division, as was previously thought.

      The paper places particular emphasis on the fate of maternal micronemes and rhoptries during division. The data show that (1) during division in wild-type cells, maternal micronemes and rhoptries are detectable in the residual body (RB); however, the majority of these organelles are localized within the parasite body, either at the apical or basal ends of the daughter parasites (Fig. 6). (2) In the absence of the myosin motor MyoF, micronemes and rhoptries accumulate in the residual body and are not properly trafficked to the daughter cells. Upon restoration of MyoF protein levels, these organelles redistribute to the daughter cells, although in an uneven manner.

      Weaknesses:

      The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. While the authors propose that the RB acts as a central hub for recycling both organelles, the current data do not fully support this conclusion.<br /> The model that microneme and rhoptry recycling is RB-dependent relies largely on the MyoF depletion phenotype and the limited detection of maternal organelles in the RB of wild-type parasites. Alternative models remain plausible, including direct trafficking to daughter cells, with RB accumulation upon MyoF depletion reflecting impaired trafficking rather than an obligatory RB-dependent recycling pathway, as now discussed by the authors.

    4. Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughter-derived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strength:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are convincing. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Weaknesses:

      (1) In addressing the question of residual body participation in sorting of organelles, a clear definition of this structure is required including when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. The authors' definition is as follows: 'The RB originates from the collapse of the maternal parasite during daughter cell budding and occupies the space previously occupied by the mother cell.' As such, a clear marker of the mother cell 'collapse' is required, but such a marker is not identified or used in the study to separate what might be considered an active part of the mother cell during early daughter formation, and the residual body. This might seem like moot a point, but it would help to give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? The authors elegantly show that MyoF is necessary for segregation of micronemes and rhoptries into daughters, and that MyoF depletion leads to accumulation of these organelles within the residual body. Moreover, restored expression of MyoF can then recover these organelles. This clearly demonstrates the activity of the residual body as part of the syncytium space that participates in the maintenance of the vacuole. But does it imply that this space necessarily handles all inherited micronemes and rhoptries as a 'trafficking hub'? My concern with the lack of a clear definition could provide some misinterpretation or overinterpretation of the contribution residual body.

      We thank the reviewer for raising this conceptual point. We agree that, in the absence of a molecular marker that uniquely defines the nascent RB, the precise transition between posterior maternal cytoplasm and a morphologically distinct RB cannot be determined during early daughter formation. We have therefore clarified our terminology in the revised manuscript and define the RB operationally as the posterior compartment/connection between daughter parasites. We also avoid assigning early posterior trafficking events unambiguously to a fully formed RB. We further agree that the current data do not establish that all inherited micronemes and rhoptries must transit through the RB. We have therefore revised the Results and Discussion and softened terminology such as “central trafficking hub” and now conclude that the RB represents an important dynamic compartment in organelle recycling and redistribution, without implying that it is an obligatory intermediate for every inherited secretory organelle.

      (2) A further, remarkable conclusion is that maternal micronemes are evenly segregated into daughters through an active process for 'balanced microneme inheritance'. The proportion of maternal micronemes is quantified up to the 8-cell stage and shown to be not significantly different between cells. But would this result be expected with random assortment at this stage? The authors model the probability of a 32-cell stage vacuole occurring with each daughter having within 0-3 maternal micronemes and this is considered unlikely. However, the authors neither present the modelling for the 8cell stage or show quantification of 32-cell vacuoles. They do show some images of large vacuoles, but it is not possible to determine the distribution of maternal micronemes in these images. A regulated process of segregation would require a complex mechanism where some form of microneme counting would be required to create the proposed balance. It is, therefore, important to have strong data supporting such a hypothesis, but this is not currently presented.

      We agree that the previous wording implied a mechanistic conclusion beyond what can be established from the present dataset. We have therefore revised the manuscript so that the relatively even distribution of maternal micronemes at the 8-cell stage is presented as an observation that is consistent with a non-random or regulated partitioning process, rather than evidence for an established microneme-counting mechanism. The 32-cell model is now presented as supportive rather than definitive evidence, and we explicitly acknowledge that the quantitative experimental dataset was obtained at the 8-cell stage. 

      Reviewer #2 (Public review):

      (1) The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. The authors strongly argue that the RB is a central hub for recycling micronemes and rhoptries; however, this conclusion is not fully supported by the data. For example, the authors state…Thus, the model that all microneme and rhoptry trafficking is RB-dependent is based primarily on the MyoF depletion phenotype (which results in RB accumulation) together with the observation that a relatively small amount of maternal microneme and rhoptry material is detectable in the RB of wild-type parasites. Although the authors' interpretation-that recycling is RB-dependent-is one possible explanation, alternative models are not discussed. For example, an alternative possibility is that the majority of micronemes and rhoptries are trafficked directly from the apical end of the mother parasite to the daughter cells without passing through the RB. In this scenario, only a subset of the organelles would enter the residual body, perhaps reflecting imperfect trafficking efficiency rather than an obligatory recycling step. Loss of MyoF would impair this trafficking pathway, resulting in the accumulation of secretory organelles within the RB. In other words, RB accumulation could be a consequence of MyoF depletion rather than evidence that all trafficking in wild-type parasites normally proceeds through the RB. This alternative interpretation seems particularly relevant for the rhoptries, given that the authors themselves state that "M-RON2 was integrated into daughter rhoptries prior to mother cell collapse and formation of the RB."

      We agree with the reviewer that the current data do not establish obligatory transit of all maternal micronemes and rhoptries through the RB. We have revised the manuscript throughout to make this distinction explicit. In particular, we now emphasize that maternal MIC2 can be directly observed entering the RB, whereas most maternal RON2 is incorporated into daughter rhoptries before mother-cell collapse and formation of a morphologically distinct RB. This observation leaves open the possibility that a substantial fraction of maternal rhoptries is transferred directly from the mother to developing daughters. We have also revised our interpretation of the MyoF-depletion phenotype. The accumulation of maternal MIC2 and RON2 in the RB following MyoF depletion demonstrates that MyoF is required for efficient redistribution of both organelle populations, but does not by itself demonstrate that both normally follow an identical spatial route through the RB. We now explicitly state that their precise trafficking routes and timing may differ.

      We nevertheless retain the conclusion that the RB is a dynamic compartment involved in organelle recycling because maternal MIC2 can be directly observed entering and leaving this compartment, and material accumulated there following MyoF depletion can subsequently be redistributed after restoration of MyoF.

      (2) Figure S10C. To determine whether microneme degradation occurs in the RB, the authors quantified the fluorescence intensity of individual micronemes in control parasites and following auxin washout, showing that after redistribution the fluorescence intensity of individual vesicles is unchanged. However, this is not the appropriate analysis to address the question being asked. To conclude that micronemes are not degraded, the authors would need to quantify the total fluorescence intensity within the entire vacuole. For example, if half of the micronemes were degraded, the remaining micronemes would be expected to retain the same fluorescence intensity as those in the control parasites. Thus, unchanged fluorescence intensity of individual vesicles does not exclude the possibility that degradation has occurred.

      We agree with this criticism and have revised the interpretation of the experiment accordingly. We no longer conclude that the analysis excludes microneme degradation. We now explicitly acknowledge that analysis of individual recovered micronemes cannot exclude degradation of a fraction of the total microneme population during RB retention. Thus, the experiment supports preservation of MIC2 signal in the recovered organelles but is no longer presented as evidence that no microneme degradation occurs.

      Reviewer #3 (Public review):

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body. 

      We agree and have incorporated this limitation into the revised interpretation. Our live imaging demonstrates that maternal MIC2 can enter the RB and subsequently redistribute to daughter parasites, but it does not establish that every individual maternal microneme follows this route. We thank the reviewer for highlighting this distinction, which has helped us clarify the model presented in the Results and Discussion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Carruthers and Sibley, 1997, is still the cited ref for Tic20 in the apicoplast despite the authors saying that would correct this to a van Dooren publication.

      Corrected

      (2) In Figure 2B three biological replicates were used however there are no error bars shown. The only reason for performing replicates is the observe the variance in the data, so if this is not shown the replicates are effectively meaningless. I strongly advice that error bars are given to indicate this seeing that the data in this figure form the basis of the major conclusions of the study. It might be necessary to show this in supplemental forms with fewer proteins if the error bars are too difficult to see in the combined figure.

      We agree and have corrected Figure 2B to display the variability between the three independent biological replicates. Error bars now represent the standard deviation.

      (3) Line 136: Can you conclude that these inheritance patterns are 'organelle-specific' when each organelle is only sampled with one or two proteins. Isn't it better to conclude that these are protein-specific, with the hypothesis that they might represent the orgnalle as a whole. I imagine that some proteins in organelles such as the apicoplast have shorter half-lives than others, and therefore some apicoplast proteins might behave like 'Group 3' proteins.

      We thank the reviewer for raising this point. We agree that individual proteins within the same organelle may differ in their turnover kinetics and that analysis of one or two markers cannot establish that every molecular component of an organelle behaves identically. However, we do not think that describing the observations exclusively as protein-specific inheritance would fully reflect the biological process investigated here. The proteins analysed are established markers of defined organelles, and our conclusions are based not only on changes in fluorescence intensity, but also on the localization, morphology, partitioning, and spatial relationship between maternally inherited and newly synthesized organelle populations.

      This is particularly evident for micronemes and rhoptries, where maternal and de novo material remain spatially separated and individual organelles can be followed during inheritance.

      In addition, the microneme phenotype observed with MIC2 was confirmed using AMA1, MIC4, and MIC8.

      We therefore retain the terminology of organelle inheritance, while acknowledging that individual proteins within a given organelle may exhibit different turnover kinetics and that the markers analysed may not represent the behaviour of every molecular component of the organelle.

      (4) Line 222: It is an odd phrase to suggest that the Golgi, ER etc 'bypass' the residual body, which suggests an active avoidance mechanism. Would the authors also conclude that the nucleus 'bypasses' the RB? Moreover, the ER and mitochondria are actually known to be present in the RB forming continuous organelles between daughters in a vacuole. So again, this might be an overstatement that mispresents how the RB participates in vacuole functions.

      We agree and have removed the term “bypass.” The revised text now states only that we did not observe comparable accumulation of the analysed Golgi, ER, or apicoplast markers in the RB during inheritance. This avoids implying an active avoidance mechanism and is compatible with the known continuity of ER and mitochondria through the RB.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      Figure 8F and Video S6: The authors should specify the time point after IAA washout at which live imaging was initiated. Does time 0 in the video correspond to the point at which IAA was removed?

      We have clarified this in the Results and Methods. Auxin was removed after 24 h of replication, and live imaging was subsequently initiated. Time 0 in Figure 8F and Video S6 corresponds to the first acquired frame after auxin washout.

      Figure S10 should read auxin, not auxine.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      I have no further suggestions. Congratulations to the authors for a lovely study.

    1. eLife Assessment

      This important study shows that nutrient resorption efficiency in the widespread wetland grass Phragmites australis varies with population origin and remains largely stable under experimentally imposed salt stress, supporting limited short-term plasticity and a role for genetic differentiation. The findings suggest that predictions of wetland nutrient cycling under increasing salinization should account for intraspecific variation and phylogeographic composition, the evidence is compelling, based on a common-garden experiment involving 110 genotypes, paired control and salinity treatments, and convergent metabolomic, ionomic, and whole-plant evidence confirming substantial physiological stress. The revised manuscript clarifies the partial overlap between ecotype and phylogeographic lineage and appropriately qualifies the modest explanatory contribution of latitude, conclusions remain restricted to one species, one growing season, and a single moderate salinity treatment; an ecotype-specific nitrogen resorption response also indicates that plasticity is not entirely absent and the physiological mechanisms and responses to chronic, stronger, or multigenerational salinity exposure remain unresolved. The study will interest researchers working on plant functional ecology, nutrient cycling, and wetland responses to global change.

    2. Reviewer #1 (Public review):

      Summary:

      This study demonstrates that nutrient resorption efficiency (NuRE) in Phragmites australis is genetically canalized rather than plastic to salt stress. Using 110 genotypes in a common garden, the authors show that intraspecific variation in NuRE is explained by phylogeographic lineage, ecotype, and latitude, not by effective salinity. Element specific regulatory strategies further reveal how N, P, and K resorption are differentially controlled. At the population level, this is an important study that fundamentally advances our understanding of plant functional trait evolution and its implications for ecosystem nutrient dynamics under global change.

      Strengths:

      This study is the first to demonstrate genetic determination of a key nutrient conservation trait under effective salt stress in a widespread macrophyte, directly testing the 'plastic acclimation versus inherent conservatism' paradigm in a non-nutrient stress context. The experimental design is rigorous: each genotype was paired across control and salt treatments, and multilevel stress effectiveness (metabolomics, biomass, Na accumulation) was confirmed before evaluating NuRE. The large sample size of a macrophyte and dual classification (phylogeography + ecotype) allow robust disentangling of genetic versus plastic sources of variation.

      The analysis comprehensively tests three resorption control hypotheses using appropriate SMA regression, revealing element specific and condition dependent patterns. The latitudinal gradient and variation partitioning provide strong evidence that genetic origin and geographic context outweigh short term plasticity, with important implications for predicting ecosystem nutrient cycling under global change. This study provides a clear empirical demonstration that a key nutrient conservation trait can remain homeostatic under non nutrient stress, and that intraspecific variation is primarily a product of population differentiation rather than short term plasticity.

      Weaknesses:

      First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study's main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermine the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

      Comments on revised version.

      The author carefully revised the parts that might cause confusion.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript shows that nutrient resorption efficiency in Phragmites australis is largely genetically canalized rather than plastic under effective salt stress, with variation mainly associated with phylogeographic lineage, ecotype, and latitude.

      Strengths:

      The common-garden experiment with 110 genotypes, paired control and salt treatments, multi-level stress validation, and element-specific analyses provides compelling support for the central conclusions.

      Weaknesses:

      The generality is limited by the focus on a single species and a single growing season, and some mechanisms are inferred indirectly.

    4. Author response:

      The following is the authors’ response to the original reviews.

      The major revisions include:

      (1) Conceptual framing and scope: defined canalization in the Introduction and clarified that our conclusions are restricted to limited NuRE plasticity during one growing season under a single moderate salinity treatment.

      (2) Methods and classification: clarified the substrate composition and elemental measurements, specified that the ecotype analysis included only Chinese populations, and explained the partial association between ecotype and phylogeographic group.

      (3) Interpretation: expanded the discussion of K resorption and inverted nutrient limitation and tempered the interpretation of latitude and the substantial unexplained variation.

      (4) Robustness and presentation: added Supplementary Figure S7 showing that carbon standardization did not alter the main conclusions, added significance symbols to Table 1, corrected the unit in Figure 2b, and revised repetitive wording in the Discussion.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      (R1-P1) First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long-term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study’s main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      We agree that our evidence is limited to the absence of NuRE plasticity during one growing season under the imposed salinity treatment and does not resolve chronic or multigenerational responses or their physiological basis. We therefore revised Discussion 4.1 to delimit the canalization inference, identify the proposed mechanisms as untested, and specify the longer-term and mechanistic studies needed to distinguish among them.

      Discussion 4.1, fourth paragraph, inserted immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied ”.

      “Our inference of canalization is therefore limited to the absence of a plastic NuRE response during one growing season under the imposed salinity treatment. Chronic, more severe, or multigenerational salinity exposure may produce acclimatory, epigenetic, or transgenerational responses that cannot be evaluated here. Moreover, although altered phloem loading, disruption of senescence-associated remobilization, and reallocation towards osmotic adjustment are plausible explanations for the observed response, we did not directly measure these mechanisms. Long-term experiments combined with targeted molecular and transport measurements are needed to distinguish among these possibilities.”

      (R1-P2) Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermines the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

      We agree and have clarified all three evidential limits. First, the resorbed N: P and N: K analyses are now described as indirect evidence consistent with nutrient-limitation control, not as a causal test; direct nutrient-addition experiments would be required for causal inference. Second, metabolomics is identified as validation of physiological stress rather than a genotype-specific mechanistic analysis. Third, because ecotype and phylogeographic group are partly associated, they were fitted in separate models. The genotype random effect accounts for paired measurements but does not remove confounding between the classification schemes; ecotype differences are therefore interpreted as complementary rather than independent evidence.

      Discussion 4.2, third paragraph.

      “Within this context, the consistent ‘inverted’ nutrient limitation pattern (i.e., a slope significantly >1 for the relationship between log Resorbed N:P and log Green N:P) provides indirect evidence consistent with nutrient-limitation control at the intraspecific level in P. australis, but it cannot establish causal nutrient limitation. Direct factorial nutrient-addition experiments would be required to determine whether the observed resorption patterns are driven by the relative limitation of N, P, or K.”

      Discussion 4.1, first paragraph, inserted immediately after the sentence ending “providing a robust foundation to evaluate NuRE responses”.

      “In this study, metabolomic profiling was used primarily to confirm that the salinity treatment induced broad physiological stress, rather than to resolve genotype-specific metabolic mechanisms underlying NuRE variation. Integrating metabolite profiles with genotype-level NuRE responses would be a valuable direction for future mechanistic research.”

      Methods 2.4.

      “Phylogeographic group and ecotype were analysed in separate linear mixed-effects models because ecotype classifications were available only for Chinese populations and were partly associated with phylogeographic structure. Each model included salinity treatment and either phylogeographic group or ecotype as fixed effects, with genotype fitted as a random effect to account for the paired experimental design in which each genotype was exposed to both control and salt conditions.”

      Discussion 4.4, first paragraph, inserted immediately after the sentence ending “governed by geographic origin (phylogeographic group and ecotype)”.

      “The separate-model approach avoids including the two correlated classification schemes as simultaneous independent predictors, but it does not fully disentangle deep phylogeographic history from recent habitat-associated differentiation. We therefore interpret the ecotype analysis as complementary evidence of habitat-associated differentiation rather than as an effect independent of phylogeographic history.”

      Reviewer #2 (Public review):

      (R2-P1) The experiment covers only one growing season, with salinity applied in June and measurements in December. While the stress is clearly effective, longer-term or multi-year stress might reveal acclimation or epigenetic effects that are not captured. Given the author team’s expertise in parental and transgenerational effects in clonal plants, this limitation is particularly relevant and warrants more thorough discussion in the manuscript.

      We agree. This concern overlaps with Reviewer #1’s temporal-scope comment. We have revised Discussion 4.1 to state explicitly that our inference is restricted to the absence of a plastic NuRE response during one growing season. We also acknowledge that chronic or multigenerational exposure could induce acclimatory, epigenetic, or transgenerational responses that were not captured by the present design.

      See the full revised text under R1-P1 above.

      (R2-P2) The salinity treatment uses a single moderate level of 10 ppt, which does not allow assessment of whether more extreme stress might trigger a plastic response. A dose-response design across a gradient would have provided stronger inference about the threshold at which NuRE canalization might be overcome. Additionally, the ecotype analysis in Figure 4 applies only to Chinese populations, as classification was not available for non-Chinese populations, which should be stated more explicitly in the Results.

      We agree with both points. We have revised Discussion 4.1 to acknowledge that the single 10 ppt treatment does not exclude the possibility of a plastic NuRE response at higher salinity or along a broader dose-response gradient. We therefore frame the identification of a possible response threshold as a priority for future experiments.

      We have also revised Results 3.2 to state explicitly that the ecotype analysis in Figure 4 included only Chinese populations because ecotype classifications were unavailable for non-Chinese populations.

      Discussion 4.1, fourth paragraph, at the same revision point as R1-P1: immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied”.

      “Because only one moderate salinity level (10 ppt) was tested, our results do not exclude plastic NuRE responses at higher salinity or along a dose-response gradient. Future experiments should determine whether a threshold exists beyond which the apparent stability of NuRE is overcome.”

      Results 3.2, second paragraph, inserted immediately after the sentence reporting the ecotype effects and ending “(Figure 4; Figure S5)”.

      “It should be noted that the ecotype analysis here was based only on Chinese populations, as ecotype classification was not available for non-Chinese populations.”

      (R2-P3) The variation partitioning shows latitude as a significant predictor, but the R<sup>2</sup> values are relatively low, indicating that much variance remains unexplained. The manuscript should avoid overinterpreting latitude’s explanatory power and more openly acknowledge the role of unmeasured factors. The interpretation of slopes greater than 1 for the resorbed N:P versus green N:P relationship, labeled as “inverted limitation”, also needs further explanation regarding its functional significance.

      We agree that the original wording overemphasized latitude. Results 3.4 and Discussion 4.3 now describe latitude as the largest contributor among the measured predictors while emphasizing its modest individual R<sup>2</sup> and the substantial unexplained variation. We also expanded Discussion 4.2 to explain that slopes greater than 1 indicate disproportionate recovery of P or K relative to N and to present nutrient balance, P conservation, K mobility, and constraints on N remobilization as non-exclusive hypotheses rather than established mechanisms.

      Results 3.4, first paragraph, replacing the sentence beginning “Furthermore, variation partitioning analysis indicated that latitude”.

      “Variation partitioning indicated that latitude had the largest individual contribution among the measured predictors, but its contribution was modest for N, P, and K resorption (individual R<sup>2</sup> = 0.092, 0.057, and 0.094, respectively; Table 1). The very small contributions of green-leaf P concentration (R<sup>2</sup> = 0.006) and stoichiometry (R<sup>2</sup> < 0.001) further indicate that most variation in P resorption was associated with factors not represented in the present models.”

      Discussion 4.3, first paragraph.

      “Although latitude explained more variation than the other measured predictors, its individual contribution remained modest. The substantial unexplained variance indicates that additional climatic, edaphic, demographic, or genetic factors also contribute to NuRE variation. We therefore interpret latitude as a significant but limited correlate of NuRE rather than as a dominant determinant.”

      Discussion 4.2, first paragraph.

      “Functionally, slopes greater than 1 indicate that changes in green-leaf N:P or N: K are accompanied by disproportionate changes in the corresponding resorbed ratio, consistent with relatively greater recovery of P or K than of N across the observed nutrient gradient. This pattern may contribute to maintaining internal N:P:K balance during regrowth. It may also reflect stronger conservation of P, the high mobility of K, or constraints on the remobilization of N retained in structural or metabolic compounds. Because the experiment did not include nutrient additions or direct measurements of remobilization costs, these explanations remain hypotheses.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) Clarify the concept of “canalization” early. Consider adding one sentence in the Introduction or Discussion explicitly defining canalization in this context (i.e., NuRE varies among populations but this variation is genetically determined and shows little plasticity to salinity). This will help readers less familiar with evolutionary biology terminology.

      We thank the reviewer for this helpful suggestion. We agree that canalization should be defined when it is first introduced. We have therefore added a sentence to the Introduction, immediately after introducing the evolutionary canalization hypothesis, to clarify how this concept is used in our study.

      Introduction, fourth paragraph.

      “Here, canalization denotes genetically based differences in NuRE among populations, coupled with limited phenotypic plasticity of this trait under the short-term salinity treatment.”

      (R1-R2) Discuss potassium more deeply. In the Discussion (section 4.2 or 4.4), speculate on why K resorption lacks concentration control. Does Na<sup>+</sup> accumulation functionally substitute for K in osmotic adjustment, thereby decoupling resorption from green leaf K concentration?

      We appreciate this suggestion and have expanded Discussion 4.2. We now propose partial functional substitution of K<sup>+</sup> by accumulated Na<sup>+</sup> during osmotic adjustment as one possible explanation for the weak coupling between green-leaf K concentration and K resorption. We explicitly frame this explanation as tentative because it was not directly tested in the present experiment.

      Discussion 4.2, second paragraph.

      “One possible explanation is the partial functional substitution of K<sup>+</sup> by Na<sup>+</sup> during osmotic adjustment. Na<sup>+</sup> can replace part of the nonspecific vacuolar osmotic function of K<sup>+</sup> in plants and has been reported to become a major osmoticum in P. australis from higher-salinity habitats (Wakeel et al., 2011; Zhao et al., 1999). Increased Na<sup>+</sup> accumulation may therefore reduce reliance on K<sup>+</sup> for osmotic adjustment, potentially contributing to the weak relationship between green-leaf K concentration and K resorption efficiency observed here.”

      Wakeel, A., Farooq, M., Qadir, M., & Schubert, S. (2011). Potassium substitution by sodium in plants. Critical Reviews in Plant Sciences, 30(4), 401–413. https://doi.org/10.1080/07352689.2011.587728

      Zhao, K. F., Feng, L. T., & Zhang, S. Q. (1999). Study on the salinity-adaptation physiology in different ecotypes of Phragmites australis in the Yellow River Delta of China: Osmotica and their contribution to the osmotic adjustment. Estuarine, Coastal and Shelf Science, 49(Supplement 1), 37–42. https://doi.org/10.1016/S0272-7714(99)80006-7

      (R1-R3) Address the potential confounding of ecotype and phylogeography. The dual classification (Figure 3 vs 4) is elegant. However, Chinese ecotypes (freshwater, coastal, inland saltmarsh) may be partially confounded with phylogeographic lineages. A brief sentence explaining how the analytical approach (separate models, random effects) helps separate deep evolutionary history from recent local adaptation would strengthen the interpretation.

      We agree. This concern is addressed in detail under R1-P2. We clarified why phylogeographic group and ecotype were fitted in separate models, what the genotype random effect accounts for, and why the ecotype results cannot be interpreted independently of phylogeographic history.

      See R1-P2, Locations C1 and C2.

      (R1-R4) In section 2.1, specify whether the soil mixture ratio (2 soil: 1 peat moss: 1 river sand) is by volume or by mass.

      We thank the reviewer for identifying this ambiguity. The ratio was based on volume, and we have revised Methods 2.1 accordingly.

      “The plants were planted in barrels (total volume 25 L; top diameter 32.5 cm, bottom diameter 28.4 cm, height 38.5 cm) containing 20 L of a substrate composed of soil, peat moss, and river sand in a 2:1:1 volume ratio (Figure S1).”

      (R1-R5) In the Materials and Methods section, the element potassium (K) is described twice. Specifically, in line 176, the sentence “Besides C, N and P, other eight elements (K, Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” should be revised to “Besides C, N, P and K, other seven elements (Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” to avoid redundancy.

      We thank the reviewer for identifying this redundancy. We have corrected the sentence in Methods 2.2.

      “Besides C, N, P and K, seven additional elements (Cu, Zn, Fe, Mn, Mg, Si, and Na) were quantified in leaf tissues.”

      (R1-R6) Citations should be formatted and ordered alphabetically or by year of publication.

      We thank the reviewer for pointing out this issue. We standardized multi-reference citations, reordered the reference list alphabetically by first-author surname, and verified correspondence between in-text citations and reference entries.

      Full-manuscript citation and reference-list audit.

      Reviewer #2 (Recommendations for the authors):

      (R2-R1) In the discussion of the “inverted” nutrient limitation, elaborate on why P and K are recovered more relative to N. Consider whether this reflects a strategy to maintain an optimal N: P: K ratio or arises from higher costs or lower availability of N.

      We agree. This issue is addressed in detail under R2-P3, where we explain the meaning of slopes greater than 1 and present the possible functional mechanisms as hypotheses rather than established explanations.

      See R2-P3, Location C.

      (R2-R2) Expand Table 1 to include confidence intervals or p-values for individual effects. The extremely low R<sup>2</sup> for concentration and stoichiometry on P resorption should be more explicitly noted as evidence for the dominance of latitude.

      We agree and have expanded Table 1 by adding significance symbols (*) to the individual R<sup>2</sup> values. Significance was assessed for the corresponding fixed effects in the full linear mixed-effects models using Type III tests with Satterthwaite’s approximation for degrees of freedom. For P resorption, latitude had the largest individual contribution and was significant (R<sup>2</sup> = 0.057, p = 0.005), whereas the contributions of green-leaf P concentration and stoichiometry were very small and nonsignificant (R<sup>2</sup> = 0.006 and < 0.001, respectively). We revised the Results to emphasize this contrast while acknowledging that latitude explained only a modest proportion of the total variation.

      See R2-P3, Locations A and B, for the corresponding Results and Discussion revisions.

      (R2-R3) Standardize the abbreviation to “NuRE” throughout, correcting the use of “NRE” in the introduction.

      We thank the reviewer for identifying this inconsistency. We standardized the abbreviation to NuRE throughout the manuscript.

      (R2-R4) Acknowledge in the discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response, justifying future dose-response experiments.

      We agree. This limitation is addressed under R2-P2, where we state that a single 10-ppt treatment cannot exclude plastic responses at higher salinity or along a broader dose-response gradient.

      See R2-P2, Location A.

      (R2-R5) In the Results section, explicitly state that the ecotype analysis in Figure 4 is based only on Chinese populations, not all 110 genotypes, because ecotype classification was not available for non-Chinese populations.

      We agree. This clarification is provided under R2-P2, where Results 3.2 is revised to state that the Figure 4 ecotype analysis includes only Chinese populations.

      See R2-P2, Location B.

      (R2-R6) Consider adding a supplementary figure comparing raw and carbon-standardized NuRE values to show whether the correction altered main conclusions.

      We agree and have added Supplementary Figure S7 comparing raw and carbon-standardized NuRE. The two estimates were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999), and analyses using raw NuRE retained the same conclusions for phylogeographic group, ecotype, salinity, their interactions, and latitude. Thus, carbon standardization slightly shifted the absolute values without altering the main conclusions.

      Methods 2.3.

      “Raw NuRE was calculated without carbon standardization and compared with carbon-standardized NuRE using Pearson correlations; the main linear mixed-effects analyses were also repeated using raw NuRE.”

      Results 3.4, inserted after the existing paragraph reporting the latitude effects and referring to Figure 6 and Table 1.

      “Raw and carbon-standardized NuRE were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999; Figure S7), and analyses using raw NuRE did not alter the conclusions for phylogeographic group, ecotype, salinity, their interactions, or latitude.”

      Supplementary Figure S7 legend

      “Figure S7 Comparison of raw and carbon-standardized nutrient resorption efficiency (NuRE) for N, P, and K. Points represent genotype-by-treatment observations (n = 206), coloured by treatment. Grey dashed lines indicate the 1:1 relationship, and black lines show ordinary least-squares fits. Pearson’s r and the mean standardized-minus-raw difference (Δmean, percentage points) are shown.”

      (R2-R7) In Figure 2b, check the y-axis label; the text reports mg/kg, but the axis shows g/kg, which needs correction.

      We thank the reviewer for identifying this unit discrepancy. We corrected the y-axis label in Figure 2b from g/kg to mg/kg so that it matches the units reported in the text.

      (R2-R8) In the Discussion, rephrase the sentence “This indicates that NuRE is a conservative trait…” to avoid repetition with the Results summary, for example, “This finding highlights the conservative nature of NuRE.”

      We thank the reviewer for this helpful wording suggestion. We have rephrased the sentence to avoid repetition.

    1. eLife Assessment

      Dohi et al. asked what role the dorsal hippocampus and medial prefrontal cortex play during different temporal epochs of a memory-guided navigation task, a long-standing question for neuroscientists studying hippocampal-prefrontal contributions to working memory. This useful study used a delayed, cue-guided T-maze task in mice and reported impaired choice accuracy when silencing occurred early in the central-arm run but not during the delay. However, the evidence for the study's claim is incomplete in its current form: the silencing windows are not duration-matched across epochs, key negative findings rest on three to four mice without power analysis or reported effect sizes, no non-mnemonic control task distinguishes disrupted memory-guided behavior from a general action-selection deficit, and no neural recordings accompany the causal manipulations to verify the manipulations' assumed mechanistic effects. The reported perseveration also reflects an increased directional bias rather than repetition of the previous choice, and the language should be revised accordingly.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors trained mice to perform a memory-guided navigation task, in which they must navigate to a previously cued arm after a delay period. They optogenetically inhibited the dorsal hippocampus or mPFC (targeting PL) in different task epochs. They found that for both regions, inactivating at the beginning of the navigation epoch impaired performance and induced mice to revert to habitual side biases. Inactivating during other epochs, including a delay period before the navigation phase, had little or no impact on behavior. The relationship between trial duration and behavioral performance was differentially impacted by hippocampal and mPFC inactivation, suggesting that the nature of the deficits was somewhat different.

      Strengths:

      The effects of perturbations are robust across animals and generally convincing. The lack of effect at some task epochs serves as a nice internal control. The finding that hippocampus and mPFC inactivation produced subtly different effects is interesting.

      Weaknesses:

      The simplicity of the behavior makes it difficult to resolve how exactly the hippocampus and mPFC contribute to working memory. Also, the language does not always reflect the trends in the data: the authors claim that optogenetic perturbations cause mice to repeat previous choices, but the data show that perturbations increase the likelihood of choosing a preferred side (which is left for most mice). A side bias is not the same as choice repetition. This has implications for interpreting the nature of the behavioral effects.

    3. Reviewer #2 (Public review):

      Summary:

      The study uses transient optogenetic silencing of the dorsal hippocampus or prefrontal cortex in mice using a delayed response task in a T-maze, with a temporal delay of 1s after a visual cue, followed by a central stem run period before choice execution. Silencing of either the dorsal hippocampus or the dorsal prefrontal cortex is executed for varying time periods during the stem/ central arm running epoch or the 1s temporal delay epoch by targeting either PV+ or Dlx interneurons in a block-wise or random-trial design. The main result reported is that silencing of either region during long periods of stem running impaired choice behavior, and silencing during the temporal delay period did not have an effect.

      Strengths:

      The major strength of the study is using the optogenetic silencing strategy to target different temporal periods of the task.

      Weaknesses:

      (1) A major weakness of the study is the lack of a balanced design with equal time periods of silencing during the stem running period and temporal delay period in many of the animals, which precludes any conclusion about distinct functional roles of the regions during these two phases of the task. The central question of this study is not new, with many previous studies investigating distinct and overlapping roles of hippocampus and prefrontal cortex in spatial working memory tasks and memory-guided navigation, using inactivation of one or both regions, crossed inactivation approaches, as well as targeting direct and indirect connections between the regions (PMIDs: 20074655, 9030646, 17045348, 30179661, 27511010, 10491611, 26017312, etc.), in addition to several physiology studies. A key extension for the current study would have been to show a distinction between roles in the temporal delay period after the cue and the stem-running working memory period. However, the inactivation period during stem running shows effects only for long inactivation periods, 2s initial periods and later ~1.6s periods (run after 0.8s, >2/3 total running periods), whereas the temporal delay period inactivation is 1s for the large majority of animals, which is a clear mismatch in inactivation periods, obviating this conclusion of distinction.

      (2) It is not clarified why such short delay periods were used compared to long periods of ~10s in T-maze spatial alternation tasks with delay, and whether the temporal delay period of 1 sec is strictly distinct from the spatiotemporal delay period during stem running in terms of short-term memory function. The choice of time windows for inactivation needs to be better justified, which currently appears to be rather random (2s initial running period, run after 0.8s, run after 1.6s; for an average reported running period of ~2.4-2.5s). The ideal design clearly would have been to use a 2s temporal delay period so that the inactivation time in this epoch matched the initial 2s run period. Only a subset of 4 animals were run with a longer temporal delay period, and too with only with hippocampal inactivation (Figure 4d). The main conclusion of distinction between temporal delay and stem running delay periods is therefore not adequately tested for the prefrontal cortex, and the statistics in terms of number of animals for this important control for hippocampal inactivation are also not comparable to the main experiment.

      (3) The mixture of PV-Cre animals (5 animals), Dlx targeting (2 animals), and one WT animal is also suggestive of a fragmented approach, and inactivation efficacy cannot be assumed to be similar for different animals. Importantly, there is no physiological evidence for confirmation of suppression in the optogenetic experiments, even in exemplar animals.

    4. Reviewer #3 (Public review):

      The authors sought to determine when the dorsal hippocampus and mPFC are causally required during a delayed spatial working memory task. Using temporally precise optogenetic silencing in mice performing a delayed cue-guided T-maze task, they tested the effects of transient perturbations during distinct behavioral epochs. Contrary to the common view that these regions are primarily required during the delay period to maintain working memory representations, they report that silencing during the delay had little effect on performance, whereas perturbation during the early phase of central-arm traversal consistently impaired performance. The authors conclude that hippocampal and prefrontal contributions to memory-guided behavior are dynamically engaged during active navigation rather than passive maintenance of information.

      The study has several strengths. The behavioral paradigm is carefully designed to dissociate cue, delay, and movement epochs, allowing temporally specific causal manipulations. The systematic comparison of multiple task epochs represents a major strength and provides compelling evidence that the behavioral effects of perturbation are epoch-dependent. The authors also include several important controls, including stimulation during multiple task phases, a longer-delay condition to dissociate task epoch from elapsed time since cue presentation, and analyses of movement trajectories and perseverative behavior that provide additional insight into the nature of the behavioral deficits. Together, these experiments convincingly demonstrate that transient dorsal hippocampal and mPFC perturbations have markedly different behavioral consequences depending on when they occur within a trial.

      The evidence is generally solid and supports the primary finding that perturbations during early navigation produce larger impairments than perturbations during the delay period. However, some aspects of the broader interpretation are less well supported. Most notably, the study lacks a non-memory control task, such as a visually guided version of the maze, making it difficult to determine whether the observed deficits specifically reflect disruption of memory-guided behavior or more general impairments in action selection, behavioral flexibility, or movement planning. The observed increase in perseverative responding and delayed commitment to a turn are consistent with either interpretation. In addition, several experimental conditions rely on relatively small numbers of animals, limiting confidence in some negative findings, particularly for the longer-delay and later-run manipulations. Finally, while the Discussion proposes that hippocampal-prefrontal circuits become engaged during the transformation of stored information into action, this mechanistic interpretation remains speculative because no neural recordings accompany the causal manipulations.

      Overall, the authors achieve their primary aim of demonstrating that the behavioral consequences of dorsal hippocampal and mPFC silencing depend strongly on task epoch. The data convincingly support the conclusion that these structures are more vulnerable to perturbation during early navigation than during the brief delay period used in this task. The broader conclusion that these findings redefine when hippocampal-prefrontal circuits support working memory should be interpreted more cautiously, as alternative explanations involving action selection or behavioral state remain plausible in the absence of additional control tasks.

      The findings are potentially important because they challenge the common assumption that hippocampal and prefrontal contributions to delayed-response tasks are centered on delay-period maintenance. Instead, the work supports the idea that these circuits may be recruited when remembered information is translated into goal-directed behavior. This framework is broadly consistent with recent distributed models of working memory and provides an interesting perspective that may help reconcile previous studies reporting effects during different task phases. The behavioral paradigm and temporally precise perturbation approach should also be useful for future studies aimed at dissecting the dynamic contributions of hippocampal-prefrontal circuits during memory-guided behavior.

    5. Author response:

      On the eLife Assessment. We appreciate the assessment’s recognition that this study addresses a long-standing question concerning hippocampal and prefrontal contributions to working memory. We think the broader significance lies in constraining how causal manipulations in working-memory tasks are interpreted. Working memory encompasses many processes distributed across a trial and showing that a region is required for a delayed-response task does not establish when its contribution is necessary. The observation that hippocampus and mPFC are required while navigating to a goal, but not during stationary delay, challenges a common assumption and, in our opinion, has implications across the broader field of working-memory research.

      We agree that the original manuscript did not adequately report effect sizes or convey the uncertainty associated with some smaller samples. However, the dataset does include duration-matched stationary and running conditions, including a hippocampal long-delay control matched to early-running stimulation in both duration and elapsed time after cue onset. The new effect-size and within-animal analyses support a robust hippocampal epoch difference, while the corresponding mPFC comparison is less precisely estimated and should be interpreted more cautiously.

      We therefore think the central finding remains well supported: hippocampal function, and potentially mPFC function, is required during the active navigation period of this memory-guided task but not detectably during the stationary delay. This does not establish the specific computation disrupted during running. Neural recordings would certainly provide further insight into the underlying mechanism, but their absence does not detract from the value of the behavioral result itself.

      Reviewer #1 (Public review):

      The simplicity of the behavior makes it difficult to resolve how exactly the hippocampus and mPFC contribute to working memory.

      The task was designed to combine the temporal precision of cue-based delayed-response paradigms with the behavioral richness of freely moving navigation. Few tasks combine a fixed cue-presentation period, an explicit delay, and a subsequent navigation phase involving extended running. This structure creates well-defined behavioral epochs that can be targeted with temporally precise perturbations, allowing us to ask when hippocampal and mPFC contributions are required within an ongoing memory-guided behavior. How these regions contribute is the harder question, and one we are pursuing next. Identifying when perturbations disrupt behavior is an important step toward understanding how these regions support memory-guided navigation

      Also, the language does not always reflect the trends in the data: the authors claim that optogenetic perturbations cause mice to repeat previous choices, but the data show that perturbations increase the likelihood of choosing a preferred side (which is left for most mice). A side bias is not the same as choice repetition. This has implications for interpreting the nature of the behavioral effects.

      We agree that our results do not clearly distinguish a directional bias from a tendency to repeat the previous choice. To examine whether mice consistently favored a particular direction, we compared their side preferences during silencing across sessions. Mice generally favored the same side across silencing conditions, although some switched direction in individual sessions (Author response image 1a). Within sessions, the preferred side was maintained from no-stimulation to stimulation trials in 13 of 19 cases and reversed in six (Author response image 1b). These observations are consistent with a directional preference that can sometimes reverse during stimulation, and cannot be disambiguated from perseveration. In the revision, we will describe the effect as increased directional bias and revise the language concerning choice repetition and perseveration throughout the manuscript.

      Author response image 1.

      Silencing increases directional bias. (a) Fraction of choices made to the right in each session, for every mouse (rows) and each silencing condition (symbols). Open symbols, no-stimulation trials; filled symbols, stimulation trials from the same session; blue and red denote a left or right preference during silencing. Filled symbols falling predominantly on the same side of 0.5 within a row indicate that a mouse generally favored the same direction across silencing conditions, although some mice switched direction. One session per mouse and condition, hippocampal silencing only; T2 and T6 did not perform the 0.8 s condition. (b) The same sessions expressed as signed bias, from no stimulation to silencing. Black lines mark the six sessions in which the preferred side reversed; grey lines the thirteen in which it was maintained.

      Reviewer #2 (Public review):

      A major weakness of the study is the lack of a balanced design with equal time periods of silencing during the stem running period and temporal delay period in many of the animals, which precludes any conclusion about distinct functional roles of the regions during these two phases of the task. The main conclusion of distinction between temporal delay and stem running delay periods is therefore not adequately tested for the prefrontal cortex, and the statistics in terms of number of animals for this important control for hippocampal inactivation are also not comparable to the main experiment.

      We agree that matching stimulation duration is an essential control and recognize that the relevant comparisons and statistics were not sufficiently clear in the original manuscript. Four hippocampal conditions used closely matched stimulation durations of 2 s: Cue+Delay, long delay, early running, and running with a 0.8 s onset (Author response image 2a). Crucially, the long-delay control matched both stimulation duration and elapsed time after cue onset to the early-run condition, while the mouse remained stationary.

      Author response image 2b–c shows the estimated impairment and its 95% confidence interval for each condition. In the hippocampal experiments, the duration-matched stationary conditions showed effects close to zero, whereas the running conditions showed large impairments. Despite the smaller sample, the upper confidence limit for the long-delay impairment was approximately 10 percentage points, substantially below the observed early-running impairment. These estimates establish that despite the smaller sample, the data support a lack of effect compared to early running.

      Author response image 2.

      Duration-matched stimulation produces different behavioral effects across task epochs. (a) Stimulation timing relative to cue onset. Numbers within bars indicate calculated median stimulation duration in seconds; black ticks indicate door opening. Bold labels identify conditions with 2 s stimulation. (b-c) Mean impairment in choice accuracy for hippocampal and mPFC manipulations. Impairment is accuracy during baseline minus accuracy with stimulation, in percentage points. Error bars show 95% confidence intervals across animals; numbers indicate mice. Open circles denote single-animal observations, and arrows indicate confidence intervals extending beyond the plotted range.

      For mPFC, the duration-matched Cue+Delay condition likewise showed an effect close to zero, whereas early-running stimulation produced substantial impairment. However, the long-delay condition included only one mouse. The later-running effects in both regions were also less precisely estimated. We will distinguish these limitations from the more informative stationary-condition results.

      To directly test whether the duration-matched effects differed across epochs, we next compared impairment within the same mice, including only animals tested in both conditions (Author response image 3). For the hippocampus, every mouse showed greater impairment during early running than during either duration-matched stationary condition, and the confidence intervals for both paired differences excluded zero. These within-animal comparisons support an epoch-dependent effect that cannot be explained by stimulation duration alone. The mPFC comparison showed the same direction of effect, although the confidence interval for the Cue+Delay versus running difference narrowly included zero. We will therefore distinguish the stronger evidence for the hippocampal epoch difference from the more limited evidence for mPFC.

      Author response image 3.

      Within-animal comparisons of duration-matched stimulation effects. (a–b) Impairment during Cue+Delay or long-delay stimulation compared with early-running stimulation for hippocampal (a) and mPFC (b) manipulations. Points represent individual mice, lines connect observations from the same mouse, and black bars indicate mean. Annotations report the mean paired difference in impairment (running minus stationary), and its 95% confidence interval calculated using the t distribution. Positive differences indicate greater impairment during running. The single-mouse mPFC long-delay comparison is descriptive. Blue indicates stationary epochs and orange indicates running.

      The central question of this study is not new, with many previous studies investigating distinct and overlapping roles of hippocampus and prefrontal cortex in spatial working memory tasks and memory-guided navigation, using inactivation of one or both regions, crossed inactivation approaches, as well as targeting direct and indirect connections between the regions (PMIDs: 20074655, 9030646, 17045348, 30179661, 27511010, 10491611, 26017312, etc.), in addition to several physiology studies.

      We agree that hippocampal and prefrontal contributions to spatial working memory have been extensively studied. That’s precisely why we find these results impactful when placed in the rich context of the field. The requirement for these regions in delayed working memory tasks has often been interpreted in terms of their contributions during the delay period. However, a requirement during a delay-based task does not itself demonstrate a requirement during the delay. Previous manipulations have not isolated delay periods of waiting from the subsequent navigation within a trial.

      Our experiments extend this work by separately targeting cue presentation, the delay, and different portions of navigation within the same task. This allows us to test whether the behavioral consequences of perturbation depend on the particular epoch in which it occurs. This distinction is important: knowing that a region is required for a memory-guided task does not establish when its contribution is needed. Identifying those periods constrains how we interpret the deficits produced by longer-lasting inactivation. We believe that this is an important result that should be considered when interpreting these broader findings. We will ensure that the appropriate literature and discussion are included in the revision.

      It is not clarified why such short delay periods were used compared to long periods of ~10s in T-maze spatial alternation tasks with delay, and whether the temporal delay period of 1 sec is strictly distinct from the spatiotemporal delay period during stem running in terms of short-term memory function.

      The 1 s delay was chosen to maintain reliable task performance, as some mice could not perform the task with longer delays. Indeed, one reason the longer-delay condition includes fewer animals is that two mice could not perform reliably with the 3 s delay (one additional mouse was not tested in this condition). Delays on this timescale have also been used in rodent cued delayed-response tasks, including a 0.5 s delay in Kopec et al. (2015), a 1.3 s delay in Guo et al. (2014), and a 1.2 s delay in Inagaki et al. (2019). Like these tasks, our paradigm requires mice to remember an externally presented cue specifying the upcoming response, rather than their own previous arm choice as in spatial alternation. It therefore combines spatial navigation with a cued delayed-response requirement, and the delay durations tolerated in alternation tasks are not necessarily directly comparable. We will clarify this rationale in the manuscript.

      We agree that the stationary delay and the subsequent run both require retention of information after cue offset. Our experiments distinguish these behavioral epochs, but do not establish that they involve separate short-term memory processes. The different effects of perturbation suggest that the contribution of these regions changes as the animal moves from waiting to navigating. Determining what accounts for this change is an important direction for future work.

      The choice of time windows for inactivation needs to be better justified, which currently appears to be rather random (2s initial running period, run after 0.8s, run after 1.6s; for an average reported running period of ~2.4-2.5s).

      We sought to target different portions of the central-arm run. Given the typical traversal time of approximately 2.4–2.5 s, stimulation onsets at 0, 0.8, and 1.6 s sampled the beginning, middle, and later portions of the run. Our setup allowed precise control of stimulation timing, and, as shown in Figure 4b, these onset times correspond approximately to the start, middle, and end of the central arm. Each condition used a nominal 2 s stimulation window, truncated if the mouse reached the choice point sooner. The windows therefore overlap, with the later-onset condition generally producing shorter stimulation. We will clarify this rationale and the distinction between stimulation onset and duration in the revised manuscript.

      The mixture of PV-Cre animals (5 animals), Dlx targeting (2 animals), and one WT animal is also suggestive of a fragmented approach, and inactivation efficacy cannot be assumed to be similar for different animals. Importantly, there is no physiological evidence for confirmation of suppression in the optogenetic experiments, even in exemplar animals.

      The Dlx animals were included to improve regional specificity through local viral expression and to test whether the behavioral effects were consistent across targeting approaches. Both approaches produced comparable impairments during running, supporting their inclusion in the same analysis. We will clarify this rationale in the manuscript.

      Optogenetic activation of inhibitory interneurons is an established approach for suppressing local principal-cell activity, with physiological validation in previous studies (Guo et al., 2014; Li et al., 2019, Zutshi et al., 2022). In our experiments, running-period stimulation produced robust behavioral impairments that were consistent across animals and targeting approaches. Furthermore, comparisons across epochs were performed within animals, using the same preparation and stimulation parameters. Differences in efficacy between animals therefore cannot readily account for the observed epoch dependence. Although direct recordings would establish the magnitude and spatial extent of suppression in our preparation, the central behavioral finding is supported by these within-animal comparisons.

      Reviewer #3 (Public review):

      The findings are potentially important because they challenge the common assumption that hippocampal and prefrontal contributions to delayed-response tasks are centered on delay-period maintenance.

      We thank the reviewer for describing our findings as “potentially important” and for highlighting their implications for how hippocampal and prefrontal contributions to delayed-response tasks are understood. We appreciate the constructive suggestions and address the public comments below.

      Most notably, the study lacks a non-memory control task, such as a visually guided version of the maze, making it difficult to determine whether the observed deficits specifically reflect disruption of memory-guided behavior or more general impairments in action selection, behavioral flexibility, or movement planning. The observed increase in perseverative responding and delayed commitment to a turn are consistent with either interpretation.

      We agree that leaving the cue on throughout the trial would provide an important control for distinguishing memory-specific effects from broader effects on action selection or movement planning. The senior author is currently setting up a new laboratory, so implementing this control may take some time. We hope to include it in the revised manuscript.

      The running-period deficit indicates a disruption of processes that enable the animal to act on a remembered cue. Whether this reflects disruption of memory itself, movement planning, or another component of translating the cue into a choice requires further clarification but does not detract from the observed dependence on task epoch. Uncertainty about the mechanism of the running-period deficit also does not change the observation that the same manipulation produced no detectable impairment during the stationary delay, when the cue was absent and still had to be remembered. This will be clearly discussed in the revision.

      In addition, several experimental conditions rely on relatively small numbers of animals, limiting confidence in some negative findings, particularly for the longer-delay and later-run manipulations.

      We agree that small sample sizes limit the interpretation of some negative findings. We now report animal-level effect estimates and 95% confidence intervals for each condition (Author response image 2b–c; Author response table 1).

      Of the nine conditions with no detectable impairment and more than one mouse, seven had confidence intervals that excluded effects as large as the observed mean early-running impairment in the same region. This included the hippocampal long-delay condition, despite its smaller sample. These results argue against similarly large impairments in these conditions, although smaller effects remain possible. The two later-run conditions remained too uncertain to exclude such impairments. These and the two single-animal conditions are marked in Author response table 1 and will be interpreted cautiously.

      Author response table 1.

      Animal-level impairment estimates and uncertainty

      Finally, while the Discussion proposes that hippocampal-prefrontal circuits become engaged during the transformation of stored information into action, this mechanistic interpretation remains speculative because no neural recordings accompany the causal manipulations.

      We will clarify in the Discussion that the proposed transformation of stored information into action remains speculative. Nevertheless, we believe this is an exciting possibility raised by our findings that warrants further investigation. Neural recordings would help test this interpretation but are beyond the scope of the current paper. We plan to explore this question in future work.

      References

      Guo ZV, Li N, Huber D, Ophir E, Gutnisky D, Ting JT, Feng G, Svoboda K (2014). Flow of cortical activity underlying a tactile decision in mice. Neuron 81:179–94. doi:10.1016/j.neuron.2013.10.020. PMID 24361077.

      Kopec CD, Erlich JC, Brunton BW, Deisseroth K, Brody CD (2015). Cortical and subcortical contributions to short-term memory for orienting movements. Neuron 88:367–77. doi:10.1016/j.neuron.2015.08.033. PMID 26439529.

      Inagaki HK, Fontolan L, Romani S, Svoboda K (2019). Discrete attractor dynamics underlies persistent activity in the frontal cortex. Nature 566:212–217. doi:10.1038/s41586-019-0919-7. PMID 30728503.

      Li N, Chen S, Guo ZV, Chen H, Huo Y, Inagaki HK, Chen G, Davis C, Hansel D, Guo C, Svoboda K (2019). Spatiotemporal constraints on optogenetic inactivation in cortical circuits. eLife 8:e48622. doi:10.7554/eLife.48622. PMID 31736463.

      Zutshi I, Valero M, Fernández-Ruiz A, Buzsáki G (2022). Extrinsic control and intrinsic computation in the hippocampal CA1 circuit. Neuron 110:658–673.e5. doi:10.1016/j.neuron.2021.11.015. PMID 34890566.

    1. eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRSIPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. The analysis of reciprocal daughter-cell pairs provides compelling evidence for SCE events, providing evidence consistent with CDK1-TTF2-TRAIP mediated cell-cycle regulated CMG helicase disassembly and fork cleavage at unreplicated regions.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9-induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs.

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

    3. Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed large-scale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations.

    4. Reviewer #3 (Public review):

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are « genetically silent ». Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly « permissive » for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRISPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole-genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. However, the evidence supporting the proposed involvement of under-replicated region/replication termination-zone resolution and TRAIP/URR-like pathways is currently incomplete and could be strengthened with an increased number of reciprocal daughter-cell pairs and by genetic or molecular perturbation, or alternatively, this can be addressed by changing the discussion.

      We appreciate the editor’s and the reviewers’ recognition of Cas9-induced SCE as an important previously invisible repair outcome and of RDCP analysis as a notable feature of the study. As proposed, we incorporated the two recent Science studies and now tone down the Discussion of our RDCP observations as consistent with, rather than definitive evidence for, the TRAIP-dependent pathway. We also detail how our single-cell genomic observations complement and extend these two studies in terms of the biological significance of the CDK1-TTF2-TRAIP axis.

      Specifically, these changes are in:

      Discussion. We changed

      “Two recent studies revealed how the CDK1-TTF2-TRAIP axis is cell-cycle regulated to trigger mitotic CMG helicase disassembly and fork cleavage: one study showed a two-fold SCE reduction in mouse ES cells (Fujisawa and Labib, 2026), while the other showed that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions (Can et al., 2026). Our observation of the "WWC-or-WCC/deletion pair" signature in wild-type cells provides, to our knowledge, the first genetic evidence of linking a deletion with SCE and revealing both W and C unreplicated template strands present in the reciprocal daughter cell, consistent with this mechanism at single-cell genomic resolution (illustrated in Fig.3), although this is limited by the observation of only one RDCP.”

      While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells.

      Fig.3. Title and legends. We changed

      “Haplotype-aware analysis of observed RDCP (Pair 4, chr1) shows SV patterns at the SCE junctions consistent with the predicted RDCP signature of SCE mediated by URRs or replication termination zones (green shaded area), although the lagging strands, rather than the leading strands, must be resolved to generate these mitotic breaks.”

      Reviewer #1 (Public review):

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest. The evidence for structural complexity associated with some induced SCEs is intriguing, but the mechanistic interpretation should either be tested directly or presented more cautiously.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs. Given that potential, the current manuscript would benefit greatly from any experiments characterizing this sub-population: are these cells in a particular cell cycle state, experiencing changes in gene expression, or do they have other unique biological properties?

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

      Thank you very much for this assessment.

      Weaknesses:

      The number of informative RDCPs is limited, and the mechanistic interpretation of the "WWC-orWCC/deletion" signature is more suggestive than definitive. In particular, the manuscript invokes (even though only in the Discussion section) URR or replication-termination-zone resolution and discusses TRAIP-dependent CMG unloading, nuclease cleavage, and polymerase theta-mediated joining, but these pathway components are not directly tested herein. A more conservative conclusion that some Cas9-associated SCEs coincide with structural alterations is more appropriate, particularly in the Discussion and Conclusion. For example, the statement that this work provides "direct genetic evidence" for a URR-type mechanism is overstated unless supported by additional experiments or a more extensive analysis of alternative models. Similarly, while the authors explain the limitations of acute Cas9 disruption of LIG3, LIG4, XRCC1, and XRCC4, the manuscript should clarify what biological questions this experiment can and cannot answer.

      Please see response to the eLife Assessment as this is a common point raised by multiple reviewers.

      Additionally, we clarified what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) by adding the following text in the “Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell” section:

      “One additional complication is that, although Cas9 RNP achieves >90% knockout efficiency in bulk assays and is therefore used as a substitute for siRNA, sorting BrdU-labeled cells in the subsequent G1 enriches for cells that escaped frameshift editing, particularly for essential genes; thus, 100% knockout in 90% of cells is not equivalent to 90% knockdown in every cell, representing a unique challenge for single-cell assays.”

      Reviewer #1 (Recommendations for the authors):

      (1) Temper the mechanistic claims about URR/TRAIP-type resolution.

      The RDCP data support the conclusion that some Cas9-associated SCEs are accompanied by structural alterations and may arise through non-classical mechanisms. However, claims about TRAIPdependent CMG unloading, URR resolution, or polymerase theta-mediated joining should be framed as a model unless directly tested.

      We cited mechanistic dissection of the CDK1-TTF-2-TRAIP axis, which was published since the review of the paper. While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells but qualified that this observation is in only one RDCP. We further tempered our claims by changing “direct evidence” to “consistent with” as we did not perturb genes involved in these processes.

      (2) Clarify the impact of the small RDCP sample size.

      The manuscript would be strengthened by explicitly stating how many total RDCPs were analyzed, how many SCE events were informative, and how much confidence can be placed in the estimated fraction of SCEs associated with SVs. A short table summarizing RDCP counts, SCE counts, copy-neutral events, and SV-associated events would be helpful.

      A total of 15 RDCPs were recovered from close to 4,000 single cells analyzed across all conditions. We added Tab.S3 detailing SCEs in all 15 RDCPs in addition to SCE and SV breakdown in Tab.S2 (originally Tab.S1). The new Tab.S3 is cited in the “RDCP analysis reveals large-scale SVs on chromosomes with induced SCE, as well as structural alterations at Cas9-induced SCE junctions” section.

      (3) Provide more detail on the "rescued" SCE calls.

      Because the central conclusions rely on SCE detection, the criteria for breakpoint R-based calls versus rescued calls should be explained clearly in the main text or methods. It would be useful to know how sensitive the main conclusions are to the inclusion or exclusion of rescued calls. Is this laid out in greater detail in an additional manuscript?

      We previously included an “On-target SCE identification” section in the Methods, where we described in detail the rescue of on-target SCEs missed by the initial breakpoint R calls. We also depicted calls and calls+rescues for all the conditions in Fig.S1B.

      We agree with the reviewer and now expanded the description of rescued SCE calls in the main text (under the “A single Cas9 DSB induces potent local SCE” section). In brief, 50-92% of SCEs (typically >70%) across the four single-targeting sites were directly called rather than rescued, with the exception of LIG4, where only 30% were direct calls. This is because LIG4 is located only 6 Mb from the telomere and is therefore particularly prone to missed breakpointR calls in low-coverage cells. On-target rescue at individual sites is self-contained in this manuscript because our lab primarily focuses on spontaneous SCE, for which there are no expected SCE sites. However, the rescue methodology was previously implemented in the original development of sci-L3-Strand-seq to identify SCEs at centromeres.

      (4) Clarify the biological interpretation of the repetitive-target enrichment.

      The high-SCE subset analysis is interesting, but the manuscript should explain whether these cells have evidence of higher RNP uptake, altered cell-cycle state, greater DNA damage, or lower sequencing quality. If these possibilities cannot be distinguished, the text should state this clearly.

      We agree with the reviewer. The high SCE subset does not have lower sequencing quality by coverage or background (0.3% coverage for high-SCE vs. 0.28% coverage overall, and 3% background for both high-SCE and overall). However, our current data do not allow us to distinguish among biological explanations for the elevated SCE. The original manuscript acknowledged this limitation (“Whether this reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment, remains to be determined.”). To make this limitation more explicit and to address the possibility of sequencing quality raised by the reviewer, we have revised the text as follows: “The high-SCE subset did not show evidence of lower sequencing quality, based on either sequencing coverage (p=0.17) or background SCE levels (p=0.13). However, we cannot distinguish whether the elevated on-target SCE reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment.”

      This question may be better explored by future co-assays with sci-L3-Strand-seq; currently we cannot enrich for cells with high SCEs to characterize the molecular features of this subset of the cells using other omics approaches.

      (5) Reconsider the framing of the DNA repair gene targeting experiment.

      The current data do not strongly test whether LIG3, LIG4, XRCC1, or XRCC4 regulate Cas9-induced SCE, because functional protein loss is delayed and essential-gene targeting introduces selection. This section may be better framed as a negative/control observation rather than as a pathway analysis.

      We agree and please refer to Public Reviews for a single-cell assay-specific explanation.

      (6) Consider including some additional control experiments, for example, Cas9 without sgRNA, nontargeting sgRNA, or mock-transfected cells to make sure that some phenotypes (for example, cell-cycle arrest) directly result from DNA cleavage rather than from the transfection procedure.

      We thank the reviewer for this suggestion. We have carefully considered these additional controls but have chosen not to add further experiments. Our existing Cas9 nickase experiments provide a control that directly addresses whether the observed arrest is attributable to DSB formation rather than RNP delivery/transfection. Both the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9, but neither produced the cell-cycle arrest observed following DSB induction by wild-type Cas9. Thus, these experiments control for Cas9 RNP delivery while altering the nature of the DNA lesion and support the interpretation that the observed arrest is associated specifically with Cas9-induced DSBs rather than the transfection procedure itself.

      (7) Figure 1: the fonts should be increased. The majority of the labels are impossible to read in a printed copy of this manuscript.

      We thank the review for pointing this out. We enlarged Fig.1 fonts.

      (8) Figure 1C. The pileup plots should be described and interpreted in a clear way. In its present form, it is unclear how the interpretations and conclusions are made.

      We added explanation of the pileup analysis immediately following mentioning the Fig.1C pileup: “We next examined … SCEs using genome-wide pileup analysis (Fig.1C, Fig.S1B), in which we plot the total number of SCEs detected across all single cells within each 1 Mb window.”

      Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed largescale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Thank you very much for this assessment.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations. The language and logic in the paper can be improved, and some of the claims seem incorrect. For example, the abstract reads "A single Cas9 cut at a unique genomic locus led to strong local enrichment of SCE at the break site, reaching up to 41% in the same cell cycle and 17% in the subsequent division, indicating that DSB repair frequently engages non-local inter-sister repair." The evidence that only a single Cas9 cut was made is lacking (see my earlier comment); it is not clear how local enrichment or non-local inter-sister repair are defined.

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We also agree that novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome these limitations, perhaps by using vfCas9 but more importantly, if new approaches to turn off Cas9 are developed. We have added a brief discussion on this limitation (in the Limitation section) and the resulting constraints on extrapolating mechanisms of DNA instability and repair from the observed genomic rearrangements. We thank the reviewer for pointing out the distinction between “a single Cas9 cut” vs. "Cas9 targeting of a single genomic locus." We went through the manuscript and revised where cutting only once was implied. We also explicitly acknowledge the possibility of multiple rounds of cutting at the same sites.

      We thank the reviewer for pointing out that “non-local repair” is a non-standard term. We use it operationally to distinguish repair confined to the broken chromatid (e.g., fill-in synthesis or end joining in cis) from repair involving exchange between sister chromatids. We have added a schematic (Fig. S1A) illustrating this distinction and cited this figure immediately before where we operationally defined SCE as a “reciprocal strand switch between sister chromatids, without implying a single mechanistic pathway.” This distinction is important because, particularly for two-ended Cas9 DSBs, an SCE-like outcome could potentially arise through either HR-mediated crossover or NHEJ of DNA ends across sister chromatids; the latter may involve different genetic requirements from classical NHEJ at least in end-tethering. We have revised the manuscript to define “non-local repair” explicitly at its first use.

      Reviewer #2 (Recommendations for the authors):

      References to relevant earlier studies using Strand-seq to study SCEs are missing (PMID: 27185886 and PMID: 29348659).

      We thank the reviewer for pointing this out. We added these references in the 3rd paragraph of the Introduction where we briefly review Strand-seq methods.

      Reviewer #3 (Public review):

      Summary:

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are « genetically silent ». Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly « permissive » for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

      Strengths:

      This is an interesting paper that molecularly explores sister chromatid exchanges, which represent an important challenge in molecular biology since they are genetically silent.

      Thank you very much for this assessment.

      Weaknesses:

      A complexity of the current paper is that it heavily relies on a recently published paper (Chovanec et al 2026, NAR) describing the powerful but complex technique sci-L3-Strand-seq. Knowledge of this paper is a prerequisite to understanding the current manuscript because no reminder is provided. In addition, the current manuscript presents the use of the sci-L3-Strand-seq technique in the study of SCE after Cas9-induced DSBs, while a companion study is referred to several times for containing results about SCE in XRCC1 KO. At some point, one questions the relevance of splitting the use of sci-L3-Strandseq in different papers instead of making a single integrated one.

      We appreciate this concern. The original sci-L3-Strand-seq study is an extensive methodology paper that establishes and validates various computational framework, whereas the companion study focuses on the genetic regulation of spontaneous SCE. The experimental designs and biological questions of the companion study and the present work are therefore distinct, although we draw on selected results from the companion study where they provide useful comparisons and contrasts between spontaneous and Cas9-induced SCE. The present study addresses a distinct biological question, the response to Cas9-induced DSBs, and we therefore believe that combining these studies would make the resulting manuscript unnecessarily broad and obscure their different biological questions.

      We nevertheless agree that the present manuscript should be understandable without requiring detailed knowledge of either paper. We have therefore added a brief description of the sci-L3-Strandseq approach (3rd paragraph of Introduction, Fig.S1A legends, and Fig.1B legends) and clarified the relevant methodological concepts where they are first introduced. We hope to improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract

      "Identical sisters ": redundant

      "non-local" inter-sister repair: the meaning is not clear. Do the authors refer only to "inter-sister" and therefore "non-local" is redundant, or do they imply something specific by "non-local", in which case it needs to be clarified?

      "237 repetitive targets": at least a slight description of this target is needed. Is it a "random" repeat, a satellite sequence, a sequence related to a transposable element ?...

      We agree with the reviewer that “identical sisters” is technically redundant. However, we have retained “identical” here to emphasize the distinction between sister chromatids vs. homolog, as SCE is sometimes misconstrued as exchange between homologs and as potentially causing loss of heterozygosity. We prefer the slight redundancy here for conceptual clarity.

      We thank the reviewer for pointing out that the meaning of “non-local” was unclear. As discussed in our response to Reviewer #2, we use “non-local repair” operationally to distinguish repair confined to the broken chromatid in cis from repair involving exchange between sister chromatids. We have added a schematic (Fig.S1A) illustrating this distinction and explicitly define the term at its first use in the revised manuscript. Please see our response to Reviewer #2 above for the detailed rationale.

      We thank the reviewer for asking us to clarify the nature of the 237 repetitive targets. The sgRNA targets an Alu sequence and was selected from a larger screen of >20,000 sgRNAs targeting repetitive sequences occurring at >200 genomic sites. In that screen, cellular toxicity did not simply scale with the number of predicted target sites; we therefore selected this sgRNA because its intermediate phenotype allowed us to introduce a large number of programmed DSBs without either minimal perturbation or excessive loss of cells. Thus, the 237-site guide was not an arbitrarily selected Alu-targeting sgRNA. The full repetitive-element screen is beyond the scope of the present study, but we have clarified in the Abstract that these 237 sites are Alu targets and added a brief description of the guide selection in the Methods.

      (2) Introduction:

      "non-local outcome / non-local repair processes": The use of "non-local" is not standard and is obscure for the reader. Specify if it has any meaning or remove it.

      Please see our response above to both Reviewers #2 and #3 regarding our definition and use of “nonlocal repair”.

      The authors mention that replication through a DSB generates four broken ends. However, in case the DSB is reached by one replication fork before the converging one, there are only three broken ends for at least the time required for the converging fork to reach the DSB from the other side. This may influence the repair outcome.

      We agree with the reviewer. If one replication fork encounters the DSB before the converging fork, a transient three-ended intermediate can exist before the second fork reaches the break. This temporal asymmetry could influence repair pathway choice, including engagement of HR, end joining, or BIR-like repair. Our assay captures the resulting SCE outcome but cannot distinguish the order in which replication forks encounter the DSB or the repair pathway engaged at these intermediate stages. We have revised the text (2nd paragraph of the Introduction) to clarify that four broken ends represent the eventual configuration after replication through the DSB, rather than necessarily a simultaneous intermediate.

      (3) Results

      Cell cycle arrest experiment: it seems that a control condition with no Cas9 is missing to conclude better about what looks like a G2-M arrest, but that is not clearly mentioned.

      Please see our response to Reviewer #1, Recommendation 6, regarding additional controls for the cell-cycle arrest experiment. Briefly, the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9 but did not produce the cell-cycle arrest observed following DSB induction, providing an internal control for RNP delivery/transfection and supporting the association of the arrest with Cas9-induced DSBs. We would also like to clarify that the observed cell-cycle arrest is primarily a G1/S, rather than G2/M, basing on the FACS signal (see revised Fig.1B legend). This is consistent with the strong G1/S checkpoint in mammalian cells and the predominantly G1 cell-cycle distribution of BJ-5ta cells.

      Note that the font size in Figure 1 is too small for readability.

      We have enlarged the font sizes throughout Figure 1 to improve readability.

      Figure 1B, D10A and H840A conditions:

      The authors mention that nicks can be converted into DSBs through the passage of the replication fork, but do not see any cell cycle defect in the conditions tested. Is it possible that the absence of effect results from the fact that the analysis is done prior to nicks being converted into DSBs? This remark notably applies to the 237 target sites experiment. It seems that controlling for cell cycle delays for longer times is needed to conclude clearly about this aspect. In case a clear absence of cell cycle delay is observed in the 237 target sites in the Cas9 nicking condition, this would suggest that replication born DSBs behave differently from "classical" two-ended DSBs and do not trigger cell cycle arrest.

      We agree that the timing of nick conversion during replication could contribute to the absence of a detectable cell-cycle delay under the conditions examined. However, extending the duration of Cas9 nickase treatment or labeling would not necessarily resolve this question, because persistent Cas9 activity permits repeated rounds of nicking across successive cell cycles, making it difficult to relate a later cell-cycle phenotype to a defined replication-born lesion. More generally, we believe that the relationship between replication-associated nicks, SCE formation, and cell-cycle progression is better addressed in the context of spontaneous SCE, which is the focus of our companion study. The present study is focused on SCE following programmed Cas9-induced DSBs, and analysis of replication-born nick lesions would require precise temporal control (ideally vfCas9 nickases that can be turned off) of individual nicking events relative to replication, for which an appropriate experimental system is not currently available to us.

      Figure 1C should mention somewhere the genomic location of the four targets to clearly show that they correspond to the four major SCE peaks. In addition, there is no legend for the vertical pink stripes. Finally, it might be wise to keep the same y-axis scale for better comparisons.

      Figure 1 overall: it might be wise to clearly show a no Cas9 condition to clearly set the SCE baseline and show that it is independent of Cas9 induction. As of now, it is not clear whether the non-targeted SCE comes from a specific cleavage of Cas9 or not. Such an aspect could benefit from putting Figure S1C in the main Figure 1. Alternatively, results from Chovanec et al 2026 (NAR) should be better restated because the reader does not necessarily have them in mind.

      We thank the reviewer for these suggestions. The expected Cas9 target positions were already indicated by vertical bars in Fig. 1C; however, we agree that this was not sufficiently clear. We have therefore revised the figure legend to explain that the vertical bars indicate the expected Cas9 target positions. The Chovanec et al. (2026, NAR) study focused entirely on spontaneous SCE, which we simultaneously map here as the background signal, rather than the on-target SCEs induced by Cas9. We hope that explicitly identifying the target locations in the revised legend makes this distinction clear and ensures that prior knowledge of the NAR study is not necessary to interpret Fig. 1C.

      We have retained the individual y-axis scales because the magnitude of SCE enrichment differs substantially among conditions. Using a common y-axis scale would make several of the on-target SCE peaks difficult to visualize.

      The section « Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell » is questionable in the results section for the following reasons:

      (i) The DNA repair genes are used here as target sites for Cas9 cleavage, but are not the object of the study, but may be the object of a companion paper. This aspect is slightly misleading.

      (ii) As first mentioned in this section, there is evidence strongly suggesting that inactivating DNA repair genes will not affect SCE, and this is what the authors observed.

      (iii) As an alternative, one could put the emphasis on the fact that the effect of Cas9-mediated inactivation of DNA repair genes (ie LIG3) starts to be detectable only in the washout condition ie after at least one cell cycle. But in this case, this is addressing the role of DNA repair genes in Cas9-induced SCE, which is not the point of the current paper.

      We thank the reviewer for this comment and agree that the original framing of this section could give the impression that these experiments were intended to test the functions of the targeted DNA repair genes in SCE. This was not our intent. Rather, these genes provided defined genomic target sites for Cas9 cleavage, and the primary purpose of the experiment was to characterize SCE associated with Cas9-induced DSBs at these loci.

      As discussed in our response to the Editor Assessment, there are important limitations to using these experiments to infer the consequences of loss of the targeted proteins, including the delay between Cas9 cleavage and depletion of pre-existing protein and, and particularly in the single-cell assay, selection for cells that escape disruptive editing at essential genes. We have added text to the Results explicitly describing these limitations.

      We therefore agree with the reviewer that the delayed effects observed under the washout condition should not be interpreted here as establishing a role for individual DNA repair genes in Cas9-induced SCE. We have revised the section title to “Cas9 targeting of DNA repair gene loci did not immediately alter overall SCE frequency per cell” to clarify the scope of this experiment and to avoid implying that testing the functions of the targeted DNA repair genes is a major objective of the present study.

      The conditions in Table 1 need to be homogenized and better explained:

      - 237 sites and 237 cuts are used: homogenize?

      We thank the reviewer for spotting this. We revised both to be “237 sites”.

      - May explain better the rationale for putting BrdU simultaneously with Cas9 or after 24 h and a wash.

      For the single-targeting sites, we observed more SCE when BrdU was added simultaneously with the Cas9 for the same 24 hours, compared to adding BrdU in the subsequent division after a wash. Therefore, for the 237 sites, we analyzed both conditions.

      - Typo in the text: 237cuts_24ws_40BrdU instead of 237cuts_24ws_BrdU

      We apologize for the lack of clarity in these labels and have substantially revised the Table 1 legend. In brief, the “40” is not a typo. In the 237 sites experiments, wild-type Cas9 considerably prolonged the cell cycle. Therefore, rather than labeling with BrdU for 24 hours as in the other conditions, we extended BrdU labelling to 40 hours in the last two conditions to allow more cells to progress into the subsequent G1 for successful Strand-seq analysis. We clarified that “237 sites 24 + 16hrs BrdU” refers to the condition in which Cas9 RNP and BrdU were added simultaneously. After 24 hours of Cas9 RNP treatment, BrdU labeling was continued for an additional 16 hours (a total of 40 hours of BrdU). The “237sites 24ws40BrdU” condition is the corresponding washout condition, in which Cas9 RNP was removed after 24 hours and cells were then labeled with BrdU for 40 hours post-washout.

      - 237 cuts: Are some sites more enriched in SCE than others?

      Yes, some sites showed greater SCE enrichment than others. We tested whether this variation correlated with chromatin accessibility but found no significant association. This was not unexpected, as the sgRNAs predominantly target Alu elements.

      - Table 1: There is a difference between 237 sites 24ws24BrdU and 237cuts24ws40BrdU, with a significant enrichment of on-target SCE for the latter condition only. Could the increase in SCE rise even more with longer BrdU exposure? In other words, does the low enrichment in SCE at target sites in the 237 sites experiment result from a non-optimal timing for the analysis?

      Yes, this is possible. We did not systematically test additional treatment or labeling durations. A 24-hour Cas9 RNP treatment is typically used for Cas9 RNP-mediated knockout experiments, and we therefore initially used this duration to assess gene-editing outcomes. For Strand-seq, BrdU labeling is ideally limited to approximately one cell division. Because BJ-5ta cells have an approximately 24-hour cell cycle, extending BrdU labeling substantially beyond 40 hours could allow some cells to undergo a second round of replication and become double-labelled. We therefore did not extend BrdU labeling beyond 40 hours. Thus, the lower enrichment in the 24ws24BrdU condition may in part reflect the timing of the assay.

      - Figure 2 / RDCP analysis:

      Interpretation of this figure relies exclusively on the 2026 NAR paper from the authors. This, at least, should be mentioned to help the reader understand it. Once the legend restates, this figure misses clear identification of the SCE and other genomic rearrangements. For readability, maybe the full genome should be kept for the supplementary data, and only the rearranged chromosomes should be kept in the main figure so that the rearrangements are clearly visible and annotated.

      We thank the reviewer for this suggestion. To make the Strand-seq plots interpretable without relying on our 2026 NAR paper, we have added an explanation of Strand-seq orientation in the third paragraph of the Introduction and in Fig.S1A. We have revised Fig.2 legends to improve readability. We have retained the whole-genome view because Strand-seq data are conventionally presented in this format and it provides important genome-wide context for interpreting the observed events. The rearranged chromosomes and events were annotated in Fig.S3.

      (4) Discussion

      - Most DSB never formed or did not produce SCE: how to understand this better? What would be the argument in favor of one or the other possibility?

      We agree that these are two possible explanations that cannot be distinguished by the current experiment. The absence of an SCE at a targeted site could reflect either inefficient DSB formation or repair of a DSB through a pathway that does not generate an SCE. Distinguishing these possibilities would require direct measurement of cutting efficiency at individual target sites, which was beyond the scope of this study.

      - The conclusion about the effect of the Cas9 nickases needs to be toned down as long as the proper timing for SCE analysis has not been performed (see comment above).

      We agree and have toned down this conclusion by specifying that no significant on-target SCE enrichment was detected under the conditions tested and acknowledging that we cannot exclude SCE formation at other time points (Discuss, first paragraph).

      - As much as possible, avoid the use of non-conventional acronyms like URR.

      We agree and have reduced the use of non-conventional acronyms where possible. We have retained URR (under-replicated region), as the term appears seven times throughout the manuscript, but have ensured that it is clearly defined at first use.

      - The discussion about the RDCP analysis in the second paragraph of the discussion should refer to Figure 3.

      Thank you for pointing this out. We added this reference to Fig.3

    1. eLife Assessment

      This study presents a valuable tool for comparing immune receptor data across multiple samples while properly accounting for statistical uncertainty and receptor similarity. The evidence supporting the tool is solid overall. However, some key concerns regarding potential sequencing artifacts in one of the validation datasets and an unclear strategy for false discovery rate control in the proposed framework remain to be addressed.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors present ClustIRR, a computational tool that analyzes multiple TCR repertoires together, instead of one at a time. It builds one shared similarity graph across all the repertoires, then finds communities (CJs) on that graph that can be compared across samples. Building one shared graph across repertoires, rather than clustering each repertoire on its own, is a real improvement over existing tools. However, there are a few concerns that need to be addressed.

      (1) The Introduction motivates ClustIRR by contrasting it with beta-binomial regression, Fisher's exact test, and scCODA/tascCODA (lines 62-77), but none of these are actually run on the same data in the Results. Could the authors include a direct comparison, e.g., applying a standard beta-binomial or Fisher's exact test to the Dataset 1 CJ occupancy matrix, to show the Bayesian model gives a lower false-positive rate or better-calibrated intervals than the alternatives it's positioned against?

      (2) Prior predictive and posterior predictive checks are both reported as "(data not shown)" (lines 644, 728). Since the manuscript's central claim is rigorous, uncertainty-aware inference, it would help to include these diagnostic plots, along with Rhat and ESS values, in the supplement rather than stating they were checked.

      (3) With 7,505 CJs tested simultaneously for differential occupancy in Dataset 1 alone (Figure 1 legend), what is the expected false discovery rate under the non-overlapping-HDI criterion used throughout? A short discussion of multiple-comparisons correction, or an argument for why it isn't needed under this framework, would strengthen the statistical claims.

      (4) Line 115 states the Dataset 1 joint graph produced 10,301 CJs, of which 3,038 were singletons, leaving 7,263 non-singleton CJs. The Figure 1 legend reports 7,505 CJs used for the β modeling. Could the authors clarify how these two numbers relate - whether some singletons were included in the model, or a filtering step was applied that isn't described in Methods?

      (5) The Methods section states that archival pretreatment tumor tissue was available for four patients (Pt4, Pt32, Pt36, Pt38; line 531), but the T+/T- DCJ analysis in Fig. 3B-C is only shown for Pt4. Was this analysis attempted in the other three patients? Extending it, even partially, would substantially strengthen the claim that contracting DCJs are enriched for tumor-infiltrating TCRs, which is currently based on a single patient.

      (6) Dataset 1 was generated by deliberately stimulating T cells with EBV or MART1 antigen, so recovering EBV/MART1-annotated CJs from VDJdb is closer to a positive control than a blinded validation. Do the authors have, or could they obtain, any independent confirmation (e.g., tetramer data or an orthogonal cohort) for the CJs with large β that lack VDJdb annotation (orange dots, Figures 1B-C)?

    3. Reviewer #2 (Public review):

      Summary:

      The study confronts a major obstacle in repertoire analysis. Given that individual TCR sequences are diverse and sparsely detected across repertoires, identifying sample-specific enrichment of TCRs based on sample-to-sample comparison of exact clonotype sequences can be intractable. This paper attempts to build on insights that sequence-similar TCRs can share antigen recognition, such that aggregating similar sequences derived from multi-sample joint-graph communities (i.e. clusters of tightly connected nodes) could reduce sparsity and boost signal.

      Strengths:

      The study is a well-motivated effort to address a need in the field. The paper takes a unique approach. The core method is modeling community occupancy with a hierarchical Dirichlet-Multinomial model that attempts to account for the high level of overdispersion present in repertoire sampling, a technique that has been previously applied to compositional microbiome data.

      The authors apply this framework to both single-cell paired-chain and bulk single-chain TCR data. They develop a set of examples from public and synthetic data, with the most promising real-world application shown in reanalysis of longitudinal data during treatment of cancer patients with checkpoint inhibitors.

      The manuscript is well structured and cogent. The authors are to be commended for contributing a well-documented, open-source R/Bioconductor package and for providing the underlying analysis datasets in a well-organized repository. In the joint graph construction step, the authors opt to use an existing implementation of the BLAST algorithm on CDR3 sequences, which is a slightly odd choice since it ignores potential contributions of other V-gene germline-encoded CDRs, but the authors also envision that their statistical package could be extended to include community graphs developed with other established TCR clustering tools. This will allow others to potentially explore the utility of Bayesian hierarchical Dirichlet-Multinomial models for differential occupancy analysis of immune receptors under varied clustering criteria.

      The methods described here were demonstrated on relatively small datasets from 2-5 samples, and future work is likely needed to extend the joint graph differential occupancy concept to larger datasets. The authors are transparent about this and some of the other limitations in their current tool, most notably the computational cost of graph construction based on an all-versus-all sequence alignment to construct a joint sequence graph and the challenge of Bayesian parameter estimation as the number of subgraph entities scales with input data size. Since efficient approximate methods exist to find edges between similar text strings, the underlying idea of applying uncertainty-aware statistical inference to subgraph communities is promising, and the paper advances its primary goal.

      Weaknesses:

      The paper proposes the utility of the joint-graph community occupancy framework through three examples. I discuss potential weaknesses apparent in each separate example in turn.

      (1) Weaknesses in Example 1

      A broad weakness of the first results section ("Detecting EBV- and MART1-antigen reactive T cell communities from single cell datasets") is its reliance on a single vendor-generated dataset generated by the company ParseBio with no published experimental protocols and limited, if any, prior peer review. For reasons I will explore in greater detail below, the EBV-sample data may be particularly prone to chimeric pairings that confound the authors' primary analysis goals, and, at the very least, may not reflect physiologically realistic conditions for identifying antigen-reactive TCRs in other contexts.

      Let us first consider Dataset 1 in more detail. The authors compare 2 antigen-stimulated repertoires with 3 unstimulated controls. To improve on single clonotype-level comparisons, the authors propose comparing the cell count aggregated across cells within joint graph communities constructed from paired CDR3 sequences across all the samples. Thus, one of the most relevant questions one hopes the authors answer in this section is whether the resolved graph communities are made up of many distinct clonotypes (i.e., are they polyclonal), allowing the method to function as intended by aggregating across multiple clones with putative shared antigen-reactivity.

      Supplementary Figure 1B shows the size of all the communities with callouts for the putative EBV-expanded communities. The authors listed communities strongly enriched in the EBV-stimulated sample as e1, e2, e3, and e5. Each contains {greater than or equal to}100 clonotypes, and the authors note they contain at least one clonotype with a CDR3 sequence matching an EBV-annotated clone in VDJdb - a database of TCRs with some experimental evidence of epitope-reactivity. Community "e2" is notable for its remarkable size, including 1,339 unique clonotypes. At first glance, this seems promising for a method attempting to boost signal through community detection. However, it is worth re-investigating the individual clone sizes and sequences within this extraordinary community.

      In the EBV-associated community "e2", a look at the data provided by the authors on the paper's GitHub repository suggests a single clonotype (clonotype_9; TRAV12-3 CATQGSNDYKLSF / TRBV9 CASSTGQVATNEKLFF) comprises 26,917 cells. As such, it makes up 29% of the sample with a total of 90,588 cells. If one clone supplies most of a graph community's cell counts, the community-level posterior estimate of β (Figure 1B) effectively tracks a single-clone estimate, and the premise of borrowing statistical power across a polyclonal expansion in this example is hard to assess.

      There is also considerable evidence to believe that the apparent mega-polyclonality of cluster e2 may be partially an illusion, stemming from an artifact of this hyper-expanded clone's massive size and the experimental method used to assign TCRα-TCRβ chain pairings. In fact, the same α-chain CDR3 (CATQGSNDYKLSF) appears in ~1,219 clonotypes paired to distinct β chains, generating much of e2's remarkable 1,339-clonotype count. I believe two features warrant caution here. First, a single clonotype making up ~29% of total T cells in the sample is highly unexpected in ex vivo repertoires, suggesting intense non-physiological expansion conditions unlikely to generalize to other settings. That is, one would almost never expect to see a signal this strong.

      Second, one TCR-α chain paired to ~1,220 distinct and diverse β chains in one sample is also biologically unexpected given what we know of the best-characterized epitope-specific responses for EBV, including to the well-known HLA-A*02 EBV BMLF-1 epitope, which recruits a tetramer-stained repertoire with conserved CDR3 motifs in both chains and at least some constraint on favored V-gene/α-β pairing (See Extended Data Figure 5 in Dash et al., Nature 2017). This raises the concerning possibility of barcode collision and mispairing against a hyperexpanded clone in the ParseBio split-pool method, unlikely to be robust to a clone occupying a third of the sample. Most of the ~1,219 β chains paired to the dominant α have a cell count of 1, further raising the concern of artifactual pairing versus genuine convergence. The second largest community "e1" also seems to suffer from the same issue, with a TCRβ sequence from one super clone making up 7% of the sample potentially being artifactually over-paired to >500 rare single-cell-count TCR α chains.

      Taken together, these observations suggest that the authors' first positive control example passes but probably for the wrong reason since at least some of the "antigen-specific communities" are strongly anchored by a single hyper-clone. This could be remedied by repeating the same type of analysis on an ex vivo single-cell repertoire following more modest stimulation or natural infection (e.g., yellow-fever vaccine, influenza, or SARS-CoV-2 single-cell TCR datasets). I would advise future work using an alternative data source with better-documented experimental protocols, given the concerns above.

      (2) Weaknesses in Example 2

      Example 2 explores the application of Bayesian methods for identifying sample-enriched joint graph communities found in longitudinal data from many participants at two time points and longitudinal data from 1 person (Pt4) at 5 time points. A potential weakness of Case Study 2 is that it yields limited additional biological insight compared to what was previously shown by the authors of the underlying input data. Previously, Formenti et al. 2018 showed that the number of expanded clones after treatment strongly reflects responder status in this cohort, greater in patients with CR/PR versus SD or PD (See Figure 2b of the study). This 2018 primary analysis showed that tracking individual clonotypes was sufficient to reveal biological insight without the need for the computational demands of constructing a massive sequence similarity graph, finding communities on joint graphs, or Bayesian statistical inference. Thus, the impact of Case Study 2 in proving the unique utility of ClustIRR is somewhat diminished.

      It is not clear how this study's result is more "robust" than the original analysis. Perhaps the authors could further clarify what is learned from the uncertainty-aware approach that could not be learned from exact clone tracking in time series.

      Thus, example 2 shows that a complex method recapitulated the finding of a much simpler method for analyzing longitudinal TCR data where a strong signal of expansion was already present at the single-clonotype level. Since the Dirichlet method is sensitive to absolute counts, the large expanding clone in each community at 22 days may alone have carried most of the signal, which is not fully explored.

      In this section, the authors make an interesting observation that CDR3β detected in both PBMC and patient-matched tumor samples were enriched in contracting communities (7/14) versus expanded communities (1/41). The authors do not indicate whether a similar enrichment of tumor-infiltrating lymphocytes (TILs) matched Day 0-22 contracting communities in the other 4 patients with tumor-matched samples, which, if consistent, would strengthen their finding.

      (3) Weaknesses in Example 3

      The final example explores "convergent repertoire differences between species." This is intriguing in principle but less informative due to the use of synthetic data with somewhat predictable properties that some may reasonably consider baked in by the data-generating process. That is, some of the results might be guaranteed by the way the data is constructed using the OLGA/IGoR generative model. For instance, the authors observe a positive correlation between CJ community size and Pgen of constituent clonotypes, stating: "This indicates that CDR3 sequences with high Pgen are statistically more likely to be generated, leading to convergence of similar sequences into public CJs." I may be mistaken, but this conclusion is almost guaranteed by the way the data-generating OLGA model outputs more similar high-Pgen sequences and fewer lower-Pgen sequences.

      An interesting finding in this section is shown in Figure 4B, where community-level aggregation allowed for discrete clustering of human samples away from mouse samples that was not possible by comparing cosine similarity of a sparse clonotype occurrence matrix, a result that would be higher impact if it could be shown to separate real repertoire samples from humans with differential serology, vaccination status, or HLA backgrounds.

      (4) Weaknesses in General

      More generally, one aspect of the method that seems under-emphasized is the fact that many of the nodes in a multi-sample joint sequence similarity graph may have no edges. These zero-degree nodes would probably frequently occur in only one sample but be absent in other samples. It is not strongly emphasized in the paper how the model would infer whether such a singleton found in only one sample in the graph could be reliably inferred to be sample-specific enriched (see, for example, the large single node in Figure 3D, the orange node labeled "GQYF" in the far-right position of the lowest row in panel D). Presumably the number of cell counts represented in this single-sequence node is so great at sampled timepoint Day 22 as to yield a statistically strong signal in the multinomial model; however, the authors may wish to comment on how, for such singleton sequences, the power to detect sample-specific enrichment differs from prior single-clonotype-based methods.

      With any large effort to find statistically significant features from a large candidate set examined all at once, a reader might be concerned with the potential for false discovery. Throughout, the authors seem to assign statistical significance when the 95% high-density interval (HDI) of the posterior estimate excludes zero or when the 95% HDIs of two features do not overlap. There is little discussion of how this implicitly handles multiplicity adjustment via shrinkage, which the authors could address more directly and explain more clearly to a broad audience, including many non-statisticians, who will read this paper.

    4. Reviewer #3 (Public review):

      Summary:

      Analysis of immune receptor repertoires (IRR) needs to take into account the underlying diversity of the repertoires analysed, and the limitations inherent to the technologies used to measure IRRs: limited sampling depth relative to total number of cells and clonotypes, and experimental noise. In this work, Kitanovski and colleagues present ClustIRR. ClustIRR proposes to improve the analysis of immune receptor repertoires, specifically TCRs in the presented applications, by performing two steps: (1) consistent and comparable sequence clustering across repertoires, to account for sparsity of sampling, and (2) estimation of sequence clusters of interest using a Bayesian approach. They showcase ClustIRR performance in 3 scenarios: detection of antigen-specific T cells in a peptide-stimulation, detection of T cells responding to immunotherapy in the context of lung cancer, and analysis of mouse and human T cell repertoires.

      Strengths:

      (1) The cluster occupancy calculation presents an important conceptual framework which would be of interest and useful to the TCR repertoire field as it smoothly integrates information over a set of experimental conditions or time points. The application to longitudinal TCR sequencing datasets is particularly interesting, and could easily be extended to BCR sequencing datasets. Moreover, the calculation can be performed with any user-defined grouping of TCR clones, which allows for usage of other existing methods as the user wishes.

      (2) The results of the human and mouse repertoires provide a very insightful argument for the use of metaclones as opposed to single clones for analysis of repertoires compared to single clones.

      Weaknesses:

      (1) While ClustIRR provides an interesting framework to analyse TCR sequencing datasets, it is not clear whether ClustIRR can perform more informative sequence clustering than state-of-the-art methods. A comparison of obtained clusters with existing methods would provide a useful benchmark. Moreover, computation time scales quite fast with the number of sequences included. This is a major limitation, as the authors state that a time of 2.5 hours is required for clustering of ~100,000 sequences, a number of clones that can easily be reached when analysing multiple samples together.

      (2) The authors claim in the abstract that ClustIRR is integrated with gene expression data. However, in the results presented, the gene expression and TCR sequencing data are analysed separately, and the results are simply correlated. No real integration in the analysis exists for these two data types. The claim should be removed from the abstract. Moreover, the differential gene expression section, while it presents interesting results, lacks clarity and transparency. Presentation of the data in more transparent ways (such as showing violin plots or clustering on the UMAP) would increase the strength of the claims.

      (3) The score calculation does not seem to be normalized to take into account the underlying diversities of CDR3a and CDR3b. While still a useful metric for sequence clustering, I worry about the impact of the lack of normalisation on the conclusion that the "alpha chain drives functional convergence through germline bias". The observation that clustering is mostly driven by the J gene is known and expected (https://pmc.ncbi.nlm.nih.gov/articles/PMC5553937/), as the J gene has lower diversity and greater overlap with the definition of CDR3. Because CDR3a is lower diversity generally from CDR3b, it will likely dominate the sequence similarity signal. Thus, it will appear that Ja drives the signal. The authors do try to address this by looking at clusters driven by CDR3b similarity. However, a low number of CDR3b-driven clusters is consistent with the similarity definition. I wonder if instead the appropriate control for this analysis would be to run the same analysis disregarding CDR3a sequence altogether, and quantify whether similar species-specific DCJ are identified when only CDR3b similarity is used? Absence or reduction of species-specific clustering would confirm that the effect is driven by the CDR3a sequence.

    1. eLife Assessment

      This meta-analysis provides a valuable contribution by integrating findings from dozens of macaque electrophysiology studies to reconcile discrepancies in choice probability and identify factors that robustly influence choice signals in the visual cortex. Such systematic synthesis across studies is rare in this field, and the analysis provides solid evidence for several conclusions, including the cross-study consistency of the relationship between choice probability and sensitivity and the distinctiveness of V1.

    2. Reviewer #1 (Public review):

      This meta-analysis addresses long-standing questions about the reliability and interpretation of choice probability in macaque visual areas, and provides some important findings (e.g., the cross-study consistency of the CP-sensitivity relationship, V1 distinctiveness). However, the evidence for several claims is incomplete: the analysis does not consider the statistical dependence of data points from the same studies and monkeys, and both the bistable-stimulus effect and the stimulus-duration effect rely on interpretive assumptions.

      Strengths:

      The paper's transparency about its own limitations is a genuine strength. Several sections of the paper and the supplement report null results (task exposure, lapse rate, eccentricity) rather than omitting them. This kind of self-scrutiny is uncommon in meta-analyses and substantially increases confidence in the parts of the analysis that do hold up.

      Weaknesses:

      (1) No mixed/hierarchical statistical models for nested data. The paper considers 150 data points from 59 studies and treats them as independent samples, though many share monkeys and brain areas. This reduces the confidence in the reported p-values. A standard way of dealing with this would be to use mixed-effect models with random intercepts rather than OLS.

      (2) Evidence for one of the main findings in the abstract ("First, CPs were higher in tasks involving bistable percepts, reinforcing the link between CP magnitude and subjective perception.") is weak. This effect relies entirely on studies using bistable rotating cylinder stimuli performed in a single lab (lines 702 - 711). I would suggest making this more explicit in the abstract / discussion and in Figure 8b,c.

      (3) The interpretation of main drivers of CP is unclear. The discussion summarizes the 4 main drivers of CP as "four systematic drivers of this variability: neuronal sensi758tivity, brain area, stimulus duration, and task type." The independent contribution of task type is however, questionable. In line 623, it is stated that the difference between coarse and fine discrimination can be entirely explained by the difference in sensitivity (explained possibly by differences in optimizing the stimuli). Again, detection tasks (line 658) show a trend for higher CP because most studies used tailored stimuli from single-recording experiments. Bistable task: see point 2. Thus, all "task effects" can be attributed to confounds, and the claim of the "four drivers of CP" should be revised.

      (4) The paper could be improved by a Discussion that synthesizes the results in a concise manner. Now it seems more like a re-iteration of the results. Overall, I appreciate that the paper is thorough and discusses many of the caveats. However, those are somewhat buried in the long subsections, and I fear that the quick reader may walk away with a stronger impression of "four robust independent drivers" than the text, read carefully, actually supports.

      (5) Datapoints are not weighted according to their standard error (common practice in meta-analysis is inverse-variance weighting). The concern is that underpowered studies with high variance (e.g., due to a low number of recorded neurons) have the same impact as well-powered studies, and this may change some of the estimates. For example, Supplementary Figure 4 shows that mean CP values decrease with statistical power of the study, consistent with the concern. If SEMs are not available, could the authors show the robustness of the results by weighting by sample size as a partial check?

      (6) Interpretation of feed-forward vs. feedback origin of CP [Disclosure: I am an author of Wimmer et al. 2015.]. This paper presents a mechanistic network model of area MT and a decision area that decomposes CP into two components with distinct time courses, and, directly relevant to Section 2.6, shows how a combination of an early feedforward and a late feedback component can produce a roughly time-invariant (flat) CP. This is a specific, quantitative instance of the "sustained plateau" pattern the authors themselves note is inconsistent across studies (lines 480-486) but don't develop further. Engaging with this model in Section 2.1/2.6 would let the authors contrast their duration-effect interpretation against an explicit dynamical model rather than the generic feedforward/feedback dichotomy in Figure 7a.

      (7) Reaction-time experiments. I am worried that differences in CP in RT vs. fixed duration tasks (Supplementary Figure 12) could have an influence on the main regression analysis (because RT experiments are mostly from detection tasks, and because RT experiments presumably include less of a post-decision period). Could this factor be included in the main analysis?

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript by Pletenev et al. provides a meta-analysis of 59 published studies on decision-related activity in the macaque visual cortex. The work does not contain original research material, but by conducting extensive analyses and compilations of previously published data, it provides several new insights not already conveyed in recent reviews on this topic.

      Strengths:

      The work is scholarly and helps organize and integrate a broad set of findings. A difficulty in such an undertaking is making sure that the original studies are accurately characterized and tabulated. The lead authors should be commended for the rigorous approach they've taken; the inclusion of many of the authors of the original studies as co-authors provides additional reassurance that trends visible across studies reflect an accurate quantification of what each study has shown.

      Weaknesses:

      My sole scientific concern is on the 'duration' section (lines 487 to 536). Specifically, the authors relate the dependence of choice probability (CP) on time to two competing (but not mutually exclusive) views of how CPs arise-the 'feedforward' vs 'feedback' views. The basis for the predictions in this section was not clear to me. In particular, it was not obvious that the predictions fully considered all the relevant factors. For instance, did the feedforward predictions consider how the response covariance depends on duration and how this would affect CP values (Equation 1)? (See for example Figure 4 of Kang and Maunsell, 2012, JNP.)

      I would suggest either explaining the basis of the predictions much more carefully or moving the predictions to a supplementary section where they can be presented in more detail (i.e., more convincingly). Alternatively, since the section is inconclusive in the end (there is not strong evidence in favor of FF or FB), the authors could mention the theoretical predictions in passing only, i.e., much more briefly, just to make the reader aware that the different theories can provide predictions about the duration dependence.

    4. Reviewer #3 (Public review):

      Summary:

      This study presents a comprehensive meta-analysis of choice probability (CP), a classic metric of the relationship between single-neuron responses and an animal's perceptual judgment. The authors compiled data from 59 macaque electrophysiology studies and identified several factors that consistently influence CP magnitude, including neuronal sensitivity, brain area, stimulus duration, and task type. These results provide evidence that helps settle several long-standing debates about how CP should be interpreted.

      Strengths:

      This work is a rare example of meta-analysis in macaque electrophysiology, focusing on an important and long-debated metric, choice probability (CP), in the study of sensory and decision-making mechanisms. CP has been measured across many studies, but its interpretation remains contentious because it depends on numerous task and recording factors in addition to sensory and decision-making mechanisms. Individual macaque studies also typically include few subjects, and CP effects are generally small, which prevents any single study from drawing strong conclusions about general patterns. The authors identified an ideal use case for meta-analysis and combined fragmented findings from individual primate studies into a coherent picture of which factors matter most for CP and which do not. This work can also serve as a guide for future meta-analyses of macaque electrophysiology data.

      Weaknesses:

      While the survey and discussion of CP's interpretation are comprehensive, the paper would benefit from a clearer conclusion on why measuring CP remains important and what future directions could make CP more useful for revealing sensory and decision-making mechanisms.

    5. Author response:

      We thank the Editors and Reviewers for their encouraging evaluation and constructive feedback. We are glad that they recognized the value of this systematic synthesis in reconciling disparate findings across many studies on an important question, and in providing new insights that help address long-standing debates around the interpretation of CP values. They also acknowledged our rigorous data curation involving many original study authors, and our transparency in reporting limitations and null results.

      To address their constructive recommendations, we will provide additional hierarchical regression analyses where possible, state some limitations more explicitly, and condense the Discussion section to improve focus and readability.

      Regarding the hierarchical nature of the data, we distinguish three potential hierarchical levels: studies, monkeys, and neuronal samples. Because individual monkeys contribute roughly one observation per study and cannot be tracked across publications, animal-level random effects are statistically unidentifiable. We will address the rare cases where identical neuronal pools were evaluated across multiple task conditions by providing a sensitivity analysis restricted to one data point per unique neuronal sample. At the study level (median 2, range 1–7 observations per study), we will present linear mixed-effects models with random study intercepts.

      On the bistability findings, we agree with Reviewer 1's concern about the limited number of studies. We noted in the Results and Discussion that this effect currently relies on rotating-cylinder paradigms and emphasized the need for CP to be quantified with other forms of bistable stimuli. We will make this limitation explicit in the Abstract and Figure 8 caption. We will revise the Discussion to emphasize the three primary drivers (neuronal sensitivity, brain area, and stimulus duration) and treat the bistable stimulus effect separately as a distinct finding.

      Regarding Reviewer 1's concern about reaction-time experiments: because reaction-time (RT) paradigms are heavily confounded with detection tasks in the existing literature (89% of detection tasks are RT tasks, and 67% of RT tasks are detection ones), including RT as a separate factor introduces near-complete collinearity. We will make this constraint and the inability to statistically disentangle them explicit in the main text.

      Addressing Reviewer 2's concern regarding the CP–duration predictions, we agree that they depend on specific modeling assumptions. While our predictions—for the feedforward framework in particular—reflect standard models from the literature, we will explicitly acknowledge that alternative feedforward assumptions—such as duration-dependent response covariance—could alter the expected relationship.

      In response to Reviewer 3’s comments on the rationale and future utility of measuring CP, we will revise the Discussion to emphasize that while the interpretation of CP has evolved from feedforward readout to include feedback mechanisms, our results confirm that CP remains a robust neural correlate of subjective perception. Although the exact mechanisms linking CP to perception remain unresolved, this ambiguity does not justify abandoning the metric; rather, CP remains an indispensable tool, provided it is supplemented with additional analyses as outlined in our recommendations.

    1. eLife Assessment

      This study presents an open-source reinforcement learning framework for the real-time, closed-loop optimization of spatiotemporal electrical stimulation in engineered neuronal networks. Using single-spike-resolution activity as continuous feedback, the authors provide solid evidence that their platform can identify stimulation patterns that drive specific activity motifs within a structurally constrained four-node circuit. While this reproducible system offers a valuable and accessible tool for interacting with biological neural networks in an adaptive manner, further validation is needed to determine how well these stimulation strategies generalize to larger, unstructured, or more conventional network architectures.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript presents several reinforcement learning (RL) approaches to control activity in small, ring-shaped neural networks. The paradigm is to find temporally and spatially patterned stimuli applied to axons that maximize the length of activity propagation along the network. Several RL versions were compared. Some RL-designed stimulation patterns worked better than others.

      Strengths:

      The work is technically and statistically very sound, and important controls were done that are missing in some earlier work along this line of research. The statistics and technical solutions seem solid, though I am no specialist in RL. The figures are well done and informative, with matching quality of the captions. The work presented is of value mainly for someone who wants to build a good closed-loop stimulation system to experiment with neuronal networks in-vitro.

      Weaknesses:

      The manuscript appears undecided about whether it wants to be about RL control, about a technical implementation of long-term stimulation in vitro, or about the properties of neuronal networks and interaction with them. The introduction is well written, comprehensive and insightful, focusing on biological aspects. Methods are then extensively about RL algorithms, without explaining why several were used or why these in particular, but with specialist language hard to understand for neuroscientists. The results then quantitatively compare the performance of the RL but do not really explain what this teaches us about neuroscience or what we learn about the networks beyond that they can be stimulated for longer propagation patterns. Extensive supplementary material almost advertises the hardware built by the team. Some figures suggest, though I'm not 100% certain about this, that different algorithms find different optimal stimulation patterns in the same network - which I find puzzling. What then does this tell us about the stimulation patterns and the variability of the responses? There is some speculative mechanistic reasoning, but no data to support this further. The conclusions hardly address the initial motivation of the manuscript. To me, it is not clear what can be learned that was not, in one way or another, presented previously, with as well as without RL.

    3. Reviewer #2 (Public review):

      The authors developed a reinforcement learning framework to identify stimulation patterns that produce a predefined activity pattern in neuronal networks cultured on a microelectrode array. The main contribution of this study lies in the development and validation of the experimental framework, rather than in providing new insight into neuronal response properties or the biological mechanisms underlying network activity.

      A major strength of the work is that the framework is implemented using open-source software. Electrical stimulation through microelectrode arrays is already widely used, and the proposed approach is therefore likely to be of interest to researchers seeking to automate the exploration and optimization of stimulation parameters. The authors also provide the software, hardware configuration, and experimental data, which substantially improves transparency and should facilitate reproduction and further development of the system by other laboratories.

      The experiments provide a useful proof of concept showing that the framework can search for stimulation patterns associated with the predefined task in the tested neuronal network. The evidence is solid for demonstrating the feasibility of the approach within this specific experimental configuration. In particular, the closed-loop integration of stimulation, recording, evaluation of the neuronal response, and subsequent selection of stimulation patterns is clearly implemented and experimentally tested.

      An important limitation is that the framework was evaluated using a relatively small and highly structured network. Four stimulation electrodes were positioned around a single network, and stimulation was delivered to microchannels in which axons were concentrated. This configuration is well suited to the initial demonstration, but it remains uncertain whether the same approach will perform similarly in conventional monolayer dissociated cultures, larger networks, or systems with different numbers and spatial arrangements of electrodes. Additional validation across a broader range of network structures and experimental configurations would therefore be needed to establish the general applicability of the framework.

      Overall, this study presents a valuable and reproducible methodological framework for the closed-loop optimization of electrical stimulation in cultured neuronal networks. Its likely impact lies primarily in providing an accessible experimental and computational platform that can be adapted for studies requiring systematic exploration of stimulation patterns, although the extent to which the current findings generalize beyond the tested network configuration remains to be determined.

    4. Reviewer #3 (Public review):

      Summary:

      This study integrates living neuronal circuits with a real-time reinforcement-learning framework to enable closed-loop optimization of directional firing sequences at single-spike resolution. By using the spatiotemporal structure of neuronal firing as feedback, the system adaptively searches for stimulation patterns that increase a predefined propagation reward. The work represents an exciting step toward systematic and adaptive control of living neuronal networks.

      Strengths:

      The study presents a cutting-edge closed-loop platform that combines neuronal cultures, high-temporal-resolution electrophysiology, and reinforcement learning with millisecond-scale latency. The ability to optimize directional activity propagation at single-spike resolution is particularly novel. More broadly, the work provides a compelling framework for interacting with biological neural networks in an adaptive rather than purely predefined manner.

      Weaknesses:

      Several aspects of the analysis and interpretation require further clarification. In particular, the definition and computation of directional propagation and reward are not always clear. The generalizability of the optimized stimulation patterns across cultures also remains unclear.

    5. Author response:

      eLife Assessment

      This study presents an open-source reinforcement learning framework for the real-time, closedloop optimization of spatiotemporal electrical stimulation in engineered neuronal networks. Using single-spike-resolution activity as continuous feedback, the authors provide solid evidence that their platform can identify stimulation patterns that drive specific activity motifs within a structurally constrained four-node circuit. While this reproducible system offers a valuable and accessible tool for interacting with biological neural networks in an adaptive manner, further validation is needed to determine how well these stimulation strategies generalize to larger, unstructured, or more conventional network architectures.

      We thank the editors and reviewers for their kind and insightful words, as well as their constructive feedback and assessment. We agree that in its current stage, the manuscript does not convincingly argue, that specific stimulation strategies applied to one network generalise well to other network architectures. We intend to address this problem by adjusting the focus of the manuscript to be more on the platform than on any specific neuroscience claim, as we believe that this is the more useful angle for the community at large. We also intend to discuss the transferability of our results between the presented culture system and other neuroscience systems.

      Based on the reviewer comments, we will further give a more approachable introduction to reinforcement learning to make the concept more accessible to a wider audience. We will also motivate the algorithm selection in more depth.

      In the following, we will discuss the comments put forward by the reviewers and how we intend to adapt our manuscript to address them. In this provisional response, we will only focus on major concerns. Minor comments, where we are following the suggestions made by the reviewers directly, will not yet be discussed.

      Reviewer #1 (Public Review):

      The manuscript appears undecided about whether it wants to be about RL control, about a technical implementation of long-term stimulation in vitro, or about the properties of neuronal networks and interaction with them. The introduction is well written, comprehensive and insightful, focusing on biological aspects. Methods are then extensively about RL algorithms, without explaining why several were used or why these in particular, but with specialist language hard to understand for neuroscientists. The results then quantitatively compare the performance of the RL but do not really explain what this teaches us about neuroscience or what we learn about the networks beyond that they can be stimulated for longer propagation patterns. Extensive supplementary material almost advertises the hardware built by the team. Some figures suggest, though I’m not 100% certain about this, that different algorithms find different optimal stimulation patterns in the same network - which I find puzzling. What then does this tell us about the stimulation patterns and the variability of the responses?

      We thank the reviewer for pointing out these issues with our manuscript. The goal of our work is the hardware framework. We plan to highlight this more in the revised version. The RL part will be revised with less specialist language and will get less emphasis in the next version of our manuscript. We will also discuss why different algorithms seem to find different solutions, which can be traced back to having different networks and agents finding different local maxima.

      By clearly setting the scope of our manuscript on the hardware aspects and revising the RL part to be seen more as one possible closed-loop control paradigm, we further intend to address the concerns put forward by the reviewer regarding our focus on RL in other parts of the manuscript.

      Reviewer #2 (Public Review):

      An important limitation is that the framework was evaluated using a relatively small and highly structured network. Four stimulation electrodes were positioned around a single network, and stimulation was delivered to microchannels in which axons were concentrated. This configuration is well suited to the initial demonstration, but it remains uncertain whether the same approach will perform similarly in conventional monolayer dissociated cultures, larger networks, or systems with different numbers and spatial arrangements of electrodes. Additional validation across a broader range of network structures and experimental configurations would therefore be needed to establish the general applicability of the framework.

      We thank the reviewer for their feedback. They are right to point out that the manuscript here focuses purely on small and highly structured networks. With the presented experiments, our manuscript does not and cannot make any biological claims about how neurons communicate. However, with this manuscript, we also do not intend to do so, as such an endeavour would be out of scope. In the next version of our manuscript, we will better highlight the focus of our manuscript, which lies on the hardware framework itself. Furthermore, we will discuss in the outlook the generalisability and scalability concerns raised by the reviewer in more detail.

      Reviewer #3 (Public Review):

      Several aspects of the analysis and interpretation require further clarification. In particular, the definition and computation of directional propagation and reward are not always clear. The generalizability of the optimized stimulation patterns across cultures also remains unclear.

      We thank the reviewer for their insightful feedback. We agree with them and will discuss (1) the currently implemented reward based on directional propagation of activity and (2) the limitations and aspects influencing reward algorithm selection in more detail in our revised version of the manuscript. We further believe that we cannot make any claims about generalisability between cultures or when changing the experimental paradigm.

    1. eLife Assessment

      This important study provides convincing evidence that somatosensory relay nuclei remain engaged during attempted hand movements following chronic cervical spinal cord injury, including in individuals with severe motor impairment. The combination of functional and quantitative MRI with physiological controls supports the main findings and may have implications that are of importance in understanding sensorimotor processing after spinal cord injury. However, the interpretation of this activity as specifically reflecting top-down, and particularly corticocuneate, signaling may be overstated. The origin and anatomical route of the observed activity appears to remain less firmly established.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigated somatosensory processing along the afferent somatosensory pathway (cuneate nucleus, thalamus, S1) in a group of spinal cord injury patients and a group of controls. They propose that reduced motor function in SCI patients would reduce bottom-up activity; thus, recorded activity in SCI patients would reflect top-down modulation of overt or attempted movements.

      Strengths:

      (1) Strong methods.

      (2) Experimental and control groups.

      (3) Strong writing.

      (4) Results well presented.

      (5) Appropriate statistics.

      Weaknesses:

      Some results (or lack of) cast doubt about the ability of the used technique (3T fMRI) to detect the desired effects (bottom-up vs top-down activity).

    3. Reviewer #2 (Public review):

      Summary:

      This study addresses a question that has been essentially inaccessible in humans: whether the early somatosensory relay nuclei are engaged by anything other than peripheral drive. Using functional and quantitative MRI in individuals with chronic cervical spinal cord injury, the authors show that the cuneate nucleus and VPL are robustly engaged during overt or attempted hand movement, and that this engagement persists in a participant with complete hand paralysis and no detectable muscle activity. They further report structural degeneration of the cuneate nucleus that is unrelated to the preserved task-evoked activity, and interpret the residual activity as reflecting top-down corticocuneate signalling.

      Strengths:

      The work is well conceived and clearly written, and the demonstration that the cuneate nucleus and VPL are robustly engaged during (attempted) hand movement after chronic cervical spinal cord injury is, to my knowledge, novel at this level of anatomical resolution. The finding is convincing. The dissociation between preserved task-evoked activity and marked structural degeneration of the cuneate nucleus is a highly interesting result with clear implications for rehabilitation. The manuscript is straightforward to follow, and the imaging protocol is carefully executed.

      Weaknesses:

      My main reservation concerns the inferential step from "not peripheral" to "corticocuneate". The data establish the former convincingly; the latter may not.

      (1) Attribution of the observed activity to corticocuneate projections.

      The central claim rests on an argument by elimination: because bottom-up drive is excluded in PT01, the residual activity must be top-down and, by extension, corticocuneate. Two distinct gaps should be addressed. First, "top-down" is not equivalent to "direct corticocuneate". Descending influence could reach the cuneate nucleus through multiple indirect pathways. Second, the activity observed at the three levels (cuneate, VPL, S1) need not be serially propagated, since layer 6 corticothalamic projections, for example, could drive VPL independently of any cuneate contribution. I would ask the authors to either provide evidence bearing on the routing, or to consistently use a route-neutral term (e.g. "descending" or "top-down") and reserve "corticocuneate" for the discussion of candidate mechanisms.

      (2) Afferent input arising above the lesion level.

      The EMG control in PT01 was restricted to hand and forearm muscles. Musculature innervated above C4 (cervical paraspinals, trapezius, and to a variable extent the shoulder girdle) remained available to this participant, and attempted hand movement is frequently accompanied by increased proximal co-contraction, postural stabilization, and altered respiratory effort. Afferent to the upper cervical cord is known to project to the ipsilateral cuneate nucleus, and its activity would produce lateralized, ipsilaterally dominant cuneate input - that is, precisely the pattern reported. This alternative is not excluded by the present control and should be addressed directly, ideally with proximal EMG in PT01 (and, if possible, in the other participants), or at minimum with an explicit discussion. Relatedly, the authors recorded respiratory and cardiac signals: please report whether respiratory volume or heart rate differed between movement and rest blocks, and between groups, since the dorsal medulla lies adjacent to cardiorespiratory nuclei.

      (3) Functional significance of the preserved top-down signal.

      The discussion establishes that top-down input persists but says relatively little about why it should. If the principal role of descending input to the cuneate nucleus is the gating of incoming afferent traffic, then in the absence of afferents there is nothing left to gate, and one might have expected the signal to be lost. Its persistence is the most interesting aspect of the finding and deserves fuller discussion. Candidate accounts the authors may wish to consider include: an efference copy or predictive signal delivered to a comparator that no longer receives its input, in the framework the authors already invoke (references 27, 28); engagement of the non-lemniscal outputs of the dorsal column nuclei (e.g., cuneocerebellar, cuneo-olivary projections); attempted movement engages motor imagery and attention, in which case the relevant question becomes what distinguishes these from movement-related gating. A related interpretational point: in behaving primates, movement-related modulation of cuneate transmission is bidirectional and includes prominent suppression (refs. 7/12). Note also that BOLD increases are compatible with increased inhibition, so they do not indicate facilitated throughput.

    4. Author response:

      Reviewer #1:

      Some results (or lack of) cast doubt about the ability of the used technique (3T fMRI) to detect the desired effects (bottom-up vs top-down activity). 

      As the reviewer notes, 3 T MRI alone cannot partition the relative contributions of bottom-up and top-down signals, since both are present during movement in an intact system. This reflects the premise of our study design but is also an important caveat when interpreting the group-level results, which characterise the net task-related response. Our central inference therefore focuses on the persistence of activation in PT01, in whom hand movement was absent. PT01 therefore lacks bottom-up signals, and any observed activity must be driven by top-down processes. In the revised manuscript, we provide measures of signal quality and activation magnitude to better characterise the sensitivity of these measurements and will draw on existing evidence that our approach resolves task-specific responses within these nuclei.

      Reviewer #2:

      (1) Attribution of the observed activity to corticocuneate projections: My main reservation concerns the inferential step from "not peripheral" to "corticocuneate". The data establish the former convincingly; the latter may not. The central claim rests on an argument by elimination: because bottom-up drive is excluded in PT01, the residual activity must be top-down and, by extension, corticocuneate. Two distinct gaps should be addressed. First, "top-down" is not equivalent to "direct corticocuneate". Descending influence could reach the cuneate nucleus through multiple indirect pathways. Second, the activity observed at the three levels (cuneate, VPL, S1) need not be serially propagated, since layer 6 corticothalamic projections, for example, could drive VPL independently of any cuneate contribution. I would ask the authors to either provide evidence bearing on the routing, or to consistently use a route-neutral term (e.g. "descending" or "top-down") and reserve "corticocuneate" for the discussion of candidate mechanisms. 

      We agree with the reviewer that our previous attribution of the observed brainstem effects to corticocuneate processing was speculative and should have been presented as such. Our findings support a non-peripheral, top-down contribution but do not allow us to attribute this descending influence to a specific anatomical route. We have therefore revised the manuscript throughout to use “top-down” processing as a more route-neutral term, reserving the corticocuneate pathway for discussion of possible candidate mechanisms. We have also clarified in the revised discussion that activity observed across the cuneate nucleus, VPL, and S1 does not necessarily imply serial propagation through these structures.

      (2) Afferent input arising above the lesion level: The EMG control in PT01 was restricted to hand and forearm muscles. Musculature innervated above C4 (cervical paraspinals, trapezius, and to a variable extent the shoulder girdle) remained available to this participant, and attempted hand movement is frequently accompanied by increased proximal co-contraction, postural stabilization, and altered respiratory effort. Afferent to the upper cervical cord is known to project to the ipsilateral cuneate nucleus, and its activity would produce lateralized, ipsilaterally dominant cuneate input - that is, precisely the pattern reported. This alternative is not excluded by the present control and should be addressed directly, ideally with proximal EMG in PT01 (and, if possible, in the other participants), or at minimum with an explicit discussion. Relatedly, the authors recorded respiratory and cardiac signals: please report whether respiratory volume or heart rate differed between movement and rest blocks, and between groups, since the dorsal medulla lies adjacent to cardiorespiratory nuclei.

      The reviewer raises an important point. To test whether proximal muscle activity could account for the cuneate response in PT01, we collected additional EMG data during attempted hand movements, focusing on muscles innervated above the lesion level. These included the anterior (AD) and middle deltoid (MD), upper (UT) and middle trapezius (MT), and cervical paraspinals (CP), along with two of the original distal recordings (thenar eminence, TE; extensor digitorum, ED). Nonetheless, we found no significant difference in EMG activity in any recorded muscle during attempted left- or right-hand movement versus rest (Supplementary Figure 3A and 3B). To confirm that the chosen electrode montage could detect proximal muscle activity, we further instructed the participant to perform left or right shoulder shrugs during the same session. This produced a clear increase in activity during movement across the trapezius, cervical paraspinal and middle deltoid recordings (Supplementary Figure 3C and 3D).

      In addition, we analysed the cardiac and respiratory recordings acquired during fMRI to determine whether movement-related physiological changes could explain the observed brainstem activity. Importantly, any physiological change in heart rate during movement is global and therefore cannot explain the hand-dependent lateralisation of the cuneate response. Furthermore, cardiac and respiratory nuisance regressors were included in all first-level models. Heart rate showed a small but significant increase during movement compared with rest (controls: +0.31 bpm; SCI: +0.69 bpm; main effect of condition F(1,33) = 10.22, p = 0.003, η<sup>2</sup></sub>p</sub> = 0.24, BF<sub>10</sub> = 8.44), whereas respiratory volume per time (RVT) did not differ between conditions (F(1,33) = 0.35, p = 0.56, η<sup>2</sup></sub>p</sub> = 0.01, BF<sub>10</sub> = 0.24). Given the autonomic consequences of cervical injury, we also tested whether these changes differed between groups. Neither measure showed a Group × Condition interaction (heart rate: F(1,33) = 1.41, p = 0.24, BF<sub>10</sub> = 0.56; RVT: F(1,33) = 1.60, p = 0.22, BF<sub>10</sub> = 0.61), suggesting that they cannot account for the group differences we report.

      We have added the additional EMG control and the cardiorespiratory analyses to the Supplementary Material of the revised manuscript. Together, these controls suggest that neither proximal muscular nor cardiorespiratory factors explain our findings.

      (3) Functional significance of the preserved top-down signal: The discussion establishes that top-down input persists but says relatively little about why it should. If the principal role of descending input to the cuneate nucleus is the gating of incoming afferent traffic, then in the absence of afferents there is nothing left to gate, and one might have expected the signal to be lost. Its persistence is the most interesting aspect of the finding and deserves fuller discussion. Candidate accounts the authors may wish to consider include: an efference copy or predictive signal delivered to a comparator that no longer receives its input, in the framework the authors already invoke (references 27, 28); engagement of the non-lemniscal outputs of the dorsal column nuclei (e.g., cuneocerebellar, cuneo-olivary projections); attempted movement engages motor imagery and attention, in which case the relevant question becomes what distinguishes these from movement-related gating. A related interpretational point: in behaving primates, movement-related modulation of cuneate transmission is bidirectional and includes prominent suppression (refs. 7/12). Note also that BOLD increases are compatible with increased inhibition, so they do not indicate facilitated throughput. 

      We thank the reviewer for their comment and agree that the persistence of this descending signal despite profound loss of peripheral input is a very interesting aspect of the findings, and that our manuscript will benefit from a more extended discussion of this result. We will expand on this and the candidate accounts raised in the revised discussion. We will also clarify that movement-related modulation of cuneate processing may include both facilitation and suppression, and that our finding of increased BOLD activity does not necessarily imply facilitated sensory throughput.

    1. eLife Assessment

      This valuable study examines the relationship between simultaneous LC single-unit recordings and pupillometry, both within and across baseline and evoked epochs, finding that cross-epoch relationships are unreliable. The evidence is solid and suggests that care should be taken when interpreting pupil responses as reflecting LC activity. This manuscript will be interesting to basic and clinical researchers working in cognitive and decision neuroscience as well as computational psychiatry.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses a valuable dataset of simultaneous LC single-unit recordings and pupillometry in awake monkeys to examine one aspect of the relationship between LC activity and pupil diameter: whether baseline LC activity predicts evoked pupil and whether baseline pupil predicts evoked LC. These relationships are largely absent in the present dataset, and the authors conclude that this type of prediction is not reliable.

      Strengths:

      This is a valuable dataset of simultaneous LC single-unit recordings and pupillometry in awake monkeys, the within-modal and within-epoch analyses are sound. The new results further caution the use of pupil diameter to infer LC activity.

      Weaknesses:

      There is an obvious rationale for asking whether baseline LC relates to baseline pupil or evoked LC relates to evoked pupil. It is also obvious to ask about the relationship between baseline and evoked LC activity, and between baseline and pupil responses. However, the rationale for the cross-modal and cross-epoch analysis is not clear. Why should one expect baseline LC to predict evoked pupil, or baseline pupil to predict evoked LC? What is the biological importance of such predictions? These analyses simultaneously change the modality and the temporal domain.

      Furthermore, some recent studies, which are not referenced in this manuscript, showed variability in the relationship between pupil and LC within the same epoch, and that tonic and phasic pupil responses could be regulated by different inputs to LC. Given that the pupil-LC coupling is readily imperfect and potentially controlled by different cellular and circuit mechanisms, it is not surprising that their cross-epoch relationship is even more variable.

      One main analysis that correlated baseline pupil with evoked LC showed statistically different results between the two monkeys, raising the question of to what extent a general claim for the cross-modal cross-epoch analysis can be made.

    3. Reviewer #2 (Public review):

      The manuscript by Thompson and Gold reported that, although baseline and evoked LC activity were positively correlated with baseline pupil size and evoked pupil dilation, respectively, there were no reliable cross-epoch relationships between LC activity and pupil size - that is, between baseline LC activity and evoked pupil dilation or between baseline pupil size and evoked LC activity. A major strength of the study is its large dataset, comprising recordings from more than 100 LC single units collected across 83 recording sessions in two monkeys. The authors further strengthened their conclusions by performing several control analyses to rule out potential confounds, including whether the findings were influenced by (1) the use of residuals in the main analyses, (2) the possibility that the auditory stimulus failed to evoke the full dynamic range of LC activity and pupil responses, and (3) nonlinear cross-epoch relationships between LC activity and pupil size.

      These findings are important because they remind researchers to exercise caution when assuming that non-illuminance-mediated fluctuations in pupil size can reliably serve as a proxy for LC activity under all experimental conditions

      With that being said, I have two suggestions that I believe would further strengthen the manuscript.

      (1) Previous studies have suggested that the LC is organized into functionally distinct subpopulations (e.g., PMID 28920933 and PMID 40770025). It is therefore possible that a subset of the LC neurons recorded in this study does not participate in controlling pupil size. I encourage the authors to discuss this possibility in the Discussion. In addition, this possibility could be tested by repeating the cross-epoch relationship analyses using only LC units that exhibit a strong correlation with pupil size (the darker points in the last column of the top two rows of Figure 1?).

      (2) The manuscript would benefit from a clearer visualization of the analyses addressing the possibility that the auditory stimulus failed to evoke the full dynamic range of LC activity and pupil responses. A supplemental figure illustrating the distributions or ranges of LC activity and pupil responses, together with the corresponding control analyses, would help readers better understand this important point.

    1. eLife Assessment

      This study addresses a key question concerning whether neurofeedback can enhance the neural representation of a selected speaker during competing continuous speech and whether such enhancement translates into behavioral benefits. The solid findings provide important insights into the online modulation of auditory attention and the extent to which selective listening can be voluntarily controlled. Although the observed effects were relatively small and did not consistently persist beyond the feedback period, the study advances understanding in an under-explored area and provides a helpful foundation for future research on sustained learning and behavioral transfer.

    2. Reviewer #1 (Public review):

      Summary:

      The authors asked whether neurofeedback during competing continuous speech can help to modulate the attention-related N1-component in the temporal response function (TRF), which is an event-related-response-like estimate of the phase-locked EEG activity following the envelope. The research question is relevant because it asks to what degree the strength of attention can be controlled beyond the binary decision to attend or ignore something, and whether this control is beneficial for the behavioral outcome.

      Strengths:

      (1) Sample size of 56 participants.

      (2) Control group with sham feedback.

      (3) Novelty: Under-explored field of neurofeedback in selective speech tracking.

      (4) Pragmatic and reasonable methodological decisions.

      (5) Transparent results not hiding the fact that effect sizes are small.

      Weaknesses:

      Besides some need for clarification, I could only find one methodological weakness, which the authors discuss anyway:

      (1) Overall, speech tracking-based neurofeedback may lead to more robust results, because the N1-extraction does not have to be handcrafted and all components would be taken into account. As the authors state, the P2-component has been related to effort, and this may provide more "room to play" for voluntary modulation.

      The following "weaknesses" are related to the impact of the results:

      (2) Non-translating effects to post-training trials, neither neurally nor behaviorally.

      (3) Neurofeedback-related Modulation of N1

    3. Reviewer #2 (Public review):

      Summary

      This manuscript investigates whether neurofeedback based on the N1 component of the temporal response function can be used to modulate neural responses during selective attention to continuous speech. Participants listened to two competing audiobooks and were instructed to attend to one of them. In the neurofeedback group, trial-by-trial N1 responses were converted into visual feedback, whereas the sham-feedback group received replayed feedback from other participants. The authors found a significant interaction between group and block for the N1 response to target speech over a small fronto-central cluster, with larger N1 responses during feedback blocks in the genuine neurofeedback group but not in the sham group. No neurofeedback effect was found for the distractor response. The authors also reported exploratory post-training effects and an association between changes in N1 and speech-comprehension performance at right fronto-central electrodes.

      Overall, the study is conceptually interesting and novel. The online, trial-by-trial estimation of neural responses from continuous speech is an attractive development for auditory neurofeedback, and the inclusion of a randomised sham-feedback group is an important strength. However, the manuscript provides stronger evidence for modulation of a neural response during feedback than for learning or training of selective attention. Some aspects of the analysis and interpretation also require further consideration/clarification, particularly the use of group-specific N1 time windows, the spatial confound between target and distractor streams, the absence of artefact correction in the signal used for feedback, and the relatively weak behavioural evidence.

      Strengths

      The main strength of this study is its novel use of neurofeedback during continuous competing speech. Rather than providing feedback based on a general measure of brain activity, the authors targeted a specific neural response associated with selective auditory attention. The online implementation is technically impressive, allowing neural responses to be estimated from 22-second speech segments and converted rapidly into feedback. The inclusion of a sham-feedback group is another important strength, as it helps distinguish effects of genuine neurofeedback from nonspecific effects such as task engagement or motivation. The relatively large sample for a neurofeedback study and the use of natural continuous speech also increase the robustness and ecological relevance of the work.

      Weaknesses

      (1) The effect was present during the feedback blocks but did not increase across training blocks, and the post-training effect was not found at the same electrodes used for feedback. The evidence therefore supports online modulation more strongly than learning or lasting self-regulation, and claims about successful training or persistent learning should be interpreted cautiously.

      (2) The offline analysis used different N1 time windows for the neurofeedback and sham groups. This introduces a potential bias in the group comparison because the dependent measure was defined differently between groups. Confirmation of the main result using a common, independently defined N1 window would strengthen the evidence.

      (3) The target speech was always presented from the front and the distractor from behind. Differences between target and distractor responses therefore cannot be attributed entirely to attention because spatial location is also different. This limits the interpretation of target-versus-distractor differences and may also contribute to the weaker reliability of the distractor response.

      (4) Feedback blocks always contained two speakers, whereas half of the baseline trials contained only one speaker. Since the presence of a distractor altered the neural response, it is important that the neurofeedback comparison is based on acoustically matched multi-speaker baseline trials. If this were not the case, differences between baseline and feedback could partly reflect differences in the acoustic condition rather than neurofeedback.

      (5) No artefact correction was applied to the signal used for online feedback. Because feedback was derived from fronto-central electrodes, eye or muscle activity could potentially contribute to the measured signal. An offline demonstration that the main neural effect remains after appropriate artefact control would strengthen the interpretation that the effect reflects neural modulation rather than systematic changes in non-neural activity.

      (6) The behavioural evidence is weaker than the neural evidence. There was no significant overall improvement in speech comprehension in the neurofeedback group. The reported behavioural effect is instead based mainly on associations between changes in the neural response and changes in comprehension, and some of these effects were weak before the whole-scalp analysis. Therefore, these findings are better interpreted as exploratory associations rather than evidence that neural modulation directly caused improved comprehension.

      (7) Adding the distractor did not significantly reduce comprehension performance. The absence of a significant distractor effect on comprehension suggests that the listening condition may not have produced a strong behavioural cocktail-party difficulty in this sample. The large spatial separation between speakers and the nature of the behavioural task may have reduced sensitivity to distraction, which could limit the strength of the conclusions regarding improvement of speech understanding in challenging listening conditions.

      (8) The study is described as double-blind, but the manuscript provides limited detail on how blinding was maintained, particularly when the experimenter manually checked the N1 estimate. In addition, participants' belief in or perceived control over the feedback was not formally assessed. These factors make it difficult to determine how effectively expectancy or motivation-related effects were controlled.

      Overall assessment:

      This study provides a valuable methodological and conceptual advance by showing that online, trial-by-trial neurofeedback can modulate the neural response to attended speech during competing speech. The evidence is solid for an immediate neural effect during feedback, and it was supported by comparison with a sham-feedback group. However, it remains incomplete for broader claims about learned self-regulation, persistent effects, distractor suppression, and improved speech understanding.

    1. eLife Assessment

      This study presents a fundamental molecular resource, offering subtype-specific insight into the composition of ribosome-associated protein complexes in the developing cerebral cortex. The evidence is compelling in terms of data quality and is strongly supported by the results, given the rigorous technical execution. While primarily descriptive in nature, this resource will be of great use to the field.

    2. Reviewer #1 (Public review):

      This work provides a valuable toolkit for endogenous isolation of projection neuron subtypes. With further validation, it could present a solid method for low-input ribosome affinity purification using a ribosomal RNA (rRNA) antibody. The experimental evidence for the distinct ribosomal complexes is limited to this method and indirect support from complementary analyses of pre-existing data. However, with additional experimental data to support the specificity of ribosomal complex pulldown and confirmation of the putative ribosomal complex proteins of interest, the study would provide compelling evidence for translation regulation of neuronal development through compositional ribosome heterogeneity. This work would be of interest to neuroscientists, developmental biologists, and those studying translational networks underlying gene regulation.

      Strengths

      (1) This in vivo labeling of specific projection neurons and ribosomal rRNA affinity purification method accommodates a low input of <100K somata per replicate, which is useful for the study of neuronal subtypes with limited input. In principle, this set of techniques could work across different cell types with limited input depending on the molecule used for cell type labeling.

      (2) The authors are also able to isolate endogenous neurons with minimal perturbation up to the point of collection, preserving the native state for the neuron in vivo as long as possible prior to processing.

      (3) This study identified over a dozen potential non-ribosomal proteins associated with SCPN ribosomal complexes, as well as a ribosomal protein enriched in CPN.

      Limitations

      (1) In this study, the authors address the advantages of their ribosomal complex isolation method in SCPN and CPN against RPL22-HA affinity purification. While this does show more pull down of the ribosomal RNA by the Y10B rRNA antibody, the authors claim this method identifies cell-type specific ribosomal complex proteins without demonstrating a positive control for the method's specificity. There are very limited experiments to truly delineate how "specific" this method is working and whether there could be contamination from other complexes bound by the antibody. I see this as the major limitation that should be addressed. To boost their claims of capturing cell-type specific ribosomal complexes, the authors could consider applying their rRNA affinity purification pipeline to compare cell-types with well-characterized ribosome-associated proteins, like mouse embryonic stem cells and HELA cells. The reviewer can completely appreciate the elegance in the neural characterization here, but it seems there needs to be a solid foothold on the specificity of the method, perhaps facilitated by cell types that can be more readily scaled up and tested.

      (2) The authors followed up on their differentially enriched ribosomal complex proteins by analyzing ribosome association of these proteins in external datasets. While this analysis supports the ribosome-association of these proteins, there is limited experimental validation of physical association with the ribosome, much less any functional characterization. The reciprocal pulldown of PRKCE is promising; however, I would recommend orthogonal validation of several putative ribosomal complex proteins to increase confidence. Specifically, the authors could use sucrose gradient fractionation of SCPN and CPN, followed by western blot to identify the putative interaction with the 80S monosome or polysomes. This would also provide evidence towards the pulldown capturing association with mature ribosome species, which is currently unclear. This experiment would provide substantial evidence for the direct association of these non-ribosomal proteins with subtype-specific ribosomal complexes.

      (3) The authors state interest in learning more about the differences underlying translational regulation of projection neuron development. This method only captures neuronal somata, which will only capture ribosomes in the main cell body. There are also ribosomes regulating local translation in the axons, which may also play a critical role in axonal circuit establishment and activity. These ribosomal complex interactions may also be rather transient and difficult to capture at only one developmental stage. Therefore, this method is currently limited to a single developmental snapshot of ribosomal complexes at P3 within the main cell body. It would be exciting to see extended utility of this method to sample neurites and additional developmental stages to gain further resolution on the developmental translation regulation of these projection neurons.

      Likely impact of the work on the field, and the utility of the methods and data to the community

      The authors introduce a unique pipeline of techniques to identify cell-type specific ribosomal complex compositions. With more validation, there is certainly potential for those studying neuronal translation to leverage this method in limited primary cells as an alternative to existing methods that do not rely on ribosomal protein tagging, such as ARC-MS (Bartsch et al., 2023), RAPIDASH (Susanto and Hung et al., 2024), and RAPPL (Nature Communications, 2025).

      Comments on revised version.

      We thank the authors for their thorough response to our comments. The revised manuscript satisfactorily addresses most reviewer comments through clarification and expanded discussion, although we believe some important limitations remain. In particular, we continue to view experimental validation of the identified ribosome-associated proteins as an important component of introducing this methodology to the field, rather than work that falls beyond the scope of the study. The authors have acknowledged that these experiments are future directions, and it is clear they plan to pursue additional validation outside of the current manuscript. Given their emphasis that the primary contribution is the development of a methodological framework, we believe the work may be appropriately framed as a Tool and Resource article. Such positioning would better align reader expectations with the manuscript's strengths as a technical advance while recognizing that further validation and functional characterization will be needed in future studies. Despite these limitations, I believe the manuscript makes a valuable methodological contribution.

    3. Reviewer #2 (Public review):

      Summary:

      The study by presents a sophisticated molecular dissection of ribosome-associated complexes (RCs) in two well-defined cortical projection neuron subtypes (ScPN and CPN) during early postnatal development. The authors develop and optimize an rRNA immunoprecipitation-mass spectrometry (rRNA IP-MS) workflow to recover RCs from FACS-purified, retrogradely labeled neurons, achieving remarkable subtype specificity and biochemical resolution. Through proteomic profiling, they reveal both shared and distinct ribosome-associated proteins between ScPN and CPN, with a focus on non-core RC components and their potential functional relevance. The work advances our understanding of cell-type-specific translation regulation, moving beyond the transcriptome to explore the proteome-level complexity in neuronal subtypes.

      Strengths:

      This work stands out for its technical sophistication and innovation. The authors combine retrograde labeling, FACS purification, and an optimized rRNA IP-MS approach (low input) to isolate ribosome-associated complexes from highly specific neuronal subtypes in vivo, a challenging issue that they execute with impressive rigor. The methodological pipeline is both elegant and well controlled, yielding high-quality, reproducible data. The depth of proteomic coverage is remarkable, with nearly all known cytoplasmic ribosomal proteins identified, along with hundreds of ribosome-associated proteins (RAPs), including translation factors, chaperones, and RNA-binding proteins. The analysis not only reveals shared components between ScPN and CPN RCs but also uncovers subtype-specific differences in associated proteins.

      Particularly notable is the integration of this new proteomic dataset with previously published transcriptomic and ribosome footprinting data, which helps to validate the specificity and relevance of the findings. Overall, the clarity of the writing, the robustness of the data, and the transparency of the methods make this a strong and compelling contribution.

      Weaknesses:

      Despite the depth and high quality of the dataset, the study remains descriptive. While the identification of subtype-specific RC components is intriguing, the current version of the manuscript does not explore their functional roles or biological consequences of their alterations. There is no perturbation, causal testing, in vitro or in vivo manipulation to demonstrate whether these proteins are necessary for ScPN or CPN identity, specific axonal targeting, metabolism or synaptic function.

      One important point that is also highlighted by the authors in their discussion and that is critical to establish the subtype specificity of the identified protein. One important point highlighted by the authors in the discussion - and critical for establishing the subtype specificity of the identified proteins-is that some ribosomal complexes may be specialized for specific developmental stages, rather than exclusively for the subtype-specific needs of projection neuron development. The work presented here provides a valuable starting point for further investigation into such RC specialization. However, it will be essential to determine to what extent these RCs exhibit true subtype specificity, independently of their temporal maturation context.

      As a result, key mechanistic insights remain a bit speculative. Although several of the identified proteins have known roles in processes like synaptogenesis or metabolism, their relevance to the specific neuronal subtypes under study is not experimentally addressed. That said, given its rich content and the comprehensive early postnatal dataset, the manuscript represents an extremely valuable resource for the community. While primarily exploratory, it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work provides a valuable toolkit for endogenous isolation of projection neuron subtypes. With further validation, it could present a solid method for low-input ribosome affinity purification using a ribosomal RNA (rRNA) antibody. The experimental evidence for the distinct ribosomal complexes is limited to this method and indirect support from complementary analyses of preexisting data.

      However, with additional experimental data to support the specificity of ribosomal complex pulldown and confirmation of the putative ribosomal complex proteins of interest, the study would provide compelling evidence for translation regulation of neuronal development through compositional ribosome heterogeneity. 

      This work would be of interest to neuroscientists, developmental biologists, and those studying translational networks underlying gene regulation.

      Strengths

      (1) This in vivo labeling of specific projection neurons and ribosomal rRNA affinity purification method accommodates a low input of <100K somata per replicate, which is useful for the study of neuronal subtypes with limited input. In principle, this set of techniques could work across different cell types with limited input, depending on the molecule used for cell type labeling.

      (2) The authors are also able to isolate endogenous neurons with minimal perturbation up to the point of collection, preserving the native state for the neuron in vivo as long as possible prior to processing. 

      (3) This study identified over a dozen potential non-ribosomal proteins associated with SCPN ribosomal complexes, as well as a ribosomal protein enriched in CPN.

      We appreciate the reviewer's thoughtful and detailed review. We especially appreciate the positive evaluation of its strengths, including the use of rRNA affinity purification to access ribosomal complexes in low-input neuronal subtypes in vivo with minimal perturbation, and the resulting identification of distinct ribosomal complexes in SCPN and CPN with associated non-ribosomal proteins. We are also pleased by the recognition of its significance in advancing our understanding of neuronal subtype-specific post-transcriptional gene regulation. We have carefully addressed the limitations below.

      Limitations

      (1) In this study, the authors address the advantages of their ribosomal complex isolation method in SCPN and CPN against RPL22-HA affinity purification. While this does show more pull-down of the ribosomal RNA by the Y10B rRNA antibody, the authors claim this method identifies cell-type-specific ribosomal complex proteins without demonstrating a positive control for the method's specificity. 

      There are very limited experiments to truly delineate how "specific" this method is working and whether there could be contamination from other complexes bound by the antibody. I see this as the major limitation that should be addressed. To boost their claims of capturing cell-typespecific ribosomal complexes, the authors could consider applying their rRNA affinity purification pipeline to compare cell types with well-characterized ribosome-associated proteins, like mouse embryonic stem cells and HELA cells.

      The reviewer can completely appreciate the elegance in the neural characterization here, but it seems there needs to be a solid foothold on the specificity of the method, perhaps facilitated by cell types that can be more readily scaled up and tested.

      We thank the reviewer for the opportunity to further clarify how our experimental design addresses the question of specificity of ribosomal complex pulldown. The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or non-specific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.

      (2) The authors followed up on their differentially enriched ribosomal complex proteins by analyzing the ribosome association of these proteins in external datasets. While this analysis supports the ribosome-association of these proteins, there is limited experimental validation of physical association with the ribosome, much less any functional characterization.

      The reciprocal pulldown of PRKCE is promising; however, I would recommend orthogonal validation of several putative ribosomal complex proteins to increase confidence. 

      Specifically, the authors could use sucrose gradient fractionation of SCPN and CPN, followed by a western blot to identify the putative interaction with the 80S monosome or polysomes. This would also provide evidence towards the pulldown capturing association with mature ribosome species, which is currently unclear. This experiment would provide substantial evidence for the direct association of these non-ribosomal proteins with subtype-specific ribosomal complexes.

      We thank the reviewer for the feedback and suggested future directions for candidate validation. We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.

      We appreciate the reviewer's suggestion of sucrose gradient fractionation for polysome profiling. We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph.

      (3) The authors state interest in learning more about the differences underlying translational regulation of projection neuron development. This method only captures neuronal somata, which will only capture ribosomes in the main cell body. There are also ribosomes regulating local translation in the axons, which may also play a critical role in axonal circuit establishment and activity. These ribosomal complex interactions may also be rather transient and difficult to capture at only one developmental stage. Therefore, this method is currently limited to a single developmental snapshot of ribosomal complexes at P3 within the main cell body. It would be exciting to see the extended utility of this method to sample neurites and additional developmental stages to gain further resolution on the developmental translation regulation of these projection neurons.

      We thank the reviewer for the opportunity to further highlight the foundational significance of this work in translational regulation of projection neuron development. Here, we identified subtype-specific differences in ribosomal complex composition within somata at a critical developmental time window for two PN subtypes. This work provides foundation for future investigation of local translation and its regulation in growth cones and axons. It will also enable direct comparison with potential future results from axons and growth cones, developmental subcellular specializations at axon tips that implement pathfinding and circuit formation (such work is not yet feasible due to exceptionally low available input). Our lab has significant ongoing work regarding subtype-specific axon and growth cone biology, including recent investigations of growth cone-localised RNA and protein molecular machinery that regulate circuit formation, maintenance, and function of distinct cardinal PN subtypes (Poulopoulos*, Murphy* et al. Nature 2019; Engmann*, Hatch* et al. Nature Prot 2022; Itoh et al. Cell Rep 2023; Veeraraghavan*, Engmann* et al. Nature Neurosci 2026; Durak*, Kim* et al. bioRxiv 2023; Veeraraghavan*, Tillman* et al. bioRxiv 2025; Tillman et al. bioRxiv 2026). Combining subtype-specific growth cone purification with ribosomal complex investigation represents a logical and exciting future direction. We appreciate the reviewer's encouragement of this future line of investigation and now highlight it in the revised Discussion.

      Likely impact of the work on the field, and the utility of the methods and data to the community:

      The authors introduce a unique pipeline of techniques to identify cell-type-specific ribosomal complex compositions. With more validation, there is certainly potential for those studying neuronal translation to leverage this method in limited primary cells as an alternative to existing methods that do not rely on ribosomal protein tagging, such as ARC-MS (Bartsch et al., 2023), RAPIDASH (Susanto and Hung et al., 2024), and RAPPL (Nature Communications, 2025).

      Reviewer #2 (Public review):

      Summary:

      This study presents a sophisticated molecular dissection of ribosome-associated complexes (RCs) in two well-defined cortical projection neuron subtypes (ScPN and CPN) during early postnatal development. 

      The authors develop and optimize an rRNA immunoprecipitation-mass spectrometry (rRNA IPMS) workflow to recover RCs from FACS-purified, retrogradely labeled neurons, achieving remarkable subtype specificity and biochemical resolution. Through proteomic profiling, they reveal both shared and distinct ribosome-associated proteins between ScPN and CPN, with a focus on non-core RC components and their potential functional relevance. The work advances our understanding of cell-type-specific translation regulation, moving beyond the transcriptome to explore the proteome-level complexity in neuronal subtypes.

      Strengths:

      This work stands out for its technical sophistication and innovation. The authors combine retrograde labeling, FACS purification, and an optimized rRNA IP-MS approach (low input) to isolate ribosome-associated complexes from highly specific neuronal subtypes in vivo, a challenging issue that they execute with impressive rigor. The methodological pipeline is both elegant and well-controlled, yielding high-quality, reproducible data. The depth of proteomic coverage is remarkable, with nearly all known cytoplasmic ribosomal proteins identified, along with hundreds of ribosome- associated proteins (RAPs), including translation factors, chaperones, and RNA-binding proteins. 

      The analysis not only reveals shared components between ScPN and CPN RCs but also uncovers subtype-specific differences in associated proteins. Particularly notable is the integration of this new proteomic dataset with previously published transcriptomic and ribosome footprinting data, which helps to validate the specificity and relevance of the findings. Overall, the clarity of the writing, the robustness of the data, and the transparency of the methods make this a strong and compelling contribution.

      Weaknesses:

      Despite the depth and high quality of the dataset, the study remains descriptive. While the identification of subtype-specific RC components is intriguing, the current version of the manuscript does not explore their functional roles or the biological consequences of their alterations. There is no perturbation, causal testing, in vitro or in vivo manipulation to demonstrate whether these proteins are necessary for ScPN or CPN identity, specific axonal targeting, metabolism, or synaptic function. One important point highlighted by the authors in the discussion - and critical for establishing the subtype specificity of the identified proteins - is that some ribosomal complexes may be specialized for specific developmental stages, rather than exclusively for the subtype-specific needs of projection neuron development. The work presented here provides a valuable starting point for further investigation into such RC specialization. 

      However, it will be essential to determine to what extent these RCs exhibit true subtype specificity, independently of their temporal maturation context. As a result, key mechanistic insights remain a bit speculative. Although several of the identified proteins have known roles in processes like synaptogenesis or metabolism, their relevance to the specific neuronal subtypes under study is not experimentally addressed. 

      That said, given its rich content and the comprehensive early postnatal dataset, the manuscript represents an extremely valuable resource for the community. While primarily exploratory, it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.

      We thank the reviewer for their excellent summary and for their very positive assessment of our work. We are pleased that the methodological rigor, proteomic depth, and integrative analyses were well-received. We thank the reviewer for highlighting that our “work presented here provides a valuable starting point for further investigation into such RC specialization” and that “it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.” This is exactly how we view this contribution – as a foundation to share with broader field so such functional investigation and investigation of developmental dynamics can be pursued by multiple groups in the broader related field.

      We again thank the reviewer for the very positive and insightful comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) As listed in the limitations, I would recommend that the authors consider applying their rRNA affinity purification to additional cell lines to confirm the specificity of the method as a positive control, where just demonstrating the technology may be easier to carry out than with more limited samples.

      As we noted in our response above to Limitation 1: “The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or nonspecific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.”

      (2) Also, as listed in the limitations, I recommend that the authors provide orthogonal experimental evidence for the putative SCPN and CPN ribosomal complex proteins of interest (e.g., sucrose gradient > western blot).

      As we noted in our response to Limitation 2 above: “We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.”

      Regarding sucrose gradient fractionation (for polysome profiling), we also noted in our response to Limitation 2: “We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph”.

      (3) The authors are interested in preserving the native state of the projection neurons to isolate ribosomal complexes; therefore, they may consider using biotin-conjugated CTB (Thermo Fisher) sorting through anti-biotin MACS columns (Miltenyi Biotech) as opposed to FACS in the future. While it does not allow for the same visualization as using a CTB-FP, this slight pipeline adjustment could help save time and physical processing of the projection neurons, helping preserve their endogenous state.

      While we appreciate the reviewer’s constructive suggestion for this theoretically alternative approach, we respectfully submit that this approach is unlikely to effectively isolate projection neurons from the living brain as effectively as FACS approach employed here. We considered this approach. We respectfully offer that CTB standardly enters neurons by binding GM1 gangliosides at the axon terminal, after which it is internalized and undergoes retrograde transport to the soma, several millimeters or more away depending on the projection. By the time CTB reaches the soma, we further respectfully offer that it is standardly fully internalized and is no longer surface-exposed for the theoretically suggested antibiotin capture for MACS. We used magnetic bead-based pull-down for the ribosomes themselves and find those molecular approaches very beneficial. Despite our use and openness to magnetic-conjugate-based separation approaches, we judge that fluorophore-conjugated CTB and FACS remain the most efficient and feasible approach for isolation of these exceptionally polarized projection neurons while maintaining cell viability. 

      (4) While the authors provided a comparison of overall protein detection levels between SCPN and CPN, I recommend an additional analysis and potential normalization for the average core ribosomal protein abundance across samples. I will note that this is more accessible when samples are prepared using TMT-labeling methods, which might be considered for future experiments.

      We thank the reviewer for these suggestions. We appreciate the opportunity to address them together, to further clarify our deeply considered choice of MS-based proteomic analysis and corresponding primary normalization approach. To complement our initial normalization approach, we have now also implemented the reviewer’s suggested approach of normalization to the average intensity of core ribosomal proteins. Notably, these new results agree with those from our initial normalization approach. This insightful suggestion has further strengthened the paper’s results and interpretation. 

      Here, we employed label-free quantification (LFQ), wherein samples are assayed sequentially rather than simultaneously, to most rigorously establish which proteins are truly present in some neuronal subtypes but absent in others. As the reviewer is aware, this capability to determine absence vs. presence of MS-detectable peptides distinguishes LFQ from approaches that assay samples simultaneously, such as isobaric tandem mass tag (TMT) labeling, which standardly offers advantages in relative quantification studies. Advances in sample preparation, instrumentation, and data analytical algorithms (from recent developments in single-cell proteomics and related approaches) now enable application of LFQ in quantitative differential analysis, even in the ultra-low-input regime. We now include both detailed discussion of these points and relevant citation in the text. 

      As the reviewer is also aware, in LFQ, normalization is crucial to ensure the protein quantification is accurate and comparable across sequential runs. For quantitative differential analysis of proteins detected in both subtypes, we have implemented a primary normalization approach employing “median-of-ratios” normalization across all detected proteins for robustness against outliers and technical variability. This primary approach results in similar overall distributions of protein abundances across CPN and SCPN samples (Figure S2A), providing confidence in quantitative comparison between SCPN and CPN in the ultra-lowinput regime. 

      Following the reviewer’s suggestion, to further ensure rigor of identification of differential proteins, we have also implemented a second normalization approach, rescaling each sample to the average intensity of its core ribosomal proteins alone (new Figure panels S2B, C). Subsequent differential analysis reveals equivalent CPN > SCPN enrichment of RPS30/eS30, GUCY1A1, and CELF3 (proteins identified as CPN-enriched with the primary normalization approach). These three proteins rank among the five proteins with lowest p-values, though false-discovery-rate-corrected significance is reduced. This confirmation by a second normalization approach further strengthens the findings.

      We have included clarifications in the main text, added the second normalization approach and subsequent analysis in both the main text and Figure S2. In addition, we have now noted in the discussion that TMT labeling with correspondingly appropriate normalization approaches might better define relative quantitative differences between functional candidates present in multiple subtypes.

      Recommendations for improving the writing and presentation:

      (1) I recommend this as a Tools or Resource article, seeing as the biological conclusions are limited.

      We respectfully submit that this work investigated biological questions and identified biological answers beyond pure development of Tools or offering a dataset as a Resource. We further respectfully submit that the question of differential neuronal subtype-specific translation of shared transcripts has become an emerging area of interest in regulation of precise neuronal and circuit development, maintenance, and function, as well as the neurobiological basis of disease. This has been quite hard to study, and this paper brings a first level of answers to that biological question. Of course, the biological results of this paper are not the complete answer, but as with all biological discovery papers, it provides a foundation for many further studies by multiple labs. 

      (2) In the rationale for studying ribosomal complex machinery, it may be helpful to say that ribosome composition and associated proteins that are present in the cytoplasm provide a way in which ribosomes can tune translation rapidly. This is especially important, seeing as this affords post-mitotic neurons the opportunity to remodel and repair by using readily available proteins while also avoiding the energetic demands of producing new ribosomal complex proteins.

      We thank the reviewer for this insightful comment and fully agree. We have now added this rationale to the Introduction. 

      (3) Is there a need for a CTB injection control? Does GM1 binding affect translation pathways? Please list citations, if possible.

      We respectfully submit that a CTB injection control is not necessary. As noted in our response to Limitation 1 in the Public Review portion, both PN subtypes underwent retrograde labeling with CTB in this work. While it remains unknown whether CTB-GM1 binding affects translation, potential effects of CTB are expected to apply equivalently to both subtypes and would therefore not confound these between-subtype comparisons. Please also see our response to recommendation 4 immediately below, in which we further clarify that retrograde tracing with CTB is a long-standing, well-accepted method shown to cause minimal damage and not interfere with continued neuronal development.

      (4) Do the traced/labeled PNs keep developing normally? In other words, does retrograde tracing inhibit proper PN development? Please list citations, if available.

      We thank the reviewer for encouraging us to further clarify that these are longstanding and well-accepted methods in the field, found to cause minimal damage and not to interfere with continued neuronal development. These and related retrograde labeling methods have led to the identification of the field's cardinal regulatory genes and molecules of axonal connectivity. This includes substantial work from our own lab (PMID in parentheses): Arlotta*, Molyneaux* et al. Neuron, 2005 (15664173); Lai*, Jabaudon* et al. Neuron, 2008 (18215621); Molyneaux*, Arlotta* et al. J. Neurosci., 2009 (19793993); Galazo et al. Neuron, 2016 (27321927); and more recently Sahni et al. Cell Rep., 2021a (34686320); and Sahni et al. Cell Rep., 2021b (34686337). These methods have also employed by other groups, such as Bin Chen (e.g. McKenna et al. PNAS, 2015 (26324926)) and Marta Nieto (e.g. De León Reyes et al. Nat. Commun., 2019 (31591398)). We have now clarified this in the text and included references for the benefit of the readers.

      (5) It may be helpful to mention that RPS30 associates with immature ribosomes during biogenesis (PMID: 25706898).

      We thank the reviewer for the opportunity to further clarify background knowledge of RPS30. As the reviewer is aware, RPS30/eS30 is definitively part of the mature 80S ribosome, as established, e.g., by cryo-EM of human 80S ribosomes (PMID: 25901680). Like multiple other ribosomal proteins, RPS30/eS30 also associates with immature ribosomes during biogenesis (PMID: 25706898). Intriguingly, RPS30/eS30 is produced by cleavage of a fusion protein comprising ubiquitin-like FUBI and RPS30/eS30, with cleavage recently identified as a late step in cytoplasmic 40S maturation (PMID: 34318747). As noted in the text, we confirmed that the peptides used for RPS30/eS30 identification appropriately map only to the amino acid sequence of RPS30/eS30 and not FUBI. We now mention that it is a core component of the mature 80S ribosome and its immature ribosomal association. 

      (6) Please clarify how the rRNA-IP is pulling down mature ribosomes. If not, this should be incorporated into the discussion.

      We thank the reviewer for raising this interesting point. As the reviewer notes, it is well established in the ribogenesis field that ribosomes are continuously produced and therefore exist at various stages of maturation. Our protocol removes a major source of immature ribosomes by subjecting FACS-purified cells to two centrifugal spins that remove the nucleus, the site of ribogenesis and early steps of maturation. In pilot experiments, nuclear removal was confirmed by the absence of a contaminating genomic DNA peak on Bioanalyzer electropherograms of total RNA extracted from input samples (without genomic DNA removal) immediately prior to rRNA-IP. However, ribosomes also undergo cytoplasmic maturation steps, and various functional states of ribosomes have been found to be present in the cytoplasmic fraction. For these reasons, we have referred to what we pulled down as "ribosomal complexes" throughout the manuscript. We now explain this nuclear/immature ribosome depletion and cytoplasmic ribosomal enrichment in the text.

      (7) Would have been interested to see some discussion of the most enriched CPN RAPs or why these might not exist in most replicates (inter-subtype heterogeneity?)

      The reviewer asks an interesting question. As the reviewer is aware, when considering a single sample in isolation, absence of mass spectrometry-based detection is not definitive proof of absence, especially not within this work’s ultralow input regime. One of us (B. Budnik) has substantial experience with ultra-low-input samples across multiple cell types outside of the nervous system, and identifying a protein in three of four identical samples is not uncommon. We therefore used detection in three or four samples as an indicator of presence, while absence across all samples was taken to indicate true absence.

      That said, the reviewer is correct that further diversity and heterogeneity within both CPN and SCPN subtypes additionally might be involved. This is an interesting question for future research, and we have added relevant text to the Discussion. 

      (8) It may be helpful to mention that RPL22, while not stoichiometric, is known to have extraribosomal functions and can pull down independently from assembled ribosomes (PMID: 17381311, PMID: 28575669. Moreover, RPL22/eL22-3xFLAG has been previously used as a control for ribosome affinity-based pulldowns and could be added to citations for Figure S1 (PMID: 28625553, PMID: 28625553).

      We thank the reviewer for this helpful recommendation. We have added relevant text and citations. 

      Minor corrections to the text and figures:

      (1) Want to confirm that in Figure 2, P adj is <0.1 is correct? 

      Yes

      (2) It is not necessary to show MS spectra in Figure 2.

      While we understand the spectra are not strictly necessary, we respectfully submit that they enhance the figure and aid readers in assessing data quality. 

      Reviewer #2 (Recommendations for the authors):

      To strengthen the impact and interpretation of the authors' findings, we encourage consideration of the addition of functional validation experiments for at least one (or more) of the ribosomeassociated proteins that are differentially enriched in ScPN. This could include genetic manipulation (e.g., knockdown or overexpression) to test whether these proteins influence subtype-specific features or neuronal function. Even a limited set of perturbation experiments, such as targeting PRKCE, which is particularly interesting due to its known role in synaptogenesis, would help move the study from descriptive to a more mechanistic nature of the work.

      There appear to be no issues related to data availability, ethics, or compliance, assuming all raw proteomic data and associated code for differential analysis are made publicly available.

      We again thank the reviewer for this encouragement and highlighting that our work provides the foundation for further functional and developmental dynamic investigations. We view this work as providing that foundation for further investigation by multiple labs in the broader fields, beyond the scope of this paper.

    1. eLife Assessment

      This important study addresses how listeners learn the statistical properties of acoustic spaces, combining well-designed psychophysics in virtual rooms with non-invasive brain stimulation. The evidence is solid: the behavioral results are strong and show that speech understanding is best in rooms with everyday levels of reverberation, and the stimulation data are consistent with a contribution of dorsolateral prefrontal cortex, though the effects are modest and the spatial precision of TMS is inherently limited. The work will be of interest to researchers in auditory neuroscience and psychoacoustics, helping to understand how the brain copes with reverberant environments.

    2. Reviewer #3 (Public review):

      Summary:

      This manuscript presents a well-designed study examining human adaptation to room acoustics, building on prior work. The psychophysical results are convincing and add meaningful knowledge to our understanding of reverberation learning. The transcranial magnetic stimulation (TMS) component shows a role of prefrontal cortex in this listening task, targeting dorsolateral prefrontal cortex (dlPFC). Cautious interpretation of the TMS results is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, and the limited ability of TMS to precisely target dlPFC in individual subjects. A surprising and interesting finding in the study is that listeners performed the speech recognition task more poorly in anechoic conditions than in those with naturalistic levels of reverberation. This is likely due to contributions of spatial release from masking provided by reverberation acoustics, which may counteract the detrimental effects of reverberation on speech perception itself. Overall, the experiments are well performed and clearly presented, improving our understanding of how the brain copes with reverberant environments.

      Strengths:

      (1) Well designed acoustical stimuli and psychophysical task.

      (2) Comparisons across room combinations is well conducted.

      (3) Virtual acoustic environment is impressive and applied well here.

      (4) Timely study with interesting behavioural results.

      (5) Causal evidence of a role for dlPFC in reverberation learning.

      Weaknesses:

      (1) Poorer performance in anechoic environments than some reverberant environments suggests and interplay of spatial release from masking and speech intelligibility that are not fully unpicked here. This could be controlled in future experiments, for example by comparing monaural and binaural listening conditions.

      (2) Lack of evidence for targeting TMS to dlPFC in individual participants. This is simply a limitation of the technique which the reader should keep in mind.

      (3) Most interesting effect of TMS is a null result compared to a weak statistical effect for "meta-adaptation"

    3. Reviewer #4 (Public review):

      The authors use a d' defined for 2-alternative forced choice experiments, but their data are 4-alternative (for color) and 8-alternative (for number) forced-choice. So, the d' is not computed correctly. For mAFC experiments, the authors should use the Hacker-Ratcliff (1979) method, also defined in chapter 10 of the Macmillan & Creelman textbook.

      Normalization of the stimuli was arbitrary, and consequently the unexpected improvement in performance in reverberation compared to anechoic condition is still not explained. A natural normalization across different environments is to take the direct portion of the BRIR (and HRTF for the anechoic condition) and make sure that that is scaled identically across the different simulated rooms (with the reverberant tails scaled naturally). That corresponds to the situation when the sources are emitting the sound at the same level in each environment. The current study scaled the overall levels. As a minimum, it should be reported how this scaling boosted/attenuated the targets and maskers in each environment.

      The potential that the listeners are tuning to individual voices, as opposed to rooms, has not been eliminated. The authors suggest that a lack of interaction with different voices is evidence that that is not the case. This is not correct: lack of this interaction just means that there are no differences in tuning between the voices. But that does not mean that the same amount of tuning is happening for each voice, as observed in previous studies. Unless the authors provide a follow-up data with randomly varying voices within each trial, these claims should be tuned down.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of an experiment that demonstrates a disruption in statistical learning of room acoustics when transcranial magnetic stimulation (TMS) is applied to the dorsolateral prefrontal cortex in human listeners. The work uses a testing paradigm designed by the Zahorik group that has shown improvement in speech understanding as a function of listening exposure time in a room, presumably through a mechanism of statistical learning. The manuscript is comprehensive and clear, with detailed figures that show key results. Overall, this work provides an explanation for the mechanisms that support such statistical learning of room acoustics and, therefore, represents a major advancement for the field.

      Strengths:

      The primary strength of the work is its simple and clear result, that the dorsolateral prefrontal cortex is involved in human room acoustic learning.

      Weaknesses:

      A potential weakness of this work is that the manuscript is quite lengthy and complex.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how listeners adapt to and utilize statistical properties of different acoustic spaces to improve speech perception. The researchers used repetitive TMS to perturb neural activity in DLPFC, inhibiting statistical learning compared to sham conditions. The authors also identified the most effective room types for the effective use of reverberations in speech in noise perception, with regular human-built environments bringing greater benefits than modified rooms with lower or higher reverberation times.

      Strengths:

      The introduction and discussion sections of the paper are very interesting and highlight the importance of the current study, particularly with regard to the use of ecologically valid stimuli in investigating statistical learning. However, they could be condensed into parts. TMS parameters and task conditions were well-considered and clearly explained.

      Weaknesses

      (1) The Results section is difficult to follow and includes a lot of detail, which could be removed. As such, it presents as confusing and speculative at times.

      (2) The hypotheses for the study are not clearly stated.

      (3) Multiple statistical models are implemented without correcting the alpha value. This leaves the analyses vulnerable to Type I errors.

      (4) It is confusing to understand how many discrete experiments are included in the study as a whole, and how many participants are involved in each experiment.

      (5) The TMS study is significantly underpowered and not robust. Sample size calculations need further explanation (effect sizes appear to be based on behavioural studies?). I would caution an exploratory presentation of these data, and calculate a posteriori the full sample size based on effect sizes observed in the TMS data.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a well-designed and insightful behavioural study examining human adaptation to room acoustics, building on prior work by Brandewie & Zahorik. The psychophysical results are convincing and add incremental but meaningful knowledge to our understanding of reverberation learning. However, I find the transcranial magnetic stimulation (TMS) component to be over-interpreted. The TMS protocol, while interesting, lacks sufficient anatomical specificity and mechanistic explanation to support the strong claims made regarding a unique role of the dorsolateral prefrontal cortex (dlPFC) in this learning process. More cautious interpretation is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, the imprecise targeting of dlPFC (which is not validated), and the lack of knowledge about the timescale of TMS effects in relation to the behavioural task. I recommend revising the manuscript to shift emphasis toward the stronger behavioural findings and to present a more measured and transparent discussion of the TMS results and their limitations.

      Strengths:

      (1) Well-designed acoustical stimuli and psychophysical task.

      (2) Comparisons across room combinations are well conducted.

      (3) The virtual acoustic environment is impressive and applied well here.

      (4) A timely study with interesting behavioural results.

      Weaknesses:

      (1) Lack of hypotheses, particularly for TMS.

      (2) Lack of evidence for targeting TMS in [brain] space and time.

      (3) The most interesting effect of TMS is a null result compared to a weak statistical effect for "meta adaptation"

      Reviewer #4 (Public review):

      Summary:

      Several behavioral experiments and one TMS experiment were performed to examine adaptation to room reverberation for speech intelligibility in noise. This is an important topic that has been extensively studied by several groups over the years. And the study is unique in that it examines one candidate brain area, dlPFC, potentially involved in this learning, and finds that disrupting this area by TMS results in a reduction in the learning. The behavioral conditions are in many ways similar to previous studies. However, they find results that do not match previous results (e.g., performance in anechoic condition is worse than in reverberation), making it difficult to assess the validity of the methods used. One unique aspect of the behavioral experiments is that Ambisonics was used to simulate the spaces, while headphone simulation was mostly used previously. The main behavioral experiment was performed by interleaving 3 different rooms and measuring speech intelligibility as a function of the number of words preceding the target in a given room on a given trial. The findings are that performance improves on the time scale of seconds (as the number of words preceding the target increases), but also on a much larger time scale of tens to hundreds of seconds (corresponding to multiple trials), while for some listeners it is degraded for the first couple of trials. The study also finds that the performance is best in the room that matches the T60 most commonly observed in everyday environments. These are potentially interesting results. However, there are issues with the design of the study and analysis methods that make it difficult to verify the conclusions based on the data.

      Strengths:

      (1) Analysis of the adaptation to reverberation on multiple time scales, for multiple reverberant and anechoic environments, and also considering contextual effects of one environment interleaved with the other two environments.

      (2) TMS experiment showing reduction of some of the learning effects by temporarily disabling the dlPFC.

      Weaknesses:

      While the study examines the adaptation for different carrier lengths, it keeps multiple characteristics (mainly talker voice and location) fixed in addition to reverberation. Therefore, it is possible that the subjects adapt to other aspects of the stimuli, not just to reverberation. A condition in which only reverberation would switch for the target would allow the authors to separate these confounding alternatives. Now, the authors try to address the concerns by indirect evidence/analyses. However, the evidence provided does not appear sufficient.

      The authors use terms that are either not defined or that seem to be defined incorrectly. The main issue then is the results, which are based on analysis of what the authors call d', Hit Rate, and Final Hit rate. First of all, they randomly switch between these measures. Second, it's not clear how they define them, given that their responses are either 4-alternative or 8-alternative forced choice. d', Hit Rate, and False Alarm Rate are defined in Signal detection theory for the detection of the presence of a target. It can be easily extended to a 2-alternative forced choice. But how does one define a Hit, and, in particular, a False Alarm, in a 4/8-alternative? The authors do not state how they did it, and without that, the computation of d' based on HR and FAR is dubious. Also, what the authors call Hit Rate, is presumably the percent correct performance (PCC), but even that is not clear. Then they use FHR and act as if this was the asymptotic value of their HR, even though in many conditions their learning has not ended, and randomly define a variable of +-10 from FHR, which must produce different results depending on whether the asymptote was reached or not. Other examples of usage of strange usage of terms: they talk about "global likelihood learning" (L426) without a definition or a reference, or about "cumulative hit rate" (L1738), where it is not clear to me what "cumulative" means there.

      There are not enough acoustic details about the stimuli. The authors find that reverberant performance is overall better than anechoic in 2 rooms. This goes contrary to previous results. And the authors do not provide enough acoustic details to establish that this is not an artefact of how the stimuli were normalized (e.g., what were the total signal and noise levels at the two ears in the anechoic and reverberant conditions?).

      There are some concerns about the use of statistics. For example, the authors perform two-way ANOVA (L724-728) in which one factor is room, but that factor does not have the same 3 levels across the two levels of the other factor. Also, in some comparisons, they randomly select 11 out of 22 subjects even though appropriate test correct for such imbalances without adding additional randomness of whether the 11 selected subjects happened to be the good or the bad ones.

      Details of the experiments are not sufficiently described in the methods (L194-205) to be able to follow what was done. It should be stated that 1 main experiment was performed using 3 rooms, and that 3 follow-ups were done on a new set of subjects, each with the room swapped.

      We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and insightful comments. We greatly appreciate the time and expertise invested in reviewing our work. The feedback has been invaluable in improving the clarity, rigor, and overall presentation of the manuscript.

      In response to the reviewers’ comments, we have carefully revised the manuscript throughout. The revisions include clarification of the study hypotheses, re-analysis of the TMS data using a mixed ANOVA framework, additional methodological details regarding the TMS procedures and behavioural analyses, expanded justification of the statistical modelling approach, clarification of the room-acoustics paradigm, revision of figures and figure legends, additional discussion of study limitations, and a more balanced interpretation of the TMS findings. We have also substantially revised the Discussion section and improved the overall structure and readability of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is understood that this topical area is necessarily detail-heavy, but if there are ways to streamline the manuscript to more quickly arrive at key results (Figure 3?), the work might have an even greater overall impact.

      We appreciate the reviewer’s feedback and have carefully revised the manuscript to address all comments from all reviewers. However, we have retained the existing order of figures and results to maintain consistency and avoid extensive structural changes that could compromise the clarity and flow of the manuscript.

      (2) Minor point: I believe Equation 1 should be d' = z(H) – z(F).

      Thanks for noticing this, we have fixed the equation.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 115: Towards the end of the introduction, the hypotheses for the current study remain unclear. Please explicitly outline each hypothesis before the Methods section.

      We have now specified the hypotheses tested at the end of the Introduction.

      (2) Line 229: Please state the minimum MEP amplitude criterion used during TMS thresholding (usually 50 µV). Please also reference the EMG hardware and software used to measure MEPs.

      We are grateful to the reviewer for noticing that a few details regarding our TMS procedure were missing in the “Continuous theta-burst stimulation” section of the Methods. We have extensively revised this section to include the required details. We did not, however, use MEP amplitude as a criterion for estimating motor thresholds. For single-pulse TMS-induced motor threshold determination, we used visual observation of the first dorsal interosseous (FDI) muscle twitch (Varnava et al., 2011).

      (3) Line 248: The Coordinate Response Measure corpus contains combinations of callsigns, colours, and numbers. What was the rationale in asking participants to only identify the colour and the number spoken, and not the callsign?

      The rationale for reporting only Color and Number was to ensure an equal number of keyword identifications across phrase lengths—for example, CP0 and CP1 do not contain a callsign. This has now been clarified in the Procedure section of the Methods.

      (4) Lines 327–334: The analyses outlined here are unclear – it appears that there are multiple different statistical tests being conducted on the same outcome variable in this study. Given that the design includes a between-subjects factor (TMS condition) and two within-subjects factors (Rooms, Carrier Length Phrase), could the analyses be simplified by employing a mixed ANOVA as opposed to separate repeated-measures and univariate ANOVAs? If this is, in fact, the analysis that was conducted, please improve the wording. Clarification/correction is also needed surrounding the use of the term “univariate” ANOVA, as this can refer to different statistical tests.

      We appreciate the reviewer pointing this out. We have re-analysed the TMS data using a mixed ANOVA design and reported it as such in the Results section. Although the numerical values have changed, the significant findings and conclusions remain the same.

      (5) Line 340: Whilst the authors identify that an alpha value of 0.05 and Bonferroni corrections were used for statistical inference in two-tailed t-tests, there is no indication of the alpha value used in the interpretation of the ANOVA results. Please include this before the Results section. Furthermore, given that multiple tests are being conducted in this project, the alpha value used to infer statistical significance should be corrected in accordance with the number of hypotheses, to reduce Type I error rates (e.g., 4 hypotheses would result in an alpha inference criterion of α = 0.0125).

      We appreciate the reviewer’s observation. The corrected alpha value was not explicitly reported because all statistical analyses were conducted using IBM SPSS Statistics for Windows, Version 29.0.2.0 (IBM Corp., Armonk, NY; RRID: SCR_002865). In SPSS, Bonferroni corrections are applied by adjusting the p-values rather than the alpha threshold itself. For transparency, we have now clarified in the manuscript the factors included in each statistical comparison to make the tested hypotheses fully explicit.

      (6) Line 346: The sample size calculation used in the current study could be improved. It is unclear why a sample size estimation of ≥18 is used when the actual sample recruited is significantly greater than this (62). Is this because 62 participants were divided across multiple experiments in this paper? Further clarification is needed. The alpha value used in this calculation should also be corrected to account for multiple statistical models.

      We thank the reviewer for noting this. We have clarified that the initial sample size estimation (n = 18) referred to individual ANOVA analyses. In the revised manuscript, we specify in each experimental section (Identity of Sound Environments and Continuous Theta Stimulation) the exact number of participants recruited per experiment, which together sum to a total of 74 participants across all experiments (also clarified in the Participants section of the Materials and Methods).

      (8) Line 438: Many of the statistical tests presented in the Results section have not been outlined in the Methods section or had the rationale explained in the Introduction – this makes the analyses feel confusing and speculative. Explicitly identifying the core hypotheses earlier in the manuscript and clearly stating which hypotheses are confirmatory or exploratory would be an important improvement.

      We appreciate the reviewer noticing this. We have clarified the hypotheses tested in the Introduction (4th and 5th paragraphs) and provided a more detailed account of the statistical analyses in the Materials and Methods (Statistical Analysis section).

      (9) Line 565: It is unclear how the 62 participants recruited in the study were divided across each of the experimental conditions. I would recommend expanding the Participants section in Methods to outline the number of participants involved in each stage of the study.

      Thanks for noticing this discrepancy—this calculation was indeed confusing and incorrect. We ultimately tested a total of 74 participants. We have specified the sample size per experiment in both the “Identity of selected sound environments” and “Continuous theta-burst stimulation” sections of the Methods, and reiterated this in the Results section to prevent confusion.

      (9) Line 792: The Results section as a whole is incredibly complex and lacks structure. This could be condensed significantly – details regarding previous research should be removed from Results, as this should already be outlined in the Introduction as rationale for the current project. It would be clearer to outline each confirmatory and exploratory hypothesis in the Introduction, then identify at each stage in the Results section where it is being tested.

      We appreciate the reviewer’s point. We have rewritten the 4th and 5th paragraphs of the Introduction to outline the confirmatory and exploratory hypotheses. While we have streamlined parts of the Results section to enhance clarity, we have retained contextual information for each analysis to help readers follow the logic of the findings, given the complexity and scope of the study.

      Reviewer #3 (Recommendations for the authors):

      (1) Introduction – Overinterpretation and Hypothesis Clarity. The final paragraph of the Introduction (page 4) discusses the experimental findings and their interpretation, which belongs in the Discussion. This section should instead clearly state hypotheses for both the behavioural and TMS experiments. In particular, the TMS experiment lacks a clear rationale: what mechanism is being tested, and what behavioural outcome is predicted? Please revise this section to focus on the theoretical motivation, clearly defined hypotheses, and expected results, and less on summarizing the results.

      We appreciate this point and have accordingly deleted the final paragraph of the Introduction, as it is already incorporated into the Discussion section. We have stated the hypotheses/rationale more clearly in the Introduction.

      (2) TMS – Mechanism, Timeframe, and Clarity. The manuscript does not adequately explain how TMS produces long-lasting effects relevant to the task, which occur minutes (or possibly longer) after stimulation.

      (a) What is the specific timeframe between stimulation and behavioural testing?

      The timeframe between stimulation and behavioural testing has been clarified in the Methods section, e.g.: “Participants exposed to ‘real’ or ‘sham’ TMS completed the familiarization and behavioural task right after cTBS procedures.”

      (b) What is the evidence that TMS to the prefrontal cortex affects function on this timescale?

      We have added the following paragraph to the Discussion to address the timescale of continuous theta-burst stimulation (cTBS): cTBS, as used in our study, typically induces aftereffects lasting 20–50 minutes (Huang et al., 2005; Wischnewski & Schutter, 2015). While these effects are well established in the motor cortex—with motor-evoked potential changes persisting for up to one hour—recent evidence suggests that similar durations of cortical modulation can also occur in the prefrontal cortex (Taylor et al., 2025). Specifically, studies applying inhibitory rTMS to the dorsolateral prefrontal cortex (dlPFC) during cognitive tasks have demonstrated functional effects lasting up to one hour in healthy participants (Wagner et al., 2006). Furthermore, Tupak et al. (2013) showed that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels, reflecting decreased cortical activity, for at least 45 minutes—the same duration as the experimental task in our study. Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), the fNIRS-measured alterations in local cerebral blood oxygenation provide an indirect but reliable indicator of TMS-induced neural modulation within this timescale.

      (c) Can post-stimulation effects be objectively measured or confirmed?

      Although no objective post-stimulation measures were collected in the present study, we acknowledge this as a limitation. However, previous research has shown that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels—reflecting decreased cortical activity—for at least 45 minutes (Tupak et al., 2013). Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), these findings indicate that fNIRS can serve as an indirect but reliable method for confirming TMS-induced neural modulation. We plan to incorporate such objective measures in future studies.

      We thank the reviewer for raising these important points and have addressed them by updating the Methods and Results sections and adding a paragraph to the Discussion. We would like to clarify, however, that the TMS effects observed in our study are not weak: the statistically significant differences between sham and TMS conditions were accompanied by large effect sizes, indicating that bilateral inhibitory stimulation of the dlPFC produced a robust and consistent effect across participants—specifically, a reduction in performance, reflecting decreased improvement in speech understanding with increasing exposure to the reverberant environment.

      (3) Figure 3A – Anatomical Specificity and Interpretation. Figure 3A implies precise stimulation of dlPFC and its projections to auditory cortex (A1), but the authors cannot actually target dlPFC or its connections specifically with this approach. Rather, the TMS protocol disrupts an undetermined region of PFC, with diffuse downstream effects. This should be clearly acknowledged in the figure legend and main text.

      We appreciate this comment and have acknowledged this point in the figure legend and Discussion, e.g., Figure 3 legend: “Although the TMS protocol was intended to target the dlPFC, it likely affected adjacent prefrontal regions, leading to diffuse downstream effects that may have included modulation of A1.”

      (4) Additionally, the Discussion overstates the evidence for a specific dlPFC → AC role in reverberation learning. The weak and poorly localized TMS effect does not support strong claims about this pathway. Please scale back this interpretation and focus more on the robust psychophysical results, which are the manuscript's stronger contribution.

      We thank the reviewer for this comment. We have revised our interpretation to clarify that the proposed dlPFC–auditory cortex link is speculative, and have added caveats regarding the limited spatial precision of TMS targeting and individual variability in its effects, in the Discussion section “A role for dlPFC in statistical learning of room acoustics.”

      (4a) Line 152: Extra comma after “of”? Also, why are there square brackets around “callsigns” etc.?

      Fixed.

      (4b) Line 156: “RRID:SCR_001622” is unexplained and likely unnecessary—consider removing.

      Removed.

      (4c) Line 159: Why was no ramping applied at the end of the noise? Please clarify.

      Similar to Brandewie & Zahorik (2013), no ramping was applied. This has been clarified in the Methods (“Acoustic Stimuli”).

      (4d) Line 207: Methods do not describe the sham TMS protocol—please add this information.

      Thanks for noticing this. Information on the sham TMS protocol has been added.

      (4e) Line 352: Fix bracket formatting.

      Fixed.

      (4f) Line 382: Sentence is grammatically incorrect—please revise.

      Fixed.

      (4g) Line 394: Unclear use of square brackets—clarify or standardize.

      Fixed.

      (4h) Line 401: It is unclear how interleaving the talker and length ensures a different room each trial. Aren't these variables independent?

      The reviewer is correct: the only variable that was pseudorandomized was room order, to prevent carry-over effects, similar to Brandewie & Zahorik (2013). This has been clarified in the Methods section.

      (4i) Figure 1: Clarify that AI-generated images refer only to the room images, not other components.

      Fixed.

      (4j) Figure 2A: Confirm that “overall” includes all speakers and durations—clarify in legend.

      Fixed.

      (4k) The interesting duration effects in Figure 1D are not discussed in the text and appear before overall room effects in Figure 2A—please reorder and comment on these results.

      Fixed.

      (4l) Supplementary Figure 2: Caption contains a typo (“Lecture Room/Open-Plan Office”).

      Typo has been fixed.

      Also, I recommend adding this result to the main figure set—e.g., include overall d′ for all six talkers in Figure 2 alongside rooms (2A) and lengths (2D).

      We thank the reviewer for the suggestion. Including overall performance for all six talkers in Figure 2 would require substantial restructuring and risk making the figure crowded. We have therefore retained these results in the Supplementary Materials (Supplementary 2 and 3), as originally presented.

      (4m) Line 555: Phrase “to better understand” could be clearer—consider rewording.

      This section has been reworded.

      (4n) Lines 587–592: The lack of main effect of TMS is helpful, but more important is whether interactions between TMS and room/length variables occur. Please report these interactions, as they are central to interpreting the TMS effects.

      We appreciate the reviewer highlighting this. We have reviewed this analysis and reported the interaction Condition x CP length as follows: “A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], and no significant interaction Condition x CP length was observed: [F (3,93) =0.48, p=0.69, ŋp2 = 0.01]; confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population”. 

      (4o) Line 601: Reiterate the timeframe of the TMS-behaviour gap. Is there supporting evidence that TMS can affect behaviour over this duration? Could null effects reflect fading TMS efficacy?

      We appreciate the reviewer pointing this out. We have clarified the timeframe of TMS stimulation in both the Methods and Results, e.g.: “The procedure began with TMS manipulation, and although the behavioural task lasted 45 minutes, the inhibitory effects of TMS extended for at least 60 minutes post-stimulation (Huang et al., 2005; Gamboa et al., 2010; Hoogendam, Ramakers, & Di Lazzaro, 2010; Romero et al., 2022).” However, we cannot dismiss individual differences in the duration of TMS effects, nor differences in efficacy duration between anatomical areas (motor cortex vs. dlPFC). We have noted this in the Discussion section “A role for dlPFC in statistical learning of room acoustics” (Pallant, 2011).

      (4p) Figure 3I: The “meta-adaptation” effect is marginal in both Exp 1 (p = 0.03) and Exp 2 sham (p = 0.04). These should be interpreted cautiously, given their statistical fragility.

      We appreciate this comment. We have now calculated effect sizes for all Wilcoxon signed-rank tests (Pearson’s r) and report them. For the two comparisons noted by the reviewer, the effect sizes are medium (Exp 1) and large (Exp 2). We are therefore confident that, even where the p-values are not extremely low, the statistical differences are reliable.

      (4q) Line 696: Reverberation is described as “common,” but it is nearly universal. Consider rephrasing to reflect this.

      We appreciate this suggestion and have rephrased this line.

      (4r) Line 816: The authors state that TMS reduced overall performance, but the earlier ANOVA (lines 587–592) shows no such effect. Please correct this discrepancy.

      We appreciate the reviewer noticing this. This section has been clarified: the lack of statistical significance at lines 587–592 relates to the comparison between a subset of ‘no-TMS-exposed’ listeners and ‘sham’-TMS-exposed listeners, made only to demonstrate the absence of placebo effects in the sham sample. Following the reviewers’ suggestions, we also re-analysed the data using a one-way ANOVA; this slightly changed the numerical values of the reported main effect but did not change the statistical outcome.

      Reviewer #4 (Recommendations for the authors):

      (1) Lines 201–202: It's not clear what is meant by combination and by carrier length here.

      This section has been rewritten for clarity.

      (2) Line 330: What is meant by “Univariate” here? I think this was a mixed ANOVA, with a betweensubject factor of TMS exposure and the remaining factors within-subject.

      We appreciate this suggestion; we have re-analysed this section to use a mixed-ANOVA design. The numerical results differ, but the statistical outcome remains the same.

      (3) Lines 335–343: This is impossible to follow if one does not understand that there were 3 followup experiments.

      Thank you for highlighting that this section was confusing. We have rewritten it to clarify the following: “Three follow-up experiments were performed (univariate ANOVA) to assess whether speech understanding was affected by room context (i.e., the third room in which Open-Plan Office and Underground Car Park were learnt), with one between-subjects factor: room context (levels: Anechoic Room, Living Room, Lecture Room, and Highly Reflectant Room).” We have also added the following earlier in the Methods: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).”

      (4) Lines 414–416: The review of Tsironis et al. (2024) (doi:10.1177/23312165241273399) does not provide strong evidence that there are multiple scales for adaptation to room reverberation (most adaptation effects stabilize within 1 sec).

      We apologize for this mistake, which arose from an issue with our reference manager. It has been corrected to: Robinson, Harper, & McAlpine (2016), Nature Communications, and Simpson, Harper, Reiss, & McAlpine (2014), Journal of Neuroscience.

      (5) Line 427: The supplementary figure shows that many subjects did not achieve asymptotic performance.

      We appreciate the reviewer pointing this out. The fittings have been extensively reviewed; please see the Methods and Results for the new fitting analysis. Indeed, some participants, although very slowly, keep improving over time without reaching clearly asymptotic behaviour. This section, however, referred specifically to the point at which performance stabilised within ±10% of final performance.

      (6) Line 428: The ±10% statistic is random (as discussed below). And why switch to HR now? And what is its meaning when the HR did not converge by the end of the run?

      We appreciate the reviewer raising the inconsistent use of d′ versus HR. d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis; time-course analysis does not allow us to calculate d′ at each trial or time point, owing to the lack of HR and FA values for single trials. We have clarified this in the Methods: “Given how d′ was calculated for |Color| and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data for each |Color| and |Number| affecting the temporal resolution of any generated curve. We therefore analysed the development of individual and average performance in each acoustic environment using cumulative hit rates, applying a 5-point moving average (~7 seconds) to each trace and plotting performance as a function of mean cumulative exposure time.”

      We have also extensively reviewed our fits, following these steps: (i) we compared single- vs. double-exponential fits across 22 participants (Supplementary Figure 1), which showed that double exponentials provided a better R<sup>2</sup> for the majority of participants; (ii) we re-ran all analyses forcing the fits to the HR endpoint; (iii) characterizing taus was not informative for our sample, given flat-like performance for some participants (Supplementary Figure 1)—in these cases, taus do not aid understanding of how performance stabilizes over time, particularly given the use of two taus; and (iv) cutting initial points differs by participant.

      (7) Line 434: Or that they learned/adapted to other characteristics that were fixed.

      Thank you—we have added “adapted to” in the sentence.

      (8) Supplementary tables often show differences, but the actual values are not shown. Also, the tables and figures randomly switch between d' and HR.

      Supplementary tables are intended only to show additional detail not reported in the main text or figures, to avoid redundancy; means (referred to in the Supplementary tables) are always shown in the main figures. We appreciate the reviewer raising the inconsistent use of d′ versus HR, addressed above, and have clarified this throughout the Methods and Results.

      (9) Lines 436–442: There seem to be a lot of issues with the fitting shown in Supplemental Figure 1 and Figure 2B:

      (a) It does not seem to converge, especially for the green line. So, presumably, the asymptotic value obtained for tau_slow was the upper bound set to 2000 s for many subjects' conditions. But those values are never shown—they should be in Supplemental Figure 1.

      (b) Then the FHR value, derived from that, is completely dependent on what the bound was set to, and is therefore arbitrary. And its value of 10% is also arbitrary. Why do this when tau itself of an exponential model represents the time it takes to reach 67% of the asymptotic value, from which one can derive whatever time it should take to reach the final 10%?

      (c) Even the use of the model specified by Equation 2 seems arbitrary. Average data in Figure 2B do not provide strong evidence for two time scales. If the authors are worried about the instability of the data at the beginning, a simple exponential with a weighted fit that prioritizes the later portions seems sufficient.

      (d) Lines 430–434: This conclusion seems wrong, based only on the arbitrary measure chosen for “global likelihood learning.” Looking at Figure 2B, there is no evidence that the green graph reached any asymptote, while for the yellow and blue it appears to have. The authors should try fitting a simple exponential function to it to show that tau is larger.

      We appreciate the reviewer’s comments and have significantly revised these sections of the Methods and Results. In summary, we fitted the data with single- and double-exponential functions. Double exponentials were fitted to the full time course. Single exponentials were fitted to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) to mitigate initial variability (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments. Compared against the truncated single-exponential fit, the double-exponential model retained a lower mean AIC (−420.7 vs. −399.1) and remained the preferred model for approximately half the subjects. The double-exponential model was preferred not only for its automated nature (requiring no manual truncation) but also for the magnitude of improvement: when the single model was superior, the advantage was marginal (ΔAIC = 9.7 ± 1.4), whereas the double model’s advantage was substantial (ΔAIC = 52.8 ± 8.9). We therefore used the double-exponential fit for further analysis.

      (10) Lines 436–446: How can this analysis be performed if asymptotic performance was not achieved in any of the conditions (nothing has plateaued in Figure 2B)? Also, the 10% FHR measure is dependent on the FHR estimate; correlating two measures based on the same measure is, by definition, expected to be correlated. This result seems to reflect that if one's learning is faster within a fixed number of trials (150), one has more opportunity to reach a higher final PCC even if asymptotic performance is identical.

      To clarify, the variables being correlated are not the FHR and ±10% of the FHR values themselves, but rather the time points at which each participant reached ±10% of their individual FHR during the task. This analysis therefore does not involve two measures derived directly from the same estimate. The timing of reaching ±10% of the FHR reflects the learning-settling trajectory rather than the FHR magnitude, so there is no a priori reason for the two measures to be intrinsically correlated. While asymptotic performance was not reached within 150 trials for some participants, the estimated FHR still provides a consistent individual marker of learning rate, allowing comparison of relative learning dynamics across participants and conditions.

      (11) Lines 448–461: Brandewie & Zahorik (2013) show that a large portion of that improvement is due to tuning to the voice and location of the speaker. Also, in the current study, there are some issues with the anechoic condition (see below).

      We thank the reviewer for this comment. This section refers specifically to results related to carrier phrase length, not to speaker identity (addressed separately below) or location, both of which were fixed in our study and therefore unlikely to account for the observed effects. We address the reviewer’s concerns about the anechoic condition in our responses below.

      (12) Line 491: What were the average trial numbers for the steady and initial trials? Also for the anechoic condition?

      We appreciate the reviewer raising this. We analysed the average trial number at which initial and steady trials occurred across a total of 360 trials: for all 22 participants, initial trials mean = 6 ± 4 and steady trials mean = 37 ± 9; sham TMS: initial trials mean = 5 ± 4, steady trials mean = 34 ± 6; real TMS: initial trials mean = 10 ± 9, steady trials mean = 38 ± 12; anechoic condition: initial trials mean = 6 ± 3, steady trials mean = 36 ± 8; and the 11 randomly selected subjects: initial trials mean = 6 ± 5, steady trials mean = 38 ± 7. This information has been added to the relevant Results sections.

      (13) Line 498: It's still not clear when the anechoic condition was performed. Lines 200–205 talk about combinations in which the Lecture Room was swapped, but it's impossible to follow when and how often that occurred. Given that the anechoic room was not included in the same way as the main three rooms, the conclusion at lines 500–504 is questionable.

      Thank you for noticing this. We have rewritten the relevant section of the Methods (“Identity of sound environments”) as follows: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics—i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).” The anechoic room was therefore explored in the same manner as the main three rooms.

      (14) Lines 508–521: Tuning to the talker's voice/location would not predict that the effect would be different for a different voice.

      We thank the reviewer for this point. If the improvement in speech understanding were due to tuning to a specific talker’s voice or location, we would expect the effect to differ across talkers. However, our analysis across six talkers (three female, three male) showed no significant interaction between talker, carrier phrase length, and room (F(30,630) = 0.73, p = 0.85, ηp<sup>2</sup> = 0.034). Although overall performance differed across talkers (main effect of talker: F(5,105) = 27.19, p < 0.001, ηp<sup>2</sup> = 0.56), these differences did not modulate the carrier phrase effect. We therefore conclude that the improvement in speech understanding with increasing carrier phrase length is consistent across talkers.

      (15) Lines 523–536: Neither of these tests addresses the question directly. That would require switching the talker randomly between the carrier and target (or throughout the sentence).

      We thank the reviewer for this comment. We respectfully disagree that our analyses fail to address the question. While our experiment was not specifically designed to test the effect of switching talkers between the carrier and target segments, we examined whether adaptation to a talker could explain the improvement in performance with increasing carrier phrase length through three complementary analyses: (1) a repeated-measures ANOVA testing for interactions between talker and carrier phrase length (see response above); (2) an analysis of potential carry-over effects across consecutive same-talker trials; and (3) an assessment of talker-learning effects in the absence of reverberation (anechoic condition). As detailed in the Results section “Improvements in performance are explained by exposure to the environment, not talker idiosyncrasies,” none of these analyses revealed evidence that talker identity influenced the observed improvement in speech understanding. We therefore conclude that the performance improvements with increasing carrier phrase length are better explained by adaptation to the acoustic environment than to specific talkers.

      (16) Line 544: Why is FHR used in this measure when d' is used for the standard analysis in Figure 2D?

      As noted above, d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis. Given how d′ was calculated for | Colour | and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data affecting temporal resolution. We therefore used cumulative hit rates, with a 5-point moving average (~7 seconds), plotted against mean cumulative exposure time. This has been clarified in the Methods.

      (17) Also, why is FHR, as opposed to HR (which I assume is really PCC), computed across the whole experiment?

      FHR refers to the Final Cumulative Performance. This naming was used to distinguish it from trial-by-trial Hit Rate used in the time-course analysis. FHR is the final data point of the cumulative hit rate—i.e., after all responses have been accumulated in that listening environment. This has been clarified throughout the manuscript.

      (18) Still worse, it's also not clear when these anechoic trials were measured.

      We appreciate the reviewer noting a lack of clarity here. We performed three follow-up experiments in different, naïve groups of listeners, assessing performance across combinations of three rooms, including Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners), assessed in the same way as Brandewie & Zahorik (2013). The anechoic room was therefore explored in the same manner as the main three rooms; this has been clarified in the Methods and reiterated in the Results.

      (19) And specifically, from Supplemental Table 4, it looks like the improvement was considerable (up to 15%), supporting that the effect is occurring. Also, note that there seems to be something numerically wrong in Supplemental Table 4: the improvement CP0–CP1 is −6, CP1–CP2 is −9.667, and CP2–CP3 is −2.5. Based on this, CP0–CP2 is expected to be −15.667 (which it is), but CP0–CP3 is expected to be −18.167, yet it's stated as −13.167.

      We appreciate the reviewer pointing this out. Our statistical analysis does not match the calculations the reviewer derived from Supplemental Table 4. For transparency, we report below the means for each carrier phrase in the anechoic room, exported directly from SPSS, which are the values reported in the manuscript. We have reviewed this section to improve clarity and have included a link to the raw supplemental data.

      CP0: Mean 40.500, SE 4.548, 95% CI [30.212, 50.788]

      CP1: Mean 46.500, SE 5.296, 95% CI [34.519, 58.481]

      CP2: Mean 56.167, SE 3.777, 95% CI [47.624, 64.710]

      CP3: Mean 53.667, SE 3.966, 95% CI [44.695, 62.638]

      (20) Lines 555–557: This sentence seems grammatically incorrect.

      It has been corrected.

      (21) Lines 559–560: The sentence “a brain region implicated in listening performance in noise (Houtgast & Steeneken, 1973; Knudsen, 1929; Lochner & Burger, 1961)” seems to imply that the cited studies support dlPFC being the brain region implicated in hearing in noise. None of these studies does that.

      Thank you for noticing this—this was an error introduced by our reference manager and has been corrected.

      (22) Lines 586–592: What was the “overall performance” measure—d′, HR, or PCC? Also, what is “univariate” analysis here? A mixed ANOVA with a between-group factor of condition and withingroup factors of room and CP length would be appropriate, and the whole group of 22 subjects should be used for the “no-exposure” group, rather than a random selection of an 11-subject subgroup.

      We appreciate this comment and we have revised this analysis to include a mix ANOVA as suggested by the reviewer. It reads as follows in the Manuscript: “Given the potential placebo effects of a ‘sham’ TMS stimulation, we first tested whether our sample of 11 ‘sham’ TMS participants exhibited similar behavioural performance to the larger sample of 22 participants who had not been exposed to any TMS manipulation. A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population.”

      (23) Lines 600–601: By “univariate ANOVA” is meant one-way ANOVA? And why wasn't it a two-way ANOVA with factors of room and sham/real TMS? More importantly, the 10% of FHR measure is arbitrary and should be replaced by standard fitting, as discussed earlier. Looking at Figure 3B, the black line appears near an asymptote while the red one is still growing toward the end, and that should be reflected in tau.

      The ANOVAs in this and other sections have been revised based on the reviewers’ suggestions; mixed ANOVAs have instead been performed and reported, yielding similar results. The fittings and 10% FHR calculations have also been extensively revised. We now show that double exponentials are better suited to our dataset, and that two-tau parameters are not informative about when performance reaches a stable point during the task.

      (24) Also, why is the exposure time on the x-axis different in Figure 3B from Figure 2B (150 vs 500)? And it would be good to see where the across-room average no-TMS data would lie here (or show the equivalent average in Figure 2B).

      We appreciate the reviewer noticing this mismatch. Figure 2B shows the time course for each environment (150 s of exposure to each), whereas Figure 3B shows all environments collapsed (150 s × 3). This is because, for the 22 listeners without TMS exposure, a Rooms main effect was observed, justifying separate time courses per room; however, for listeners exposed to sham and real TMS, no Rooms × TMS interaction was observed, so separating time courses per room was not statistically justified. As the only significant effect was TMS condition, we grouped the time spent across all environments by TMS condition.

      (25) Lines 606–617: This analysis and Figure 3C have the same issues as described for Figure 2C— asymptotic performance was not achieved for many conditions, so the 10% measure is arbitrary, as is the resulting correlation.

      We appreciate the reviewer raising these fitting issues. This part of the manuscript has been extensively revised, including new analyses and figures, although our results have not changed. Additional detail has been added to the Methods (“Speech Performance Analysis and Timecourse Fittings of Mean Cumulative Hit Rates”), and the following summary has been added to the Results (“Statistical learning of reverberant environments occurs over long and short time courses”): we fitted data with single- and double-exponential functions; double exponentials were fitted to the full time course, and single exponentials to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments, and remained preferred when compared against the truncated single-exponential fit (−420.7 vs. −399.1, preferred for roughly half the subjects). The double-exponential model was preferred for both its automated nature and the magnitude of improvement (marginal ΔAIC = 9.7 ± 1.4 when the single model won, versus substantial ΔAIC = 52.8 ± 8.9 when the double model won). We therefore used the double-exponential fit for further analysis.

      (26) Lines 619–620: Was d′ really calculated using FHR (the final value) and a non-final False Alarm Rate? This would be arbitrary. It is still unclear how HR and FAR are defined here. There is no apparent benefit to switching between HR (Figure 3B/C, presumably overall percent correct, PCC), d′ (D, E, F), and back to HR (H, I).

      We appreciate the reviewer raising this. We have clarified in the Methods (“Speech performance analysis and time-course fittings of mean cumulative hit rates”) how Hit Rate, False Alarm Rate, and d′ were calculated.

      (27) Line 655: Figure 3E should be Figure 3F.

      Corrected.

      (28) Lines 661–670: Why was Number only analyzed for initial trials, while Color was analyzed for both initial and steady trials? Also, the choice of trials 9–10 for “steady” is arbitrary and should be shown somewhere in Figure 3B.

      We analysed performance for |Number| on initial trials (1–2) only, for CP0, because performance for this speech token could only improve if positively influenced by short-term, within-trial accumulation of information (acknowledging that | Colour | precedes |Number|). To determine how much knowledge accumulated over repeated exposures — i.e., metaadaptation—we instead needed to compare performance on a speech token whose improvement could only stem from knowledge gained across previous trials, not within a single carrier phrase. We compared | Colour | performance on CP0 between initial trials (1–2) and later, steady trials (9–10). If this hypothesis is supported, it suggests that | Colour | performance for CP0 benefits from meta-adaptive information conveyed across trials as knowledge of the environment’s global structure accumulates—our proxy for meta-adaptation (Figure 3G). The choice of trials 9–10 as “steady” follows work on animal models of meta-adaptation (Robinson, Harper, & McAlpine, 2016), which described a faster adaptation rate after the eighth presentation of an environment. This has been clarified in the manuscript.

      (29) Lines 724–728: This description is confusing. The main effect of “Lecture Room” vs. “Highly Reflectant” context is that performance is very good in the Lecture Room (green line) and poor in the Highly Reflectant Room (purple). Averaging that with OPO and CP and reporting “mean difference = 20.09” (in what units?) as “overall performance” distracts from the main point. Moreover, how can that be entered into an ANOVA when the room contexts differ (LR+OPO+CP vs. HR+OPO+CP)? That ANOVA seems incorrect; it should only be performed on OPO+CP across the two contexts.

      This section has been revised and re-analysed as suggested. Redundant and unnecessary statistical comparisons were removed, retaining only those that show the effect of context on OPO and CP when comparing the different contexts in which these common environments were learned.

      (30) Lines 740–750: Again, it is not surprising that when Living Room replaces Lecture Room—and performance in Living Room is worse than in Lecture Room—the average of LiR+OPO+CP is lower than LER+OPO+CP, if OPO+CP performance is unchanged. The interesting question is whether anything changed in OPO+CP performance, as suggested for the previous point.

      This section has been revised as suggested by the reviewer.

      (31) Lines 752–771: Again, the same issue—the main effect is that performance in Anechoic trials is worse than in Lecture Room or Living Room trials, which alone explains the group difference. I am also sceptical of the finding that Anechoic performance is worse than reverberant performance, contrary to typical spatial-release-from-masking results, where reverberation degrades performance by adding noise energy at the better ear and reducing binaural benefit through decorrelation. This may be an artefact of how target and noise levels were normalized after convolution with HRTFs/BRIRs (or the use of Ambisonics); no acoustic analysis of the stimuli is provided. At minimum, the total received level at the two ears for target and masker in every environment should be reported. Brandewie & Zahorik (2013), using equivalent anechoic and reverberant conditions, never observed reverberant performance to exceed anechoic, contrary to what is stated here (lines 754–755).

      We appreciate the reviewer raising these points. The statistical analysis in this section has been revised: only the common rooms across the three-room conditions (Open-Plan Office and Car Park) were directly compared. Performance in Living Room and Lecture Room was not statistically different (Results, paragraph 4, “Statistical learning of room acoustics is tuned to universally experienced reverberation times”). However, Anechoic and Lecture Room performance remained significantly different (mean difference = 11.08, t(9) = 2.66, p = 0.013, Cohen’s d = 0.84), as did performance in the common rooms when learned in the context of Lecture Room versus other contexts (mean difference = 9.7, F(1,41) = 13.24, p < 0.001, ηp<sup>2</sup> = 0.26).

      While this setup resembles many masking studies, Brandewie & Zahorik (2013) tested four rooms simultaneously, whereas we tested three-room conditions explicitly designed to test environment-mix adaptation. We observed a synergistic relationship between performance in ‘good’ reverberant rooms (Lecture Room, Living Room) and the common but less favourable rooms (Open-Plan Office, Car Park, with longer RT60): in the absence of a ‘good-reverb anchor,’ performance in the common rooms improved less over time, possibly because participants had less to leverage in anechoic environments.

      We agree that verifying at-ear acoustic levels is critical to ruling out a normalization artefact. Our stimuli were normalized in the 41-channel sound field, not at the listener’s ears: source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses, the 41-channel noise energy was scaled to a target of 70 dB, and the 41-channel speech field was scaled to the target SNR; the 41-channel signals were then rendered to two channels using a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related acoustic effects such as head shadow.

      To verify that this did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (see supplied table, Summary Reverb Data). In the anechoic condition, the noise (positioned to the left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, raising noise level at the right ear to 55–56 dB depending on room, substantially lowering the ear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming this is not a normalization artefact but rather a genuine perceptual spatial release from masking, likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added the at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2, and updated the Discussion to clarify this mechanism.

      (32) Lines 755–759: Describing anechoic spaces as “rare” and as rooms whose “walls are treated” states the facts backwards. Open spaces (e.g., a grass lawn) are largely anechoic fields, and people spend considerable time in such environments. An anechoic room may be artificial, but an anechoic (or near-anechoic) space is very common, and the room is simply an attempt to simulate that within an enclosure.

      We appreciate this point and have rewritten this section to reflect it.

      (33) Lines 808–812: This sentence appears incorrect. It refers to “the ability to correctly report keywords spoken in environments with the more extreme—lower or higher—RIRs,” presumably meaning OPO and CP, but these are not the environments with extreme RIRs; or, if referring to An and HR, those were not “encountered in experimental blocks also containing the moderately reverberant Lecture Room or Living Room.”

      Thank you for noting this. We have rephrased this section to refer only to the extreme high-RIR environments encountered.

      (34) Line 843: Should Fig 3Ai be Fig 3A? Also, in that figure there are arrows between dlPFC and A1, and between A1 and (the cerebellum?)—it's unclear what these represent.

      The arrows were intended to represent feedforward and feedback information flow to lower auditory brain centres. We acknowledge they were confusing and have removed them from the figure.

      (35) Lines 929–949: The authors did not account for listeners tuning to voice and location (as now cited via Best et al.), and their own and Brandewie & Zahorik's data show improvement due to carrier phrase even in the anechoic case (with the inconsistency in Supplemental Table 4 noted earlier). A direct test—switching the environment between carrier phrase and target phrase, as in Brandewie & Zahorik and Vlahou et al.—would be needed to fully attribute the effect to reverberation rather than other factors.

      We thank the reviewer for this detailed comment and agree that directly manipulating the environment between carrier and target phrase would provide the most direct test of environment-specific adaptation. While our study did not implement this manipulation, our data provide converging evidence: (1) listeners showed improvement with longer carrier phrases even in the anechoic condition, consistent with previous reports, but this improvement did not interact with talker identity, carrier phrase length, or room, indicating it is not driven by tuning to specific voices or locations; and (2) regarding Supplemental Table 4, the calculations suggested by the reviewer do not match our statistical analysis—we have reported the SPSS-exported means directly (shown above) and reviewed this section for clarity, including a link to the raw supplemental data. Taken together, while we cannot fully quantify the proportion of adaptation attributable to reverberation versus other factors without the direct environment-switch manipulation, our results indicate that the observed improvements primarily relate to exposure to the environment rather than talker-specific effects.

      (36) Lines 951–953: It is unclear what about “understanding speech in background noise” distinguishes this study from previous studies of adaptation to reverberation, many of which also examined speech in noise (as reviewed in Tsironis et al., 2024). Rather than reviewing pertinent studies on adaptation to reverberation for speech tokens, the authors cite abstract noise-texture studies that are only partially relevant, given the prevalence of speech in everyday listening (lines 955–958).

      We appreciate the reviewer raising this point. Our intention was to refer specifically to statistical learning of implicit environmental acoustic features such as reverberation, rather than to speech-in-noise perception per se. We have revised the text accordingly: “A key feature of our study, which distinguishes it from previous investigations of statistical learning of acoustic features in human listeners, is the use of an ethologically valid listening task—understanding speech in background noise while listeners implicitly learn repeated acoustic features.”

      (37) Lines 978–980: In what way? For environments with large T60, a simpler explanation than “ecological validity” is that there is more late reverberant energy in the target acting as a masker, predictable from DRR.

      This section has been rewritten to clarify that it is the decline in performance at longer RT60 that is reminiscent of the decline observed under rTMS.

      (38) Line 981: What does “the better to understand speech in reverberant background noise” mean?

      This sentence has been revised.

      (39) Lines 987–988: When did “performance decline over the course of an experimental session”? Figures 2B, 3B, and 4A all show performance improving over the session.

      We have rephrased this sentence to refer to a decline in overall performance.

      (40) Lines 990–992: Many previous studies report better adaptation to reverberation for some rooms than others (e.g., Brandewie & Zahorik, 2010; Vlahou et al., 2021), but none have reported decreased performance for an anechoic space relative to a reverberant one. This anomaly should be explained and reconciled with the existing literature before invoking ecological explanations such as “ethologically relevant environments.”

      We thank the reviewer for raising this important point. We agree that verifying the at-ear acoustic levels is critical to ruling out a normalization artefact, particularly given our finding that reverberant performance exceeded anechoic performance. To address this directly: our stimuli were normalized in the 41-channel sound field, not at the listener’s ears. The source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses; the 41-channel noise field was scaled to a target of 70 dB and the 41-channel speech field scaled to the target SNR; the 41-channel signals were then rendered to two channels via a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related effects such as head shadow because normalization preceded binaural rendering.

      To verify that this sound-field normalization did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (Summary Reverb Data table). In the anechoic condition, the noise (positioned left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, increasing right-ear noise level to 55–56 dB depending on room, substantially worsening the atear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming the finding is not a normalization artefact but instead reflects a genuine perceptual spatial release from masking—likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added these at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2 to clarify this mechanism.

      (41) Discussion: Given the questions about the results, the discussion might need to be rewritten to only discuss claims that are actually supported.

      The Discussion section has indeed been extensively revised.

      (42) The hippocampus and other areas have been proposed for statistical learning, and studies also show that disruption of DLPFC can boost statistical learning (https://doi.org/10.1016/j.jml.2020.104144).

      We appreciate the reviewer raising this. We have cited Ambrus et al. (doi:10.1016/j.jml.2020.104144) in the Discussion (line 888) as evidence of opposing effects of dlPFC stimulation on statistical learning, and have further revised lines 891–903 of the Discussion to more clearly describe the known projections and functional interactions between dlPFC, hippocampus, striatum, and basal ganglia that support implicit and statistical learning.

      (43) Line 1739: What is “cumulative” here?

      “Cumulative” has been deleted.

    1. eLife Assessment

      This study presents an important study into the molecular function of AT-HOOK MOTIF NUCLEAR LOCALIZED 15 (AHL15), a member of the AHL protein family, identifying it as a potential regulator of three-dimensional gene-loop organization within transcribed gene bodies. The authors support this claim with compelling genome-wide evidence, integrating AHL15 binding profiles with transcriptional and chromatin accessibility changes, as well as demonstrating overlap with genes known to form loops across transcribed regions. The evidence supporting the claims of the authors is convincing. Collectively, these findings will be of broad interest to biologists seeking to understand the fundamental regulatory mechanisms underlying gene expression.

    2. Reviewer #1 (Public review):

      The study by Luden et al. seeks to elucidate the molecular functions of AHL15, a member of the AT-HOOK MOTIF NUCLEAR LOCALIZED (AHL) protein family, whose overexpression has been shown to extend plant longevity in Arabidopsis. To address this question, the authors conducted genome-wide ChIP-sequencing analyses to identify AHL15 binding sites. They further integrated these data with RNA-sequencing and ATAC-sequencing analyses to compare directly bound AHL15 targets with genes exhibiting altered expression and chromatin accessibility upon ectopic AHL15 overexpression.

      The analyses indicate that AHL15 preferentially associates with regions near transcription start sites (TSS) and transcription end sites (TES). Notably, no clear consensus DNA-binding motif was identified, suggesting that AHL15 binding may be mediated through interactions with other regulatory factors rather than through direct sequence recognition. The authors further show that AHL15 predominantly represses its direct target genes; however, this repression appears to be largely independent of detectable changes in chromatin accessibility.

      In addition to the AHL protein family, the globular H1 domain-containing high-mobility group A (GH1-HMGA) protein family also harbors AT-hook DNA-binding domains. Recent studies have shown that GH1-HMGA proteins repress FLC, a key regulator of flowering time, by interfering with gene-loop formation. The observed enrichment of AHL15 at both TSS and TES regions, therefore, raises the intriguing possibility that AHL15 may also participate in regulating gene-loop architecture. Consistent with this idea, the authors report that several direct AHL15 target genes are known to form gene loops.

      Overall, the conclusions of this study are well supported by the presented data and provide new mechanistic insights into how AHL family proteins may regulate gene expression.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Luden et al. investigates the molecular function and DNA-binding modes of AHL15, a transcription factor with pleiotropic effects on plant development. The results contribute to our understanding of AHL15 function in development specifically and transcriptional regulation in plants more broadly.

      Strengths:

      The authors developed a set of genetic tools for high-resolution profiling of AHL15 DNA binding and provide exploratory analyses of chromatin accessibility changes upon AHL15 overexpression. The generated data (CHiP-Seq, ATAC-Seq and RNA-Seq is a valuable resource for further studies. The data suggest that AHL15 does not operate as a pioneer TF, but is likely involved in gene looping.

      Weaknesses:

      The authors have extended the motif analysis to the top 1,000 shared peaks, addressing part of my previous concern. It did not erase my worries about overclaiming completely, but I think it can be considered as a terminological issue, rather than technical one. Therefore, it could be addressed through minor revision, without further analyses.

      Specifically, the absence of a predominant enriched motif does not establish that AHL15 binds non-specifically or lacks sequence preferences. Low motif prevalence limits the proportion of peaks it could explain but does not exclude a genuine binding preference; a binding motif need not be rare in the genomic background (binding can be supported by other factors). Please revise the interpretation in lines 188-194, and the conclusion in lines 204-206 to state that the present analysis did not identify a strongly enriched motif accounting for a substantial proportion of AHL15-associated regions. Any equivalent claims of non-specific binding elsewhere in the manuscript should be toned down.

      Additional minor points:

      (1) Please provide the exact HOMER background-selection procedure, and definition of the 50-bp search windows.

      (2) Figures 2B-C and the related supplementary figures, add the x axis label.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigated the role of AHL15 in regulation of gene expression using AHL15 overexpression lines. Their results do show that more gene are downregulated when AHL15 is upregulated and its binding is not affecting the chromatin accessibility. Further, they investigated AHL15 binds in regions depleted in histone modifications and other epigenetic signatures. Subsequently, they investigated the presence of AHL15 in the gene chromatin loops. They found overlaps with both upregulated and downregulated genes. The methods are appropriately described, but could be improved to include the analysis of self-looping gene boundaries.

      Strengths:

      Their study clearly showed lack of any specific sequence enrichment in the AHL15 binding sites, other than these being AT-rich, suggesting that AHL proteins do not recognize a specific DNA sequence but are recruited to their AT-rich target sites in another way. The study does suggest significant enrichment of AHL15 binding sites at TSS and TES, and AHL15 sites are depleted of any histone marks. They also identified that AHL15 binding sites overlap with self-looping gene boundaries. The authors have addressed the comments raised in the revised manuscript.

      Comments on revised version.

      The authors have addressed the comments raised, in the revised manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study by Luden et al. seeks to elucidate the molecular functions of AHL15, a member of the AT-HOOK MOTIF NUCLEAR LOCALIZED (AHL) protein family, whose overexpression has been shown to extend plant longevity in Arabidopsis. To address this question, the authors conducted genome-wide ChIP-sequencing analyses to identify AHL15 binding sites. They further integrated these data with RNA-sequencing and ATAC-sequencing analyses to compare directly bound AHL15 targets with genes exhibiting altered expression and chromatin accessibility upon ectopic AHL15 overexpression.

      The analyses indicate that AHL15 preferentially associates with regions near transcription start sites (TSS) and transcription end sites (TES). Notably, no clear consensus DNA-binding motif was identified, suggesting that AHL15 binding may be mediated through interactions with other regulatory factors rather than through direct sequence recognition. The authors further show that AHL15 predominantly represses its direct target genes; however, this repression appears to be largely independent of detectable changes in chromatin accessibility.

      In addition to the AHL protein family, the globular H1 domain-containing high-mobility group A (GH1-HMGA) protein family also harbors AT-hook DNA-binding domains. Recent studies have shown that GH1-HMGA proteins repress FLC, a key regulator of flowering time, by interfering with gene-loop formation. The observed enrichment of AHL15 at both TSS and TES regions, therefore, raises the intriguing possibility that AHL15 may also participate in regulating gene-loop architecture. Consistent with this idea, the authors report that several direct AHL15 target genes are known to form gene loops.

      Overall, the conclusions of this study are well supported by the presented data and provide new mechanistic insights into how AHL family proteins may regulate gene expression.

      However, it is important to note that the genome-wide analyses in this study rely predominantly on ectopic overexpression of AHL15 at developmental stages when the gene is not usually expressed. Moreover, loss-of-function phenotypes for AHL15 have not been reported, leaving unresolved whether AHL15 plays a physiological role in regulating plant longevity under native conditions. It therefore remains possible that longevity control is mediated by other AHL family members rather than by AHL15 itself. In this regard, the manuscript's title would benefit from more accurately reflecting this broader implication.

      The ahl15 loss-of-function phenotype has previously been described in Karami et al., 2020 (Nat. Plants), Rahimi et al., 2022a (New Phyt.), and Rahimi et al., 2022b (Curr. Biol.), showing that ahl15 loss-of-function among others results in accelerated vegetative phase change and flowering, a reduced number of leaves produced by axillary meristems in short day grown plants and reduced secondary growth in the inflorescence stem. The dominant-negative ahl15 delta-G allele, expressing a mutant protein lacking the conserved G motif in the PPC domain, shows these phenotypes more clearly in the heterozygous ahl15 +/- background, and is embryo lethal in the homozygous ahl15 background (Karami et al., 2021, Nature Comm.). In addition, we recently show that leaf senescence is significantly accelerated in the ahl15 loss-of-function mutant (Luden et al., 2025, BioRxiv). These results show that AHL15 is involved in several aspects of ageing in Arabidopsis, and we have adjusted the introduction to discuss these previous findings more explicitlyWe agree with reviewer 1 on the possibility that multiple AHLs could have an effect on longevity, which is partially supported by the delayed flowering time observed in the AHL20, AHL27, or AHL29 overexpression lines (Karami et al., 2020, Street et al., 2008). However, the induction of the AHL15-GR fusion alone by DEX shows a clear delay of developmental phase transitions and the aging process in general, indicating that AHL15 by itself is able to extend longevity as other AHLs are not affected by DEX treatment (proven by the fact that their expression is not significantly changed in our RNA-seq analysis of DEX-treated 35S:AHL15-GR seedlings).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Luden et al. investigates the molecular function and DNA-binding modes of AHL15, a transcription factor with pleiotropic effects on plant development. The results contribute to our understanding of AHL15 function in development, specifically, and transcriptional regulation in plants, more broadly.

      Strengths:

      The authors developed a set of genetic tools for high-resolution profiling of AHL15 DNA binding and provided exploratory analyses of chromatin accessibility changes upon AHL15 overexpression. The generated data (CHiP-Seq, ATAC-Seq and RNA-Seq is a valuable resource for further studies. The data suggest that AHL15 does not operate as a pioneer TF, but is likely involved in gene looping.

      Weaknesses:

      While the overall message is conveyed clearly and convincingly, I see one major issue concerning motif discovery and interpretation. The authors state that because HOMER detected highly enriched motifs at frequencies below 1%, they conclude that "a true DNA binding motif would be present in a large portion of the AHL15 peaks (targets) and would be rare in other regions of the genome (background)."

      I agree that the frequency below 1% is unexpectedly low; however, this more likely reflects problems in data preprocessing or motif discovery rather than intrinsic biological properties of the transcriptional factor that possesses a DNA-binding domain and is known to bind AT_rich motifs. As it is, Figure 2 cannot serve as a main figure in the manuscript: it rather suggests that the generated CHiP-Seq peakset is dominated by noise (or motif discovery was done improperly) than that AHL15 binds nonspecifically.

      Since key methodological details on the HOMER workflow are missing in the M&M section, it is not possible to determine what went wrong. Looking at other results, i.e. the reasonably structured peak distribution around TSS/TTS and consistent overlap of the peaks between the replicas, I assume that the motif discovery step was done improperly.

      Therefore, I recommend redoing the motif analysis, for example, by restricting the search to the top-ranked peaks (e.g. TOP1000) and by using an appropriate background set (HOMER can generate good backgrounds, but it was not documented in the manuscript how the authors did it). If HOMER remains unsuccessful, the authors should consider complementary methods such as STREME or MEME, similar to the approach used for GH1-HMGA (https://pmc.ncbi.nlm.nih.gov/). If the peakset is of good quality, I would expect the analysis to identify an AT-rich motif with a frequency substantially higher than 1%-more likely in the range of at least 30%. If such a motif is detected, it should be reported clearly, ideally with positional enrichment information relative to TSS or TTS. It would also be informative to compare the recovered motif with known GH1-HMGA motifs.

      If de novo motif discovery remains inconclusive, the authors should, at a minimum, assess enrichment of known AHL binding motifs using available PWMs (e.g. from JASPAR). As it stands, the claim that "our ChIP-seq data show that AHL15 binds to AT-rich DNA throughout the Arabidopsis genome with limited sequence specificity (Figure 2A, Figure S2-S4)" is not convincingly supported.

      Another point concerns the authors' hypothesis regarding the role of AHL15 in gene looping. While I like this hypothesis and it is good to discuss it in the discussion section, the data presented are not sufficient to support the claim, stated in the abstract, that AHL15 "regulates 3D genome organization," as such a conclusion would require additional, dedicated experiments.

      The motifs discovered by HOMER are ranked by their enrichment over background, of which the highest-scoring motifs are very rare in the AHL15-bound targets, but even rarer in the background, which is why they score highly on the percent enrichment score. As expected by reviewer 2, we identified AT-rich motifs that were present in a larger percentage of AHL15 targets (found in 3-18% of targets, depending on the motif, see for example motif #5 in figure S4A), which can be seen at the right tail of the histograms shown in figures 2B-C and figures S2-S4 B-C. However, these motifs were also common in the background and were therefore not considered as significantly enriched in the AHL15-bound regions, with a target:background ratio of <2. As most of these motifs were flagged by HOMER as possible false-positives, and to limit the size of the (supplemental) figures, we did not show each of the motifs identified by HOMER in table form, but the full tables of de novo motifs identified by HOMER, including possible false-positive results are included in Additional file 3.

      Although the identification of AT-rich motifs shows that AHL15 (and very likely most other AHL proteins as well) binds AT-rich regions, it does not sufficiently explain the binding of AHL15 to its target genes, as these motifs are found at almost equal frequencies in non-AHL15-bound regions. In addition, a sequence found at this frequency in the genomic background is, in our view, too unspecific to be considered as a transcription factor binding site. Based on this, we concluded that AHL15 lacks a specific binding motif that can define the genes it binds.

      We have updated the methods section to include more details on the HOMER analysis and have also run the analysis in the top1000 shared peaks as suggested by reviewer 2 for both AHL15 and AHL29 ChIP-seq peaks, which showed that unlike in AHL29, a clear AT-rich motif cannot be found for AHL15 (Additional file 1: Figure S5).

      Reviewer #3 (Public review):

      Summary:

      This study investigated the role of AHL15 in the regulation of gene expression using AHL15 overexpression lines. Their results do show that more genes are downregulated when AHL15 is upregulated, and its binding does not affect the chromatin accessibility. Further, they investigated AHL15 binds in regions depleted in histone modifications and other epigenetic signatures. Subsequently, they investigated the presence of AHL15 in the gene chromatin loops. They found overlaps with both upregulated and downregulated genes. The methods are appropriately described, but could be improved to include the analysis of self-looping gene boundaries.

      Strengths:

      Their study clearly showed a lack of any specific sequence enrichment in the AHL15 binding sites, other than these being AT-rich, suggesting that AHL proteins do not recognize a specific DNA sequence but are recruited to their AT-rich target sites in another way. The study does suggest significant enrichment of AHL15 binding sites at TSS and TES, and AHL15 sites are depleted of any histone marks. They also identified that AHL15 binding sites overlap with self-looping gene boundaries.

      Weaknesses:

      The claim that AHL15 acts as a repressor and genes regulated by it are downregulated needs to be investigated based on AHL15 binding sites, to show enrichment/ depletion of AHL15 binding sites in overexpressing genes and repressed genes. The authors should provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title. Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression may be helpful to understand the significance of the study. Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful. It is not clear how the overlap of AHL15 peaks with self-looping genes has been carried out.

      A metagenome plot of AHL15 binding around genes that are differentially expressed upon DEX treatment can be found in Figure 3F. This analysis shows that AHL15 binding near differentially expressed genes is more pronounced compared to all AHL15-bound genes, and that AHL15 binding near the TSS is especially enriched for upregulated genes.

      As also suggested by reviewer 2, we ran a motif enrichment analysis on the differentially expressed genes that are bound by AHL15 to see if any motifs are enriched compared to the background and overrepresented in the AHL15-bound genes. Again, this did not reveal an AT-rich motif nor a motif that was conserved between up- and downregulated AHL15-bound genes (Additional file 1: Figure S6).

      Plant longevity in 35S:AHL15-GR Arabidopsis plants treated with DEX has been reported previously.. DEX treatment extended vegetative development after flowering resulting in polycarpy (Karami et al., 2020, Nature Plants), enhanced secondary growth resulting in woody stems (Rahimi et al., 2022, Current Biol.) and recently we showed that it delays leaf senescence in Arabidopsis (Luden et al., 2025, bioRxiv). All these observations have now been incorporated in the results section where the p35S::AHL15-GR plants are first presented. In addition, we show that 35S:AHL15-GR plants treated a single time with DEX at 10 days after germination show a significantly delayed flowering time in figure 4C-D of this manuscript.

      The enrichment of AHL15 ChIP-seq peaks in self-looping genes will be analyzed as suggested and compared to a random set of genes as a control, and the methods section will be updated to clarify how the analyses on self-looping genes were carried out.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors need to correct Line 28 on page 10:

      By comparing the AHL15 ChIP-seq data with 1766 previously identified self-looping genes -> By comparing the AHL15 ChIP-seq data with 1792 previously identified self-looping genes (Liu et al., 2016)

      Note: 1792 genes with self-loops were first reported by Liu et al. (2016)

      Liu, C., Wang, C.M., Wang, G., Becker, C., Zaidem, M., and Weigel, D. (2016). Genome-wide analysis of chromatin packing in at single-gene resolution. Genome Res 26, 1057-1068.

      These numbers will be corrected in the text.

      Reviewer #2 (Recommendations for the authors):

      (1) The newly generated datasets have been deposited only as raw sequencing reads. For reusability and reproducibility, the authors should also provide processed data accompanied by detailed metadata.

      We will upload the processed data and corresponding metadata to GEO.

      (2) Figures 4B and 5E are of low quality; can they be improved?

      We have submitted the original high-quality images to the publisher, which should resolve the issue.

      (3) Supplementary Table 5 lacks description (columns do not have names).

      This has been fixed.

      Reviewer #3 (Recommendations for the authors):

      Suggestions:

      (1) Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful.

      This analysis will be performed and included in the revised manuscript as Additional file 1: Figure S6.

      (2) Investigate AHL15 binding sites to show enrichment/ depletion of AHL15 binding sites in overexpressing genes than repressed genes.

      This analysis has been done, please see figure 3F.

      (3) Provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title.

      For the effect of AHL15-GR induction by DEX on vegetative phase change and flowering time, please see figure 4C-D and Rahimi et al., (2022, New Phytologist). For other phenotypic changes induced by DEX treatment of 35S:AHL15-GR plants, please see Karami et al. (2020; Nature Plants), Rahimi et al., (2022, Current Biology) and Luden et al., (2025; BioRxiv). Text has been added to the results section where the 35S:AHL15-GR line is first introduced to refer to these previous publications.

      (4) Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression, may be helpful to understand the significance of the study.

      This analysis has been done and included in the revised manuscript as Additional file 1: Table S1.

      (5) Describe how the overlap of AHL15 peaks with self-looping genes has been carried out.

      The methods section has been updated with detailed information on this analysis.

    1. eLife Assessment

      This important study reveals the roles of two lytic transglycosylases in the progression of spore formation in the research model species Myxococcus xanthus. Solid evidence is provided for the roles of these two enzymes in spore formation and for interplay between their functions and peptidoglycan synthetic systems. These findings may have broader implications for studies of peptidoglycan metabolism across a range of species.

    2. Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to all points of the previous reviews.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to most points of the previous review.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted in the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. The conclusions concerning the PG degradation during sporulation needs to be clarified, as described below. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation". This needs to be reflected also in title of the paper.

      We have addressed the reviewer’s concerns and updated the text. 

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 100-125: I am still concerned about clarity in the description of the muropeptides from spores. The claim is that 90% of the recovered muropeptides are anhydro LTG products. If LTGs had truly cleaved so many of the glycosidic bonds such that 90% of the muramic acid was now in the anhydro form, then all of the PG in the spores would be in small fragments (dimers and trimers containing 4-6 sugars and 8-12 amino acids), and these would be soluble and lost during purification. Whatever the form of the PG, it must not be easily soluble, so it must be either larger than that or bound to something else. The fact that the muropeptides are solubilized by muramidase indicates that there are some NAM-NAG bonds remaining, and cleavage of these should release some non-anhydro muropeptides. The spore muropeptide chromatograms have a large, late "mound" of UV-absorbing material that was released by muramidase digestion, and two of the identified anhydro products are present in this mound. It is not clear how these two anhydro products were quantified within this mound and what other (presumably) muropeptide species might be present in this mound. Some explanation of how these two muropeptides were quantified is needed.

      We thank the reviewer for this careful point. As the reviewer notes, our purification procedure recovers only sedimentable PG, and any PG fragments solubilized by LTG activity would be lost during the washing steps and therefore not represented in our analysis. We have now explicitly acknowledged this limitation and discussed its implications in the revised text:

      “Consequently, any PG fragments solubilized by LTG activity during sporulation would be lost at this stage, and the muropeptides we detect derive from the more highly crosslinked material that survives the procedure.

      “Despite the overall decline in identified muropeptides, anhydro-muropeptides a minor component of vegetative PG were enriched in both spore types (Figure 1).”

      I feel that the authors need to address some of this uncertainty in the results and discussion. They might be able to say that anhydro-muropeptides represent 90% of the "identified muropeptides" but need to acknowledge that there is a great deal of unidentified material released by the muramidase digestion. The data might also indicate that the LTG activity solubilizes much of the PG, which is lost, and results in recovery of only highly cross-linked muropeptides that might survive the PG purification process.

      We agree that two of the identified anhydro species elute within a broad, unresolved region of the chromatogram that accounts for a substantial fraction of the muramidase-released material. Because reliable assignment and quantification of individual components within this region is not possible, we concluded that expressing anhydro-muropeptides as a fraction of the total identified muropeptides could be misleading. We have therefore removed the quantitative statement and now describe the enrichment of anhydromuropeptides in spores relative to vegetative cells qualitatively:

      “The identified muropeptides, however, represent only a part of the material released by muramidase from spores: a substantial, as a late-eluting portion of the chromatogram could not be assigned, and its composition remains unknown.”

      This would not eliminate the conclusion that "their abundance in spores indicates that certain LTGs must play essential roles (perhaps change to "might play important roles") in M. xanthus sporulation", which leads to the remaining studies in the paper.

      Following the reviewer's suggestion, we have softened the conclusion of this section to state that LTGs "may play important roles" in sporulation.

      (2) The authors have changed the statement about a "new mechanism of sporulation" at the beginning of the discussion, but this language is still in the paper title. Something more like "Programmed peptidoglycan degradation plays an important role in Myxococccus sporulation"

      Following the reviewer's suggestion, we changed the title to “A novel mechanism for morphological change during bacterial sporulation based on programmed peptidoglycan degradation”.

    1. eLife Assessment

      This valuable study provides convincing evidence for deficits in aversive taste learning and taste coding in a mouse model of autism spectrum disorders. Specifically, the authors found that Shank3 knockout mice exhibit behavioral deficits in learning and extinction of conditioned taste aversion, and calcium imaging of the gustatory cortex identified impaired neuronal responses to taste stimuli. This paper will likely be of interest to researchers studying how learning and sensory processes are affected by genetic causes of autism spectrum disorders.

    2. Reviewer #1 (Public review):

      Summary:

      The study from Wu and Turrigiano investigates how disruption of taste coding in a mouse model of autism spectrum disorders (ASDs) affects aversive learning in the context of a conditioned taste aversion (CTA) paradigm. The experiments combine 2photon calcium imaging of neurons in the gustatory portion of the anterior insular cortex (i.e., gustatory cortex) with behavioral training and testing. The authors rely on Shank3 knockout mice as a model for ASDs. The authors found that Shank3 mice learn CTA more slowly and extinguish the memory more rapidly than control subjects. Calcium imaging identified impairments in taste evoked activity associated with memory encoding and extinction. During memory encoding, the authors found less suppressed neuronal activity and increased correlated variability in Shank3 mice compared to control. During extinction, they observed a faster loss of taste selectivity and degradation of taste discriminability in mutants compared to controls.

      Strengths:

      This is a well-written manuscript that presents interesting findings. The results on the learning and extinction deficits in Shank3 mice are of particular interest. Analyses of neural activity are well conducted and provide important information on the type of impaired cortical activity that may correlate with behavioral deficits.

      Weaknesses:

      The authors did an excellent job addressing the weaknesses highlighted in my first assessment.

    3. Reviewer #2 (Public review):

      Summary

      Wu and Turrigiano investigate how Shank3 loss affects experience-dependent changes in sensory representations during conditioned taste aversion learning and extinction. Using longitudinal two-photon calcium imaging in the gustatory cortex, the authors show that Shank3 knockout mice acquire taste aversion more slowly but, after additional conditioning, reach an aversion comparable to wild-type mice; this learned aversion then extinguishes more rapidly. At the neural level, knockout mice exhibit reduced stimulus-evoked suppression and increased correlated variability; while learning and extinction are accompanied by changes in the reliability, selectivity and population-level discriminability of taste representations. The revised manuscript more clearly distinguishes baseline genotype-dependent differences in cortical activity from learning-associated changes and appropriately frames the relationship between neural activity and behaviour as associative rather than causal.

      Strengths

      A major strength of the study is the combination of longitudinal cellular-resolution imaging with a behavioural paradigm that allows cortical population activity to be followed across acquisition, retrieval and extinction. This provides a rich description of how sensory representations evolve as learned value changes, and how these dynamics differ following Shank3 deletion. The observation that knockout mice eventually acquire a robust aversion but subsequently extinguish it more rapidly is particularly useful because it separates impaired acquisition from subsequent instability of the learned association.

      The revised manuscript has substantially addressed several concerns raised in the original review. Importantly, the authors now show that reduced stimulus-evoked suppression is already evident during pre-learning habituation and that increased coactivity therefore appears to reflect, at least in part, a pre-existing network property rather than a consequence of learning. They additionally report that the level of coactivity at the beginning of conditioning correlates with subsequent behavioural acquisition, providing a useful link between individual variation in cortical activity and learning performance. This analysis strengthens the association between cortical network state and behaviour without establishing causality.

      Another useful addition concerns the potential contribution of licking behaviour to taste decoding. Because sampling behaviour necessarily differs as animals acquire an aversion, separating sensory representations from movement-related activity is difficult in this paradigm. The authors now perform decoding during the ten-second post-sampling period and find above-chance decoding after licking has ceased, making it less likely that differences in licking alone explain the principal population-decoding results.

      The manuscript is also clearer in its anatomical and conceptual terminology. The recordings are now appropriately described as being from gustatory cortex rather than implying coverage of the broader anterior insular cortex, and the conclusions have been restricted primarily to conditioned taste aversion rather than generalised to cognitive flexibility more broadly. The latter is particularly important because whether these findings generalise to reversal learning, probabilistic learning, or other forms of adaptive behaviour remains unknown.

      Weaknesses

      The principal remaining limitation is mechanistic. The experiments establish a robust association between Shank3 deletion, altered cortical activity and altered learning dynamics, but they do not establish the causal relationships among these observations. The new correlation between early coactivity and subsequent learning is informative, but manipulating the relevant network property would ultimately be required to determine whether increased correlated variability contributes directly to slower acquisition. The authors now acknowledge this limitation and have appropriately removed language implying causality.

      Similarly, the cellular or circuit origin of the altered correlated variability remains unresolved. Reduced inhibition is an interesting potential explanation, but the authors do not directly measure interneuron function or inhibitory transmission here. Consequently, the proposed relationship between Shank3 loss, altered inhibition, increased correlated activity, and impaired sensory encoding should remain a hypothesis emerging from the results rather than a demonstrated mechanism.

      A further limitation is the absence of a full conditioned-stimulus-only Shank3 knockout control group. The newly analysed habituation recordings partly address this issue by demonstrating reduced suppression and a tendency toward increased coactivity before learning, and the cross-session decoding analysis suggests that naïve knockout cortical populations can nevertheless distinguish water from saccharin. These analyses considerably improve interpretation of the existing experiment, although they are not equivalent to longitudinal comparison with a knockout control group undergoing the complete protocol without aversive conditioning.

      Finally, the interpretation of population activity in terms of taste identity, learned value and their interaction remains necessarily limited by the task design. The results clearly demonstrate experience-dependent changes in population discriminability, but because taste identity, learned value, and sampling behaviour covary during conditioned taste aversion and extinction, the present experiments cannot fully separate the precise information represented by these population changes.

      Overall assessment

      The revision has addressed the major interpretational concerns raised in the previous review and has strengthened the manuscript through several useful additional analyses. In particular, distinguishing pre-existing cortical abnormalities from learning-associated changes, relating early coactivity to subsequent behaviour, controlling more carefully for licking-related activity, and removing causal language better align the conclusions with the evidence. The study therefore provides solid evidence for altered experience-dependent sensory coding and learning dynamics following Shank3 loss, while the mechanisms connecting these phenomena remain an important question for future work.

    4. Reviewer #3 (Public review):

      In this study Wu & Turrigiano investigate an ethologically relevant form of associative learning (conditioned taste aversion-CTA) and its extinction in the Shank3 KO mouse model of ASD. They also examine the underlying circuits in anterior insular cortex (AIC) simultaneously, using two-photon calcium imaging through GRIN lens. They report that Shank3 KO mice learn CTA slower and suggest that this is mediated by a reduction in tastant-stimulus activity suppression of AIC neurons and reduced signal-to-noise ratio due to increased noise correlations in AIC neurons. Interestingly, once Shank3 KO mice do acquire CTA, they extinguish the aversive memory more rapidly than wild-type. This accelerated extinction is accompanied by a faster loss of neuronal and population-level taste selectivity and coding in the AIC compared to WT mice.

      This is an important study that uses in vivo methods to assess circuit dysfunction in a mouse model of ASD, related to sensory perception valence (in this case taste). The study is well executed, the data are of high quality, and the analyses procedures are detailed. Furthermore, the behavioural paradigm is well thought, particularly the approach for assessing extinction through repeated retrieval sessions (T1-T5), which effectively tests discrimination between saccharin and water rather than relying solely on lick counts or total consumption as a measure of extinction. Finally, the statistical tests used are appropriate and justified.

      Comments on revised version.

      The authors have addressed all comments satisfactorily.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The experiments rely on three groups: CS-only WT, CTA WT, and CTA KO. Can the authors provide a rationale for not having a CS-only KO group?

      We did not include the CS-only (KO) group because longitudinal in vivo recordings in behaving animals are technically demanding, and our primary goal was to follow and compare the dynamics across learning in WT and Shank3 KO mice. However, we recorded an additional habituation water session one day before CST1, which allows us to address (1) the decoding performance in naïve KO animals, and (2) whether higher correlated noise is already present in KO animals before learning. To this end, we first trained and tested the classifier on cross-registered data from the habituation (HAB) water session and the CST1 saccharin session across the CSonly (WT), CTA (WT), and CTA (KO) groups. Because this comparison is made across sessions, it is not exactly comparable to discrimination within a session, but it does show that GC responses in naïve KO animals can discriminate water from saccharin at baseline. These data are now included in Figure 6 – figure supplement 2 and are described in lines 329-338.

      Interestingly, we also observed a significant reduction in the amplitude of suppressed responses during the habituation water session in KO compared to WT animals, as well as a trend toward higher stimulus-evoked neuronal coactivity (data included in Figure 2 – figure supplement 2). This suggests that reduced suppression is already present in GC of KO animals before learning, potentially contributing to slower CTA acquisition (now described in lines 209-215).

      (2) The authors design an effective behavioral paradigm comparing consumption of water and saccharin and tracking extinction (Figure 3). This paradigm shows differences in licking across distinct behavioral conditions. For instance, during T1, licking to water strongly differs from licking to saccharin for both WT and KO. During T2, licking to water strongly differs from licking to saccharin only for WT (much less for KO), and licking to saccharin in WT differs from that in KO. These differences in taste sampling across conditions could contribute to some of the effects on neural activity and discriminability reported in Figures 5 and 6. That is, sucrose and water trials may be highly discriminable because in one case the mouse licks and in the other it does not (or licks much less). The author may want to address this issue.

      This is an important point. As noted by the reviewer, active licking can modulate neuronal activity in the gustatory cortex independent of taste identity (Neese et al., 2022). Because our paradigm required animals to voluntarily sample tastants of different valences, motivated differences in licking are inherently tied to taste value, making it difficult to fully disentangle taste-evoked responses from lick-related activity.

      However, taste exposure is known to induce prolonged neural responses that persist beyond the sampling phase (Juen et al., 2024). In our recordings, we included a 10second post-sampling epoch. We thus trained and tested classifiers using calcium traces during this post-sampling period; in particular, we divided the 10-second duration into five 2-second bins, matching the length of the tastant delivery phase, and analyzed the classifier built within each bin. We found that in both WT and KO animals, decoding performance was consistently above chance throughout the postdelivery period (now added to Figure 6 – figure supplement 1, lines 321-329), suggesting that decoding accuracies during sampling likely reflect taste rather than licking.

      (3) Are there any omission trials following CTA? If so, they should be quantified and reported. How are the omission trials treated with regard to the analyses?

      On the day following each CST session, animals underwent a water-only session (i.e., saccharin was omitted) to minimize context–malaise association. During these sessions, animals resumed licking both in the total lick counts and in the number of trials they engaged in, to levels comparable to pre-conditioning behavior. We did not observe significant differences between the WT and KO groups during these omission sessions. This point has been mentioned in the Methods section of the revised manuscript (lines 544-547).

      (4) The authors describe the extinction paradigm as "alternative choice". In decision-making, alternative choice paradigms typically require 2 lateral spouts to report decisions following the sampling from a central spout. To avoid confusion, the authors may want to define their paradigm as alternative sampling.

      We have revised this terminology to “alternative sampling” to avoid confusion with the classical alternative-choice paradigms.

      (5) Figure 4 reports that CTA increases the proportion of neurons that consistently respond to saccharin and water across days. While the saccharin result could be an effect of aversive learning, it is less clear why the phenomenon would generalize to water as well. Can the authors provide an explanation?

      Water and saccharin activated an overlapping population of neurons in GC. When we further quantified their tuning properties in the lifetime plots (Figure 4), we found that neurons responsive to both stimuli showed the most stable responsiveness across days, compared to neurons that responded only to saccharin or only to water (Author response image 1). Because the water-responsive and saccharin-responsive groups in Figure 4 both include this subset of dual-responsive neurons, this likely explains why both plots show increased reliability. This effect on reliability of single-cell responses is thus distinct from changes in the ability to discriminate between tastants at the population level (Fig. 6).

      Author response image 1.

      GC neurons responding to both water and saccharin are more stable during CTA extinction. Lifetime plot showing significant responses of the same neurons responding to only water (blue), only saccharin (magenta), and to both saccharin and water (gold) across test sessions (T1-5) in the CTA (WT) group.

      (6) The recordings are performed in the part of the anterior insular cortex that is typically defined as "gustatory cortex" (GC). Given the functional heterogeneity of the anterior insular cortex (AIC) and given that the authors do not sample all of the anteroposterior extent of AIC, I would suggest being more explicit about their positioning in GC. Also, some citations (e.g., Gogolla et al, 2014) refer to the posterior insular cortex, which is considered more inherently multimodal than GC. GC multimodality is typically associative in nature, as only a few neurons respond to sound and light in naïve animals.

      Our stereotaxic coordinates targeted the conventional gustatory region within AIC (see revised manuscript Methods section, lines 489-490). We have revised the terminology throughout the manuscript to more explicitly reflect this anatomical positioning.

      (7) It would be useful to add summary figures showing the extent of viral spread as well as GRIN lens placement.

      Revised Manuscript Figure 1B shows a representative example of confirmed GRIN lens placement and the viral spread of GCaMP. In most cases, GCaMP expression is confined to GC, with minimal spread to the piriform cortex and along the injection track. Depth and GCaMP expression in GC were further validated during two-photon imaging.

      (8) I encourage the authors to add Ns every time percentages are reported. How many neurons have been recorded in each condition? Can the authors provide the average number of neurons recorded per session and per animal?

      We now included these numbers in the revised manuscript (lines 157-158, 162-163, 253, 268, 271-272).

      (9) It looks like some animals learned more than others (Figure 1E or Figure 3C). Is it possible to compare neural activity across animals that showed different degrees of learning?

      We thank the reviewer for this suggestion – we now show a significant correlation between the magnitude of CTA and the coactivity metric in Figure 1 Figure supplement 3; we elaborate in our Response to Reviewer #3 Public Review 1.

      Reviewer #2 (Public review):

      (1) Causality: The paper infers that increased correlated variability causes learning deficits, but no causal tests (e.g., optogenetic modulation of inhibition or interneuron rescue) are presented to confirm this.

      Although we now provide data showing that correlated variability prior to learning is significantly correlated with the magnitude of CTA (see Response to Reviewer #1 Public Review 1above), we agree that we cannot infer causality without additional manipulations. While it might be possible to manipulate correlated variability by targeting inhibition within GC, optogenetic and chemogenetic manipulations of inhibition are likely to impact behavior through multiple mechanisms; for example, enhancing PV-interneuron activity in visual cortex profoundly impairs vision-dependent learning (Bissen et al. 2026, Leman et al. 2025). Thus, testing this would require finding a paradigm that specifically restores synchronization to WT levels without over-inhibiting the network, which is beyond the scope of the current study. We have rewritten the manuscript throughout to remove the inference of causality, and instead describe these two findings as being “associated” (see e.g. lines 95, 219-220, 379-382).

      (2) Behavioural scope: The study focuses exclusively on taste aversion; generalisation to other flexible learning paradigms (e.g., reversal or probabilistic tasks) is not addressed.

      Our study is focused on conditioned taste aversion (CTA) acquisition and extinction, which provides a well-established model for examining the formation and updating of aversive associative memories. We agree that cognitive flexibility encompasses a broad range of behavioral paradigms, and in the revised manuscript have sought to confined our conclusions to CTA. Whether the mechanisms identified here extend to other forms of flexible learning, such as reversal or probabilistic learning, or even to other sensory-stimulus-guided behavior, will require future investigation.

      (3) Mechanistic insights: While providing interesting findings of altered sensory perception and extinction of learning-related signals in AIC, it offered nearly no mechanistic insights. This makes the interpretation, especially on how generalisable these findings are, difficult. Also, different reported findings are "potentially" connected, but the exact relation between increased correlated variability and faster loss of taste selectivity cannot be assessed.

      In a new analysis we find that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3). This new piece of data (now added to revised manuscript lines 215218) provides a link between these two findings, and suggests that baseline coactivity levels in GC can influence the speed of CTA learning. We agree that this association does not imply causation, and have taken pains to avoid stating this.

      Reviewer #3 (Public review):

      (1) The authors don't make a causal link between the behaviour and AIC neurophysiology, both the percentage of suppressed cells and the coactivity measurements. For the % of suppressed cells, it seems that both WT and KO cells are suppressed in the transition between CST1 and CST2 (Figure 1L), yet only the WT mice exhibit CTA (at least by CST2). For the taste-elicited coactivity measure, it seems that there is an increase in coactivity from CST1 to CST2 in WT (Figure 2C - blue, although not statistically tested?), but persistently higher coactivity in KO. Is this change of coactivity in WT important for the expression of CTA? Plotting behavioral performance (from Figure 1G) against coactivity (from Figure 2C) for each animal would be informative.

      This is a good suggestion (also made by the other reviewers), and we now show that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3).

      (2) Shank3 KO cells already show an increase in baseline coactivity (Figure 2- figure supplement 1), and the authors never examine CS-only responses in the KO group, therefore making it difficult to determine whether elevated coactivity and noise correlations reflect a generalized AIC abnormality in Shank3 KOs (perhaps through impaired PV-mediated inhibition in insular cortex - Gogolla et al, 2014) that is not directly responsible/related to CTA?

      We agree that the increased coactivity in KO animals likely reflects a general network defect in AIC before CTA learning. This baseline change is unrelated to the taste stimuli, because (1) the coactivity is already elevated before the animals receive the taste stimuli (that is, a baseline abnormality), (2) the neuronal responses to taste delivery during CST1 is indicative of CS-only responses, as it occurs before the injection of LiCl, when the associative learning process is initiated, and (3) a trend toward higher coactivity is already present in the habituation water session before CST1 (Figure 2 and Figure 2 – figure supplement 2).

      We have clarified our description of these findings to avoid claiming that the increased coactivity “causes” poor learning performance (lines 379-382).

      (3) How do the authors interpret the large range of lick ratios (Figure 1G) for WT (almost bi-modal distribution)? Is there a within-subject correlation with any of the neurophysiological measurements to suggest a relationship between AIC neurophysiology and behavioural expression of CTA?

      See response to Point 1 above.

      (4) Indeed, CTA appears to be successfully achieved for Shank3 KO mice delayed by 1 day, as the level of saccharin aversion during the first retrieval session (T1) is comparable between Shank3 KO and WTs. In this context, not extending the first part of the paradigm to include CST3 seems to be a missed opportunity. Doing so would have allowed for within-cell and within-subject comparison of taste-elicited pairwise correlation across the learning and to investigate the neural mechanism of delayed extinction in KOs more effectively.

      We did not include a third CST session because when we analyzed the lick counts, KO animals already formed robust CTA after CST2 that was indistinguishable from WT animals. This suggests that the faster loss of CTA memory during extinction is due to a faster extinction process, rather than a weaker CTA memory from the outset. Adding a third CST could potentially lead to a memory that is harder to extinguish. Whether Shank3 KO mice would exhibit faster loss of memory in this scenario is an open question that would be interesting to explore in a future study.

      (5) How to interpret Figure 5F: Absolute discriminability is lower for T5 for CTA WT and CTA KO compared to CS-only? Why would AIC neurons have less information on taste identity by the end of extinction than during the unconditioned (CS-only) condition? And if that is the case, how is decoding accuracy in Figure 6C higher in T5 for CTA WT vs CS-only?

      We appreciate the reviewer's confusion about the discrepancy between our single-cell and population-level discriminability results. We speculate that in the CS-only state, individual AIC neuronal responses mostly reflect taste identity. However, after learning (in the CTA group), these neurons develop “mixed selectivity” (Tye et al., 2024), encoding not only identity but also the learned valence and extinction history. The lower single-cell discriminability after extinction (T5) in Figure 5F suggests that, although taste identity may remain constant, the learned history (e.g., "this taste used to be dangerous, but now it's safe") has shifted. This mixing of information makes each cell a weaker discriminator on its own.

      However, the higher population decoding accuracy in Figure 6C demonstrates that the entire population of neurons can work together more effectively. The learning process could reorganize the neural ensemble in such ways that our support vector classifier (SVC) is able to identify and combine the relevant signals within the population, even when the valence of taste stimuli has changed, to better decode stimuli and outperform the non-learned state. This suggests that the brain shifted to a more robust, population-based coding strategy for complex, learned information, which is resistant to changes in selectivity at the single-cell level. The finding that population coding is robust to single-neuron variability has also been reported in other cortical regions (Montijin et al., 2016).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Mechanistic experiments: Consider inhibitory neuron-specific imaging or manipulation (e.g., optogenetic enhancement of interneuron activity) to test whether restoring inhibition rescues learning flexibility.

      We have addressed the limitation and potential issues for manipulating cortical inhibition in Response to Reviewer #2 Public review 1.

      (2) Clarify limitations: Explicitly acknowledge the correlational nature of neuralbehavioural relationships in the Discussion.

      We have removed language that implies a causal relationship throughout, and have emphasized the correlational nature of our findings in the Discussion section of our revised manuscript (lines 379-382)

      (3) Enhance clarity: Simplify some dense methodological sections and expand figure legends to guide interdisciplinary readers.

      We have adjusted the Methods section and figure legends as needed for better readability.

      Individual Comments for Authors:

      (1) L83-90: Confusingly written, not easy to understand for someone not knowing the paradigm in detail.

      - What are the different stages? Memory encoding? Leaning? Extinction

      - More reliable in taste responsiveness - what does that mean?

      We have emphasized the behavior stages where each finding was observed in the revised manuscript (lines 82-91)

      (2) L112: Not sure if these references support the "crucial", since they do not seem to be causal.

      We have reworded this for accuracy (line 111).

      (3) Figure 1: panels h and i in the heat maps, it looks like that in the KO animal, activity is more suppressed from CST1 to 2?

      Panels j, l, m, and Figure 2: Neuronal suppression is already higher in CST1; therefore, there is no CTA effect but a general "perceptual" issue in the Shank3 model. The only effect seems to be a potential reduction in activation in CST2 in KO animals.

      This point has been discussed in Reviewer #3 Public Review 2.

      (4) Clarify in text. Especially with the sentence in the next paragraph, it might be confusing: "We wondered what other features of AIC activity during CTA acquisition might differ between WT and Shank3 KO mice."

      We have rewritten this in the revised manuscript (lines 169-170).

      (5) Clarify which are CTA-dependent and which are general (e.g., if writing suppression during CTA acquisition, it implies that it is related. But these changes were present before CTA.

      We have clarified this in the revised manuscript (lines 209-215).

      (6) Figure 4: Mainly shows a CTA-related increase in reliability in their taste responsiveness. This is not addressed anywhere else in the document and is not taken up in the discussion. How could it be related to the other findings, and what is its relevance? Please elaborate (e.g., in the discussion) or potentially remove?

      We measured response reliability, as stabilization of stimulus-evoked responses has been reported in other sensory cortices across different learning tasks. Yet, it remained unclear whether CTA learning would induce similar changes in AIC. We took advantage of our longitudinal recording to address this question and believe that this piece of evidence will contribute to the research community that studies taste and learning in general. In addition, what is striking to us is that while the taste selectivity is degraded faster in KO animals, their response reliability is largely preserved. This suggests that these two sensory stimulus-related neuronal properties may involve distinct cellular and/or circuit mechanisms.

      (7) Figure 5: Problematic to compare T5 between both groups, since T5 is lower than T4 in WT (against the trend) and T4 is an outlier in KO. e.g., if compared at T5, completely different results? Or why is there significance between T1 and T2 but not between T1 and T4 in KO? Could the authors address this point?

      In Figure 5B, the slightly lower average for WT animals at T5 was driven by a single outlier, and there was no statistically significant difference between T4 and T5 (corrected post hoc t-test, WT, T4 vs. T5, p = 0.4097). Therefore, it does not contradict the trend toward an overall increase in nonselective neurons during CTA extinction. For KO animals, the lower average at T4 than T5 (corrected post hoc ttest, KO, T4 vs. T5, p = 0.0082) was intriguing, and one possible explanation is that neurons in the KO group might undergo more dynamic and variable changes in their responsiveness during CTA extinction, fluctuating before finally stabilizing.

      Comparing T5 instead of T4 thus ensures that neuronal responsiveness is stabilized and reflects an “extinct” CTA memory more truly.

      General Comments:

      (1) While changes in SNR were observed in Shank3 models, the mechanism underlying decreased correlated variability has not been reported to date. Since decreased variability is usually associated with improved SNR ratio, it might be worth highlighting the distinction between "signal" and "noise" as separated in your analyses to make it more understandable for the reader.

      We have described in the Results section what signal and noise correlations indicated and how they were separated in our analyses in both the Results and Methods section of the revised manuscript (lines 185-193, lines 673-681).

      (2) What is the origin of the increased correlated variability?

      We have discussed that reduced cortical feedback inhibition could be a potential source of increased correlated variability in the Discussion section of our revised manuscript (lines 370-375).

      (3) Is the variability generally increased between trials (bigger fluctuations between trials for each neuron), or is the variability of each neuron similar, but they are just more correlated (more synced)?

      Our pilot analysis did not detect any evident changes in the response variability for each neuron across trials; thus, we think that in KO animals, neuronal responsivity becomes more correlated and synchronized.

      Reviewer #3 (Recommendations for the authors):

      (1) Point in line 422-424: Rephrase the closing statement of the discussion as you have shown that mutant mice are actually able to update their behaviour (in fact faster) when the valence of the sensory input changes.

      The “reduced ability to update behavior when the valence of a sensory input changes” refers to the finding that KO animals learned CTA more slowly; i.e., they were unable to timely adjust their behavior after malaise. We have rephrased this for clarity (line 448-449)

      (2) The Figure 6 legend does not correspond to panels D and E in the figure. Νο I, J in figure.

      We have fixed this mismatch in the revised manuscript.

      Minor concerns:

      (1) Cue/lick/taste-responding neurons greatly overlap and are not exclusively selective (Figure 1- figure supplement 2). Is there a genotype difference for the % of selective neurons (i.e., ones that only respond during cut/lick/taste) or the % of overlap?

      When we quantified the stimulus responsivity in KO animals, we also identified neurons that were activated by cues, licks, or tastes. Their respective percentages and overlap did not differ significantly from those in the WT group, indicating that the modality of KO neurons across different sensorimotor cues is not compromised in the KO condition (Author response image 2).

      Author response image 2.

      Neurons in WT and Shank3 KO animals show comparable responsiveness to sensorimotor stimuli during conditioning. (A) Percentage of neurons activated by the cue (left), lick movement (middle), and the tastant (right) in the CTA (KO) group (B) during the first conditioning session (CST1). (B) Venn diagram showing the overlaps among cue-, lick-, and tastant responsive neurons in (Figure 1 - figure supplement 2 C) and (A).

      (2) For Figure 1: The authors could also express consumption as a % of consumed (trial-averaged licks) over the number of trials. It is mentioned that mice undergo daily training sessions consisting of 'approximately 30 trials' (line 114). This can give an indication of how strong the learning is between cta1 and cta2 and how strong the genotype difference is.

      We are not sure if dividing trial-averaged licks over the number of trials would provide additional information, as the trial-average lick is already normalized to the number of trials.

      (3) Figure 4: Why is there a different number of neurons in C vs G?

      The figures B, C, D showed neurons that were activated by saccharin, and the figures F, G, H showed neurons that were activated by water. In all experimental groups, the numbers of neurons responsive to saccharin and water were different (i.e., B vs. F, C vs. G, D vs. H). The exact numbers were included in the corresponding figure legends in the revised manuscript.

      (4) Figure 5B: The grey background box is moved to the left.

      We kept the current figure format, as it effectively presents the mean, fitted mean, error bars, and individual animal data.

      (5) In line 142: (1-2), (2-3), (3-4), the numbers in parentheses are confusing.

      We have relabeled this as epoch 1-2, epoch 2-3, and epoch 3-4 in both text and figures for clarity (lines 145-146).

      (6) Line 188: Do the authors mean noise correlations?

      Rosenbaum et al. and Khoury et al. indeed measured correlated variability (noise correlation) in their study. On the other hand, Rothschild et al. did not specifically separate the noise from signal activities, which more likely reflect the coactivity measured in our case. We have rewritten this for accuracy (line 196).

      (7) Where mentioning in the CS-only group, please explicitly state the CS-only WT group.

      We have relabeled this throughout our revised manuscript.

      (8) In lines 273-274: if the comparison is the reduction in discriminability being faster for the KO animals that had CTA, the correct comparison should be CSonly KO vs CTA KO.

      We think that the better comparison to test how fast taste discriminability is reduced would be to perform post-hoc tests comparing T1 vs T2 within genotypes. We did not see significant changes between T1 and T2 in either genotype, which was reported in the figure legends of the reviewed preprint (lines 1140-1141).

    1. eLife Assessment

      This important study uses diffusion magnetic resonance imaging to non-invasively map the white matter fibres connecting the zona incerta and cortex in humans. The authors present compelling evidence to indicate that these connections are organized along a rostro-caudal axis. The findings will be of interest to researchers interested in neuroanatomy and cortico-subcortical connectivity.

    2. Reviewer #1 (Public review):

      Summary:

      This is a study which used 7T diffusion MRI in subjects from a Human Connectome Project dataset to characterize the zona incerta, an area of gray matter whose involvement has been demonstrated in a broad range of behavioral and physiologic functions. The authors employ tractography to model white matter tracts that involve connections with the ZI and use clustering techniques to segment the ZI into distinct subregions based on similar patterns of connectivity. The authors report a rostral-caudal organization of the ZI's streamlines where rostrally-projecting tracts are rostrally-positioned in the ZI and caudally-projecting tracts are caudally-positioned in the ZI.

      Strengths:

      The paper presents robust findings that demonstrate subregions of the human ZI that appear to be structurally distinct using a combination of spectral clustering and diffusion map embedding methods. The results of this work can contribute to our understanding of the anatomy and structural connectivity of the ZI, allowing us to further explore its role as a neuromodulatory target for various neurological disorders.

      Weaknesses:

      There should be further discussion of the clustering methods employed and why they are appropriate for the pertinent data. Additionally, the limitations of analyzing solely the cortical connections of the zona incerta should be addressed, as anatomical studies of the ZI have shown significant involvement of the ZI in tracts projecting to deep brain regions.

      Comments on the latest version:

      I reviewed the file and am more than satisfied with the authors responses and edits.

    3. Reviewer #2 (Public review):

      Summary:

      Haast et al. investigated the organization of the zona incerta (ZI) in the human brain based on its structural connectivity to the neocortex. They found that the ZI is organized according to a primary rostro-caudal gradient, where the rostral ZI is more strongly connected to the prefrontal cortex and the caudal ZI to sensorimotor cortex. They also found that the central region of the ZI is differently connected to neocortex compared with the rostral and caudal regions and could be important as a deep brain stimulation target for the treatment of essential tremor.

      Strengths:

      I think the overall quality of this work is great, and the results are presented in a very clear and organized manner. I particularly appreciate the effort that the authors put into validating the results using 7T and 3T data, as well as test-retest data.

      Weaknesses:

      The initial version of the manuscript left me with a couple of minor concerns that the authors addressed in the revised version. These were related to the clinical relevance of the work (now addressed with the lower-resolution data), and to the initial emphasis on a dorso-ventral gradient that the data did not strongly support.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study that used 7T diffusion MRI in subjects from a Human Connectome Project dataset to characterize the zona incerta, an area of gray matter whose involvement has been demonstrated in a broad range of behavioral and physiologic functions. The authors employ tractography to model white matter tracts that involve connections with the ZI and use clustering techniques to segment the ZI into distinct subregions based on similar patterns of connectivity. The authors report a rostral-caudal organization of the ZI's streamlines where rostrally-projecting tracts are rostrally-positioned in the ZI and caudally-projecting tracts are caudally-positioned in the ZI.

      Strengths:

      The paper presents robust findings that demonstrate subregions of the human ZI that appear to be structurally distinct using a combination of spectral clustering and diffusion map embedding methods. The results of this work can contribute to our understanding of the anatomy and structural connectivity of the ZI, allowing us to further explore its role as a neuromodulatory target for various neurological disorders.

      Weaknesses:

      There should be further discussion of the clustering methods employed and why they are appropriate for the pertinent data. Additionally, the limitations of analyzing solely the cortical connections of the zona incerta should be addressed, as anatomical studies of the ZI have shown significant involvement of the ZI in tracts projecting to deep brain regions.

      We are grateful to the reviewer for recognizing the strengths of our study, as well as for providing constructive suggestions to further strengthen the manuscript.

      In response to the reviewer’s feedback, we have expanded our discussion of the clustering methods employed, including the rationale for using spectral clustering in combination with diffusion map embedding, and clarified why this approach is well-suited to connectivity-based parcellation of the ZI.

      Additionally, we have expanded the Discussion to address the limitations of focusing exclusively on cortical connections. As the reviewer correctly notes, anatomical studies have demonstrated that the ZI has extensive connections with deep brain regions, and our approach therefore represents only a partial view of its connectivity. We have previously demonstrated the feasibility of reconstructing subcortical pathways using in vivo diffusion MRI (Kai et al., NeuroImage, 2022), providing a foundation for extending the present framework beyond cortical connectivity. However, as iterated below in our specific response to reviewer 1, we believe this deserves a separate thorough investigation. Nonetheless, we now explicitly discuss this limitation in the revised Discussion and outline directions for future work incorporating subcortical connectivity analyses.

      Reviewer #2 (Public review):

      Summary:

      Haast et al. investigated the organization of the zona incerta (ZI) in the human brain based on its structural connectivity to the neocortex. They found that the ZI is organized according to a primary rostro-caudal gradient, where the rostral ZI is more strongly connected to the prefrontal cortex and the caudal ZI to the sensorimotor cortex. They also found that the central region of the ZI is differently connected to the neocortex compared with the rostral and caudal regions, and could be important as a deep brain stimulation target for the treatment of essential tremors.

      Strengths:

      I think the overall quality of this work is great, and the results are presented in a very clear and organized manner. I particularly appreciate the effort that the authors put into validating the results using 7T and 3T data, as well as test-retest data.

      Weaknesses:

      That being said, I was left with a couple of concerns after reading the paper.

      - Although the authors discussed animal evidence for a dorsal-ventral organization of the ZI, I thought that the evidence they presented for it in this paper was not so convincing. In Figure S5, the second gradient (G2) shows a clear dorsoventral pattern, but this pattern seems to primarily separate the ZI and H fields rather than show an internal topology of the ZI. This is more likely the case given that there are two bands (superior and inferior) of high G2 values surrounding a single band (middle) of low G2 values. The evidence for the rostrocaudal gradient, on the other hand, is quite convincing.

      - HCP data is still too advanced for clinical translation. Although 3T is becoming more and more prevalent for presurgical planning, the HCP 3T dataset is acquired with a voxel size of 1.25mm, which is a far higher resolution than the typical clinical scan. It would be very useful for clinical readers to see what individual subject replicability looks like if the data were acquired at the more typical voxel size of 2mm. This could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for their positive evaluation of our work and for highlighting the clarity of the results, and our validation efforts across 7T, 3T, and test-retest datasets.

      Regarding the reviewer’s concern about the evidence for a dorsal-ventral organization, we agree that the rostro-caudal gradient is more prominent and convincing in our data, while the dorsal-ventral pattern is less robust. As the reviewer points out, the second gradient (G2) in Figure S5 may primarily reflect differences between the ZI and surrounding H fields, rather than a clear internal subdivision within the ZI itself. We have revised the Discussion to clarify this interpretation, emphasizing that our evidence for a dorsal-ventral organization is more tentative and requires further validation, particularly in light of prior animal literature.

      We also appreciate the reviewer’s important point regarding clinical translation. Indeed, the HCP datasets, both at 7T and 3T, use acquisition parameters (e.g., 1.25 mm voxel size at 3T) that exceed those of typical clinical scans. We therefore assessed the replicability of our findings in data acquired at more clinically representative resolutions (i.e., 2 mm voxel size at 3T). Details concerning this analysis are outlined in our response to reviewer 2 below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) When using "spectral clustering" to segment the ZI per its structural connectivity, it is unclear how k=6 clusters were chosen. Moreover based on previous reports of rodent ZI cytoarchitecture into rostral, dorsal, ventral, and caudal regions there is disagreement with the topographic organization of six clusters presented in this manuscript. Moreover, "diffusion map embedding" was not described or cited. This technique of dimensionality reduction and how it was applied to the data should be specifically described.

      We thank the reviewer for this thoughtful comment. We agree that selecting the optimal number of clusters in data-driven approaches such as spectral clustering is inherently challenging in the absence of a definitive ground truth. To address this, we computed alternative cluster solutions across a range of k values (k=2-8), which are presented in Supplementary Figure 2B. Our decision to focus on k=6 was guided by prior cytoarchitectonic descriptions of the rodent ZI by Romanowski et al., 1985, who delineated six distinct sectors (pars rostropolaris, pars dorsalis, pars ventralis, pars magnocellularis, pars retropolaris, and pars caudalis). Thus, our approach aligns the data-driven clustering with established anatomical subdivisions. Additionally, we found that k=6 provided a meaningful level of granularity for probing location-dependent neuromodulatory effects within the ZI, as discussed in the revised Discussion section (‘Discrete subregions of the zona incerta using spectral clustering’, second paragraph).

      Finally, we acknowledge the lack of clarity concerning “diffusion map embedding” in the original submission. We have now more explicitly mentioned diffusion map embedding in the “Connectivity gradients” paragraph in the Methods section, which includes the relevant references as well as description on how it was applied to the data.

      (2) "Strong correlations with cognitive terms and cortical hierarchies were particularly evident for clusters 3-5 (Figure 7a-b), which have the highest number of connecting streamlines (Figure 3c). Cluster 5, located near the central sulcus, is significantly linked to movement related cognitive processing (Pearson's r = 0.623, p < 0.005) and CogPC1 (Pearson's r = 0.593, p < 0.005). Cluster 4, situated more anteriorly, overlaps with regions involved in working memory (Pearson's r = 0.494, p < 0.001) with a high level of expression of serotonin (5-HT1B) receptors (Pearson's r = 0.413, p < 0.005) (Beliveau et al., 2017). Cluster 3, located further towards the frontal pole, is associated with mood (Pearson's r = 0.473, p < 0.005) and impulsivity (Pearson's r = 0.473, p < 0.001), and the sensorimotor association axis (Pearson's r = 0.583, p < 0.005) (Sydnor et al., 2021). Clusters 1, 2, and 6, characterized by the least number of connecting streamlines (Figure 3c), were relatively weakly associated (i.e., low Spearman's coefficient and/or within the spatial autocorrelation range) with cognitive terms or cortical hierarchies."

      The validity of identifying correlations with the spectral clusters and functional connectivity in tasks related to "keywords", and reporting similarities between these data and the regions most strongly connected via tractography, is a bit questionable. These results are based on studies that may not report all relevant findings. There is a bias towards regions that are more commonly studied with fMRI. Moreover, claims can be made about assigning functions specific to a brain region for almost any structure/function relationship (with some exceptions).

      We agree that correlations between connectivity-defined ZI clusters and functional annotations derived from external datasets should be interpreted with caution, given potential biases in the available literature (e.g., overrepresentation of well-studied cortical regions in fMRI meta-analyses) and the inherent risk of over-assigning functions to structural subdivisions. Our intention was not to make definitive claims about the functional specialization of individual ZI subregions, but rather to provide an exploratory framework for situating the ZI within broader cortical hierarchies and functional domains. These analyses are intended to generate hypotheses and to offer preliminary insight into how connectivity-based subdivisions of the ZI may relate to cognition and behavior.

      We agree with the reviewer that future studies should be specifically designed to address these questions more directly, for example, by combining connectivity-informed parcellations of the ZI with task-based or resting-state fMRI in the same subjects. Such targeted approaches will be necessary to rigorously establish the integration of the ZI within the brain’s functional organization.

      (3) The ZI's connections to many subcortical structures have also been reported in rodents and non-human primates. Moreover, the authors describe the efficacy of DBS of the caudal ZI in alleviating symptoms in patients with essential tremor, which indicates modulation of the dentato-rubro-thalamic tract fibers that project to subcortical structures such as the VIM thalamus, red nucleus, and cerebellum. The atlas in the study was characterized per the ZI's cortical connections only. These concerns should be addressed in the discussion.

      We agree with the reviewer that incorporating subcortical connections is essential for a comprehensive understanding of ZI connectivity. Building on our prior work demonstrating the feasibility of subcortical tractography (Kai et al., NeuroImage, 2022), we propose a systematic investigation of in vivo subcortico-incertal tractography as a critical next step. We believe that a dedicated investigation is required to systematically evaluate subcorticoincertal tractography, optimize reconstruction of key pathways (e.g., the dentato-rubrothalamic tract), and determine how these subcortical connections contribute to the topographic organization of the human ZI. We now explicitly discuss these considerations and identify them as an important direction for future work. Such (currently ongoing) work will not only clarify how subcortical inputs shape the topography of the ZI, but will also enable targeted optimization of tractography parameters to maximize the reliable reconstruction of specific but key pathways, including the dentato-rubro-thalamic tract. We believe this line of investigation will be important in advancing both the anatomical characterization and translational relevance of the ZI.

      Reviewer #2 (Recommendations for the authors):

      (1) Re: data quality compared to the clinic, this could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for this valuable suggestion. While this was suggested as potential future work, we felt adding this analysis would strengthen the manuscript. In response, we repeated our analyses using diffusion MRI data that more closely approximates a clinical acquisition with a lower spatial resolution (i.e., 2 mm vs. 1.25 mm isotropic) and number of diffusion-encoding directions (i.e., 130 vs. 270, Kasa et al., NeuroImage Clin., 2022). We have included these analyses in the revised manuscript.

      Reassuringly, the principal rostro-caudal gradient of cortico-incertal connectivity was preserved, demonstrating that the dominant organizational feature of the ZI is robust even under clinically representative acquisition conditions. However, finer-grained parcellations were less consistent with the original HCP analyses. In particular, cluster solutions with larger numbers of clusters (k > 3) became increasingly variable, indicating that differentiation of subtle connectivity-defined subregions benefits from the higher spatial and angular resolution afforded by research-grade diffusion MRI.

      We believe these findings provide a more nuanced assessment of the translational potential of our approach. They suggest that the large-scale topographic organization of the ZI can be recovered using clinically realistic diffusion MRI, while also highlighting the current limitations of routine clinical acquisitions for resolving finer anatomical subdivisions.

      We have incorporated these results and their implications into the revised manuscript.

      (2) Figure 6 legend labels: (c) and (d) should be (b) and (c).

      Thank you for highlighting this discrepancy. We have corrected as proposed.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have responded critically to all issues raised in the initial round of reviews.]

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular.

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    4. Reviewer #3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful comments on the manuscript. In response to their suggestions, we have:

      Improved hardware calibration flexibility and documentation (Rev 1).

      Clarified the optical specifications of the system, including axial resolution and working distance (Rev 1).

      Updated Figure 2 and Figure S1 (Revs 1 and 2).

      Corrected typographical errors and clarified terminology throughout (Revs 1 and 2). In addition, we have a new Zapit release (v1.0.4, which includes release notes), that contains many improvements and bug-fixes including suggestions from reviewers. 

      Public Reviews:

      Reviewer 1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      We thank the reviewer for their assessment of the manuscript, particularly the reference to finding the balance between modularity and an integrated package.

      (1) Command signals

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      We thank the reviewer for this suggestion. Zapit uses a linear calibration by default, which works well for high-quality diode lasers with built-in power control. For users with EOMs or AOMs we have now implemented a feature that allows creation of the appropritate sigmoid calibration curve. This is in Zapit version 1.0.4 and the process for generating the calibration curve is documented on our GitBook doc site. and the commit containing most of the changes is here. We also include a third order polynomial fit, which we hope will help users of some cheaper lasers where the control function has non-linearities. The appropriate non-linear fit is chosen automatically. We describe this in the legend and main text associated with Fig. 7F.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      The number of grid lines and their spacing are set via GUI options. The initial galvo-toimage mapping uses a field-centred affine guess based on the user’s "scanners.voltsPerPixel" setting, and the setup process is described in the user guide. The software ships with a suitable default gain value, which is unlikely to require modification. Invert flags for X and Y are also provided. Once beam locations are recorded, a similarity transform is fitted for residual offset, rotation, and scale. The latest Zapit release includes bug-fixes associated with the centering of the initial calibration point grid in the field of view.

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      We kept the mapping fixed for simplicity in both the build instructions and the code. Zapit will require a dedicated DAQ so we do not anticipate re-wiring is a hurdle. Nonetheles, the code is open-source and such a change is possible: the settings file would need to be augmented and the functions that write the analog output waveforms modified. 

      (2) Laser and optics

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      This suggestion mirrors our own thought process, but we opted not to add a ray-tracing rendering to Figure 1, as doing so comprehensively would require illustrating additional optical principles that would detract from accessibility. However, building on the reviewers recommendation, we now provide references explaining the underlying scanning principles for interested readers (Schottdorf et al. 2025, referenced in Fig. 2). The relationship between scanner angle and beam position is also shown diagrammatically in Figure 7.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      This is an excellent point. With our specifications (0.8 mm beam diameter, 473 nm wavelength, 200 mm focal length objective), the effective NA is approximately 0.002, yielding a Rayleigh range of approximately 37.6 mm. The beam must therefore travel nearly 4 cm from focus before the point-spread function doubles in width, making the system highly insensitive to skull curvature. We have added this calculation and noted its practical advantage in the revised manuscript. (Section 2.2).

      What is the working distance?

      The Plossl scan/objective lens is housed at the end of the lens tube, giving a working distance of approximately 20 cm for the 200 mm focal length objective. We have added this information to Figure 2.

      Reviewer 2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank the reviewer for their positive assessment of the manuscript, and their vision for Zapit as an important community resource.

      Reviewer 3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

      We thank the reviewer for their detailed assessment of Zapit.

      Concerns & comments:

      (1) While the authors argue that it offers the best utility to affordability trade-off - faster than motorized drivers and require much less power than DMDs (100X) and less expensive/easier to use compared to SLMs, in the current form, the manuscript does not clearly list the limitations of the approach. At such, in my opinion, the authors should include side by side comparisons (perhaps as a table). For example, clear statements should be included with respect to comparisons in lateral (x-y), axial (z) spatial resolution, as well as temporal sequential aspect of Zappit and other photo-stimulation techniques involving DMDs or SLMs.

      Section 3.2 (Comparison to other approaches) compares the scanner-based approach to related techniques. Whilst this is brief, we believe it is adequate because the resolution is ultimately limited by tissue scattering. Indeed, we demonstrate that the radius of neural inhibition ~10 times larger at 2 mW than the lateral PSF (Figure 8D). The size of a DMD pixel on the brain will likely also be smaller than the excitation area, but it does depend on the imaged size of the DMD on the brain. Since that can vary from system to system, a comparison of even theoretical resolution is not straightforward. In terms of spatial patterning, DMDs and SLMs allow for arbitrary shapes to be created on the brain and we point this out in section 3.2. 

      (2) Is power really a limitation in terms of the laser sources? Or is this a disadvantage mainly because using less power has beneficial effects on the tissue health? It may be useful to provide metrics of comparisons along these lines between Zappit and DMD-based approaches.

      The reviewer highlights an important distinction. Too much laser power can cause phototoxicity, and can make neurons more excitable due to heating. It also results in a larger region of stimulation and increases the chance of off-target effects. However, when activating multiple sites with a galvo-based system like Zapit, dwell time goes down ~linearly as the number of “simultaneous” stimulation sites increases. Therefore, even if the same power at the sample is maintained (and the same risk of phototoxicity), peak power must go up to provide the same average power at each site. For example, stimulating 20 points at 40 Hz requires 10 times the peak power compared to stimulating 2 points at 40 Hz.

      The same is true of DMD-based approaches. For example, the Mightex recommend a 1 to 4 W laser to run their Polygon DMD-based photostimulation system over an area the size of the mouse dorsal cortex (personal communication); Kauvar, et al. 2020 used a 5 W laser to cover an area 7 mm across using a Polygon system. The cost of such a laser and the Polygon alone likely exceeds 60,000 USD. 

      Other than prices and logistics, the laser powers needed at the sample and their duty cycles are essentially the same across approaches and so there are no meaningful comparisons we can provide in this domain. However, based on the reviewer’s comments, we now clarify the effects of heating from high laser powers on neural excitability in the discussion (Section 1.1).

      (3) Arbitrary-scanning vs random scanning may be more appropriate to describe to strategy.

      We agree with the confusion surrounding “random-scanning” and have chosen the phrase “laser-scanning” rather than “random access”, which is the term used in the original pre-print. 

      Reviewer 1 (Recommendations for the authors):

      Consider noting that most other galvo scanner models can probably be used.

      We have added: "Other scanners could also be substituted, as can other lasers, lens combinations, and laser wavelengths etc.” (Section 4.3).

      The basic version of Zapit requires MATLAB (although alternatives are possible and guidance/code is provided), which is not unreasonable but may limit adoption.

      We acknowledge this point. MATLAB is widely available in academic neuroscience laboratories, and we provide a Python-based interface, but a complete conversion is beyond the scope of this manuscript. We hope that the open-source code will be adapted to other languages by the community as needed.

      Where laser power is mentioned (e.g., Discussion: "we recommend using 1-2 mW time-averaged light power..."), it is not always clear if this is at the laser or in the specimen plane; clarifications would be helpful.

      Thank you, we have clarified throughout that reported laser powers refer to measurements at the specimen plane.

      Abstract: "causally manipulating" - the "causal" part is redundant and can be dropped, or replaced with "transient" (more relevant).

      Changed to "transiently".

      Section 2.2: "The modular and open-source nature of Zapit means that exciting new configurations" - strike "exciting".

      Corrected.

      Figure 7D - y-axis label text is clipped.

      Corrected.

      Reviewer 2 (Recommendations for the authors):

      (1) In Figure S1 - Our paper (Heindorf et al.) is incorrectly listed as not making code available - the "camera controlled laser stimulation" software we used is part of Iris2p (that is freely available on SourceForge and linked as such in the paper). Granted, it's not user-friendly or easy to find in the large Iris2p software package - but it is technically "available".

      We apologise for this oversight. We have updated the Figure S1 legend to note that “available” code in this context refers to a dedicated and documented standalone package. We have also added links to both Pinto and the Keller-lab software packages.

      (2) In Figure 3: "and THE beam goes"

      Corrected.

      (3) In Figure 4: "the useR will be prompted"

      Corrected.

    1. eLife Assessment

      In this study, Yuan and colleagues carry out transcriptomic and epigenomic experiments to study open chromatin regions and transcripts that change upon larval settlement in the sponge Amphimedon. The authors present compelling evidence to show that sponge larvae prepare for receiving an environmental cue (sunset) by extensively modifying their chromatin accessibility. The study represents a fundamental advance in understanding the fine genetic control of larval settlement and has significance beyond the immediate field of sponge larval biology.

    2. Reviewer #1 (Public review):

      Summary:

      Yuan and colleagues present a thorough study of gene activation before and during metamorphosis in sponge larvae, combining in depth analyses of staged transcriptomes and chromatin accessibility profiling (ATACseq). Amongst several very interesting findings, the study reveals that the acquisition of settlement competence, which arises in response to decreasing light at sunset, is characterized by changes in chromatin accessibility that anticipate strong transcriptional shifts occurring as metamorphosis starts. Another notable finding is a set transcription factors amongst the genes strongly up-regulated at the onset of metamorphosis. In addition, larvae exposed to constant light, a condition that stalls metamorphosis, were found to activate metabolic pathways that are not normally expressed in swimming larva, Together the findings provide a rare level of understanding into how environmental conditions can promote deployment of alternative developmental programs in planktonic larvae.

      Strengths:

      This is a comprehensive and rigorous study of a phenomenon of wide interest. It will inspire researchers working on other species to look for similar, environmentally-driven "anticipatory" epigenetic mechanisms. It also provides a wealth of detailed information on genes, notably transcription factors, that are candidates for involvement in regulating specific metamorphosis transitions - and beyond. The data presented here are thus undoubtedly a rich and valuable resource.

      Weaknesses:

      It is not always straightforward to connect the conclusion statements in the text to the figures, or to grasp the aims and logic of the workflows.

    3. Reviewer #2 (Public review):

      Summary:

      It is demonstrated that sponge larvae prepare for receiving the environmental cue (sunset) by extensively modifying their chromatin accessibility in the absence of large gene expression changes - "anticipatory" chromatin remodeling. This program can be offset by modifying the cue (making light constant), leading to a novel molecular state. These are the most important novel findings. Then, around metamorphosis there is dramatic regulation of transcription factors and their targets. While this result is hardly surprising, it is nice to see it demonstrated thoroughly in a phylogenetically fundamental animal study system.

      Strengths:

      This is a top-notch study of a key life cycle transition in an organism of great phylogenetic importance, involving concurrent gene expression and chromatic accessibility profiling (to the best of my knowledge, this has never been done in non-bilaterians and likely anywhere outside Vertebrata). The result is highly non-trivial. There is also an additional experiment modifying the key environmental cue (constant light), adding additional insight.

      Weaknesses:

      The paper presents a great amount of material, which somewhat obscures results related to the major novel finding ("anticipatory" chromatin remodeling). Among them is the fact that there is no significant statistical association between regions opened during remodeling and genes subsequently regulated during metamorphosis. This lack of direct association suggests that there may be more to the story than just priming the genome for metamorphosis.

    4. Reviewer #3 (Public review):

      Summary:

      In their manuscript, Huifang Yan and colleagues perform RNA-seq (CEL-seq) and ATAC-seq experiments to profile the transcriptome and chromatin accessibility of sponge larvae across larval competence, settlement and early postlarval development. Amphimedon, the sponge species that they use, is amenable to lab experiments and can therefore be a convenient model for experimenting with this otherwise difficult to assay ecological parameters and cues. They had previously observed that light conditions (diminished light) at sunset are Sucritical for larvae to enter a pre-settlement stage and prime them for settlement and metamorphosis. In this paper they report that these conditions induce a gain of accessibility in many genes, including transcription factors, and that altering these conditions by providing continuous light at sunset affects this reprogramming event.

      Strengths:

      The above is a very interesting observation, one that the authors speculate that could have a broader significance and be a theme in many more larvae. I agree with the authors that this is an important finding and I think that the paper will be interesting for the broad readership of eLife. If this is the case, the authors open up a new theme of chromatin regulation, extensively studied in mammalian contexts, but severely understudied in pretty much every other context.

      Weaknesses:

      I think however that their paper often reports the data in a difficult to follow way, and that other sorts of analyses would have made the results more accessible for the broad readership.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Yuan and colleagues present a thorough study of gene activation before and during metamorphosis in sponge larvae, combining in-depth analyses of staged transcriptomes and chromatin accessibility profiling (ATACseq). Amongst several very interesting findings, the study reveals that the acquisition of settlement competence, which arises in response to decreasing light at sunset, is characterized by changes in chromatin accessibility that anticipate strong transcriptional shifts occurring as metamorphosis starts. Another notable finding is a set of transcription factors amongst the genes strongly up-regulated at the onset of metamorphosis. In addition, larvae exposed to constant light, a condition that stalls metamorphosis, were found to activate metabolic pathways that are not normally expressed in swimming larvae. Together, the findings provide a rare level of understanding into how environmental conditions can promote deployment of alternative developmental programs in planktonic larvae.

      Strengths:

      This is a very comprehensive, well-documented and rigorous study of a phenomenon of wide interest. It will inspire researchers working on other species to look for similar, environmentally-driven "anticipatory" epigenetic mechanisms. It also provides a wealth of detailed information on genes, notably transcription factors, that are candidates for involvement in regulating specific metamorphosis transitions - and beyond. The data presented here are thus undoubtedly a rich and valuable resource.

      We thank reviewer #1 for finding our study on sponge metamorphosis interesting and compelling, and that it is likely to inform future studies on gene regulation and activity in environmentally regulated developmental processes and metamorphosis.

      Weaknesses:

      I see no significant weaknesses; however, the documentation of the data is very compressed, with all the findings contained in 4 multi-panel figures with succinct legends. It is not always straightforward to connect the conclusion statements in the text to the figures. Although the relevant data is available in supplementary files, I would appreciate more help in navigating the data to assess the support for key conclusions, if possible, illustrating each text conclusion explicitly in the main figures.

      Thank you sincerely for these suggestions on how to better present the results. We agree that the figures and associated legends are succinct. To rectify this in the hope of improving clarity and accessibility, we have (i) created two new figures by splitting our original four figures into six, and (ii) expanded figure legends to provide more explanatory details. We also made minor additions to the main Results text to more fully explain some results (see also reviewer #3’s comments).

      Specifically, we:

      (1) Removed panel K (heat map of TF expression) from original Fig. 1 and created a new figure (new Fig. 2) that focuses solely on TF expression and emphasises the extraordinarily high expression of many TFs. We also moved into the new Fig. 2 a panel from the original SFig. 1 that documents the larval cell types that express these most highly expressed abundant TF transcripts. This new figure should provide the reader with a clearer perspective on high TF expression in the larval competence and the initiation of metamorphosis.

      (2) Removed panel F from the original Fig. 4 (now Fig. 5) to create a new, expanded Figure 6 that presents a stand-alone summary of the main findings of this work; that is, environmental regulation of competence and early metamorphosis. This allowed us to (i) incorporate the constant light experiment into the summary figure, and (ii) provide a more detail explanation in the legend.

      Reviewer #2 (Public review):

      Summary:

      It is demonstrated that sponge larvae prepare for receiving the environmental cue (sunset) by extensively modifying their chromatin accessibility in the vicinity of genes that are going to be regulated during metamorphosis, in the absence of large gene expression changes. This program can be offset by modifying the cue (making light constant), leading to a novel molecular state.

      Strengths:

      This is a top-notch study of a key lifecycle transition in an organism of great phylogenetic importance, involving concurrent gene expression and chromatic accessibility profiling (to the best of my knowledge, this has never been done in non-bilaterians and likely anywhere outside Vertebrata). The result is highly non-trivial. There is also an additional experiment modifying the key environmental cue (constant light), adding additional insight.

      We thank reviewer #2 for their efforts and for appreciating the approaches we employed to understand environmental induction of sponge metamorphosis. In addition to the phylogenetic importance of sponges, their pelagobenthic life cycle is likely shared with disparate bilaterians (but not with vertebrates and other chordates, whose metamorphoses are probably derived).

      Weaknesses:

      I have only a couple of suggestions.

      (1) Not all new pre-emptively opened OCR regions are associated with genes that are going to be regulated during metamorphosis. Is their association with such genes statistically significant? (Fisher's exact test?)

      Thank you for raising this helpful point. In following your suggestion to statistically test this, we determined that a Fisher’s Exact Test was not appropriate because that test is generally used only for small samples or tables with expected counts below 5; in our data, all four expected cell frequencies are well above 5 (minimum = 228.6) and N = 25,149. Thus we instead tested for an association between chromatin accessibility and differential gene expression using the more appropriate Pearson's chi-squared test. We found no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.18), and have added these details into the Results (lines 344-46) as follows: “Consistent with this interpretation, 62% of all genes that are differentially expressed in 1 hps postlarvae (3032) have proximal chromatin regions already accessible in competent larvae (Supplementary Tables 3 and 9), although statistically we find no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.184).”

      (2) Re: extended discussion on possible reasons for activation of specific transcription factor families. I feel it is not terribly useful since it is hardly more than guesswork. The authors should consider condensing this part to better emphasize the major (and most unexpected) large-scale regulation patterns.

      We agree with this appraisal and have modified the beginning of Discussion to highlight the large-scale and rapid changes of overall gene expression. This emphasises the regulatory processes – TF expression and chromatin state changes – that must be in place to allow such transcriptional changes to occur. We feel this addition enhances the focus on TF activation and regulation at competence and early metamorphosis, especially given the scale and level of change, with most of TFs being expressed at very high levels (i.e. top 5% of all gene expressed). As outlined above, we created a new figure focussed on highly expressed TFs (new Fig. 2) to hopefully further highlight this phenomenon. It would be of great interest to know if this is conserved amongst animals with a pelagobenthic life cycle and rapid metamorphosis.

      (3) Re: enrichment analysis based on significant genes (Figure 1H): Even though it is a common practice, there is nuance: as we all know very well, many genes pass a significance threshold not because they are highly differentially regulated (i.e., show large fold-change), but because they are more abundantly expressed overall and so the statistical power for them is greater. A good example is ribosomes - before we realized what was happening, they would show up as enriched in almost every experiment of ours, which was not very useful since their fold-change was quite trivial. I see the authors have ribosome enrichment too, and I suspect there are a few more functional groups that made it because they tend to express highly on average. Ideally, we want to see what is enriched among highly regulated genes, not among abundantly expressed genes. Because of this we moved to compute enrichment based only on fold-change, using the GO_MWU package (https://github.com/z0on/GO_MWU). I suggest authors give it a shot, to see if the enrichment results become more interpretable. GO_MWU is also very powerful to analyze enrichment in WGCNA modules, in case the authors want to try that.

      Thank you for this interesting insight and advice. We applied the GO_MWU package to our gene expression dataset. Overall, these new results corroborated the original KEGG enrichment analysis, largely identifying GO biological processes, cellular components and molecular functions consistent with the previously identified KEGG molecular and cellular processes operating at larval competence and 1 hps. These include genomic regulatory processes underlying transcriptional changes and morphogenetic processes that occur in the first hour of metamorphosis, and which are also highlighted in a recent BioRxiv paper (https://doi.org/10.64898/2026.04.23.719999) from our group.

      We have added (i) results from the GO MVU analyses to Supp. Fig. 1 and Supp. Table 4, (ii) the following statement to Fig. 1 legend: “GO-MWU analysis of upregulated genes reveal stage-specific enrichments largely consistent with the KEGG analysis (Supplementary Fig. 1 and Supplementary Table 4).”, and (iii) a brief description of this approach into the Methods.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, Huifang Yan and colleagues perform RNA-seq (CEL-seq) and ATAC-seq experiments to profile the transcriptome and chromatin accessibility of sponge larvae across larval competence, settlement and early postlarval development. Amphimedon, the sponge species that they use, is amenable to lab experiments and can therefore be a convenient model for experimenting with this otherwise difficult to assay ecological parameters and cues. They had previously observed that light conditions (diminished light) at sunset are critical for larvae to enter a pre-settlement stage and prime them for settlement and metamorphosis. In this paper, they report that these conditions induce a gain of accessibility in many genes, including transcription factors, and that altering these conditions by providing continuous light at sunset affects this reprogramming event.

      Strengths:

      The above is a very interesting observation, one that the authors speculate could have a broader significance and be a theme in many more larvae. I agree with the authors that this is an important finding, and I think that the paper will be interesting for a broad readership. If this is the case, the authors open up a new theme of chromatin regulation, extensively studied in mammalian contexts, but severely understudied in pretty much every other context.

      We thank reviewer #3 for their positive assessment of our findings, pointing out the novelty of this research and its broad relevance.

      Weaknesses:

      I think, however, that their paper often reports the data in a difficult-to-follow way, and that other sorts of analyses would have made the results more accessible for a broad readership. Here, I present some suggestions that the authors might want to take into account to improve their results.

      Reviewer #1 also commented on how the results were difficult to follow. Based on your and their comments, we have reworked parts of this section, and the figures and legends. Details of these changes are listed above and can be viewed in the new version with track changes on.

      We note that no further specific suggestions were visible to us in your review.

    1. eLife assessment

      This study presents a reaction-coupled molecular simulation framework to investigate how enzymatic post-translational modifications regulate biomolecular condensates and where reactions preferentially occur within them. These findings are important, in particular, to the emergence of spatially heterogeneous reaction fluxes and enhanced activity at condensate interfaces. The work is timely and has implications beyond the immediate subfield, especially for understanding condensates as biochemical reaction centers. The strength of evidence is solid overall: the computational framework is well motivated and broadly supports the main conclusions, although some aspects of the model implementation, parameter choices, and controls would benefit from further clarification and validation.

    2. Reviewer #1 (Public review):

      Summary:

      A well-presented computational work on how post-translational modifications take place from a thermodynamic and mechanistic point of view.

      Strengths:

      A model capable of recapitulating complex phenomena to simulate, such as phosphorylation.

      Weaknesses:

      The methodology relies on multiple user-defined parameters that alter the setup. Below are the specific concerns:

      (1) The abstract reads: 'First, reactions that weaken favorable interactions are thermodynamically suppressed within condensates. As a consequence, regulation of condensate solubility is most efficient when PTMs tune interactions to values close to the solubility threshold.' This phrasing is confusing. PTMs that weaken interactions will, in principle, shift the solubility to higher values, and therefore closer to the thermodynamic conditions that allow for condensate formation, but if those transitions are most efficient when they tune interactions to values close to the solubility threshold, they will not be suppressed? The authors have to make this statement clearer to readers, which is particularly important in the abstract of the manuscript

      (2) The computational framework used by the authors is reasonable given the coarse-grained nature of the model required to study this problem. However, there are multiple user-defined variables that require further validation in order to make their choices justifiable:

      a) NN and NK have the same interaction epsilon value, as well as K-K. This would effectively make kinase form condensates on their own if the right stoichiometry was imposed. This should be re-evaluated with a reduced K-K interaction to validate whether the conclusions remain invariant with this assumption.

      b) N, P and K beads also have the same molecular diameter. This should be justified, for instance, with solvent available surface area calculations to determine the excluded volume for the different species, or by citing other works that support this approach.

      c) The phosphorylation reaction takes place when 2 particles are found within a 1.5sigma distance; however, this value choice is not justified. The authors should prove how variations to this choice affect their conclusions.

      (3) Figure 2a is informative although not entirely intuitive to follow. It can be concluded, as the authors mention, that the capacity to form condensates decreases as phosphorylation is favored. Therefore, as the text also says, there is a higher fraction of P particles as lambdaP increases; however, the fraction NP/NT appears to decrease based on the color scale. According to the figure caption, this is meant to represent the fraction of P particles over the total, but this should not exceed 1. This should be clarified and better explained in a revised version. Moreover, the information related to this, shown in Figure S4a, is very informative, and I advise the authors to include it as part of the main set of figures, as it will potentially help many readers to follow the manuscript better. Moreover, in Figure S4a, some lines appear to be disconnected; this probably comes from trajectory merging; the authors must check this.

      (4) 'At the interface, N and K concentrations remain relatively high, but scaffold proteins experience fewer stabilizing interactions, lowering the energetic cost of phosphorylation. This leads to enhanced reaction activity specifically at the boundary between phases.' This statement perfectly explains why the density profiles of K and N do not match the phosphorylation probability curve. This probability is determined by the energetic impact of the reaction and the probability of encountering each other in space, but also on the short timescale diffusion: N and K proteins have greater access to more microstates at the interface and can access them faster, while having enough density to encounter each other. It would be interesting for the authors to prove or invalidate this argument. At the very least, it should be mentioned.

      (5) Characterizing the real impact of condensate interfaces in real size condensates is an interesting approach, nonetheless this paragraph lacks most of the necessary details to be robust and obtain any reliable conclusion out of it in its current form:

      a) There are several CALVADOS parametrizations, which one do the authors use? The force field must be cited.

      b) R is not well defined in the caption.

      c) In the rendered images, periodic boundary conditions appear not to be implemented; this should be clarified

      d) In the Intermolecular energy profiles, it seems that FUS-LC has no condensate bulk.

      e)How is the interface width calculated?

      f) Why do the authors choose the energy profile and not density? Or other observables such as the radius of gyration.

      g) The interface width will depend on the temperature, and how distant this temperature is from the critical temperature for phase separation. Currently, this information is lacking.

      h) Variations in the temperature, quantity used to define the interface, should be addressed in order to make the interfacial importance claim robust.

      (6) 'Since the size of typical cellular condensates rarely exceeds the 2 μm diameter'. This statement should be supported by multiple references.

      (7) Figure S2 should be improved: The use of 'weak' or 'strong' labels is subjective and it is unclear which parameters are being used. Moreover, 'exp reference' is not described or cited. Furthermore, this plot shows what appear to be sketched curves. The calculation of the coexistence densities to construct this phase diagram is trivial for the system studied here; the authors should provide direct estimates.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Thermodynamic principles of enzymatic regulation in biomolecular condensates from reaction-coupled molecular modeling" reports results of a very coarse-grained molecular dynamics (MD) simulation of three types of particles that undergo phase separation and chemical reactions. The study focuses on a molecule that can exist in two states (phosphorylated vs. unphosphorylated), whereas the third species is an enzyme that catalyzes the reaction in one direction of the transition. This topic of chemically active multicomponent mixtures is timely, and its connection to biomolecular condensates is well motivated in the introduction. Using their minimal setup, the authors find that driven reactions affect the composition of the dilute and dense region, and can even suppress phase separation completely. Moreover, the reaction fluxes are heterogeneous and exhibit a pronounced peak at the interface. The authors then interpret this acceleration of reactions in light of the role of condensates as reaction centers and conclude that the effect they report is relevant in cells.

      Strengths:

      A strength of the manuscript is the setup of a minimal system to understand the complex roles of enzymatically controlled, active reactions in condensates. This is a subtle topic since the physics of phase separation, implying non-ideal systems, affects the reactions, which can then no longer be described in a dilute approximation. The authors tried to take special care to ensure thermodynamic consistency, which is key to describing the interplay of the two effects accurately. The authors also perform relevant numerical tests, e.g., by determining the effect of reactions on the phase diagram, and measuring densities and reaction fluxes carefully. Moreover, they attempt to interpret the obtained quantities with intuitive pictures, although this is not always convincing.

      Weaknesses:

      The main weakness of the work is its presentation: I could not follow the detailed setup of the model since important details (such as the concrete interaction potentials and particularly the implementation of the reactions) are not described thoroughly. In particular, the implementation of the actively driven reaction is mysterious, and a proper negative control without activity is missing. Such a control is crucial since it would allow testing whether the implementation of the code ensures local detailed balance and thus thermodynamic consistency. Furthermore, such a passive null model would surely help in establishing the effects of activity, which is currently unclear. Since I don't understand the detailed setup, I cannot judge whether the major result, namely that reactions are accelerated at interfaces, is correct. It might very well be true, but I'm unable to judge this based on the current presentation. I detail my criticism in separate points below, and I hope that addressing these points helps the authors to improve their manuscript:

      (1) I find the model setup unclear, which makes it difficult for me to gauge the correctness of the results. I think the details of the model need to be explained in more detail, both in the main text and the SI. There are three different aspects that I find lacking:

      1a) The setup of the interactions in the model is unclear. First, only single interaction parameters \epsilon_i are specified, but the simulation likely needs to specify pairwise interactions. Table S2 is not particularly helpful since it uses a different notation (Is this switch of notation necessary?). Second, the statement that k_BT = 0.75 is confusing since k_BT should have units of energy. Third, the authors mention "a truncated and shifted LJ potential", but it is unclear whether this is the potential they use or not. Given this lack of details, I would not be able to immediately repeat the simulation, even without reactions.

      1b) In the main text (page 6), it is unclear whether an active or passive system is studied. The subsequent results suggest that the system is active (and thus has sustained energy fluxes), but this needs to be explained in detail. In any case, I would like to see an explicit activity parameter, so that a passive control can be added. For instance, I would expect that a passive system remains the same when reaction rates are changed, but the energetics are kept constant. In any case, since it is not explained how activity enters the system, the subsequent results are unclear to me.

      1c) It is unclear how thermodynamic consistency (i.e., local detailed balance) is ensured. For instance, how is the formation of a scaffold-enzyme complex performed (page 6)? The SI is also light on details. For instance, it is unclear whether the transitions in Equations (1-8) refer to individual particles (in this case, how are bimolecular reactions implemented?) or refer to densities or even overall particle counts. It is also suspicious that the reactions are given for the "dilute limit", whereas the main text clearly discusses condensed regions. It is also unclear how ΔU is calculated and what it means. Generally, I would expect a detailed discussion of detailed balance conditions in the SI, if not even in the main text. Here, it might help to clearly separate thermodynamic aspects (involving detailed balance) from kinetic considerations.

      (2) If the authors study an active system, reporting averaged concentrations and partition coefficients might be misleading. Generally, active systems develop gradients in the dilute and dense regions (Figure 1), so any averaged measurement (such as a density) will depend on the size of the region. It is thus necessary to clearly define the measurements in the main text (so readers know how to interpret the plots), and the discussion needs to be much more careful, particularly when comparing to phase diagrams, which typically discuss thermodynamically large systems.

      (3) The strongly increased fluxes at the interface are a bit suspicious, particularly since they are hardly visible in passive systems (Figure S6). The enhancement might originate from large reaction rates, leading to a short reaction-diffusion length scale (although I would then expect balanced reactions in the phases). Alternatively, they might originate from the microscopic details, such as the cut-off length of the kinase-interaction range. In any case, it would be important to establish a negative control using a passive setup and (based on this) explain the observed flux increase clearly.

      (4) I am not convinced by the statement about the bias of reactions in the dense and dilute phase, e.g., on page 12. I would find it extremely helpful to start with a passive system (without energy input), where I would expect balanced chemical potentials, so that all reactions are balanced at all points in the system. Already in this case, there might be accelerated reactions at the interface (due to kinetic details), but the two reactions need to be balanced to have overall homogeneous chemical potentials. Starting from such a base state, one could then discuss how activity biases reactions in a certain direction. While the current explanation correctly mentions that a transition can be hindered by energetics, it fails to account for the different abundances. In a passive system, these two aspects are perfectly balanced, leading to a linking of reaction rate constants and partition coefficients (e.g., see https://arxiv.org/abs/2202.13646). These aspects need to be discussed much more cleanly (and they are intimately linked to the setup of the model, which is currently unclear; see point 1).

      (5) I think the title is misleading. The authors do not establish new "Thermodynamic principles". I am also not sure what they mean by "reaction-coupled molecular modeling". Finally, the 
"enzymatic regulation" is hardly discussed in the main text, which instead seems to emphasize the accelerated reactions at the interface.

    1. eLife Assessment

      This study addresses a recent discovery by others that electroconvulsive therapy (ECT) generates seizure activity and spreading depolarization (SD), reflected in large increases in calcium, which can be observed through imaging fluctuations in neuronal calcium. Revisions addressed some concerns of reviewers but not the main issues. Therefore, the consensus was that the work is useful but the evidence showing that SD, rather than seizures, confers the neuroplastic and other therapeutic effects of ECT is incomplete.

    2. Reviewer #1 (Public review):

      The work shows that the ECS-induced calcium waves recapitulate several hallmarks of CSD, including hemodynamic alterations, distinct propagation patterns, and elevated Fos expression.

      Update of the weakness section.

      (1) I still have concerns about the use of Fos staining as the exclusive marker of neuroplasticity. Of course, Fos drives many forms of plasticity. But in addition to the role, Fos upregulation also reflects a recent history of elevated neuronal activity. This has been reported in numerous papers, including those cited in the article (e.g. Mahringer et al., 2019; Tyssowski et al., 2018). If the authors claim that their main finding is that ECS-induced CSD is necessary to drive Fos, then the relevant background should be presented in the Introduction.

      (2) The authors did not sufficiently address the concern regarding the bilateral suppression of EEG following unilateral calcium waves. Unilateral CSD is known to depress EEG only in the affected (ipsilateral) cortex. I suspect that the bilateral EEG suppression can be driven by the bilateral seizure (phase III oscillations) that is always coupled with the ECS-induced calcium waves.

      (3) Using the term 'phase III oscillations' instead of 'seizures' is confusing because the word "oscillations" covers a wide range of brain oscillations - from normal to pathological ones. First, the authors base the terminology change on the absence of "the ictal spikes characteristic of epileptic seizures..." during the phase. However, anesthesia can suppress ictal spiking and the phase III activity can represent an aborted seizure induced under anesthesia. Second, the authors cite the study by Brumback and Staton (1982) to justify their use of the term "phase III oscillations". However, the cited work describes the phase III as a seizure with "spike/polispike-wave activity".

      (4) Cortical SD cannot invade the hippocampus in the in vivo brain, although SD can occur in the hippocampus in response to generalized seizures. Invasion of CSD suggests its non-synaptic propagation via the contiguous gray matter. The following studies showed that CSD cannot invade the hippocampus non-synaptically: Fifkova E. Spreading EEG depression in the neo-, paleo- and archicortical structures of the brain of the rat. 1964 (PMID: 14138725); Yoshida et al. Identification of the extent of cortical spreading depression propagation by Npas4 mRNA expression. 2015 (doi: 10.1016/j.neures.2015.04.003). Concerning the papers cited in the discussion ("In mice, the ECS driven CSD likely invades hippocampus" (Mitlasóczki et al., 2025)), and response (Bahari et al., 2020; Bonaccini Calia et al., 2022), hippocampal SD was triggered by focal seizures, independent of CSD, in the experiments.

    3. Reviewer #3 (Public review):

      In this revised version of the manuscript, the authors have addressed several of the Reviewer's prior questions and comments. However, the main elements of the critique remain largely intact in the absence of additional experiments.

      Major:

      (1) The authors and Reviewer disagree as to the extent to which the main findings replicate prior work vs. represent true novelty. In their response, the authors state, "This appears incorrect. It was known that direct cortical stimulation (as was done in the Rosenthal paper) can drive a calcium event that resembles a CSD. It was also known that CSD can drive Fos expression. What was not known is that the Fos expression driven by ECS is fully explained by the calcium event (putative CSD)." First, what appears incorrect is the author's statement that Rosenthal et al. use direct cortical stimulation. So far as the Reviewer can discern, Rosenthal et al. in fact did not use intracranial or direct cortical stimulation, as stated repeatedly in the rebuttal. Instead, Rosenthal et al. use stimulating electrodes that are attached to the skull (i.e., non-penetrating), as clearly shown in Figure 1 of that paper. Describing this as "intracranial stimulation" and equating it with direct cortical stimulation used to induce CSD in Leão (1944), is incorrect and misleading.

      (2) The authors ask for guidance in improving the clarity of their messaging around this and cite the Abstract as clearly stating what they view as the novelty here. Yet, the key section of the Abstract reads: "We show that CSD drives increased expression of the immediate early gene Fos, a key marker of neuronal plasticity, and is associated with factors that predict positive ECT therapeutic outcome. Our results suggest that the therapeutic efficacy of ECT may be mediated by CSD. This challenges the seizure-centric model and implies that CSD, a currently unmonitored neurophysiological event, may serve as a more relevant biomarker for predicting and optimizing therapeutic outcomes of ECT." The Reviewer finds this quite misleading, as it suggests (at least to this Reviewer) that the authors view themselves as showing, for the first time, that ECT generates CSD and that this may be the therapeutic mechanism of ECT, while not clearly defining what is new. Instead, one could perhaps more accurately write: "Prior work has shown that ECT generates CSD and that CSD is associated with increases in the immediate early gene Fos, a key marker of neuronal plasticity. We provide data suggesting that ECT-associated CSD in fact drives this increased Fos expression. Our results imply the hypothesis that the therapeutic effect of ECT may be mediated by CSD-induced Fos expression and could serve as a biomarker of ECT efficacy."

      (3) In claiming novelty, the authors further assert that Rosenthal et al. provided "no confirmation beyond the observation of the calcium event that this is indeed a CSD." This also appears to be incorrect or misleading. Rosenthal et al. show simultaneous calcium and intrinsic optical imaging of hemoglobin dynamics associated with the propagating event. Rosenthal et al. also show DC-coupled electrophysiology demonstrating the slow potential shift characteristic of CSD using auricular ECS (the very stimulation model that the authors argue distinguishes their work from Rosenthal et al.). Hence, speculation that the "cortical seizure Rosenthal and colleagues find is a methodological artifact" is unsupported and should be removed.

      (4) In the rebuttal, the authors clarify the novelty of their manuscript as "that ECS drives CSD dependent immediately early gene expression (Fos)." The Reviewer agrees with this, yet as summarized in the bullet point, finds this novelty to be relatively limited and not particularly surprising, given that Fos is robustly induced in association with CSD in various contexts. This novelty could be markedly enhanced by linking increased Fos expression mechanistically to the therapeutic efficacy of ECT. Instead, any relevance of the CSD-dependent Fos induction is unclear, and statements to that effect, while interesting, are hypothetical, yet follow logically from Rosenthal et al. Any association between CSD occurrence and clinical outcome in human patients could establish predictive value but would also not prove causality, as suggested. Also unclear is how Fos might be used as a biomarker in humans, and how this might be valuable as a biomarker beyond CSD itself (which can actually be recorded), is unclear. The authors should state these limitations.

      Minor:

      (1) Regarding the microprism preparation, the added three-week postoperative interval is helpful toward alleviating this concern. Still the absence of spontaneous CSD does not address the fact that cortical incision and the CSD likely elicited during implantation could alter subsequent susceptibility to experimentally evoked CSD. This should simply be acknowledged as a limitation.

      (2) The Reviewer recognizes that the authors have added Tables S3 and S4 detailing stimulation parameters. However, the response that properties of the traveling calcium event do not depend on stimulation charge does not resolve the concern regarding analyses in which stimulation parameters are used to predict whether CSD occurs. Manually selecting stimulation parameters rather than selecting them randomly or systematically may confound these analyses, an issue which could simply be acknowledged.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1) The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy.

      This is correct, our claim is that ECS drives spreading depression dependent IEG expression in cortex. We are not claiming this is the only consequence of ECS. The extracortical response to ECS could contribute to the treatment, and we discuss this in the hippocampus specifically (see Discussion). In current clinical practice, cortical surface EEG is used as a biomarker for predicting treatment outcome. Given that the CSD is necessary to drive Fos expression in cortex, we think it is warranted to speculate that monitoring CSD during ECT is worth exploring as a potentially superior biomarker of treatment outcome. Especially given that this is readily feasible (Discussion).

      Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      The idea that Fos is a marker of neuronal activity is outdated. The primary correlate of Fos expression is neuronal plasticity and learning-related circuit modifications, rather than just neuronal activity – see e.g. (Bolhuis et al., 2001; de Hoz et al., 2018; Fleischmann et al., 2003; Kimpo and Doupe, 1997; Mahringer et al., 2022, 2019; Nakadate et al., 2012; Roy et al., 2016; Ryan et al., 2015; Tanaka et al., 2018; Tyssowski et al., 2018; Watanabe et al., 1996; Yassin et al., 2010). We have now referenced some of those articles in the manuscript (Introduction). But more importantly, our claim is that the CSD is necessary to drive Fos. The fact that CSD can occur in a single hemisphere allows, within a brain, to control for direct ECS response and phase III oscillations. In unilateral CSD, contralateral hemispheres displayed Fos levels in the cortex that were similar to sham mice. Investigating the exact plasticity pathway that is triggered by the CSD is an interesting question we are pursuing in follow-up work, but does not influence our conclusions here.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      The reviewer is correct in that postictal suppression is currently one of the best correlates of positive clinical outcome used in the clinic. However, note that EEG-based correlates are generally weak predictors of treatment outcomes: even recently identified correlates are unreliable between patients cohorts and account for less than 60% of the clinical outcome (Francis-Taylor et al., 2020; Scangos et al., 2019). See also our response to comment (2) of reviewer 3 on this topic.

      The primary problem in the interpretation of postictal suppression of EEG activity is that it is unclear what the source of the EEG signal is in this case. During phase III oscillation following the CSD, cortex is silent and all of the EEG is likely driven by thalamic input to cortex. Note, source localization in EEG does not help as the source of the signal is likely the thalamic axons in cortex (Huels et al., 2023) (or their postsynaptic, and in this case subthreshold, effects). Whatever the source of the cortical surface EEG may be, we know that following CSD it cannot be cortical activity, hence, whatever is driving postictal suppression is also not cortical. If we had to speculate, postictal suppression is likely an exhaustion of thalamus in attempting to drive cortex without getting any excitatory feedback. During an epileptic seizure, thalamus oscillates at delta frequencies and cortex responds at each peak of the oscillation, generating ictal spikes on the EEG (Meeren et al., 2002; Polack et al., 2009). The fact that ictal spikes are missing in an “optimal” phase III oscillation is consistent with our findings that the CSD completely silences cortex for minutes. Thus, during an ECS-driven phase III oscillation, cortex does not respond to thalamic input and at some point thalamus runs out of energy to maintain the oscillation - it is likely the combination of thalamus running out of energy and a silent cortex that gives rise to postictal suppression. Note we do find an asymmetry for the mid-oscillation amplitude (Figure 6A), which in this interpretation can be explained by the thalamo-cortical circuit attempting to compensate for the absence of cortical feedback by increasing its drive to the CSD-affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      This is a misunderstanding. We claim (and show) that the CSD is necessary to drive immediate early gene expression following ECS. Based on this we speculate that the CSD “may serve as a more relevant biomarker for predicting and optimizing therapeutic outcomes of ECT” than EEG biomarkers.

      We do not “state [that there is] a beneficial role of CSD” – we actually address the possibility that SD may not be therapeutically beneficial (Discussion). The suggestion of investigating behavioral consequences of CSD in mice is interesting, but likely not the most relevant follow-up work. This is for two reasons:

      (1) Very different from humans, most mouse behavior does not depend on cortex (Kawai et al., 2015; Pandey et al., 2026). Thus, any potential behavioral consequences of ECS in mice are difficult to interpret in the context of clinical relevance. We don’t think there is any known biomarker based on mouse behavior that has any direct translational value to psychiatric treatments. The reviewer may disagree with this assessment but just consider that no mouse behavior assay has ever been instrumental to the development of a novel psychiatric treatment.

      (2) The more direct – and simpler approach – is to directly test whether CSD correlates with treatment outcomes in patients. Our primary aim with this paper is to inspire exactly this. Note, we are currently pursuing this as well, of course. Monitoring CSDs in humans during ECT is likely possible with fNIRS or other hemodynamic-based measurements (Discussion).

      However, to reiterate – the claim of the manuscript is exactly as stated in the title. Thus, demonstrating the clinical relevance of the CSD is outside of the scope of the current manuscript, but will be trivial to prove by measuring treatment outcomes while routinely monitoring patients for CSD.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

      Reviewer #1 (Recommendations for the authors):

      The results of sham tests are shown only for the immunohistochemical data. However, results from control experiments should be provided for the EEG and calcium data.

      Sham data are included in Author response image 1. We are not sure how they support our conclusions and have left them out of the manuscript. A sham stimulation just characterizes ongoing activity under anesthesia. The more direct comparison is the difference between pre- and post-ECS activity. Baseline widefield calcium imaging data is already shown in Figure 1, 2, 6; baseline mouse EEG data is already shown in Figure 1 and 6. Baseline EEG in patients is not available in the present dataset.

      Author response image 1.

      Sham ECS does not result in phase III oscillations or a CSD. (A) Representative spectrograms (top) and raw traces (bottom) of concurrent EEG (left) and widefield calcium imaging (right) during a sham ECS session. (B) Population raster plot (top) and two example activity traces of neurons (bottom) during a sham ECS two-photon recording.

      Moreover, baseline (pre-ECS) calcium and EEG activity should be shown in Figures 1 and 6 to correctly assess the changes induced by stimulation.

      We are not sure we understand as pre-ECS data are already shown in Figures 1 and 6. We assume the reviewer may mean “more baseline data should be shown”? We have extended the time scale of the relevant panels in Figure 1, 2 and 6 to include 30 s instead of 10 s of pre-ECS data.

      Using the term 'oscillations/phase III oscillations' instead of 'seizures' throughout the text is confusing because the word covers a wide range of brain oscillations - from normal to pathological ones.

      This is indeed confusing – but the confusion arises from the often imprecise usage of the term “seizure” in the ECT literature. The EEG response to ECS is not equivalent to that of an epileptic seizure. The description of a “phase III oscillation” is a more precise description of the EEG signature. It was introduced by (Brumback and Staton, 1982) and is not our terminology. We dedicate an entire paragraph in the introduction to this distinction. We are not sure how to make this clearer in the current manuscript. Continuing to describe the EEG response to ECS as “seizure” is inaccurate, and mechanistically misleading, especially given that cortex is silent during the phase III oscillation following the CSD (i.e. cortex cannot be “seizing” during this phase of the EEG response, see our point above on postictal EEG suppression).

      The presentation of clinical data is scarce and unclear. e.g., the authors claim that the oscillation frequency is similar in mice and humans (lines 206-208). However, in Figure 1F, oscillations during the early post-ECS phase have twice the frequency (6-7 Hz) in the patient EEG recording (6-7 oscillations per 1 second) compared to the mouse EEG (3 Hz, i.e., 6 oscillations per 2 seconds during 73-75 s). It seems a bit odd because Figure 1H shows that the frequency does not exceed 5 Hz, even in patients.

      Please excuse, this is our mistake. The x-axis label in the human data of Figure 1F was incorrect. It should have read 1 to 3 s, not 1 to 2 s. Window sizes were of course matched between mice and human data and are all 2 s. The mistake is now corrected. The frequencies in patients and mice are compared in Figure 1H and are not different.

      A part of the discussion (lines 627-635) is based on factual inaccuracy: cortical SD cannot invade the hippocampus in the in vivo brain (only in slices), although SD can occur in the hippocampus in response to generalized seizures.

      It would be helpful if the reviewer would back up this claim with references. Short of this we are left to speculate - we suspect the reviewer may be referring to earlier work in the rat cortex, that claimed that a cortical SD can only invade hippocampus if glia has been impaired (Largo et al., 1997). This is now an outdated model: there are more recent reports of CSD invading the hippocampus (Bahari et al., 2020; Bonaccini Calia et al., 2022). If the reviewer has specific concerns with any of these papers, we would be happy to discuss in more detail, but the reviewers’ claim seems unfounded here.

      It is reasonable to expect that bilateral stimulation produces bilateral calcium/SD waves. In Rosenthal's experiments, this situation was most common. In the present study, bilateral ECS triggers mostly unilateral calcium waves. Do you have any idea what the reasons for the result are? As stimulation parameters (polarity, intensity, and frequency) have been shown to control the occurrence of SD waves, their unilateral pattern suggests non-uniform stimulation conditions. I am curious whether the uni- or bilateral pattern of calcium/SD waves depends on the ECS parameters.

      The main difference between our patient and mouse data is the polarity of stimulation. For patient data, the polarity of the current alternates with every pulse, while in our mouse data the stimulation was always right unipolar. This is discussed when we mention the asymmetry of SD in our data (Results, Methods) - we suspect the reviewer may have missed this.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      (1) The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      This is correct, our work confirms that the calcium event following ECS is a CSD. The main claim of our paper is that ECS drives immediate early gene expression in a CSD-dependent manner.

      However, note that the Rosenthal paper concluded that the calcium event is a CSD while providing only little evidence for that claim – mind you, we agree with their interpretation, but provide more evidence for the conclusion. We are happy to discuss the limitations of the Rosenthal paper as highlighted below more prominently in the manuscript, if the reviewer thinks this would be helpful, but we think that is likely not necessary.

      Briefly, in the data presented in (Rosenthal et al., 2025), the only support for the calcium response being a CSD is the speed of propagation and the DC shift reported in extracellular recordings. We add to this by showing that the spread of the calcium event follows the pattern expected by a CSD through cortical layers (Figure 4), results in heterogeneous returns to baseline calcium levels (Figure 4), causes vasoconstriction (Figure S5), travels at the speed expected of a CSD regardless of stimulation parameters (Figure S3) and causes Fos expression (Figure 5).

      Most importantly, the ECS as used by Rosenthal and colleagues is not a mouse model of ECT, in the sense that is not a scalp electrical stimulation. The stimulation method is fundamentally different between our two articles: (Rosenthal et al., 2025) implanted stimulating electrodes directly in the mice’s dorsal cranium. Direct cortical stimulation is well known to be able to cause CSD (Leao, 1944). However, it is unclear whether direct cortical stimulation is a useful model for ECT. We suspect that Rosenthal and colleagues were led to believe that auricular stimulation does not work because they saw no evidence of a cortical seizure in calcium recordings following stimulation. (See discussion on the confusion of phase III oscillation and seizure in the ECT literature.) We suspect that EEG recordings following direct cortical stimulation would reveal a very different EEG pattern from that observed in patients. This highlights the importance and novelty of the comparison of mouse and human EEG we present in Figure 1. Note, in our auricular stimulation preparation we do not observe any seizure-like activity in cortex that lasts beyond the stimulation (compare Rosenthal’s Figure 1E vs Figure 1F here). Given the EEG similarity we show between mice and patients, we suspect the cortical seizure Rosenthal and colleagues find is a methodological artifact. This is puzzling to us, as Rosenthal and colleagues do briefly mention a single mouse example with an extracellular electrophysiology recording compatible with CSD following auricular stimulation (Supplementary Figure 2).

      Thus, not only is it necessary to add evidence to the interpretation that the ECS-driven calcium event is indeed a CSD, but also that it can be triggered by a stimulation method that successfully replicates the known EEG response of human patients.

      (2) The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSDCa2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

      We have added additional references to the CSD-Fos literature in the discussion. Regarding the role of Fos as a marker of plasticity rather than activity, we discuss this point in our reply to comment (1) of reviewer 1. Concerning the expression of other genes, we already referenced the TRKB/BDNF pathways (Discussion). We have now added references on RNA-seq following CSD, which show that Fos is one of the most differentially expressed genes following CSD (Dell’Orco et al., 2023). Regarding the use of Fos knockout/knockdown lines, please see our reply to the reviewer’s comment (4).

      Reviewer #2 (Recommendations for the authors):

      (3) The Results and Discussion sections should be revised to better reflect the prior discovery of CSD following ECT in rodents and initial evidence in humans, as discussed in the first point in the Weaknesses section above.

      The prior discovery of the fact that ECS can drive CSD is first mentioned in the third sentence of the abstract “However, this view is challenged by the recent finding that electroconvulsive stimulation (ECS) can trigger a cortical spreading depression (CSD).” (The abstract has no references, but the introduction should make it clear what is meant).

      There is an entire paragraph of the introduction discussing the Rosenthal results.

      The first time we discuss our results (end of the Introduction) we say: “Consistent with previous work (Rosenthal et al., 2025), we observed a slow travelling calcium event that appeared to be a CSD.” The Rosenthal paper is cited in 8 times in total throughout the manuscript.

      We are unsure what the reviewer is asking us to do here. As mentioned above, if any revision is warranted regarding that reference, it should be to clarify that the Rosenthal paper used direct intracranial stimulation rather than ECS, and did not fully confirm the calcium event as a CSD – but they should be credited for finding the first preliminary evidence for CSD following ECS in patients.

      (4) To support the authors' statement that "CSD is the primary driver of plasticity is based on its role in driving Fos expression", additional experiments with Fos knockdown or knockout (or alternative interventions) are needed.

      Our main claim is that CSD causes Fos expression in the cortex following ECS. The argument that Fos can be used as a marker of plasticity follows from the literature, not from the experiments done here – this would require a form of functional plasticity measurement, which is outside the scope of this paper (and might not be the most relevant direction, see our reply to comment (3) of reviewer 1). The only observation that a Fos genetic manipulation would give us is the lack or reduction of Fos expression following CSD, which would be orthogonal to the points we make here.

      (5) It would be helpful to use a more specific term than "Ca2+ event" in Figure 2D and throughout the related Results section. It is assumed that this is the large propagating Ca2+ event attributed to CSD, but the terminology is important, as all the other events in the recording (including during Pre-ECT and the Direct ECT period) are also Ca2+ events.

      We have added a clear definition of what we mean on first usage (Results). The reason to call it a calcium event, and not a CSD, is that we did not want to jump to conclusions. We do think that the event is a CSD, but conclusive proof of that is still lacking (in both our work and that of Rosenthal). It is conceivable that the event propagates via a different mechanism than a CSD.

      (6) The authors should comment on differences among rodent models of ECT stimulation, especially with direct and ear clip methods, as discussed in the context of translational value (Theilmann et al., 2014).

      The Theilmann 2014 paper compares ECS delivered via auricular stimulation and intracranial electrodes. They compare the two stimulation methods, but use stimulation currents, total charges, and stimulation duration that were not matched. In the case of stimulation current, those used in auricular stimulation are approximately 8 fold higher than what they use for intracranial stimulation. Moreover, the stimulations were performed in awake rats. This is scientifically - and ethically, even for 2014 - questionable in light of the fact that this is aimed at developing a model for ECT, which is always done under general anesthesia. They conclude that cortical stimulation has less adverse effects and is more effective in reducing immobility in a forced swim test. The confound in the interpretation of these results is that all stimulation was performed in awake animals. The reason this is no longer done in humans is that it is extremely painful. Direct cortical stimulation is likely much less painful (for the same reason TMS is less painful). This would explain their findings of increased adverse effects with auricular electrodes. Given the differences in stimulation parameters used, the difficulty of calibrating equivalent doses of auricular and intracortical stimulations, as well as the small effect sizes reported, we don’t think the results allow for any solid conclusions as to which method is more effective in reducing immobility in a forced swim test. However, even if one would assume the intracortical ECS is more ‘effective’, this is hardly relevant, as we are interested in using mouse ECS as a model for human ECT. One could speculate that intracranial ECS might also exhibit higher clinical benefit than surface ECS in patients, but that is not the scope of our research, and probably not clinically relevant. Finally, while the authors do include EEG recordings in the rat – these were not compared to patient recordings, and from visual inspection do not resemble patient EEG recordings that we have seen. We would argue that the best rodent model of ECT stimulation is the one that triggers neuronal activity most similar to that observed in patients. We have added a brief discussion of these points to the corresponding Methods section.

      Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      This appears incorrect. It was known that direct cortical stimulation (as was done in the Rosenthal paper) can drive a calcium event that resembles a CSD. It was also known that CSD can drive Fos expression. What was not known is that the Fos expression driven by ECS is fully explained by the calcium event (putative CSD). This is particularly relevant as most people still erroneously assume Fos is a marker or neuronal activity, and prior work has come to the conclusion that ECT does not drive, but likely downregulates Fos expression (Calais et al., 2013; Morinobu et al., 1995; Park et al., 2014; Winston et al., 1990). We have added a more prominent discussion section on this point.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      The importance of Figure 1 is to demonstrate that ECS delivered with auricular electrodes in mice causes an EEG signature that is very similar to that seen in patients. We do not claim we are the first to describe characteristics of phase III oscillations in either patients or mice (we have added the Murakami reference to the manuscript). The reply to comments (1) and (6) of reviewer 2 highlights why this comparison is so important – it was not done in the Rosenthal paper the reviewer mentions below, and it is not clear whether the intracortical stimulation used there even drives a comparable EEG response (given the calcium activity shown, we suspect the answer is no). To the best of our knowledge this direct comparison is novel – but again the key novelty of our work we highlight is the one described in the title.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      It has indeed been demonstrated that ECS delivered using intracranial electrical stimulation can trigger CSD-like events (e.g. Rosenthal et al.). However, the fact that localized intracranial electrical stimulation can trigger a CSD has been shown quite a while ago already (see e.g. (Leao, 1944)). This is not the case for surface stimulation the way it is done in ECT and the way we do it. Nevertheless, note we give full credit to the Rosenthal work for making this connection. Our main contribution – as highlighted by the title – is showing that the Fos expression driven by ECS is fully explained by the CSD.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1.

      If the reviewer has references for this claim, we would be happy to add to the manuscript – we are not aware of any such work. As far as we are aware, this is still an area under active investigation – calcium signals correlate (locally) strongly with shank recordings (Wei et al., 2020), but how this translates into an EEG signal is speculative.

      It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      The reviewer may be jumping to conclusions here. It is correct that this has been shown for a CSD. The more important question (and the reason we did this experiment) is whether the calcium event triggered by ECS is indeed a CSD. We try to be careful to describe it as a calcium event in the results (mind you the calcium event in the Rosenthal is very likely a CSD as their intracortical stimulation (‘ECS’) is likely equivalent to the electrical stimulations used in the discovery of the CSD). We then perform a series of comparisons to see whether the calcium event has the known characteristics of a CSD – and we conclude everything we test is consistent with it being a CSD. Note, once again, that we do not claim novelty in any of this – the primary novelty is the link between ECS, Fos and CSD.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      That is correct. The question however is how much of the Fos expression is explained by CSD. Prior work that has looked at Fos expression in response to ECS has found that ECS results in a slight reduction of Fos expression (Calais et al., 2013; Park et al., 2014). We suspect this is the result of not triggering a CSD. We do not claim to have discovered that CSD induces Fos expression. The novel contribution, which is the main claim of the paper, is that in the context of ECS the entirety of the cortical Fos expression can be explained by the CSD. This links the relative contributions of multiple components of the ECS response to a known marker of neuronal plasticity.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      There is probably a misunderstanding here. We argue and show that there is no cortical seizure following ECS – that is why we refer to the EEG responses as phase III oscillations (characteristic of a silent cortex). And there is little prior evidence that the calcium events triggered by ECS are indeed a CSD (we think this is likely the case, but demonstrating this conclusively will require further work). If the reviewer has references for this, we would be happy to discuss. Note, the Rosenthal et al. paper just assumes (probably correctly) that they are a CSD, but does not demonstrate this. We don’t fully demonstrate this either, we just provide additional evidence. But once again, the novelty is in the title of the manuscript, and we do not claim any other novelty to the best of our reading of our manuscript. If the reviewer has a particularly misleading passage in mind, we are happy to rephrase.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      We are not sure what the reviewer means by “cFOS expression is nonspecific to plasticity”. Does the reviewer mean Fos has other roles beside driving neuronal plasticity? That is very likely correct, but it is unclear how that is relevant. Fos is a key driver of a number of neuronal plasticity pathways (Chaudhuri et al., 2000; Cohen and Greenberg, 2008; Cruz et al., 2015; Durchdewald et al., 2009). Which exact pathways are driven by ECS is an interesting question that we are currently pursuing in follow-up work. Given what we know about Fos expression, it is probably well within reasonable bounds to conclude that Fos increases result in neuronal plasticity. Fos expression directly drives network plasticity (Yap et al., 2021): "our findings indicate that Fos expression has an instructive role in orchestrating persistent circuit modifications”. Likewise, Fos induction is strictly required for experience-dependent representational plasticity during learning (de Hoz et al., 2018): “locally blocking c-Fos expression caused […] decreased cortical experience-dependent plasticity, without affecting baseline excitability or basic auditory processing.”. Calcium activity explains about 15% of the variance of Fos expression (Mahringer et al., 2022). This is likely driven by the correlation between activity and plasticity, not by a direct necessity for Fos expression to maintain neuronal activity. This is consistent with the finding that Fos as a transcription factor does not function to maintain spiking activity; it is part of the gene-regulatory machinery that converts patterned synaptic input into lasting plastic change. Fos is induced by NMDA/Ca2+, ERK, and CREB/Elk signaling rather than by firing alone (Fields et al., 1997; Xia et al., 1996), and those same pathways are required for the transcriptional program that stabilizes long-term potentiation and other durable synaptic modifications (Davis et al., 2000). As an AP-1 transcription factor, Fos drives downstream gene expression, so its appearance is better read as entry into a plasticity-related nuclear program than as a measure of ongoing excitability (Minatohara et al., 2015; Morgan and Curran, 1991; Sheng and Greenberg, 1990). That interpretation is consistent with our work showing that early Fos expression preferentially marks neurons that later undergo the strongest learning-related functional changes (Mahringer et al., 2019) and tracks functional reorganization during learning rather than simple recent activation (Mahringer et al., 2019).

      The link between CSD and therapeutic effect is a speculation we make in the manuscript, not a conclusion. We are of course in the process of performing follow-up work to test whether CSD in patients is a better predictor of treatment outcome than EEG based metrics – no experiment we can do in mice will be able to test the hypothesis that CSD is the mediator of the therapeutic benefit of ECT.

      Regarding the power of EEG metrics to predict therapeutic outcomes, we fully agree with the reviewer. The literature on the reliability of EEG metrics computed from phase III oscillation data is controversial and noisy – this is something we establish in the introduction to motivate our research into other biological processes that could explain how ECT works. Note however, this is certainly not a fringe view – see e.g. comment (2) of reviewer 1: “Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy“ We think the reason for this is that a CSD is necessary for therapeutic benefit, but only has minor effects on the phase III oscillation – note this is a hypothesis based on our results that is trivial to test in patients (which are in currently investigating).

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      Whenever possible, we use mice for multiple experiments in the lab. This is done to reduce the total number of mice used for experiments for ethical reasons. For the mice in question, ChrimsonR was injected in the retrosplenial cortex (Methods), which was originally done to stimulate locally the axons projecting in the imaging area (primary visual cortex here). Thus the labelling is very sparse, and perfectly compatible with two-photon GCaMP6f imaging. Fluorescence emission is far too weak to activate ChrimsonR.

      Qualitatively, one can mentally compare the light power we use to activate optogenetic tools, which tends to be blindingly bright (one shouldn’t look into the optogenetic stimulation laser), with the fluorescence emission from two-photon imaging of calcium indicators, which tends to be barely visible by eye.

      Quantitatively, one can estimate this as follows: At 510nm emission, each photon carries an energy of about . Assuming a neuron that strongly expresses GCaMP6f under two-photon excitation would emit a very high 10<sup>7</sup>photons per second (Har-Gil et al., 2018), its total emission power would be P = N<sub>photons</sub> * E ≈ 4pW. Assuming this is spread over the surface of the cell (sphere with 10 µm diameter), this translates to . This is several orders of magnitude below the value of 1 mW/mm<sup>2</sup> irradiance required to activate ChrimsonR modestly at peak absorption, which is 80nm away from GCaMP6f emission (Klapoetke et al., 2014). Note that this would be true even when ChrimsonR is injected at the site of imaging (Vasilevskaya and Keller, 2026).

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      We suspect that the question is driven by a misunderstanding. While it is likely that the implantation triggers a CSD (and likely so does a standard two-photon window implantation), the implantation surgery and the experiments/imaging are separated by at least 3 weeks (Methods). We have never observed spontaneous CSDs in the days and weeks following an implantation surgery.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      We have added additional details as requested by the reviewer (Methods).

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      We have added two tables (Table S3 and Table S4) that displays the stimulation parameters used for each figure, as well as the distribution of parameters for mice and patients. More importantly, the properties of the travelling calcium event do not depend on the stimulation charge (Figure S3), which removes this confounding factor and supports the idea that the calcium event is a CSD.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      That is correct – only the cortical expression of Fos depends on the cortical CSD (see our reply to comment (1) of reviewer 1). Quantification of the whole-brain Fos expression following ECS is outside the scope of the manuscript, but it is something we are currently pursuing for separate publication. We have now reworked the section describing the Fos expression to make it clear that we are only talking about cortical expression of Fos (Results).

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      That was indeed our speculation based on ongoing work on TMS in the lab – but it was poorly phrased and unnecessary. We have rephrased.

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

      It is speculative indeed – but the speculation is not unfounded and has been made previously. The primary evidence is correlative in that schizophrenia is characterized by a reduction in coupling between thalamus and frontal areas of cortex – see e.g. (Giraldo-Chica and Woodward, 2017; Vinogradov et al., 2023). We have argued in previous work that combining this with computational models of psychosis (Sterzer et al., 2018), it is not unreasonable to speculate that a reduction in the coupling between thalamus and frontal areas of cortex could explain psychosis (Keller and Sterzer, 2024). The value of that speculation here is that it forms a testable hypothesis for the mechanism of action of ECT. Nevertheless, we have rephrased the statement slightly to make it clearer that this is still speculation.

      Reviewer #3 (Recommendations for the authors):

      The authors should design/execute an experiment(s) to attempt to prove causality between CSD, calcium influx, cFOS expression, and therapeutic effect. At a minimum, the authors need to revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field.

      Regarding novelty – see discussion above.

      Regarding therapeutic effects – we are in the process of testing this in patients. Using CSD measurements during ECT to test whether CSD is a better predictor of treatment outcome. I suspect that this will, however, require many more labs to come to firm conclusions. Our aim here is to inspire these experiments. That is why we speculate about clinical relevance in the abstract and the discussion.

      References

      Bahari, F., Ssentongo, P., Liu, J., Kimbugwe, J., Curay, C., Schiff, S.J., Gluckman, B.J., 2020. Seizure-associated spreading depression is a major feature of ictal events in two animal models of chronic epilepsy. https://doi.org/10.1101/455519

      Bolhuis, J.J., Hetebrij, E., Den Boer-Visser, A.M., De Groot, J.H., Zijlstra, G.G.O., 2001. Localized immediate early gene expression related to the strength of song learning in socially reared zebra finches. European Journal of Neuroscience 13, 2165–2170. https://doi.org/10.1046/j.0953816x.2001.01588.x

      Bonaccini Calia, A., Masvidal-Codina, E., Smith, T.M., Schäfer, N., Rathore, D., Rodríguez-Lucas, E., Illa, X., De la Cruz, J.M., Del Corro, E., Prats-Alfonso, E., Viana, D., Bousquet, J., Hébert, C., Martínez-Aguilar, J., Sperling, J.R., Drummond, M., Halder, A., Dodd, A., Barr, K., Savage, S., Fornell, J., Sort, J., Guger, C., Villa, R., Kostarelos, K., Wykes, R.C., Guimerà-Brunet, A., Garrido, J.A., 2022. Full-bandwidth electrophysiology of seizures and epileptiform activity enabled by flexible graphene microtransistor depth neural probes. Nat. Nanotechnol. 17, 301–309. https://doi.org/10.1038/s41565-021-01041-9

      Brumback, R.A., Staton, R.D., 1982. The Electroencephalographic Pattern during Electroconvulsive Therapy. Clinical Electroencephalography 13, 148–153. https://doi.org/10.1177/155005948201300306

      Calais, J.B., Valvassori, S.S., Resende, W.R., Feier, G., Athié, M.C.P., Ribeiro, S., Gattaz, W.F., Quevedo, J., Ojopi, E.B., 2013. Long-term decrease in immediate early gene expression after electroconvulsive seizures. J Neural Transm (Vienna) 120, 259–266. https://doi.org/10.1007/s00702-012-0861-4

      Chaudhuri, A., Zangenehpour, S., Rahbar-Dehgan, F., Ye, F., 2000. Molecular maps of neural activity and quiescence. Acta Neurobiol Exp (Wars) 60, 403–410. https://doi.org/10.55782/ane-2000-1359

      Cohen, S., Greenberg, M.E., 2008. Communication between the synapse and the nucleus in neuronal development, plasticity, and disease. Annu Rev Cell Dev Biol 24, 183–209. https://doi.org/10.1146/annurev.cellbio.24.110707.175235

      Cruz, F.C., Javier Rubio, F., Hope, B.T., 2015. Using c-fos to study neuronal ensembles in corticostriatal circuitry of addiction. Brain Res 1628, 157–173. https://doi.org/10.1016/j.brainres.2014.11.005

      Davis, S., Vanhoutte, P., Pages, C., Caboche, J., Laroche, S., 2000. The MAPK/ERK cascade targets both Elk-1 and cAMP response element-binding protein to control long-term potentiation-dependent gene expression in the dentate gyrus in vivo. J Neurosci 20, 4563–4572. https://doi.org/10.1523/JNEUROSCI.20-12-04563.2000

      de Hoz, L., Gierej, D., Lioudyno, V., Jaworski, J., Blazejczyk, M., Cruces-Solís, H., Beroun, A., Lebitko, T., Nikolaev, T., Knapska, E., Nelken, I., Kaczmarek, L., 2018. Blocking c-Fos Expression Reveals the Role of Auditory Cortex Plasticity in Sound Frequency Discrimination Learning. Cereb Cortex 28, 1645–1655. https://doi.org/10.1093/cercor/bhx060

      Dell’Orco, M., Weisend, J.E., Perrone-Bizzozero, N.I., Carlson, A.P., Morton, R.A., Linsenbardt, D.N., Shuttleworth, C.W., 2023. Repetitive spreading depolarization induces gene expression changes related to synaptic plasticity and neuroprotective pathways. Front. Cell. Neurosci. 17. https://doi.org/10.3389/fncel.2023.1292661

      Durchdewald, M., Angel, P., Hess, J., 2009. The transcription factor Fos: a Janus-type regulator in health and disease. Histol Histopathol 24, 1451–1461. https://doi.org/10.14670/HH-24.1451

      Fields, R.D., Eshete, F., Stevens, B., Itoh, K., 1997. Action potential-dependent regulation of gene expression: temporal specificity in ca2+, cAMP-responsive element-binding proteins, and mitogen-activated protein kinase signaling. J Neurosci 17, 7252–7266. https://doi.org/10.1523/JNEUROSCI.17-19-07252.1997

      Fleischmann, A., Hvalby, O., Jensen, V., Strekalova, T., Zacher, C., Layer, L.E., Kvello, A., Reschke, M., Spanagel, R., Sprengel, R., Wagner, E.F., Gass, P., 2003. Impaired Long-Term Memory and NR2AType NMDA Receptor-Dependent Synaptic Plasticity in Mice Lacking c-Fos in the CNS. J. Neurosci. 23, 9116–9122. https://doi.org/10.1523/JNEUROSCI.23-27-09116.2003

      Francis-Taylor, R., Ophel, G., Martin, D., Loo, C., 2020. The ictal EEG in ECT: A systematic review of the relationships between ictal features, ECT technique, seizure threshold and outcomes. Brain Stimulation 13, 1644–1654. https://doi.org/10.1016/j.brs.2020.09.009

      Giraldo-Chica, M., Woodward, N.D., 2017. Review of thalamocortical resting-state fMRI studies in schizophrenia. Schizophr Res 180, 58–63. https://doi.org/10.1016/j.schres.2016.08.005

      Har-Gil, H., Golgher, L., Israel, S., Kain, D., Cheshnovsky, O., Parnas, M., Blinder, P., 2018. PySight: plug and play photon counting for fast continuous volumetric intravital microscopy. Optica 5, 1104. https://doi.org/10.1364/OPTICA.5.001104

      Huels, E.R., Kafashan, M., Hickman, L.B., Ching, S., Lin, N., Lenze, E.J., Farber, N.B., Avidan, M.S., Hogan, R.E., Palanca, B.J.A., 2023. Central-positive complexes in ECT-induced seizures: Possible evidence for thalamocortical mechanisms. Clin Neurophysiol 146, 77–86. https://doi.org/10.1016/j.clinph.2022.11.015

      Kawai, R., Markman, T., Poddar, R., Ko, R., Fantana, A.L., Dhawale, A.K., Kampff, A.R., Ölveczky, B.P., 2015. Motor cortex is required for learning but not for executing a motor skill. Neuron 86, 800– 812. https://doi.org/10.1016/j.neuron.2015.03.024

      Keller, G.B., Sterzer, P., 2024. Predictive Processing: A Circuit Approach to Psychosis. Annu Rev Neurosci 47, 85–101. https://doi.org/10.1146/annurev-neuro-100223-121214

      Kimpo, R.R., Doupe, A.J., 1997. FOS Is Induced by Singing in Distinct Neuronal Populations in a Motor Network. Neuron 18, 315–325. https://doi.org/10.1016/S0896-6273(00)80271-8

      Klapoetke, N.C., Murata, Y., Kim, S.S., Pulver, S.R., Birdsey-Benson, A., Cho, Y.K., Morimoto, T.K., Chuong, A.S., Carpenter, E.J., Tian, Z., Wang, J., Xie, Y., Yan, Z., Zhang, Y., Chow, B.Y., Surek, B., Melkonian, M., Jayaraman, V., Constantine-Paton, M., Wong, G.K.-S., Boyden, E.S., 2014. Independent optical excitation of distinct neural populations. Nat Methods 11, 338–346. https://doi.org/10.1038/nmeth.2836

      Largo, C., Ibarz, J.M., Herreras, O., 1997. Effects of the Gliotoxin Fluorocitrate on Spreading Depression and Glial Membrane Potential in Rat Brain In Situ. Journal of Neurophysiology 78, 295–307. https://doi.org/10.1152/jn.1997.78.1.295

      Leao, A.A.P., 1944. Spreading depression of activity in the cerebral cortex. Journal of Neurophysiology 7, 359–390. https://doi.org/10.1152/jn.1944.7.6.359

      Mahringer, D., Petersen, A.V., Fiser, A., Okuno, H., Bito, H., Perrier, J.-F., Keller, G.B., 2019. Expression of c-Fos and Arc in hippocampal region CA1 marks neurons that exhibit learning-related activity changes. https://doi.org/10.1101/644526

      Mahringer, D., Zmarz, P., Okuno, H., Bito, H., Keller, G.B., 2022. Functional correlates of immediate early gene expression in mouse visual cortex. Peer Community Journal 2. https://doi.org/10.24072/pcjournal.156

      Meeren, H.K.M., Pijn, J.P.M., Luijtelaar, E.L.J.M.V., Coenen, A.M.L., Silva, F.H.L. da, 2002. Cortical Focus Drives Widespread Corticothalamic Networks during Spontaneous Absence Seizures in Rats. J. Neurosci. 22, 1480–1495. https://doi.org/10.1523/JNEUROSCI.22-04-01480.2002

      Minatohara, K., Akiyoshi, M., Okuno, H., 2015. Role of Immediate-Early Genes in Synaptic Plasticity and Neuronal Ensembles Underlying the Memory Trace. Front Mol Neurosci 8, 78. https://doi.org/10.3389/fnmol.2015.00078

      Morgan, J.I., Curran, T., 1991. Stimulus-transcription coupling in the nervous system: involvement of the inducible proto-oncogenes fos and jun. Annu Rev Neurosci 14, 421–451. https://doi.org/10.1146/annurev.ne.14.030191.002225

      Morinobu, S., Nibuya, M., Duman, R.S., 1995. Chronic antidepressant treatment down-regulates the induction of c-fos mRNA in response to acute stress in rat frontal cortex. Neuropsychopharmacology 12, 221–228. https://doi.org/10.1016/0893-133X(94)00067-A

      Nakadate, K., Imamura, K., Watanabe, Y., 2012. Effects of monocular deprivation on the spatial pattern of visually induced expression of c-Fos protein. Neuroscience 202, 17–28. https://doi.org/10.1016/j.neuroscience.2011.12.004

      Pandey, A., Kang, S., Pacchiarini, N., Wyszynska, H., Masseri, Z., O’Neill, J., Honey, R.C., Fox, K., 2026. Secondary Somatosensory Cortex Is Required for Learning but Not Execution of a Tactile Discrimination. Eur J Neurosci 63, e70390. https://doi.org/10.1111/ejn.70390

      Park, H.G., Yu, H.S., Park, S., Ahn, Y.M., Kim, Y.S., Kim, S.H., 2014. Repeated treatment with electroconvulsive seizures induces HDAC2 expression and down-regulation of NMDA receptorrelated genes through histone deacetylation in the rat frontal cortex. Int J Neuropsychopharmacol 17, 1487–1500. https://doi.org/10.1017/S1461145714000248

      Polack, P.-O., Mahon, S., Chavez, M., Charpier, S., 2009. Inactivation of the Somatosensory Cortex Prevents Paroxysmal Oscillations in Cortical and Related Thalamic Neurons in a Genetic Model of Absence Epilepsy. Cereb Cortex 19, 2078–2091. https://doi.org/10.1093/cercor/bhn237

      Rosenthal, Z.P., Majeski, J.B., Somarowthu, A., Quinn, D.K., Lindquist, B.E., Putt, M.E., Karaj, A., Favilla, C.G., Baker, W.B., Hosseini, G., Rodriguez, J.P., Cristancho, M.A., Sheline, Y.I., William Shuttleworth, C., Abbott, C.C., Yodh, A.G., Goldberg, E.M., 2025. Electroconvulsive therapy generates a postictal wave of spreading depolarization in mice and humans. Nat Commun 16, 4619. https://doi.org/10.1038/s41467-025-59900-1

      Roy, D.S., Arons, A., Mitchell, T.I., Pignatelli, M., Ryan, T.J., Tonegawa, S., 2016. Memory retrieval by activating engram cells in mouse models of early Alzheimer’s disease. Nature 531, 508–512. https://doi.org/10.1038/nature17172

      Ryan, T.J., Roy, D.S., Pignatelli, M., Arons, A., Tonegawa, S., 2015. Memory. Engram cells retain memory under retrograde amnesia. Science 348, 1007–1013. https://doi.org/10.1126/science.aaa5542

      Scangos, K.W., Weiner, R.D., Coffey, C.E., Krystal, A.D., 2019. An electrophysiological biomarker that may predict treatment response to ECT. J ECT 35, 95–102. https://doi.org/10.1097/YCT.0000000000000557

      Sheng, M., Greenberg, M.E., 1990. The regulation and function of c-fos and other immediate early genes in the nervous system. Neuron 4, 477–485. https://doi.org/10.1016/0896-6273(90)90106-p

      Sterzer, P., Adams, R.A., Fletcher, P., Frith, C., Lawrie, S.M., Muckli, L., Petrovic, P., Uhlhaas, P., Voss, M., Corlett, P.R., 2018. The Predictive Coding Account of Psychosis. Biol Psychiatry 84, 634–643. https://doi.org/10.1016/j.biopsych.2018.05.015

      Tanaka, K.Z., He, H., Tomar, A., Niisato, K., Huang, A.J.Y., McHugh, T.J., 2018. The hippocampal engram maps experience but not place. Science 361, 392–397. https://doi.org/10.1126/science.aat5397

      Tyssowski, K.M., DeStefino, N.R., Cho, J.-H., Dunn, C.J., Poston, R.G., Carty, C.E., Jones, R.D., Chang, S.M., Romeo, P., Wurzelmann, M.K., Ward, J.M., Andermann, M.L., Saha, R.N., Dudek, S.M., Gray, J.M., 2018. Different Neuronal Activity Patterns Induce Different Gene Expression Programs. Neuron 98, 530-546.e11. https://doi.org/10.1016/j.neuron.2018.04.001

      Vasilevskaya, A., Keller, G.B., 2026. A functional influence based circuit motif that constrains the set of plausible algorithms of cortical function. eLife 15. https://doi.org/10.7554/eLife.110827.1

      Vinogradov, S., Chafee, M.V., Lee, E., Morishita, H., 2023. Psychosis spectrum illnesses as disorders of prefrontal critical period plasticity. Neuropsychopharmacology 48, 168–185. https://doi.org/10.1038/s41386-022-01451-w

      Watanabe, Y., Johnson, R.S., Butler, L.S., Binder, D.K., Spiegelman, B.M., Papaioannou, V.E., McNamara, J.O., 1996. Null Mutation of c-fos Impairs Structural and Functional Plasticities in the Kindling Model of Epilepsy. J Neurosci 16, 3827–3836. https://doi.org/10.1523/JNEUROSCI.16-1203827.1996

      Wei, Z., Lin, B.-J., Chen, T.-W., Daie, K., Svoboda, K., Druckmann, S., 2020. A comparison of neuronal population dynamics measured with calcium imaging and electrophysiology. PLoS Comput Biol 16, e1008198. https://doi.org/10.1371/journal.pcbi.1008198

      Winston, S.M., Hayward, M.D., Nestler, E.J., Duman, R.S., 1990. Chronic electroconvulsive seizures down-regulate expression of the immediate-early genes c-fos and c-jun in rat cerebral cortex. J Neurochem 54, 1920–1925. https://doi.org/10.1111/j.1471-4159.1990.tb04892.x

      Xia, Z., Dudek, H., Miranti, C.K., Greenberg, M.E., 1996. Calcium influx via the NMDA receptor induces immediate early gene transcription by a MAP kinase/ERK-dependent mechanism. J Neurosci 16, 5425–5436. https://doi.org/10.1523/JNEUROSCI.16-17-05425.1996

      Yap, E.-L., Pettit, N.L., Davis, C.P., Nagy, M.A., Harmin, D.A., Golden, E., Dagliyan, O., Lin, C., Rudolph, S., Sharma, N., Griffith, E.C., Harvey, C.D., Greenberg, M.E., 2021. Bidirectional perisomatic inhibitory plasticity of a Fos neuronal network. Nature 590, 115–121. https://doi.org/10.1038/s41586-020-3031-0

      Yassin, L., Benedetti, B.L., Jouhanneau, J.-S., Wen, J.A., Poulet, J.F.A., Barth, A.L., 2010. An embedded subnetwork of highly active neurons in the neocortex. Neuron 68, 1043–1050. https://doi.org/10.1016/j.neuron.2010.11.029

    1. eLife Assessment

      This study presents a valuable contribution to comparative cognitive neuroscience by directly mapping functional homologues of the human multiple-demand network in macaques using a matched spatial maze task. However, the evidence is incomplete due to methodological asymmetries in the task design that appear to affect the results in ways that are not fully analyzed and discussed. The work will be of interest to researchers studying the evolution of cognitive control and cross-species neuroimaging.

    2. Reviewer #1 (Public review):

      Summary:

      The "multiple-demand" (MD) system is a well-known finding of human brain imaging and is thought to play a central role in cognitive control. To directly compare the MD system in humans and monkeys, Mione et al. used functional magnetic resonance imaging to measure whole brain activation in a multi-step saccadic maze task. In humans, the authors found a distributed pattern of brain activity close match to the canonical MD network and extending to adjacent regions of dorsal attention and other networks. While there was good correspondence between monkey and human data, differences were also notable in lateral frontal cortex, dorsal parietal cortex, and sensorimotor cortex.

      Strengths:

      Though previous data hint at a corresponding network in the macaque, there has been no direct comparison to human data. This study provides a direct cross-species comparison with whole-brain data of fMRI, and the findings suggest an extended and strongly interconnected brain network recruited by increased cognitive challenge.

      Weaknesses:

      In previous human imaging, the MD system is defined by overlapping activation for many kinds of cognitive demand. In the present work, however, the authors used just a single task. Although there is some overlap between putative monkey MD network and canonical MD network identified in human imaging, it should be cautious to link current findings to MD system based on limited task events.

    3. Reviewer #2 (Public review):

      Summary:

      Mione et al. aim to resolve a long-standing question in comparative neuroscience: whether the macaque brain contains a functional analogue to the distributed human multiple-demand (MD) network. To address this, the authors employ a direct cross-species fMRI comparison using a multi-step saccadic maze task in humans and a simplified two-step version in macaques. By contrasting goal-directed navigation against a control condition that requires similar motor responses but no strategic planning, the study isolates the neural signatures of cognitive control across species.

      Strengths:

      The most compelling aspect of this work is its methodological alignment. Previous attempts to compare these systems often relied on comparisons of human BOLD signals and macaque single-unit recordings. By running parallel fMRI protocols, the authors establish a shared measurement basis that allows for a more direct comparison. The resulting activation maps demonstrate conserved network topology in dorsolateral and dorsomedial frontal cortex and insula.

      Weaknesses:

      However, there is concerning inter-individual variability in the macaque data, along with a notable design asymmetry between the human and monkey tasks. In the human experiment, 2-, 4-, and 6-step trials were mixed within the same blocks, whereas macaques performed only 2-step problems. This design difference likely places human participants in a state of sustained proactive cognitive control (Braver, 2012), as they must remain prepared for more demanding trials at any moment. This elevated baseline arousal may potentially inflate MD network activation even during the simpler 2-step trials in humans, making direct comparisons with the macaque data difficult.

      The authors acknowledge this concern and note that similarities between species survive despite this procedural difference. However, their own individual-level data raise doubts about how robust these "similarities" really are. A defining feature of the MD system is consistent recruitment across individuals performing the same task. In this study, however, the two monkeys show markedly different activation maps in parietal cortex. The conjunction map (Supplementary Figure 6) does not support the claim that lateral and medial parietal regions are commonly activated across both animals at any meaningful statistical threshold, even when using a permissive threshold of |t| > 1.5, corresponding to approximately p ≈ 0.13 - 0.14 (two-tailed).

      This suggests that the group-level parietal activations are largely driven by a single animal, raising questions about whether these regions can reliably be considered as part of the monkey MD network. Notably, in Premereur et al. (2018), the only other direct macaque fMRI study of a cognitively demanding task, their conjunction map between the two animals also failed to show common parietal activation during task switching. Taken together, these findings suggest that either the monkey MD network is not as similar to its human counterpart as claimed, or the monkey version of the maze task was simply not sufficiently challenging to fully engage parietal MD nodes.

      Because of the design asymmetry described above, it remains difficult to determine which of these possibilities is correct (a genuine species difference versus insufficient task demands in macaques). The authors' current claims treat the group-level activations, which appear largely driven by one animal, as reliable evidence for homology, but without consistent individual-level support, such conclusions remain tentative.

      In summary, the authors demonstrate functional correspondence between human and macaque cognitive control networks in dorsolateral and dorsomedial frontal regions. However, assertions that this extends to lateral and medial parietal cortex are not consistently supported at the individual level. While the overall conclusions are plausible, they would be significantly strengthened by demonstrating consistent activation across both animals in all claimed MD regions, ideally with a properly thresholded conjunction map.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The "multiple-demand" (MD) system is a well-known finding of human brain imaging and is thought to play a central role in cognitive control. To directly compare the MD system in humans and monkeys, Mione et al. used functional magnetic resonance imaging to measure whole-brain activation in a multi-step saccadic maze task. In humans, the authors found a distributed pattern of brain activity close match to the canonical MD network and extends to adjacent regions of dorsal attention and other networks. While there was good correspondence between monkey and human data, differences were also notable in the lateral frontal cortex, the dorsal parietal cortex, and the sensorimotor cortex.

      Strengths:

      Though previous data hint at a corresponding network in the macaque, there has been no direct comparison to human data. This study provides a direct cross-species comparison with whole-brain data from fMRI, and the findings suggest an extended and strongly interconnected brain network recruited by increased cognitive challenge.

      Weaknesses:

      In previous human imaging, the MD system is defined by overlapping activation for many kinds of cognitive demands. In the present work, however, the authors used just a single task. Although there is some overlap between the putative monkey MD network and the canonical MD network identified in human imaging, there should be caution in linking current findings to the MD system based on limited task events.

      In the Discussion, we acknowledge the limitation of using a single task. With this one task, however, the canonical MD network is clearly shown in our human data. Accompanying activation, especially of the canonical dorsal attention network, likely reflects the specific spatial demands of the maze task. In the monkey data, much of this dorsal attention activity is not seen (e.g. superior and medial parietal), likely reflecting limited power. Instead, there is distributed overlap with the previous limited monkey studies than have contrasted higher with lower cognitive demand. Though we agree that task-specific activations may contribute to our results, these arguments suggest that, in large part, our method does successfully identify a distributed set of multiple-demand regions. At the same time we acknowledge the desirability of further work to examine a wider range of task demands.

      Reviewer #1 (Recommendations for the authors):

      (1) Though the whole-brain data obtained by fMRI can provide a direct comparison between species, a single cognitive task might be insufficient to link the findings with the MD system. A cognitively challenging task likely activates multiple regions of the MD system; however, it may also recruit some task-specific regions, which do not belong to the canonical MD network. Furthermore, this is probably the reason that the dorsal attention network showed the strongest activation rather than the core MD for the current visuospatial maze task.

      In the Discussion, we acknowledge the limitation of using a single task (p. 15-16), and the likely contribution of task-specific activations to our data, especially involving the dorsal attention network (p. 15).

      (2) Ideally, a meta-analysis recruiting more fMRI studies on humans and monkeys when they perform various similar cognitive tasks may strengthen the evidence that there is a comparable MD system between species.

      Many human meta-analyses, of course, show the common MD system. For monkeys, however, at least to our knowledge, there are insufficient studies for a similar meta-analysis. Instead we discuss overlaps between the current activation findings and two previous studies of respectively antisaccades and task switching (p. 15), suggesting that, in monkey as in human, there is convergence for different kinds of demand.

      (3) Different from human subjects, monkeys usually require substantial training before fMRI scanning. More details about the training of the two monkeys should be given. In addition, training may reshape cognitive task activations. This potential impact should also be discussed.

      We now address this point in the Discussion (p. 18). As we note, similar results for the two species apparently survive even large differences in protocol. A training summary has been added to Methods (p. 27).

      (4) According to the description in the text, the two monkeys have obvious differences in cognitive task activations. It is necessary to show individual-level brain activations as well.

      Individual results and a conjunction map are shown in Supplementary Figure 6 (see accompanying text on p. 13). As expected, the conjunction map had substantial similarity to the findings from the two animals combined. Similarities and differences between animals are addressed in the Discussion (p. 17).

      (5) Based on the current research content, the title seems to be too general.

      For the reasons given above (see point 2), we think our title is reasonable.

      Reviewer #2 (Public review):

      Summary:

      Mione et al. aim to resolve a long-standing question in comparative neuroscience: whether the macaque brain contains a functional analogue to the distributed human multiple demand (MD) network. To address this, the authors employ a direct cross-species fMRI comparison using a multi-step saccadic maze task in humans and a simplified two-step version in macaques. By contrasting goal-directed navigation against a control condition that requires similar motor responses but no strategic planning, the study isolates the neural signatures of cognitive control across species.

      Strengths:

      The most compelling aspect of this work is its methodological alignment. Previous attempts to compare these systems often relied on comparisons of human BOLD signals and macaque single-unit recordings. By running parallel fMRI protocols, the authors establish a shared measurement basis that allows for a more direct comparison. The resulting activation maps clearly demonstrate conserved network topology across dorsomedial frontal, lateral, and medial parietal, and insula cortices. Combining these results with recent research on functional and structural connectivity further supports the idea that these networks evolved across species and provides a helpful starting point for future comparative studies. The findings will be highly useful for researchers investigating the evolutionary origins of domain-general cognitive control, as well as for neuroimaging methodologists developing cross-species alignment pipelines.

      Weaknesses:

      However, there are several differences in how the two groups were studied that make it harder to compare the results precisely. The human task mixed 2-, 4-, and 6-step trials within the same experimental blocks, whereas macaques performed only 2-step trials. This design difference likely places human participants in a state of sustained proactive cognitive control (Braver, 2012), as they must remain prepared for highly demanding trials at any moment. This elevated baseline arousal may artificially inflate MD network activation during the simpler 2-step trials in humans, making direct magnitude comparisons with the macaque data difficult.

      This is a reasonable concern, which we note in the Discussion. Crucially, as we point out, similarities between species appear to survive this and other differences in procedure.

      Additionally, the general linear model combined correct and error trials into a single regressor. Given that macaques exhibited substantially higher error rates, this approach risks diluting task-specific planning signals with activity related to error monitoring and reward prediction errors. The preprocessing pipeline also applied a 4 mm full-width half-maximum smoothing kernel to macaque data acquired at 1.5 mm resolution. Relative to the smaller size of the macaque brain, this kernel is quite large and likely blurs fine-grained topographical distinctions. This may partly explain why the macaque lateral frontal cortex shows a single dorsal activation patch rather than multiple discrete patches seen in humans.

      These are also reasonable concerns. To address them, we have added supplementary analyses using (a) only correct trials, and (b) a smaller smoothing kernel (see Supplementary Figure 3 and 4). In both cases, there is some loss of power, but otherwise similar results.

      Furthermore, there is concerning inter-individual variability in the macaque data. Normally, a functional network like the MD system is identified by consistent activation across all individuals. In this study, however, the two monkeys show substantially different activation maps and behavioral patterns. This lack of consistency renders the group-level results questionable, as it is unclear whether the group-level map represents a unified biological system or merely an average of disparate individual maps.

      Results for individual animals have now been supplemented with a conjunction map (Supplementary Figure 6). As expected, the conjunction map is similar to the map obtained by pooling data from the two animals.

      Finally, the subcortical activations shown in Figure 7 require more precise anatomical localization to confidently distinguish cerebellar nodes from adjacent brainstem structures.

      The slices shown in Figure 7 have been amended to better show the detail of activation outside cerebral cortex, in particular in cerebellum.

      The authors demonstrate a broad functional correspondence between human and macaque cognitive control networks, moving the field beyond speculative homology. The data suggest that an extended, interconnected network is recruited by cognitive challenge in both species; however, the strength of this claim is limited by the inter-individual variability and methodological constraints noted above. Assertions of precise topological equivalence should therefore be tempered. The absence of ventrolateral prefrontal and strong dorsal parietal activations in the macaque group analysis may reflect genuine biological differences, but could also stem from limited statistical power, excessive smoothing, or task design asymmetries. While the overall conclusions are plausible, they would be significantly strengthened by a more explicit discussion of these limitations and additional analytical clarifications regarding individual-level consistency.

      Indeed, as we note in the Discussion, there is a strong possibility that ventrolateral frontal and dorsal parietal activation are missing in our monkey data because of limited statistical power. Further work would be needed to address these possible limitations.

      Reviewer #2 (Recommendations for the authors):

      (1) Please discuss how the mixed-difficulty block design in humans may induce a state of proactive cognitive control that elevates baseline MD activation compared to the fixed 2step macaque condition. A supplementary analysis comparing early versus late block 2-step trials in humans would help clarify whether activation magnitudes reflect sustained task-set maintenance or transient trial demands.

      We now note the potential importance of this in the Discussion (p. 18), and point out that similarities between species appear to survive this difference in procedure.

      (2) Please clarify the rationale for combining correct and error trials in a single regressor.

      Given the higher macaque error rates, please provide a supplementary GLM restricted to correct trials only for the monkey data. This will demonstrate whether the core activation topography remains consistent when error-related signals are excluded.

      A new analysis addresses this point (Supplementary Figure 3). As we note, restricting analysis to correctly-completed problems somewhat reduces power but leaves major features of the results intact.

      (3) Please justify the use of a 4 mm FWHM smoothing kernel for macaque data, given cortical thickness and brain size differences. If feasible, re-run the analysis with a smaller kernel or surface-based smoothing to assess whether finer topographical distinctions emerge in lateral frontal and parietal cortices.

      This new analysis has also been run (Supplementary Figure 4), again with reduced power but major features of the results intact. We also refer to previous work indicating choice of either 3 mm or 4 mm smoothing for macaque fMRI (p. 13), approximately matching the smoothing needed to align electrophysiological and fMRI maps (Issa et al., 2013, J.Neurosci.).

      (4) Please detail how the human '2-step problems only' analysis in Supplementary Figure 1 was specified in the GLM. Explicitly state whether 4- and 6-step trials were modeled as separate regressors of no interest to prevent hemodynamic bleed-over from contaminating the 2-step beta weights.

      Indeed, 4- and 6-step trials were removed using regressors of no interest, as now specified in Methods (p. 24).

      (5) Please provide higher-resolution slices or probabilistic atlas overlays for the cerebellar and subcortical activations in Figure 7. This will help clearly distinguish cerebellar hemispheres from adjacent brainstem structures and ensure anatomical labeling is accurate.

      To address this question, additional slices have been added to a revised Figure 7, in particular adding detail to cerebellar activation.

      (6) Please address the substantial inter-individual variability in the macaque data, particularly regarding Monkey B's performance. Even after excluding five poor-performing sessions, Monkey B only reached about 65% accuracy on the first step of the maze task. This suggests that the animal was largely guessing rather than following a strategic, goal directed plan. The notably higher accuracy on step 2 (>80%) could simply be a selection bias artifact, as the task terminates immediately following step 1 errors, meaning step 2 trials only occur when step 1 was already correct. Including an animal that likely did not fully grasp the overarching task rule in a cohort of N=2 raises serious concerns about signal dilution and the reliability of the group-level maps. Please explicitly justify why Monkey B's data were retained despite these performance concerns, or consider excluding this animal and acquiring data from a third, behaviorally reliable subject. At minimum, report a conjunction map showing regions strictly active in both animals alongside the combined analysis, and present individual subject maps more prominently to improve transparency.

      Though we appreciate this concern, we do not think the data indicate that monkey B failed to use the maze goal to constrain choices. As shown in Figure 3, the great majority of errors were timing errors (mostly not waiting for go signal). Excluding these, for step 1, of choices directed to one of the two available alternatives, 88% were correct (Figure 4, compare “correct” with “wrong open location”).

      As noted above, we have added a conjunction map (Supplementary Figure 6) to our previous presentation of individual data. We note (p. 13) that, as expected, this map strongly resembles results from the combined-animal analysis.

    1. eLife Assessment

      This valuable descriptive study describes the expression of a developmentally relevant transcription factor in the adult Tribolium brain. The evidence supporting the claims is convincing and based on a very detailed and rigorous analysis of light microscopy data, which, however, lacks single-cell resolution. This neuroanatomical study is of interest to the field of insect neural development and neuroscience.

    2. Reviewer #1 (Public review):

      [Editors' note: The reviewing editor has assessed the revisions. The authors have addressed the previous minor concerns of the reviewers, added more details on the generation of the brainbow constructs and have made the image stacks available on public repositories.]

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

    3. Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a non-standard laboratory organism.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

      Weaknesses:

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects.

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single-cell labelling”, which admittedly limits precision.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function.

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an effect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately.

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question but a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative.

      Indeed, we do not reach single-cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content.

      We also think that combining our transgenic line with dopamine expression was sufficient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism.

      Thank you for this encouraging comment.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      Comments:

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect:

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/ or motor learning in motor neurons https://pubmed.ncbi.nlm.nih.gov/38779314/ or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest.

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal thresholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal-directed navigation, respectively."

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors claim that several dop-positive cell types resemble cell types described in Drosophila - to be able to better compare the two, it would be great to have the supplement and main figure pictures in one figure.

      We have added Fig. 11 from the main text to suppl. Fig. 3 for comparison

      (2) In Figure 11 and the corresponding supplement, it would be great to have some landmarks to better understand the expression patterns.

      In the legend, we have now referred to Figs. 3, 4 and 5 for depictions of these neurons within the neuropil reconstructions

      (3) For a better understanding of neurotransmitter expression, is it possible to figure out if dopamine is rather coexpressed with Glut- or ChaT-positive neurons?

      Very interesting idea. Unfortunately, the first author of the study has graduated and left the lab, such that we are unable to add this piece of information.

    1. eLife Assessment

      This valuable study uses technically challenging long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The study provides solid evidence that ORNs have a circadian pattern of spontaneous firing that is Orco-dependent although Orco transcript abundance does not itself show circadian rhythmicity, and that cAMP can modulate Orco-dependent activity. Together with the computational work, the study proposes the provocative hypothesis that an Orco-centered post-translational feedback loop generates the circadian rhythm of spontaneous firing. Nevertheless, this mechanistic interpretation will require direct testing in future work.

    2. Joint Public Review:

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

      We thank the referees and the editors for their efforts and their positive feedback. As requested, we clarified throughout our manuscript that our proposal of an Orco-centered post-transcriptionally controlled feedback loop in the plasma membrane, which we term PTFL clock, is a novel hypothesis that needs to be tested in future experiments.

    1. eLife Assessment

      This important study combines a five-year field experiment with a meta-analysis to quantify effects of grazing on ecosystem CO₂ fluxes as modified by wetness. The evidence supporting the conclusions is convincing, but several aspects of the methods and discussion would benefit from clarification to improve reproducibility and interpretability.

    2. Reviewer #3 (Public review):

      Combining a five-year field experiment with a global meta-analysis, Wu et al. investigate how grazing intensity influences ecosystem carbon dioxide (CO₂) fluxes in grasslands and how these effects are regulated by environmental conditions such as grazing duration, wetness index, and soil temperature and moisture responses.

      The authors show that the response of net ecosystem productivity (NEP) to light grazing shifts from negative to positive along a wetness gradient, whereas heavy grazing consistently suppresses NEP across wetness conditions. Importantly, this pattern is supported by both the field experiment and the meta-analysis, suggesting that may help maintain moderate levels of grazing can potentially enhance both plant productivity and carbon sequestration under favorable moisture conditions.

      The integration of experimental data with a global synthesis is a particular strength of the study, allowing the authors to evaluate grazing impacts across both temporal variability (precipitation fluctuations in the field experiment) and spatial variability (wetness gradients across global grasslands). Overall, the conclusions are well supported by the data.

      Overall, the principal conclusions are generally supported by the reported results, and the study provides useful evidence that the effects of grazing on grassland carbon cycling depend on both grazing intensity and environmental context. The comparison between field and synthesis results is potentially valuable for understanding why grazing effects vary among grassland systems. However, some aspects of the meta-analysis remain insufficiently documented. In particular, the study-selection numbers presented in the new PRISMA diagram require clarification, and the manuscript does not clearly explain how individual response ratios were weighted when estimating the pooled effect sizes. Resolving these reporting and methodological issues would improve the reproducibility and interpretation of the synthesis.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study combines a five-year field experiment with a meta-analysis to quantify the effects of grazing on ecosystem CO<sub>2</sub> fluxes as modified by wetness. The results of this study are potentially valuable, but the Methods description is incomplete and compromises reproducibility and interpretability. A major caveat of this study is that the assessment of CO<sub>2</sub> fluxes in time and space is not complete.

      We appreciate the comments. The incomplete assessment of CO<sub>2</sub> fluxes in time and space has been added in the Limitations and implications for future study Section in the revised manuscript. Method description has been revised as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      Lines 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      Lines 468-470, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that was not monitored.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study integrates long-term (7th-11th year) ecosystem CO<sub>2</sub> flux measurements from a continuously grazed typical steppe in Inner Mongolia with a global meta-analysis of grazing experiments (585 observations) to systematically investigate how the wetness index modulates the effects of grazing intensity on net ecosystem productivity (NEP) in grasslands.

      Key findings include:

      (1) Heavy grazing significantly reduced gross primary productivity (GPP) and ecosystem respiration (ER) in the typical steppe but did not significantly affect NEP.

      (2) Globally, grazing significantly reduced GPP, ER, and NEP, and a higher wetness index and aboveground biomass (AGB) enhance the positive response of NEP to grazing.

      (3) Under light and moderate grazing, the response of NEP to grazing was positively correlated with the wetness index, a relationship potentially mediated by a higher plant relative growth rate (RGR) in wetter years.

      This study holds considerable practical significance. In the context of global change, comprehending the regulatory function of water is of great importance for the adaptive management of grasslands.

      Strengths:

      (1) The study cleverly combines long-term in situ observations (revealing temporal dynamics and potential mechanisms) with a global meta-analysis (testing the generality of patterns), forming a complete evidence chain from "process understanding" to "pattern verification".

      (2) Focusing on the 7th to 11th years of grazing treatments avoids the common "initial disturbance effects" observed in short-term grazing experiments and truly captures the steady-state response of the ecosystem after it has reached a new equilibrium.

      (3) By analyzing plant relative growth rate (RGR) and aboveground biomass (AGB), this study provides mechanistic clues regarding how the wetness index modulates grazing effects. The finding that plant compensatory growth is enhanced in wetter years is a crucial pathway explaining the variation in the NEP response.

      (4) The meta-analysis systematically integrates published literature with a substantial sample size (585 observations) and broad geographical coverage (Figure 1B).

      (5) The finding that light grazing promotes carbon sinks in wet years, while heavy grazing reduces productivity even in wet years, has direct implications for adaptive grassland management under climate change.

      Thank you for the comments. We would like to express our gratitude for your positive and insightful evaluation of this manuscript.

      Weaknesses:

      Although the paper does have strengths in principle, there are weaknesses and areas for improvement in the paper. In particular:

      (1) Incomplete mechanistic chain: While RGR and AGB data offer valuable clues for mechanistic interpretation, the causal chain from "increased wetness → higher RGR → maintained NEP" remains incomplete. Key intermediate processes, such as soil moisture dynamics, nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, are absent, making the mechanistic explanation somewhat speculative.

      Thank you for the comments. The data of the soil moisture dynamics, composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain. However, the structural equation modeling could not be set up if C<sub>3</sub> and C<sub>4</sub> plant communities were included (Fig. S8). Leaf photosynthetic physiological parameters were not collected in this study. We have made following changes in the revised manuscript:

      Lines 339-341, page 11: “In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Lines 472-474, page 15: “Furthermore, key intermediate processes, such as soil nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, should also be investigated in the future.”

      (2) Inadequate exploration of heterogeneity: The global meta-analysis reveals significant effect heterogeneity. Currently, only a few moderators, such as wetness index, precipitation, and temperature, are analyzed. Potential moderators, including grazing history, grassland type (typical steppe/alpine meadow/savanna), livestock type (cattle/sheep/mixed), and soil type, are not adequately explored.

      Thank you for the comments. Grazing duration was presented in Fig. S13. The heterogeneity analysis of grassland types and livestock types was supplemented as Fig. S11 and Fig. S12 in the Supplementary documents. We have also made following changes in the revised manuscript:

      Lines 367-370, page 12: “Grazing decreased GPP, ER and NEP in desert and temperate grassland, but had no significant effect on GPP, ER and NEP in alpine grassland and savanna (Fig. S11). In addition, cattle and sheep grazing decreased GPP and NEP, but livestock mixed grazing did not show significant effect on GPP, ER and NEP (Fig. S12).”

      (3) Integration of long-term experiment and meta-analysis could be tighter.

      Thank you for the comments. Integration of long-term experiment and meta-analysis has been revised in the revised manuscript as follows:

      Lines 449-463, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term grazing field experiments revealed the mechanisms underlying the effects of grazing on ecosystem CO<sub>2</sub> fluxes. Global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analyses consistently revealed that wetness was an important factor in regulating the ecosystem CO<sub>2</sub> fluxes in response to grazing. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP (Morgan et al., 2016; Owensby et al., 2006). Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness modulated the effects of grazing on NEP in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      Reviewer #2 (Public review):

      Overgrazing in grasslands is a widespread, enormous problem, and its consequences need more research attention.

      However, I have some concerns about the methods in this manuscript.

      (1) The plots were grazed from May to September. For the CO<sub>2</sub> measurements, it states growing season, and that data was collected on sunny days from 9 to 11 am. However, annual values are reported, and no details on how these were calculated. With no measurements outside of 9 to 11 am, and only on non-sunny days, and only during the growing season, I question how accurate the annual numbers are.

      Thank you for the comments. The ecosystem CO<sub>2</sub> fluxes were measured during growing season; the description of annual CO<sub>2</sub> have been revised to “growing-season” CO<sub>2</sub>. The ecosystem CO<sub>2</sub> fluxes were measured on sunny days during the growing season (May-September) between 9:00 and 11:00 a.m. It can represent typical daytime carbon flux dynamics during the growing season (Fan et al., 2011; Li et al., 2010). In addition, previous studies showed that the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. represented the average ecosystem CO<sub>2</sub> fluxes of the day (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), so we use the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. to calculate the total ecosystem CO<sub>2</sub> fluxes of the day. We have also revised it in the limitation for future study section. We have made following changes in the revised manuscript:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      Lines 468-472, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that were not monitored. Multiple time-point sampling should be adopted; the sampling frequency and points should also be increased in the future field study.”

      (2) This study uses one chamber, 0.5 by 0.5 m, for a total area of 0.25 m<sup>2</sup>. This is pretty small, and with no replication, I would like to see more detail on how these are placed, especially when one of the dominant species is a bunchgrass. Sampling on or off a bunchgrass will give very different readings, and neither will be representative of the entire plot. The same for the soil temp and water measurements, there is no detail on how many replicates are in each plot, and working randomly does not work when there are spatial patterns caused by the bunchgrasses in the plots.

      Thank you for the comments. Two replicates of the chambers in each plot were set, and their specific sites were labeled in Fig. S2. The study area was located in a typical steppe dominated by Stipa grandis and Leymus chinensis. Therefore, we selected quadrats mainly containing these two species. We have made following changes in the revised manuscript:

      Line 186-187, page 6: “In each plot, the chambers of two replicates were measured.”

      (3) MAP and MAT are important in this study, and no mention is given if this data is from this site or a neighboring site, and if so, what the distance to this location is.

      Thank you for the comments. The data of MAP and MAT were provided in the revised manuscript as follows:

      Lines 193-195, page 7: “The precipitation and air temperature were collected from the Xilinhot Meteorological Observation Station, which was near the experimental site in this field experiment.”

      (4) Vegetation was sampled in five locations in each plot. I note no detail on the belowground biomass, how many reps, what area, or which depth? This essential data is missing. The same goes for the RGR, and here the authors talk about grazing events. Whereas earlier in section 2.2, it implies continuous grazing every day during the growing season. RGR needs more details, such as how many times, its replication, etc.

      Thank you for the comments. The detailed method for how to collect belowground biomass and how to calculate the RGR has been supplemented in the revised manuscript as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      (5) Hypothesis 2 is vague: "act as key factors". A more precise hypothesis would be better, otherwise it just leads to p-hacking.

      Thank you for the comments. The misleading description has been deleted.

      (6) More generally, neither hypothesis is really based on the introduction, as the authors report mixed results in the literature. Hence, the more specific hypothesis 1 reads like it is HARKED, based on the results that the authors found. This is not a good research practice. See the following papers on this topic:

      Murphy & Aquinis. 2019. HARKing: How bad can cherry-picking and question trolling produce bias in published results? Journal of Business and Psychology 34:1-17

      Bishop. 2019 Rein in the four horses of irreproducibility. Nature 568: 435

      Fraser et al. 2018. Questionable research practices in ecology and evolution. Plos One 13:7

      Parker et al. 2016. Transparency in ecology and evolution: real problems, real solutions. Trends in ecology and evolution 31:9

      Thank you for the suggestion. The hypotheses have been deleted in the revised manuscript to avoid HARKED mistakes.

      (7) Figure 2 mentioned n=3, which is correct. However, the variance around the means in the figures is tiny, which raises questions about the replication that is actually used here. I would like to see the entire statistics tables, including DF and sample sizes, included as an appendix, so that the reader can evaluate this much more.

      Thank you for the comments. In this study, the data presented in Fig.2 were expressed as mean± standard error (SE):

      In addition, we have provided the complete statistical table in Supplementary Table S2 according to the suggestion of the reviewer.

      (8) Figure 3, now only low, medium, and high grazing are reported as changes from the control. This can be misleading as the reader can't see how the control varies along the various gradients. I suggest including a figure with the raw data for each as an appendix. In addition, the sample size in 3d, e, f is much higher, and it looks to me like the authors used both the reps and years together. This is not good practice. In addition, full statistics tables should be included in the appendix.

      Thank you for the comments. The relationships between ecosystem CO<sub>2</sub> fluxes and wetness index have been supplemented in Fig. S6. We have also provided the complete statistical table in Supplementary Table S3.

      (9) In Table S1, since there are already a number of recent meta-analyses on this topic, I think there needs to be a stronger justification for this one.

      Thank you for the comments. The results of our five-year field monitoring experiments showed that carbon fluxes and their components were significantly influenced by the wetness index. Therefore, we explored whether such effects also occur in grazing experiments at the global scale. The results demonstrated that similar patterns indeed exist worldwide, indicating that our meta-analysis is meaningful and necessary. In addition, previous studies have addressed related topics, net ecosystem productivity has mostly been treated as an auxiliary variable rather than the primary focus of investigation (Jiang et al., 2020; Shi et al., 2022; Zhang et al., 2022; Zhou et al., 2019). In previous studies, meta-analysis literatures on grazing and NEP were limited and not adequate. Our meta-analysis is more comprehensive and reflects the reliability of the results. We have made following changes in the revised manuscript:

      Lines 491-492, page 16: “In previous studies, meta-analysis literatures of effects of grazing intensities on NEP were limited and not adequate.”

      Lines 512-514, page 16: “In summary, the meta-analysis of this study presents the first comprehensive assessment of how annual wetness index affects the response of ecosystem CO<sub>2</sub> fluxes to grazing across global grasslands.”

      (10) Regarding the grazing-induced CO<sub>2</sub> fluxes, these are only based on the plants and do not incorporate the animal CO<sub>2</sub> flux, nor the animal litter CO<sub>2</sub> flux, as they were penned at night outside the plots. Thus, the grazing impact is inflated in the data reported here and does not really represent GPP, NET, or RE. This needs to be written about in the discussion section.

      Thank you for the comments. The objective of this study was to evaluate the carbon exchange processes from vegetation and soil under grazing disturbance. Thus, the grazing effects reported in this study were the vegetation-soil ecosystem scale rather than the net carbon balance of the entire grazing system. Consequently, our conclusions regarding the effects of grazing on grassland ecosystem carbon exchange remain reliable and ecologically meaningful.

      We have made following changes in the manuscript:

      Lines 483-490, page 16: “There are also limitations in evaluating the effects of grazing on grassland-livestock ecosystem CO<sub>2</sub> fluxes in this study. Specifically, carbon emissions derived from animal respiration were not included in this study. Therefore, the results of the field experiment and meta-analysis in this study should not be interpreted as a complete carbon budget assessment of the grassland-livestock ecosystem. Our primary objective was to evaluate the carbon exchange processes between vegetation and soil under grazing disturbance. Thus, the grazing effects reported in the field experiment and meta-analysis of this study were responses at the vegetation- soil ecosystem scale rather than the net carbon balance of the entire grazing system.”

      Reviewer #3 (Public review):

      Combining a five-year field experiment with a global meta-analysis, Wu et al. investigate how grazing intensity influences ecosystem carbon dioxide (CO<sub>2</sub>) fluxes in grasslands and how these effects are regulated by environmental conditions such as grazing duration, wetness index, and soil temperature and moisture responses.

      The authors show that the response of net ecosystem productivity (NEP) to light grazing shifts from negative to positive along a wetness gradient, whereas heavy grazing consistently suppresses NEP across wetness conditions. Importantly, this pattern is supported by both the field experiment and the meta-analysis, suggesting that moderate levels of grazing can potentially enhance both plant productivity and carbon sequestration under favorable moisture conditions.

      The integration of experimental data with a global synthesis is a particular strength of the study, allowing the authors to evaluate grazing impacts across both temporal variability (precipitation fluctuations in the field experiment) and spatial variability (wetness gradients across global grasslands). Overall, the conclusions are well supported by the data.

      However, several aspects of data acquisition, analysis, and presentation could be clarified to further strengthen the reproducibility and interpretation of the results:

      (1) A PRISMA-style flow diagram would be helpful for the meta-analysis to clearly illustrate the study selection process and facilitate interpretation of the dataset.

      Thank you for the comments. We have added the PRISMA-style flow diagram in the revised supplementary materials. See Fig. S9.

      (2) Grazing intensity requires a clearer definition and, where possible, standardization between the field experiment and the studies included in the meta-analysis. Key parameters such as the number of stock per area, days per rotation or per year, and total years of grazing should be clearly defined. In addition, the criteria used to classify grazing intensity into LG, MG, and HG in the meta-analysis should be explicitly described.

      Thank you for the comments. In our meta-analysis, the classifications of grazing intensity into light, moderate, and heavy grazing were provided in previous studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the USDA criteria (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf).

      The criteria was that: light grazing: approximately equal to a maximum of 40% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; moderate grazing: approximately equal to a maximum of 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; Heavy grazing: greater than 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season (November 15).

      We have made following changes in the manuscript:

      Lines 271-277, page 9 “(e) The classifications of grazing intensity (light, moderate, and heavy grazing) were primarily based on the definitions provided in the original studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the modified Grazing Intensity Classes proposed by the USDA (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf) (Yin et al., 2023).”

      (3) In several sections of the manuscript, it is difficult to distinguish whether phrases such as "this study" or "our study" refer specifically to the field experiment or to the overall study, including both the experiment and the meta-analysis. Clearer wording distinguishing these components would improve readability.

      Thanks for the suggestion. The specific distinctions have been revised to clarify whether they refer to the field experiment, meta-analysis or the comprehensive conclusions of the results from both field experiment and meta-analysis.

      Lines 50-53, page 2: “Overall, the meta-analysis and field experiment jointly provide global perspectives on the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity and improve our knowledge of the factors influencing the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity.”

      Lines 191-193, page 6: “The annual precipitation and mean annual air temperature (MAT) for the field study area from 2019 to 2023 were obtained from the China Meteorological Data Service Centre (http://data.cma.cn/).”

      (4) The discussion attributes the non-significant annual NEP response to intra-annual precipitation variability, with grazing enhancing NEP under wet conditions but suppressing it under dry conditions. Another potential explanation may be that grazing affects GPP and ecosystem respiration (ER) at similar magnitudes (i.e., RR(GPP) ≈ RR(ER)), resulting in limited net changes in NEP.

      Thank you for the comments. Another potential explanation has been revised in the manuscript as follows:

      Lines 385-389, page 13: “Grazing decreased the responses of ecosystem CO<sub>2</sub> fluxes and plant biomass in global grasslands, but only NEP was not significantly affected by grazing in our field experiment (Fig. 4). The lack of a significant response in NEP may be because the site in this field study was managed for year-round continuous low-intensity grazing (Liang et al., 2021). In addition, grazing affected GPP and ER at similar magnitudes, resulting in limited net changes in NEP.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see from the above reviews and the specific recommendations below, the first issue you should resolve is a full description of methodological detail such that readers can, in principle, repeat your study. They have to know how you placed the chamber to have a representative measure of vegetation (averaging between high and low biomass patches). They also must know how to calculate grazing intensity and assign the values to the three classes. They want to see your raw data and statistical analyses (including formulae, replication, and degrees of freedom). Consider non-linear relationships of grazing and wetness. Try to better integrate the results from the experiment and the meta-analysis. Finally, make sure that you develop hypotheses from the prior knowledge presented in the introduction. For instance, instead of stating that NEP responses would shift from negative to positive with increasing wetness index, simply hypothesize that wetness mitigates the negative effects of grazing on CO<sub>2</sub> fluxes, even if you find that this is not true under heavy grazing.

      Thank you for the comments. The description of methodological details has been supplemented in the revised manuscript. The results from the experiment and the meta-analysis have been integrated.

      Reviewer #1 (Recommendations for the authors):

      General suggestions:

      (1) The Introduction section and the assumptions should be rewritten and improved. The logicality of the introduction should be revised to better prioritize and contextualize the research problem. The reader is lost since the links between assumptions and previous knowledge are not clear.

      Thanks for the suggestion. The Introduction section and the assumptions have been rewritten as follows:

      Lines 135-147, page 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable? In this study, we investigated the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes by combining a long-term (7-11 years) field experiment conducted in a typical steppe and a meta-analysis of global grasslands. The objectives of this study were to: (i) investigate the effects of grazing intensity with annual wetness fluctuations on ecosystem CO<sub>2</sub> fluxes (GPP, ER and NEP) covering the 7th to 11th years of a continuous grazing experiment in a typical steppe as well as the meta-analysis in global grasslands; (ii) explore how environmental factors (particularly wetness index, soil moisture and temperature, and grazing intensity) regulate the effects of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (2) In the "Materials and methods" section, you need to provide the reason why you chose the "wetness index" in this study, rather than other drought indices (such as Standardized Precipitation Evapotranspiration Index (SPEI), Aridity Index (AI)).

      Thank you for the comments. Although the standardized precipitation evapotranspiration index (SPEI) would be a more appropriate indicator for this study, its calculation requires relatively long and continuous climate data series, which were difficult to obtain in our global meta-analysis. This limitation was particularly important because our study also included a meta-analysis, for which complete climatic datasets were often unavailable from the collected literature. In contrast, the Aridity Index (AI) cannot adequately reflect interannual variability. We have supplemented this limitation in the revised as follows:

      Lines 482-483, page 15: “Furthermore, more drought or wetness indices should be investigated in future studies of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (3) For the results and discussions, the results are currently presented in parallel (long-term experiment first, then meta-analysis). It is recommended to add a dedicated integration paragraph in the discussion.

      Thank you for the comments. The dedicated integration paragraph has been supplemented in the revised manuscript as follows:

      Lines 450-464, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term experiments revealed the mechanisms underlying the effects of grazing on plant characteristics and ecosystem CO<sub>2</sub> fluxes under control conditions. And global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analysis consistently revealed that wetness was an important factor in regulating the effect of ecosystem CO<sub>2</sub> fluxes. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP. Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness determined whether grazing promoted or inhibited the carbon sink function in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      (4) Make sure that the whole manuscript has undergone professional proofreading, and check it carefully to avoid language mistakes.

      Thank you for the comments. The manuscript has been revised by professional proofreading to avoid language mistakes.

      Specific suggestions:

      (1) Lines 126-131: It is recommended to explicitly state three levels of research questions: (i) How does grazing intensity affect ecosystem CO<sub>2</sub> fluxes? (ii) Does the wetness index modulate this effect? (iii) Are these relationships globally generalizable? This will provide a clear logical thread for the paper.

      Thanks for the suggestion. We have made following changes in the revised manuscript:

      Lines 135-139, pages 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable?”

      (2) Lines 181: Briefly justify the choice of the De Martonne wetness index (Equation 1) in the introduction (why this index over other aridity indices).

      Thank you for the comments. The De Martonne wetness index is easier to obtain compared to other drought index, as it only requires annual average temperature and precipitation data for calculation, facilitating the statistics of global meta-analysis. The justification of the De Martonne wetness index has been revised in the manuscript:

      Lines 120-125, pages 4-5: “Annual precipitation is one of the climatic parameters, while the wetness index (WI) serves as a more integrative climatic indicator that incorporates both precipitation and temperature, thereby reflecting the overall water surplus or deficit (Song et al., 2019). A higher wetness index (WI > 30) indicates sufficient water availability for plant growth, whereas a lower wetness index (WI ≤ 30) suggests the water availability may be limited (De Martonne, 1926).”

      (3) Lines 173-174: Please specify the exact timing of flux measurements (e.g., "measured three times per month between 9:00 and 11:00 AM" is already stated, but add "on sunny and calm days" to ensure consistent conditions).

      Thanks for the suggestion. We have supplemented the exact timing of flux measurements in the revised manuscript as follows:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> fluxes represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      (4) The RGR and AGB data provide important clues for explaining the moderating role of the wetness index, but the mechanistic chain can be further refined. It is recommended to use the structural equation model.

      Thanks for the comments. The structural equation model has been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Suggested addition: If data on soil moisture, soil nutrients (e.g., ammonium, nitrate), leaf photosynthetic parameters (e.g., maximum photosynthetic rate, stomatal conductance), or community composition are available, please incorporate them into the analysis to test a more complete mechanistic pathway.

      Thanks for the comments. The data of the composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      (5) The current meta - analysis has established the moderating role of the wetness index, but there is room for further exploration: Test for non-linearity. Could the relationship between the wetness index and the grazing effect size (lnRR of NEP) be non-linear in the global data? Consider fitting models that include a quadratic term for the wetness index or using generalized additive models (GAMs). If a threshold is identified, report the threshold estimate and its confidence interval and discuss its management implications.

      Thank you for the comments. Following the reviewer's suggestion, we further examined whether there is a nonlinear relationship between the wetness index (WI) and the magnitude of grazing effect (NEP_RR). We compared a linear mixed-effects model including a first-order term for WI with a quadratic mixed-effects model, and set Study ID as a random effect in both models. The results indicated that the AIC value of the quadratic model (256.56) was higher than that of the linear model (241.48), suggesting that adding the second-order term did not improve the model fitting precision. Based on this, we retained the simpler linear model in the revised manuscript.

      Author response table 1.

      The comparison of linear mixed-effects model and quadratic mixed-effects model

      Note: if the AIC value was lower, the model fitting accuracy was higher.

      (6) Section 4.1: This section is quite long. Consider splitting it into 2-3 paragraphs, discussing: (i) overall grazing effects on CO2 fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) differential effects of grazing intensities.

      Thanks for the suggestion. Section 4.1 has been split into four paragraphs according to the suggestion of the reviewer. (i) overall grazing effects on ecosystem CO <sub>2</sub> fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) integrated discussion of field study and meta-analysis.

      (7) Integrating the long-term experiment and meta-analysis. Does the effect size observed in the long - term experiment (e.g., an 85.83% increase in NEP under LG in the wettest year) align with the average effect size from the global meta - analysis under similar conditions? If not, what are the potential reasons? (e.g., specificity of the typical steppe, methodological differences between chamber and eddy covariance measurements) Is the mechanism identified in the long - term experiment (e.g., increased RGR) likely to be common globally? What are the joint management implications from both parts of the study? Are there contexts where caution is needed in extrapolating the findings? (e.g., alpine meadows might be more sensitive to grazing).

      Thank you for the comments. The long-term experiment and meta-analysis have been integrated in the revised manuscript. The results of field experiment showed that light grazing increased NEP by 85.83%. However, the response of light grazing was not significant in meta-analysis. Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions. The different results in effect size and statistical significance were mainly because the long-term experiment was conducted in a typical grassland ecosystem, which may possess a strong compensatory growth capacity under moderate water conditions. Therefore, light grazing can enhance NEP by increasing plant photosynthetic rate, promoting new leaf growth, and improving community resource utilization efficiency. However, the global meta-analysis integrated different grassland types, climatic conditions and grazing durations. Consequently, the average effect of global meta-analysis may be diluted by the high heterogeneity among ecosystems. The data to calculate RGR were not available in the original studies of the meta-analysis, so we could not include RGR in the meta-analysis. Therefore, it is still unclear if the mechanism of increased RGR could be common globally. The joint management implications have been revised as follows:

      Lines 393-400, page 13: “Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions (Fig. 3B, 6B and 7H). Light grazing usually stimulates leaf regrowth following defoliation, and these new leaves often are more physiologically active than the older leaves that contribute much of leaf area in ungrazed treatment (Polley et al., 2008), which likely imply a stronger leaf photosynthesis and C sink (Reich et al., 2007). Considering factors such as different grassland types (Fig. S11), livestock grazing modes (Fig. S12), climatic conditions, and grazing durations, the future grazing studies on NEP should focus more on ANPP and RGR, and extrapolate the results cautiously.”

      (8) Figure 8: The conceptual diagram is clear and effectively summarizes the main findings. Briefly explain the meaning of the arrows in the caption.

      Thank you for the suggestion. We have revised Fig. 8 accordingly:

      “Schematic summary of the effects of grazing intensity on ecosystem carbon dioxide (CO<sub>2</sub>) fluxes and biomass in global grasslands in the meta-analysis. The blue and black arrows indicate negative effects on grazing and grazing intensities, respectively. Asterisks (<sup>*</sup>) indicate significant effects on variables at P < 0.05. GPP, gross primary productivity; ER, ecosystem respiration; NEP, net ecosystem productivity; LG, light grazing; MG, moderate grazing; HG, heavy grazing.

      (9) Please check all references for consistency with eLife style. Some entries currently have inconsistent formatting (e.g., some include issue numbers, while others do not; page number formatting varies).

      Thank you for the comments. We have checked the references one by one to revise them consistent with eLife style.

      Reviewer #3 (Recommendations for the authors):

      (1) Method citation:

      (a) Please cite the original method references rather than studies that applied the methods. For example, the original publication introducing the wetness index (WI) is: De Martonne, E. Une nouvelle fonction climatologique: l'indice d'aridité. La Météorologie 2, 449-458 (1926).

      (b) Please also cite the R packages used in the analysis. One straightforward approach is the function citation() in R.

      Thank you for the comments. We have cited the references related to the original methods and cited the R packages used in the analysis.

      (2) Several expressions would benefit from clarification:

      (a) Line 182: What has been "referred to as soil moisture"?

      Thank you for the comments. Soil moisture refers to the volumetric water content of the soil, which has been clarified in the revised manuscript as follows:

      Lines 199-200, pages 7: “Soil moisture was the soil volumetric water content.”

      (b) Line 184-185: Do you mean "soil temperature and soil moisture were measured simultaneously with ecosystem CO<sub>2</sub> flux measurements."?

      Thank you for the comments. Yes, we simultaneously measured soil temperature and soil moisture using temperature and moisture probes while measuring ecosystem CO <sub>2</sub> flux measurements. We have made following changes in the revised manuscript as follows:

      Lines 198-199, page 7: “Soil temperature and moisture at a depth of 0-10 cm were measured simultaneously with ecosystem CO <sub>2</sub> flux measurements, using the probes of the LI-8100 system.”

      (c) Line 200: Please clarify what is meant by "the caged plots"?

      Thank you for the comments. The misleading words have been deleted in the revised manuscript.

      (d) Line 242-244: Do you mean "data from non-grazed treatments were excluded when additional treatments were present"?

      Thank you for the comments. Yes, data from non-grazed treatments were excluded when additional treatments were present. We have made following changes in the revised manuscript:

      Lines 268-270, page 9: “(c) Data from non-grazed treatments were excluded when additional treatments (e.g., fertilization, experimental warming, or precipitation manipulation) were present.”

      (e) Line 263: Should this refer to RR<sub>++</sub> instead of RR? Please clarify how RR<sub>++</sub> (or lnRR++) was calculated from the reported response ratios and study weights.

      Thank you for the comments. RR is the dependent variable of the model, representing the response ratio for each observation. RR<sub>++</sub> typically refers to the pooled effect size obtained after all RRs are weighted and subjected to a mixed-effects model, which primarily corresponds to the estimated value of the model intercept β<sub>0</sub>. We have made following changes in the revised manuscript:

      Lines 293-303, page 10: “A linear mixed-effects model, with ‘study’ included as a random factor, was employed to estimate the weighted response ratio (RR<sub>++</sub>) across studies or within a specific group, fitting with restricted maximum likelihood using the ‘lmer’ function in the ‘lme4’ package (Feng et al., 2023).

      where β<sub>0</sub> is the coefficient, π<sub>study</sub> denotes the random effect associated with ‘study’ (accounting for autocorrelation among observations from the same study), and ɛ corresponds to the residual sampling error. We checked the normality of the model residuals using the ‘check_normality’ function in the ‘performance’ package. When the assumption of normality was violated, bootstrapping with 999 iterations was performed using the ‘boot’ package to derive the 95% confidence interval (CI) for each RR<sub>++</sub> (Chen et al., 2021).”

      (3) Line 222: Please list all environmental predictors considered in the analysis.

      Thank you for the comments. The environmental predictors considered in the analysis have been listed in the revised manuscript:

      Lines: 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      (4) Line 343: The non-significant response of NEP appears only under MG, while both LG and HG significantly decrease NEP (Fig. 5). It may be helpful to discuss this grazing-intensity-dependent response more explicitly.

      Thank you for the comments. The reason why the non-significant response of NEP appeared only under MG was that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP. This was consistent with the hypothesis of moderate disturbance, which suggested that ecosystem functions remain stable under moderate levels of disturbance. We have made following changes in the revised manuscript:

      Lines 390-393, page 13: “Similarly, the reason why the non-significant response of NEP appeared only under MG may be that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP.”

      (5) Line 397: The phrase "global-scale study" typically refers to experiments conducted worldwide. "Studies across global grasslands" may be more precise here.

      Thank you for the comments. We have made following changes in the revised manuscript:

      Lines: 475-477, page 15: “Our meta-analysis has limitations due to the relatively small sample size, which stems from the scarcity of studies across global grasslands exploring the effects of grazing intensity on ecosystem CO<sub>2</sub> fluxes.”

      (6) Line 403: ...have large amounts of grassland for grazing, "but rarely investigated".

      Thank you for the comments. We have made following changes according to the suggestion of the reviewer in the revised manuscript as follows:

      Lines 480-482, pages 15-16: “In addition, future studies could conduct more experiments of ecosystem CO<sub>2</sub> fluxes in South America, Africa and Oceania, which have large amounts of grasslands for grazing, but were rarely investigated.”

      (7) Please remember to cite Figure 8 in the text.

      Thank you for the comments. Figure 8 has been cited in the manuscript as follows:

      Lines 411-414, page 13: “Moderate grazing decreased GPP and NEP under lower WI, but the responses of GPP and NEP to moderate grazing were similar under higher WI in global grasslands (Figs. 6 and 8), indicating that higher wetness offset the response of GPP and NEP to moderate grazing.”

      (8) Figure S1:

      (a) The orange line in panel A appears to represent monthly temperature rather than mean annual temperature.

      (b) Same for the precipitation, the data shown here should be monthly values rather than annual means.

      (c) Consider using a color different from orange for the wetness index in panel C, unless the variable shown is temperature instead.

      Thank you for the comments. We have revised Fig. S1 according to the suggestion of the reviewer as follows:

      Supplementary page 5: “Fig. S1 The monthly total precipitation and average air temperature (A), annual precipitation (B) and wetness index (C) in the study area of the grazing intensity experiment in the typical steppe from 2019 to 2023.”

      (9) Figure S2B: Grazing intensity labels appear in Chinese in the figure. These should be translated into English, and the information on grazing intensity should be provided in the legend or the figure here, as well as in the Method section.

      Thank you for the comments. We have replaced the figures included the Chinese labels. See Fig. S2.

      (10) Figure S6B: LG significantly affects the relationships between wetness index and ER (P<0.05), but a regression line is missing from the panel.

      Thank you for the comments. We have added the regression line between wetness index and ER in Supplementary Fig. S6.

      (11) Figure S7A: "Mean" annual precipitation

      Thank you for the comments. The “Mean” has been revised in Figure S10A.

      References

      Fan, Y., Zhang, X., Wang, J., & Shi, P. (2011). Effect of solar radiation on net ecosystem CO2 exchange of alpine meadow on the Tibetan Plateau.Journal of Geographical Sciences, 21(4), 666-676.

      Jiang, Z., Hu, Z., Lai, D., Han, D., Wang, M., Liu, M., Zhang, M., & Guo, M. (2020). Light grazing facilitates carbon accumulation in subsoil in Chinese grasslands: A meta-analysis. Global Change Biology, 26(12), 7186–7197.

      Li, X., Fu, H., Guo, D., Li, X., & Wan, C. (2010). Partitioning soil respiration and assessing the carbon balance in a Setaria italica (L.) Beauv. Cropland on the Loess Plateau, Northern China. Soil Biology and Biochemistry, 42(2), 337-346.

      Niu, S., Wu, M., Han, Y., Xia, J., Li, L., & Wan, S. (2008). Water‐mediated responses of ecosystem carbon fluxes to climatic change in a temperate steppe. New Phytologist, 177(1), 209-219.

      Rong, Y., Johnson, D. A., Wang, Z., & Zhu, L. (2017). Grazing effects on ecosystem CO2 fluxes regulated by interannual climate fluctuation in a temperate grassland steppe in northern China. Agriculture, Ecosystems & Environment, 237, 194-202.

      Shi, R., Su, P., Zhou, Z., Yang, J., & Ding, X. (2022). Comparison of eddy covariance and automatic chamber‐based methods for measuring carbon flux. Agronomy Journal, 114.

      Wan, L., Liu, G., & Su, X. (2025). Global meta-analysis reveals different grazing management strategies change greenhouse gas emissions and global warming potential in grasslands. Geography and Sustainability, 6(3), 100251.

      Yin, M., Gao, X., Kuang, W., & Tenuta, M. (2023). Soil N<sub>2</sub>O emissions and functional genes in response to grazing grassland with livestock: A meta-analysis. Geoderma, 436, 116538.

      Yu, H., Wang, X., Wu, Y., Wang, C., Yan, R., Xu, D., & Xin, X. (2025). Light grazing tends to enhance ecosystem carbon sequestration and resource use efficiency in a meadow steppe of northern China. Agricultural and Forest Meteorology, 372, 110690.

      Zhang, R., Tian, D., Chen, H. Y. H., Seabloom, E. W., Han, G., Wang, S., Yu, G., Li, Z., & Niu, S. (2022). Biodiversity alleviates the decrease of grassland multifunctionality under grazing disturbance: A global meta-analysis. Global Ecology and Biogeography, 31(1), 155–167.

      Zhou, G., Luo, Q., Chen, Y., Hu, J., He, M., Gao, J., Zhou, L., Liu, H., & Zhou, X. (2019). Interactive effects of grazing and global change factors on soil and ecosystem respiration in grassland ecosystems: A global synthesis. Journal of Applied Ecology, 56(8), 2007–2019.

    1. eLife Assessment

      This important study provides convincing evidence supporting the existence of transposable element (TE)-gene chimeric transcripts in the Drosophila brain and helps reconcile previously conflicting findings. The authors demonstrate that differences in computational approaches, experimental design, and Drosophila lines can substantially influence the detection of these transcripts. The work highlights the importance of standardized approaches for studying TE-derived transcripts, although their broader biological significance remains to be established.

    2. Reviewer #1 (Public review):

      Summary:

      Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to accurately detect TE insertions from RNA-seq. In this short response, Choucri and Treiber clearly show that differences in the tools used between their study and that of Azad et al. likely explain the contrasting results, along with RT-PCR failure to design primers that match the chimeric transcript and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.

      Strengths:

      The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and to compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. The methods are clear and thorough. The discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.

      The biological function, if any, of such chimeric transcripts remains to be determined by further analysis, including chimeric transcripts with low to high overall contribution to gene expression (Figure 1B).

    3. Reviewer #3 (Public review):

      Summary:

      This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.

      Strengths:

      The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.

      Weaknesses:

      The authors should mention that combining PCR-amplified cDNA generation with short-read sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct long-read ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv) . Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that, contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to detect TE insertions from RNAseq accurately. In this short response, Choucri and Treiber clearly demonstrate that differences in the tools used between their study and that of Azad et al. likely account for the contrasting results, along with RT-PCR failure in designing primers that would match the chimeric transcript, and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.

      Strengths:

      The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. Finally, the discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.

      Weaknesses:

      I think it is necessary to add more detail to this article, for instance, the differences between TEchim and Tidal could be laid out more precisely.

      We thank the reviewer for this helpful suggestion and agree that a more explicit comparison improves the clarity of the manuscript. Briefly, TIDAL and TEChim are designed to answer different questions. TIDAL is an insertion/deletion caller that was primarily built to determine whether a TE insertion exists at a genomic locus and how it is distributed across strains or populations. It was not designed to resolve splice junctions between exons and TEs. TEChim, by contrast, is purpose-built to detect breakpoint-spanning reads that span exon junctions and putative splice sites within a TE. This difference in design has several concrete consequences in the algorithms used:

      (1) TIDAL clusters reads that support a candidate breakpoint within a window of twice the sequencing read length (e.g. 300nt for 150nt reads). This works well for calling genomic insertions, but it can be too restrictive for a splice junction, which may fall at a variable position within the gene and TE. TEChim does not rely on fixed-window clustering, but instead maps all reads, groups them to individual nt positions within the genome, and then filters for events that were detected in more than one biological replicate.

      (2) TEChim reconstructs long in-silico reads from overlapping paired-end reads, using FLASH, prior to alignment. In-silico paired-end reads are analysed, and full-read merged fragments provide single-nucleotide resolution for a breakpoint. This increases accuracy and confidence in split reads.

      (3) TEChim explicitly intersects candidate breakpoints with annotated exon/intron/UTR. Structures and canonical splice donor- and acceptor sites.

      We have now added a concise summary of the underlying principles of TEChim in the Methods section, and added details to the comparison of TIDAL with TEChim in the discussion. We hope that this addition makes it clearer why the two pipelines produce different results from the same input data.

      Regarding the roo example, one of the caveats of this family, along with others, is the presence of simple repeats. It would be important to show that the simple repeats are not interfering with the read mapping.

      We thank the reviewer for raising this important point. We agree that simple repeats can complicate read mapping and that this requires careful consideration. We have now added to the results the exact locations and lengths of the three known and annotated repeat regions within roo (Domínguez, 2021) and show that the breakpoints we report map more than 4kb away from these regions. In addition, the splice junctions we identify (at positions 5190 and 5462) recur at the same position relative to the roo consensus sequence across multiple genomic insertions, and biological replicates. If these calls were artifacts of copyspecific simple repeats that interfered with read mapping, then we would expect breakpoints to vary between with each insertions local sequence, rather than converge on the same breakpoint. We therefore consider it unlikely that simple repeats account for the observed splice junctions. We have expanded our discussion to make this reasoning clearer.

      Regarding the experiments, if we are looking for a standardized protocol, then we should have a detailed material and methods section, with every experiment, replicate, and PCR temperature clearly defined.

      We thank the reviewer for this suggestion and agree that a more detailed description of the experimental procedures will improve the reproducibility of the study. We have substantially expanded the Materials and Methods section, including the number of biological replicates, primer information, PCR conditions and other methodological details relevant for reproducing the experiments. In addition, we have written a detailed manual for the updated version of TEChim that we used here.

      Finally, and in my opinion, more importantly, the use of RT negative controls on the RT PCRs, along with DNA PCRs to show insertion presence, is mandatory for testing the presence of chimeric genes. Of course, water negative PCR controls are also needed, and unfortunately, absent from Figure 3.

      We thank the reviewer for this helpful suggestion. We have repeated the RNA extraction on 3 new samples, and this time also included minus-RT aliquots for each of the three biological replicates. We have run our PCR for the chimeric transcript between Beadex and opus on all these samples, and in addition on a water control. All these new results are now shown in Figure 3, and confirm our previous conclusions.

      We now also provide results from our DNA testing for the opus insertion in Beadex. We conduct these at regular intervals in our lab to ensure the insertion remains stable in our stock, and mentioned the results in the original version of this manuscript, but we agree that it is important to also show the raw data of this in this study. We use primers at the up- and downstream end of the opus insertion. The downstream pair gave a single band at the predicted size, which we confirmed by Sanger sequencing. This data is now presented in Figure 3 – Figure supplement 1. The PCR around the upstream end resulted in the expected band of 888bp, and two additional bands. Sanger sequencing of the 888bp band produced signals for both Beadex and opus, but we did not get a reliable signal across the precise breakpoint (see Author response image 1). We think this might partly be due to a tandem repeat at the beginning of the opus LTR, which may interfere with Sanger sequencing. Taken together, our data provides strong evidence that the opus insertion is present in our flies.

      Author response image 1.

      Sanger sequencing results of the 888bp band: Two segments of the raw Sanger sequencing trace are shown; the intervening, unmapped section is omitted. The left segment (grey) shows clean, high-confidence signal matching the Bx locus. The right segment (pink) shows the signal falling to near baseline within the opus LTR, so base calls in this region are log-confidence, and the trace does not resolve the breakpoint. Numbers above the sequence indicate position within the raw sequencing read.

      Reviewer #2 (Public review):

      Summary:

      This study by Choucri and Treiber aims to directly address a recent critique regarding the role of transposable elements (TEs) in diversifying the neural transcriptome of Drosophila. The authors seek to demonstrate that TEs are not merely genomic "noise" but are frequently and reliably "exonized" into brain-specific mRNA. By introducing an upgraded computational pipeline, TEChim, and conducting precise experimental validations, the authors set out to show that TE-mediated splicing represents a genuine biological phenomenon that expands the molecular repertoire of the nervous system.

      Strengths:

      The study's primary strength lies in its rigorous technical "forensic" analysis of previous failed replication attempts. The authors convincingly demonstrate that the lack of signal in the opposing study stemmed from a fundamental methodological mismatch: the software used by the critics (TIDAL) is logically incapable of detecting splice sites located within TE sequences. Importantly, the authors complement this computational clarification with definitive experimental evidence through an effective "experimental rescue." By employing correctly designed primers and matching the genetic backgrounds of the fly strains, thereby accounting for genomic polymorphisms, they successfully validated all seven loci that were previously reported as undetectable. This dual-pronged strategy, addressing both algorithmic bias and experimental design, establishes a more robust technical benchmark for the detection and validation of TE-derived exons in neural tissues.

      Weaknesses:

      While the technical rebuttal is highly convincing, the scope of the study remains primarily defensive. As a response to a prior critique, the work focuses on establishing the existence and detectability of chimeric TE-derived transcripts rather than exploring their broader functional consequences. As a result, there is limited new insight into how these TEmodified isoforms influence neural circuit function or organismal behavior.

      We agree with the reviewer that the primary focus of this study is to establish a robust protocol for the detection and validation of TE-derived chimeric transcripts, and to resolve discrepancies raised by Azad et al. We believe this provides an important technical and conceptual framework for future studies investigating the functional impact of TE-driven genetic variation.

      In addition, the detection and validation of these events remain technically demanding, requiring deep sequencing and specialized bioinformatic expertise, which may limit broader adoption by laboratories without dedicated computational resources.

      We agree that the robust detection and validation of TE-derived chimeric transcripts is technically demanding, requiring both high-quality sequencing data and specialised computational analyses. However, we believe that these methodological challenges are justified by the biological insights that can be gained. By providing an updated computational approach together with experimental validation, we hope to facilitate further research into this phenomenon.

      Reviewer #3 (Public review):

      Summary:

      This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.

      Strengths:

      The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.

      Weaknesses:

      The authors should mention that combining PCR-amplified cDNA generation with shortread sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct longread ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv). Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.

      We thank the reviewer for this excellent suggestion and agree that long-read RNA sequencing represents a powerful approach to characterise chimeric transcripts, because it can capture full-length transcripts and reduce ambiguity of mapping short reads onto repetitive sequences. We have now expanded the Discussion to highlight the advantages of these technologies and to cite the suggested studies. At present, long-read RNA sequencing remains challenging for Drosophila brain samples, because the total amount of input RNA is limited. As a consequence, amplification is usually required, which itself can introduce artefacts, which we showed previously (Treiber and Waddell, 2017). While we agree that long-read sequencing will be an important approach for future studies, we believe that the combination of computational analysis and targeted experimental validation presented here provides robust evidence for the existence of chimeric TE-gene transcripts.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      To maximize the impact of the work, the authors should consider adding a technical guide that outlines the implementation of the updated TEChim pipeline and provides a standardized protocol for designing chimeric RT-PCR primers. This section should include a clear software workflow and a precise primer design strategy, emphasizing junction spanning probes and the necessity of genomic confirmation, to establish these methods as the definitive technical standard for studying transposable elements in the nervous system.

      We thank the reviewer for this constructive suggestion and want to be upfront about what can and cannot be standardized here. Because TE sequences are repetitive, the specific primer pair that works best for a given locus cannot necessarily be predicted in advance, and some empirical testing of candidate primer pairs is unavoidable. What we can standardise, and now describe explicitly in the Methods section, is, firstly the use of TEChim to identify candidate splicing events, and secondly, the logic behind how we choose primers. Several candidate primer pairs are tested for a given genomic locus, and all resulting bands are confirmed using Sanger sequencing. Together with an expanded TEChim manual on GitHub, we believe this study gives other research groups a clear and reproducible starting point, while remaining honest that, as with most repeat-adjacent primer design, some locus-specific optimization remains necessary.

    1. eLife Assessment

      This valuable study demonstrates that a multi-step differentiation program in bacteria combining a bistable switch with two quorum-sensing systems is capable of generating autonomous and self-organized spatial patterns. The evidence for the core engineering system using fluorescent reporters support patterning across several conditions is convincing and has significant implications on the process of cell differentiation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data and to make interesting predictions (which require future validation).<br /> The data presented are convincing and support the claims, the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FP-expressing cells in the population) when changing the initial ratio of green:blue FP-expressing cells.

      Comments on revised version:

      The authors' replies to my comments were very satisfactory. I think the paper has been strengthened.

    3. Reviewer #2 (Public review):

      The revised manuscript by Boni et al. is substantially improved, and in my view the authors have responded constructively to the concerns raised during the first review. Most importantly, they now explicitly distinguish the bistable green/blue states generated by the toggle switch from the reversible red and yellow quorum-sensing outputs and acknowledge that the latter do not constitute irreversible differentiated states. This clarification improves the conceptual accuracy of the work.

      The other revisions are also well justified. The authors clarify that Fig. 2d is derived quantitatively from flow-cytometry data and explain how the cross-sections in Fig. 2e were obtained; they introduce the nullcline interpretation of the toggle-switch landscape and distinguish stochastic single-cell fate from predictable population-level proportions. They also provide more information on pLux engineering, report the functional forms and parameters of fitted curves, quantify the different HSL induction ranges, explain the heuristic treatment of entry into stationary phase, and clarify image selection and sender-receiver distance calculations.

      The principal strength remains the systematic engineering of a sequential multi-circuit programme in a single bacterial genetic circuitry (spread over two plasmids). The work combines bistable symmetry breaking, LuxI/LuxR-mediated communication and an orthogonal CinI/CinR layer to generate spatially organised colony-level behaviours. The DBTL engineering cycle is also convincing: substantial cross-talk in the initially tested Lux/Las quorum-sensing combination was experimentally identified and addressed by switching to the more orthogonal Lux/Cin architecture, while subsequent promoter leakiness was identified and mitigated by replacing pLux with pLuxLac.

      Overall, I consider the revisions sufficient to address all the important concerns.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggle-switch-based symmetry breaking with quorum-sensing interactions to generate colony-scale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      Comments on revised version:

      The authors have satisfactorily addressed my previously raised issues.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data.

      The data presented are convincing and support the claims; the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FPexpressing cells in the population) when changing the initial ratio of green:blue FPexpressing cells.

      We thank Reviewer #1 for the Summary and for highlighting the strengths of our work.

      Weaknesses:

      (1) No spatial pattern emerges. There are some isolated colonies that turn on the downstream FPs, but I do not see a pattern, really. Nonetheless, some colonies do differentiate (i.e. they turn on additional FPs).

      The pattern does not arise within single colonies, but when looking at groups of colonies: a green sender is surrounded by a circle of blue-red receivers, which is in turn surrounded by blue-only receivers. This organization can be seen as a collective bullseye pattern (green centre, red annulus, blue background). When multiple green colonies are present in the plate, each gives rise to its own bullseye pattern. We have now clarified this detail in the text:

      Lines 244-245 – “This motif is reminiscent of a bullseye pattern emerging not within a single colony, but in groups of colonies”

      (2) The mathematical model appears somewhat superfluous. While it can clearly reproduce the data, it is not used to make interesting predictions, changing parameters (and not initial conditions) that guide further experimental implementations.

      It is true that we presented the mathematical model primarily as being capable of recapitulating our experimental results. However, the model helped us identify the ideal set of initial conditions (i.e. inducers concentrations) for Figure 6, and to better understand the dynamics of 3O-C6-HSL and 3O-C14-HSL diffusion in our system (Supplementary Movie 6). While changing parameters would be theoretically possible (e.g. diffusion coefficients, protein production rates...) we have limited exploration in this direction, as fine-tuning a single molecular parameter, all other things being equal, is experimentally challenging.

      Since several comments highlighted this weakness, we have now employed our mathematical model to predict patterns theoretically achievable with the sequential differentiation program under substantially different experimental conditions. Please find a detailed answer to this point below (see Reviewer #1 (Recommendations for the authors), point 8).

      Future directions:

      The utility of this differentiation process (e.g. in metabolic engineering or for the study of biofilm formation and antibiotic resistance) will become clearer once the FPs are substituted with functional proteins that exert an effect on the cells.

      We agree with Reviewer #1 that, for the purposes of an application, we would have to replace the fluorescent proteins with functional proteins. In the last paragraph of the discussion, we outline several potential applications of our differentiation system, including division of labor, biocomputation, biosensing, and engineered living materials. However, expanding the system in this direction is beyond the scope of this manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) This sentence is confusing to me: ‘When cells were pre-cultured with 100 nM aTc and 1 mM IPTG, the resulting colonies were almost entirely homogeneously green and blue, respectively.’ It reads as if both inducers were added in the same culture. I would rewrite this as: ‘When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.’

      We adapted the text as suggested:

      Lines 127-129 – “When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.”

      (2) It would be good to explain why the follow-up experiments are conducted with inducers at the beginning of the differentiation process, given that, in the absence of inducers, there is spontaneous symmetry breaking as noted by the authors: ‘When cells were pre-cultured in absence of inducers, both green and blue colonies grew in the plate (Figure 2b).’ Is it to obtain consistent results across biological replicates?

      We used defined inducer concentrations to control the ratio of senders: receivers. In the absence of inducers, the population consisted of approximately equal proportions of senders and receivers. Under these conditions, nearly all blue colonies were close to a green colony and exhibited high expression of the red reporter. To enable spatial patterning, we therefore required a population strongly biased towards the receiver state. Accordingly, we used low concentrations of IPTG to shift the population to this state (as explained in lines 235-241 of the revised version). For testing and characterization of the system, we spotted senders alongside receivers (e.g. in Fig 7C and some supplementary figures). In these experiments, pre-culturing with high concentrations of aTc or IPTG ensured that the entire population was in the desired state. We have now revised the text to clarify that the choice of inducer concentrations ensured specific population ratios and the reproducibility of the differentiation assay:

      Lines 235-241 – “We therefore selected a condition that would consistently generate a population strongly biased towards the receiver state, with only a few sparse senders to produce the diffusible signal. Having characterized the TS differentiation landscape (Figure 2), we decided to pre-culture cells starting from the green state in presence of 9, 12 and 18 µM IPTG, then plated at a cell density of 500-2000 colonies per plate, in absence of any positional information. This resulted in few sparse green senders densely surrounded by blue receivers, in a highly reproducible manner.”

      (3) In the legend to Figure 4b, please indicate what ‘experimental points’ are (single colonies?)

      We adapted the figure caption as suggested:

      Figure 4b – “Hexagons, circles and triangles represent the average fluorescence intensity of single colonies.”

      (4) This sentence was confusing to me: ‘The time-lapses supported our hypothesis that the differentiation key processes (colony growth, HSL production and detection, expression of the red reporter) happened simultaneously.’ I thought the whole idea behind the work was to have a step-wise differentiation process and that HSL production had to happen first, then its detection and finally the downstream activation of the red FP expression.

      We acknowledge that this sentence was poorly phrased and may have caused confusion. Our multi-step differentiation program indeed operates in a sequential, stepwise manner at the single-cell level: a blue receiver cell can only transition to the red state after detecting 3O-C6-HSL, and red colonies can only transition to the yellow state after detecting 3O-C14-HSL. At the colony level, one might therefore expect colonies to first appear green or blue and only later acquire red fluorescence. However, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period. Experiments at the single-cell level might make the step-wise progression blue - red - yellow more obvious than colony-level time-lapses. We have revised the text to clarify this point:

      Lines 273-280 – “We therefore collected time-lapse videos of the plate assay, first inducing receiver cells with pure 3O-C6-HSL, then spotting sender cells alongside receivers, and finally performing the full differentiation assay (Supplementary Movies 1, 2, 3, 4). Whilst one might expect colonies to first appear green or blue and only later acquire red fluorescence, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period.”

      (5) The authors write: ‘The combination of spatial and temporal information allowed us to develop a mathematical model that could qualitatively recapitulate the patterning properties of the system’. It is strangely put in my opinion. One is ‘allowed’ to develop a mathematical model under other circumstances, too. I suggest rewriting. For example: ‘We developed a mathematical model combining spatial and temporal information to quantitatively ...’.

      We adapted the text as suggested.

      (6) In Figure 6b, the scale bar is missing

      Thank you for noticing this, we have now added the scale bar.

      (7) Figure 7c: Why is the green colony huge compared with the others? The formation of such huge colonies seems to occur sometimes, but not always (it is seen also in Supplementary Figure 12a and in Supplementary Figure 13, but nowhere else; by the way, in Supplementary Figure 12a there is a spatial pattern, with a clearly visible red ring in the otherwise black colony!).

      The huge green colonies occur only in the experiments where 1 µl of green senders were inoculated at specific locations. In most of the differentiation assays, each colony arises from a single cell at unpredictable locations. For Figures 7c, S13a (previously 12a) and S14 (previously S13) we wanted a single green sender colony in the centre of the plate surrounded by blue receivers. To achieve this, as we briefly explained in the methods, we decided to spot 1 µl of culture at OD=1, therefore these green colonies originate from roughly 8×10<sup>5</sup> cells, which justifies their larger size. We have now clarified this detail in the corresponding figure captions:

      Figures 7c, S13a, S14 – “The sender colonies in this figure do not originate from one single cell, but from 1 µl of a culture of sender cells at OD=1, which explains their larger size (see Methods).”

      Concerning the red rings in figure S13a (previously S12a), we hypothesise they are the effect of the leakiness of the pLux promoter. Observing expression of mCherry in the green sender colony was actually the main indication that the sequential differentiation circuit was not working as planned: if mCherry was being expressed in the sender state, then probably cinI was also being expressed, therefore the green sender colony was undesirably producing 3O-C14-HSL. We therefore replaced pLux with pLuxLac, resulting in tighter control over the mCherry-cinI operon in the green sender state. The improvement associated with this modification can be appreciated in Figure 7c, where the green sender colonies do not display any red fluorescence.

      Regarding the specific pattern of this red signal, i.e. a ring, as opposed to an homogeneous signal in the entire colony, we have different hypotheses, but we have not investigated it thoroughly. Since the colony does not originate from a single cell but from a 1 µl inoculum, the effect might partially be ascribed to a ‘coffee ring stain’ phenomenon: upon absorption of the droplet, the outer ring could feature higher cell density compared to the centre, leading to higher red signal. In other cases, we have sometimes seen ring patterns arising in homogeneous colonies due to growth and temperature effects.

      (8) It would be nice to see the model predictions being used to create a system with different properties than the actual one. Can the authors modify parameters such as diffusion of the quorum-sensing molecule or the gene expression response to it, and see how the output would change? Changing the type of quorum-sensing molecule experimentally is doable, as is the addition of, for instance, a delay module in the gene expression module. Nonetheless, I am aware that this would require some time to do, but it would, in my opinion, strengthen the paper.

      We thank Reviewer #1 for the valuable suggestions. We have now employed the mathematical model to test the pattern resulting from various diffusion coefficients combinations and added Supplementary Figure S15 and a paragraph in the main text. Choosing diffusible signals with different diffusion coefficients would indeed allow to generate more complex patterns, including concentric rings where the expression of the red and yellow reporters does not overlap. While we agree that experimentally validating this point would be of considerable interest, we believe that changing diffusion rates is challenging and beyond the scope of a typical three-month revision, as it would most likely require substantial optimization and fine-tuning.

      Concerning the delay module, experimentally it could be implemented by introducing an intermediate step with a transcription factor modulating the expression of the CinI synthase. While we agree it would technically be feasible, we did not implement it due to time constraints. We would expect the effect of a delay module to be the following: production of C14-HSL would be delayed, therefore the front wave of C14-HSL would be consistently lagging behind the front wave of C6-HSL. The distance between the two fronts would be proportional to the delay introduced by the module. We hypothesise this would generate patterns with small yellow circles inside larger red circles. Our simulations already capture these dynamics by varying diffusion coefficients, and given the absence of experimental data to constrain parameters, we decided not to include a hypothetical delay module in our mathematical model.

      See Supplementary Figure S15

      Lines 378-385 – “Our model suggests that varying the diffusion coefficients of the two diffusible signals could generate more complex patterns, particularly concentric rings where the expression of red and yellow reporters does not overlap (Supplementary Figure S15). Interestingly, the largest region of mCitrine reporter expression is predicted when the diffusion of the first signal is slow, while that of the second is fast. This suggests that sufficient local accumulation of the first signal is required to reach the threshold for production of the second signal, which can then spread further away to produce an outer blue-yellow region.”

      Reviewer #2 (Public review):

      In this manuscript, the authors implement a three-step genetic programme in E. coli that converts an initially homogeneous population into spatially structured sender, receiver, and ‘matured’ receiver colonies on agar without externally supplied positional information. They combine a TetR/LacI toggle switch for symmetry breaking, LuxI/LuxR quorum sensing for a paracrine signalling step, and CinI/CinR for an autocrine signalling-like maturation step, and complement the experiments with a mathematical model that qualitatively reproduces pattern formation over a range of initial conditions.

      While the article has many strengths such as a clear conceptual framing using Waddington landscapes, a modular and carefully optimised circuit design, thorough experimental characterisation of the toggle and quorum-sensing modules, integration of spatial modelling with experiments, and generally clear writing and figures, I think it will benefit the article to clarify the definition and stability of ‘differentiated’ states, clarify several quantitative and modelling aspects, better explain how fitted curves and promoter engineering were done, and improve some figure design and wording to avoid ambiguity.

      We thank Reviewer #2 for the summary and for highlighting the strengths of our work. Please find below our detailed responses addressing the weaknesses and gaps you identified.

      Detailed comments below:

      (1) P5-8 / and more generally: A major concern is that producing a reporter output is not, by itself, differentiation. For a state to be credibly called ‘differentiated’, it should be stable (self-maintained) over relevant timescales, ideally in the absence of the inducing context. As written, the manuscript sometimes seems to equate cell type with reporter expression. I strongly suggest adding a short subsection explicitly defining state versus output, and for each claimed state, stating whether it is stable/bistable or unstable/reversible, with evidence. Concretely, the authors should enumerate: a) Togglederived sender versus receiver: stable? under what conditions (inducer ranges, hysteresis window)? b) Paracrine-induced ‘red’ receivers: is this a stable differentiated state, or a context-dependent induction requiring proximity to senders? c) ‘Mature’ (yellow) state: does it persist after removal from the spatial signal field? If not, it should be described as an induced output programme rather than a mature lineage state.

      At present, later sections (and the ‘maturation’ language) risk over-stating what is demonstrated.

      We acknowledge Reviewer #2’s concern regarding the stability and irreversibility of our states and agree this is a limitation in our work. The hysteresis property of the toggle switch is well documented in previous studies (Litcofsky et. al., 2012; Barbier et. al., 2020), so we did not formally re-quantify it. However, the fact that pre-culturing cells without inducer yielded predictable green: blue ratios and very few colonies exhibiting both states indicates that the toggle switch is both irreversible and stable under our experimental conditions. For the second and third steps, the expression of the fluorescent proteins is maintained throughout the duration of the experiment and for several hours to days thereafter. It is indeed established that cells in stationary phase have protein half-lives on the scale of tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025). However, we agree that re-streaking these cells in the absence of sender cells would eventually lead to the cessation of fluorescent protein expression, making the quorum sensing-mediated differentiation steps reversible.

      You can find further details on this point in our reply to Reviewer #3 (Public review), point 1.

      We have now adjusted the wording throughout the manuscript, and in particular we modified the abstract to highlight that our system ‘mimics’ cell differentiation. We have also added sections to explain how red and yellow are reversible cellular programs and not stably differentiated cellular states:

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ONOFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      (2) Figure 2d: It is unclear whether this panel is intended to be qualitative (schematic/ illustrative) or generated from quantitative data. The legend should explicitly state the origin (e.g., representative image, averaged data, simulation output, schematic) and, if quantitative, what was measured, how many replicates, and how the visualisation was constructed.

      We acknowledge the potential confusion generated from this image. While its purpose is to visually illustrate the toggle switch differentiation landscape in 3D, the figure is derived from quantitative data, specifically from the same flow cytometry data used for Figure 2c. Basically, it is the density plot of the (GFP, mCerulean) events recorded for a single replicate (initial state: mixed, induced with 0.003 mM IPTG), but the density is represented with a third dimension (depth) instead of colour intensity (as we did for Figure 2e). We have now adjusted the figure caption to clarify the figure purpose and how it was generated:

      Figure 2d – “Representative 3D energy landscape: valleys represent the green and blue stable states, red line represents the separatrix. The landscape was generated from quantitative flow cytometry data from a single replicate in c (initial state: mixed, induced with 0.003 mM IPTG). We measured GFP and mCerulean fluorescence intensity from 50,000 single cells (see Methods). The depth of each (x,y) point in the landscape corresponds to the number of recorded cells with a specific (GFP, mCerulean) intensity. This specific condition was chosen for illustrative reasons, to show the landscape in an almost symmetrical condition.”

      (3) Figure 2e: The cross-sectional line is described as meant to be comparable, yet the leftmost plot appears to have a different slope from the others. The authors should explain whether this reflects a different scaling/normalisation, a different underlying dataset/condition, or simply a plotting artefact. If these are fitted trends, report the fit function (see also the comment on fitted lines below).

      The observation is correct, the cross-sectional lines do not have the same slope. They are obtained by connecting the local minima of the green and blue population, i.e. the two points with highest density in the density plots. As the mean fluorescence intensity of the two population is not the same across conditions, the resulting lines have different slopes. We chose this strategy instead of using fixed-slope cross-sectional lines to capture the maximum depth of each valley. We have now adjusted the figure caption to clarify how the cross-sectional lines were generated:

      Figure 2e – “Flow cytometry density plots of 5 representative conditions from c (initial state: mixed, inducer concentration indicated below each plot). For each density plot, we computed the coordinates of the local minima in the green and blue populations, then generated an orthogonal plane crossing them. Black lines represent the projections of these planes on the (GFP, mCerulean) plane. [...]”

      (4) Around P7-8: (saddle/separatrix description): When describing the saddle or separatrix between the two valleys, it would be helpful to briefly connect this more directly to a quantitative dynamical-systems perspective: for instance, the intersection of nullclines and how nullcline geometry changes under IPTG/aTc induction. This will make the landscape picture more complete for readers familiar with the original genetic toggle switch work (Garder et al., 2000).

      We thank Reviewer #2 for the suggestion. We have now included the concept of nullclines in our description of the differentiation landscape in the main text:

      Lines 153-159 – “Mathematically, the profile of the toggle switch landscape corresponds to the number of intersections between the nullclines (the curve where the derivative of a given species over time equals zero), which can be one or three (Gardner et. al., 2000). If they intersect only once, there is only one minimum on the potential curve, which corresponds to a single stable state. If they intersect three times, there are two minima and one maximum, which corresponds to two stable states, and the maximum represents the separatrix.”

      Lines 168-171 – “From a quantitative dynamical-systems perspective, the addition of aTc or IPTG modified the geometric shapes of the nullclines, shifting the three solutions. If the inducer concentration is large enough, the nullclines intersect only once, producing a single stable steady state (Gardner et. al., 2000).”

      (5) P9, lines 157-159: The current phrasing (‘in absence of noise, the system would be fully deterministic... in living cells, however, stochastic bursts... change the trajectory’) risks conflating predicting population-level percentages with predicting colony-level trajectories. It would help to clearly separate (i) the ability to predict the overall fraction of ON/OFF (green/blue) colonies from inducer conditions (which is largely deterministic at the population level) from (ii) the intrinsically stochastic choice of state made by any given founder cell and its colony.

      We acknowledge that this wording has generated confusion, but we are not completely sure we understood the suggestion made by Reviewer #2. If we understood correctly, the concern is that single-cell trajectories and population-level ratios are substantially different and should not be confused, and should be investigated and modelled differently.

      To clarify, our initial sentence referred to the toggle switch differentiation landscape, not to the likelihood of a cell or a colony to be in a given state. In absence of noise, the TS landscape is deterministic: one of the two states is always the stronger attractor, for the same initial conditions cells would fall 100% of the times in that valley. In living cells, however, there are sources of noise, which make it possible for two cells starting from the same initial conditions to end up in two different valleys. The stochastic bursts in gene expression can ‘push cells back up’ and beyond the separatrix. Very large noise would allow cells to end up in the green and blue valley regardless of the TS previous state or the inducer concentration.

      The differentiation induced by the toggle switch when the system is initiated close enough to the separatrix is stochastic. The probability of observing a certain ratio green: blue at the population level reflects exactly the likelihood of a single cell falling into one of the two stable states. Our mathematical model does not take into account molecular details (synthesis and degradation/dilution of proteins, affinity of transcription factors for their cognate promoter...) to mimic the stochastic choice at the cell level. Instead, it estimates the TS differentiation landscape from the population-level percentages based on thousands of single cells.

      To avoid confusion, we have removed the initial sentence and replaced it with an explanation of the link between individual cell trajectories and the resulting population-level ratios:

      Lines 159-164 – “When the system is close to the separatrix, cells can progress towards both the green and the blue destiny. The differentiation trajectory of a single cell is largely stochastic and is influenced by bursts of gene expression. Once the TS is locked in one state, the progeny of that cell maintains a memory of that state, hence a colony has the same state as the founder cell. At the population level, the ratio of green and blue colonies reflects the probability of each cell to fall into the green or the blue valley.”

      (6) P11, lines 193-195 (promoter engineering): The main text currently only refers to screening variants and choosing pLux76; I suggest briefly stating in the main text (not only in the supplement) what was changed (for example, promoter box variants, core promoter strength modifications) and what design criteria were used (reduced leakiness, increased dynamic range).

      We adapted the text as suggested:

      Lines 206-211 – we carried out a screening of pLux promoter variants, testing combinations of Lux boxes (G1 [iGem part BBa K1216007] and pLux76 (Grant et. al., 2016) and promoters with different strengths (100% and 54%), in order to identify variants with minimal leakiness and a high fold-change. We identified pLux76 (Grant et. al., 2016) as the regulatory region with the highest fold-change and sufficiently low leakiness among our candidates (Supplementary Figure S4)”

      (7) Use of fitted lines (Figures 2, 4, 5, 7): Wherever fitted curves are overlaid on data, the authors should indicate in the figure legend the explicit form of the fit as well as the fit equation/ parameters. As a reader, it is difficult to interpret what is empirical smoothing versus what is a mechanistic functional form.

      In most cases, with the exception of Figure 2c, the fitted curves have an illustrative purpose and result from empirical smoothing of the data points. Unless otherwise stated, the resulting parameters were not implemented in our mathematical model. For Figure 2c, instead, the experimental data is fitted with a mechanistic functional equation and the calculated parameters were implemented in our mathematical model.

      We have now added the equations and parameters of the lines fitting the experimental points, either directly in the figure caption or in Supplementary Information Tables (Tables VII to XII). In the latter case, the exact table is referenced in the corresponding figure caption.

      (8) P13, lines 232-235: The comparison between induction directly with C6-HSL and induction from sender colonies is qualitative (‘significantly smaller range’). The authors should provide distances (for example, in mm) for the induction range in each case and, if possible, approximate total HSL amounts or concentrations, so that the reader can appreciate the magnitude of the difference.

      We have now calculated the induction range as the distance where half-maximal induction is observed, which allowed us to compare the induction range across conditions. We adapted the text accordingly. We also provide an estimate of the amount of 3O-C6-HSL produced by a green sender colony, based on simulations of our mathematical model, but we highlight that the two conditions are substantially different and care should be used when comparing them: in one case, a fixed amount of C6-HSL is present from t=0 and simply diffuses outwards; in the second case, a source continuously produces C6-HSL that progresses as a wavefront, therefore the concentration profiles and diffusion dynamics are different.

      Lines 220-226 – “The induction range obtained with 100 picomoles of pure 3O-C6-HSL was approximately 6-8 mm (50% of maximal induction at 3.74 mm). The induction range around sender colonies was significantly smaller (50% of maximal induction at 1.8 mm, Figure 4b). Mathematical simulations yielded a similar slope when using 0.25 picomoles of 3O-C6-HSL (data not shown), even though care should be used when comparing diffusion of a fixed amount of inducer with continuous production from a growing source.”

      (9) P13, lines 259-262: The authors model the transition to the stationary phase via a monotonically decreasing sigmoid in time for biosynthetic capacity. What is the rationale or literature basis for this approach to model entry into the stationary phase? The authors should cite prior work and clarify why this form is appropriate here, versus alternatives (nutrient diffusion limitation, logistic growth with resource depletion, etc.).

      In most mathematical models involving pattern formation, entry of cells in stationary phase is simply treated as a step-wise function (Zwietering et al., 1990): the entire population is active (activity = 1) until it suddenly becomes inactive (activity = 0). However, experimental evidence suggests that the decay in activity is better represented by a smooth decreasing function (Gefen et al., 2014). While several frameworks and mathematical equations exist to describe loss of cell viability (for example under scenarios of heat inactivation (Mafart et al., 2002; Van Boekel et al., 2002), or nutrient limitation), we could not find previous works modeling the loss of metabolic activity upon entry in stationary phase. Again, experimental evidence suggest that switches in metabolic state are rather heterogeneous and occur stochastically at the single cell level (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to implement a simple heuristic approximation that captured well our experimental data. We have now added a paragraph in the section Mathematical modelling - Bacterial activity, supported by the appropriate references, to justify our reasoning:

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – In our system, production of diffusible molecules and fluorescent reporters happens on a timescale of several hours, therefore we decided to take into account the entry of cells in stationary phase. Traditionally, bacterial activity has been modelled as a step-wise inactivation function (Zwietering et al., 1990). However, experimental evidence has shown that bacteria support a low constant rate of protein expression even while growth-arrested, suggesting a low decay in bacterial activity (Gefen et al., 2014). Even when considering cell death upon heat inactivation, a Weibull frequency distribution model is preferred over a step-wise viability function (Mafart et al., 2002; Van Boekel et al., 2002). While we could not find mathematical descriptions specifically for entry in stationary phase, several single-cell metabolism studies support gradual, asynchronous state transitions (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to describe the loss of activity function in our system via a monotonically decreasing sigmoid function, that captured well our experimental data.”

      (10) Figure 6c: Are the areas of the plate shown in each column the same field of view across conditions/time, or are these simply representative regions selected per condition (possibly from different plates)? The caption/legend should clarify whether these are matched locations and how images were chosen.

      They are indeed representative regions per each condition. Each condition is a different plate, as the differentiation assays requires plating the culture homogeneously on a fresh plate without inducers. The plates used for imaging were the same used for the quantification in Figure 6b, four images per plate were collected, one representative image was chosen for Figure 6c. We have now adapted the figure caption to clarify how images were collected and chosen:

      Figure 6c – “Representative microscopy images of the spatial patterns generated by cells harbouring the 2-step system, pre-cultured with different inducer concentrations (indicated at the bottom of each column). Each column displays one representative image (from four locations imaged) of the seven plates in b. Rows (from top to bottom): GFP channel, CFP channel, mCherry channel, composite image.”

      (11) Figure 7a: The combination of solid, dashed, and dash-dot arrows/lines is visually hard to read. I suggest replacing the dash-dot line with a fully dotted line or using different colours (if consistent with journal style) to improve readability.

      Thank you for noticing this. We have now replaced the dash-dot line with a simple dotted line in Figures 3c, 7a, S12a (previously S11a) and S13b (previously S12b) to improve readability.

      (12) Figure 7e and similar analyses: The authors should explain in the Methods and/or captions how ‘distance from sender colonies’ is computed when multiple senders exist. Is the distance always measured to the nearest sender, and how are cases handled where a receiver is in the overlapping influence of several senders? This clarification is important for interpreting the fitted curves.

      The calculation of the distance in 4b, 5b and 7e was performed only in cases where sender colonies were sufficiently sparse (i.e., more than 1 cm away) to assume each receiver was under the influence of a single sender. When collecting the microscopy images, we carefully avoided to image fields of view with green senders just outside the edges of the images. Images with multiple senders (e.g., 7c) would require a non-trivial calculation to compute the relative contribution of each sender to the red intensity of each receiver. We have now clarified this detail in the Methods:

      Lines 558-560 – “When collecting microscopy images with senders and receivers, we carefully avoided to image fields of view with green senders just outside the edges of the images.”

      Lines 590-593 – “Calculation of the distance between receiver and sender colonies was performed only for images collected from plates with few sparse (i.e., more than 1 cm away) senders, ensuring that each receiver only sensed the 3O-C6-HSL produced by a single sender colony.”

      Reviewer #2 (Recommendations for the authors):

      (1) P8, ‘upon transformation’: This phrasing is ambiguous and can be misread as referring to DNA transformation. I recommend changing ‘upon transformation’ to ‘after transition’ (or similar) to avoid confusion.

      The phrasing indeed refers to the DNA transformation of circuits into the cells. The two plasmids were co-transformed into MG1655, cells recovered for approximately 1 h, then the bacteria were plated on solid medium in absence of inducers. The resulting colonies, which we refer to as ‘upon transformation’, were a mixture of green and blue (both expressed at low intensity, as highlighted in Supplementary Figure 2). We only used these colonies for Figure 2c, middle row. For all other experiments, we selected colonies that had been pre-differentiated in one of the two states via chemical inducers, and showed strong expression of the respective reporter (Supplementary Figure 2). We have now adapted the figure caption to make this detail more explicit:

      Figure 2c – “For the initial state green and blue, cells were taken from colonies that were homogeneously green or blue, while for the initial state mixed, cells were taken from a colony obtained immediately after transformation of the circuit plasmids into cells.”

      (2) More generally, consider disambiguating terminology around ‘cell type’, ‘state’, and ‘output’, since the current wording occasionally implies stable fate commitment where the data (as presented) may instead support reversible, context-driven induction.

      We went carefully through the text and adapted it appropriately, highlighting which states are stable and which are reversible. See our detailed reply to your point 1 (Public review) above.

      Reviewer #3 (Public review):

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggleswitch-based symmetry breaking with quorum-sensing interactions to generate colonyscale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      However, I think the paper needs to tone down the extent to which the system demonstrates multi-step differentiation or morphogenesis, which is not critical for making the paper valuable. Only the first step of their circuit design (Figure 1), the toggle switch, generates stable alternative states. The latter steps are mainly signal-dependent reporter activation states layered on top of the blue receiver state, rather than true fate transitions. The authors explicitly state that red expression is added without replacing the blue identity, and they also acknowledge that red cells lose their identity upon restreaking unless they remain near sender cells. That substantially weakens the differentiation analogy and makes the Waddington framing too strong.

      We acknowledge that our sequential program does not recapitulate all the features of a differentiation trajectory, and in particular we do recognize the only two stable states are the identities associated with the toggle switch, while red and yellow are outputs indicating activation of a new molecular program. A complete differentiation program would require irreversible cell fate determination, as we mention in the discussion. For the same reason, in our illustrative Waddington landscape (Figure 1) we only represent two valleys, or minima, corresponding to the green and blue states, while the activation of the red and yellow programs does not result in further valleys. We modified the text at various points to clarify this important distinction, and the fact that our system mimics some key steps happening during multicellular differentiation, without claiming that expression of a fluorescent reporter is a stably differentiated state. Notably, we modified the abstract to highlight that our system ‘mimics’ cell differentiation.

      We believe the differentiation analogy and the Waddington framing remain valuable in our work, considering the physical constraints and timescale of our system. While it is true HSL-induced molecular program would eventually deactivate in absence of signal, this would require a significantly longer amount of time than the one we use to observe patterns. Also, upon storage of Petri dishes at 4 °C, the pattern (including red and yellow) remains stable for a few weeks. Finally, our main purpose was to investigate the capacity of a sequential program to generate autonomous spatial patterns, without the need for human intervention. The removal of a colony from the local signal field, e.g. to re-streak it on a fresh plate, represents a strong disturbance of the pattern, not dissimilar from early developmental biology experiments where transplant of tissue portions would sometimes result in fate reprogramming.

      You can find further details on this point in our reply to Reviewer #2 (Public review), point 1.

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ON-OFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      A related concern is that the 3rd step does not introduce a new spatial organizing rule. The authors show that the second signal remains confined to cells already receiving the first signal, and explicitly conclude that it functions only as an autocrine cue rather than a second paracrine layer. As a result, the 3-step system seems more like an added local readout or maturation layer. Overall, the main 2-step outcome is sparse green sender colonies surrounded by red-expressing blue receivers, with distant receivers remaining blue. That is a valid engineered pattern, but it is still a local, threshold-response circuit architecture.

      The comment is correct, we indeed refer to the 3rd step as ‘maturation’ in the manuscript, as the newly emerged red population activates a new molecular program, which cannot be activated in blue receivers that have not been exposed to C6-HSL. While we had initially expected that production of a second diffusible signal could result in signal propagation, therefore generating an outer yellow ring surrounding the red ring, our experiments showed C14-HSL only acted locally, and our mathematical simulations confirmed that the front wave of the second signal was always lagging behind the first one. We have now employed our mathematical model to explore which conditions would support a new patterning rule, and added Supplementary Figure S15. If the diffusion coefficients of C6-HSL and C14-HSL were 10-100 times smaller and 2-5 times larger respectively, a yellow-only area could appear around the red area.

      Regarding the last sentence, we are unsure about the concern raised and which alternatives Reviewer #3 would recommend. We fully agree our genetic program leverages a local, threshold-response circuit architecture, it is indeed a reaction-diffusion system based on quorum sensing signalling. Many natural patterning and morphogenesis systems do rely on local, rather than global, interactions, yet they are capable of forming complex and hierarchical structures.

      The autonomy claim should be toned down and stated more precisely. The plate patterning occurs without externally imposed spatial gradients, which is a strength. However, by design, the overall system behavior depends strongly on pre-culture inducer conditions that set the sender:receiver ratio, and this externally imposed history is central to the final pattern. This property is tied to how the circuit is designed where steps 2 and 3 largely respond to symmetry breaking introduced in step 1, which is dependent on both history and initialization on the plate. In particular, currently the pattern formation process is quite variable (e.g. figure 5), depending on how different colonies flip the toggle switch, and consequently, how many become senders and how many become receivers. It would have been fascinating if they could also demonstrate the differentiation within individual colonies, leading to intra-colony patterns. This aspect should at least be discussed.

      We would like to clarify that, by ‘autonomous’ we refer to the reaction-diffusion system that starts functioning upon the seeding of the cells on the plate. All previous steps (including the cell culturing, dilution and plating) correspond to setting the initial conditions. We agree with the description of Reviewer #3 concerning the features of our system, but we argue that autonomy and dependence on the initial conditions are two different properties. We claim our system is both autonomous and sensitive to initial conditions. Upon initialization of the system on the plate, there is no further intervention, or nudging (e.g. time-dependent light stimuli, addition or removal of inducers at specific times...), therefore the system is autonomous. It is also dependent on the initial conditions, which are set homogeneously for all the cells (no positional information, each cell experiences the exact same conditions). We believe this is a strength of our system, which allows to generate a variety of diverse yet reproducible spatial patterns. Independency from initial conditions would consistently generate the same pattern, regardless of the initial inducer concentration, which might be interesting for certain applications (e.g., ensuring a fixed population ratio, provide robustness to environmental variability...), but it is not universally superior.

      In order to clarify that we only refer to the reaction-diffusion system on the plate as the autonomous component, we have now added a sentence in the main text:

      Lines 132-134 – “In our differentiation assay, cell culturing, dilution and plating correspond to setting the initial conditions for the system. Upon initialisation on the plate, no further intervention was performed, therefore the system evolved in an autonomous fashion.”

      Concerning the intra-colony patterns, that was admittedly our initial interest. We explored conditions that would consistently generate colonies with sectors, for example by inducing cells in one of the two states and then growing them on agar supplemented with the opposite inducer. However, we immediately observed that, due to the close proximity of sender and receiver bacteria, and the rapid diffusion relative to the timescale of gene expression, all blue sectors were also red. This outcome effectively eliminated the distance-dependent nature of the quorum sensing response. In order to take full advantage of the diffusible system, we focused instead on well-separated homogeneous colonies. We have now added Supplementary Figure 5, the corresponding figure caption, and we briefly discuss in the main text the occurrence of intra-colony patterns and why we did not investigate them further:

      See Supplementary Figure S5

      Lines 228-231 – “We then tested the potential of the 2-step differentiation system to generate self-organized spatial patterns. We observed rare motifs arising within the sporadic colonies showing both green and blue sectors. Due to the close proximity of the sender and receiver bacteria within a colony, the blue sectors always showed strong red signal (Supplementary Figure S5).”

      The mathematical model is useful in guiding both the characterization of parts, modules and the overall system. However, the claims around its quantitative predictive power should also be made narrower. The simulations are built from multiple fitted and partly hand-tuned components, including toggle-switch response curves, colony-growth rules, diffusion, reporter-response functions, and activity decline. This supports a calibrated qualitative reconstruction of the observed patterns, but not a strong predictive or mechanistic validation.

      We accepted the suggestion and replaced ‘predicted’ with ‘recapitulated’ or ‘simulated’ at various locations in the main text:

      Lines 105-107 – “Throughout the work, experimental results were used to develop a mathematical model, providing us with insights that guided further experimental efforts.”

      Lines 292-293 – “We developed a mathematical model combining spatial and temporal information to qualitatively recapitulate the patterning properties of the system.”

      Lines 316-318 – “We therefore employed the mathematical model to simulate the patterns generated from the 2-step differentiation system for different initial blue: green ratios. The model suggested a variety of outcomes [...]”

      Other specific points:

      (1) Given the topic of the work, the authors should cite closely relevant studies in programming pattern formation, including: Cao et al, Cell 2016 Collective space-sensing coordinates pattern scaling in engineered bacteria Rajasekaran et al, Cell 2024 A programmable reaction-diffusion system for spatiotemporal cell signaling circuit design Lu et al, BioRxiv 2024 Discovery of interpretable patterning rules by integrating mechanistic modeling and deep learning

      We have added these references to the introduction and discussion.

      (2) The model assumes identical diffusion coefficients for C6-HSL and C14-HSL despite their substantially different molecular sizes and hydrophobicities. This assumption could distort kinetic lag with differential diffusion in explaining the autocrine confinement of the third step. Its impact should at least be explored in the simulations.

      The comment is very appropriate. To the best of our knowledge, no exact values have been published for C6-HSL and C14-HSL diffusion in water, let alone for diffusion in agar. However, estimates in water range between 3 ·10<sup>−6</sup> cm/s and 5 · 10<sup>−6</sup> cm/s at 25°C, depending on the chain length (i.e. a factor of 1.7 at most). In response to your concern, we have now computationally explored the effect of using signals with different diffusion coefficients and their impact on the resulting spatial pattern (Supplementary Figure S15). Varying the diffusion coefficient of the second diffusible signal does not substantially change the pattern: lower diffusion leads to the yellow region being slightly more concentrated around the green sender and more intense. Higher diffusion has the opposite effect, with yellow areas being wider but less intense. Conversely, changing the diffusion coefficient of the first signal leads to significant changes in the final pattern, suggesting that the kinetics of C6-HSL accumulation and dispersal have a strong influence on the production of C14-HSL and therefore of the yellow signal.

      (3) The mCherry response parameters change significantly between the 2-step and 3-step systems. The authors acknowledged this change but did not provide a clear explanation.

      We believe that the elements contributing to the different mCherry response in the 2-step and 3-step systems are: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression. We have now added a paragraph in the section Mathematical modelling Bacterial activity that lists those differences. We have also added Supplementary Figure S20, containing the experimental data that was fitted to obtain the parameters listed in the second row of Table II. Finally, we highlight that for the implementation of our mathematical model we were only interested in the response function to varying 3O-C6-HSL concentrations and not in the exact values of individual parameters.

      Supplementary Material, section A. “Mathematical modelling, subsection 3. mCherry production and 30-C6-HSL sensing – The parameters fitted for mCherry production in the 2-step and 3-step systems are quite different. We hypothesize that the following elements contribute to the difference: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression.”

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – For the implementation of our mathematical model, we were not interested in analysing individual parameters values, but rather in the resulting response function to varying 3O-C6-HSL concentrations. Further analysis would be needed to determine whether these parameters are statistically significant, but this is beyond the scope of this present manuscript.”

      (4) The 3-step system is evaluated at only a single condition with no simulation comparison, in contrast to the systematic 11-condition validation of the 2-step system.

      The rationale for not repeating the 11-condition assay with the 3-step system is that the patterns would be substantially the same as Figure 6c (red colonies would also be yellow). Furthermore, the Nikon SMZ25 stereo microscope we used to take images with a large field of view (ideal for the differentiation assay in Figure 6c) would not allow us to discriminate well between green and yellow colonies, making the interpretation of the patterns difficult. However, following up on your comment, we have now used our mathematical model to produce a representative image corresponding to Figure 6d, that we included as Supplementary Figure S16, exploring the effect of varying the green: blue ratio with the 3-step system.

      Reviewer #3 (Recommendations for the authors):

      Please see above. My comments are largely about improving the rigor and clarity of the writing, particularly those related to conceptual claims, as well as some modeling analysis to strengthen their conclusions.

      We thank Reviewer #3 for the suggestions. See our detailed replies to your points above.

    1. eLife Assessment

      This important study employs a closed-loop, theta-phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats and reports that disrupting theta-timescale coordination impairs performance of challenging aspects of spatial behaviors, while sparing hippocampal replay and spatial coding in hippocampal place cells. Technically rigorous experiments were performed, and solid evidence is provided to support the claims. The findings are expected to advance theoretical understanding of learning and memory operations and to provide practical implications for the application of similar optogenetic approaches.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Joshi and colleagues demonstrates that the precise theta-phase timing of spikes is causal for CA1 hippocampal theta sequences during locomotion on a linear track and is necessary for learning the cognitively demanding outbound component of a hippocampus-dependent alternation task (W-maze), independently of replay during immobility. To reach these conclusions, the authors developed a theta-phase-specific, closed-loop manipulation that used optogenetic activation of medial septal parvalbumin (PV) interneurons at the ascending phase of theta during locomotion. This protocol preserved immobility periods, allowing a clean and elegant dissociation from SWR-associated replay.

      The manuscript is well written and was a pleasure to read. The work described if of high quality and introduces several notable advances to the field:

      a) It extends prior studies that manipulated theta oscillations by examining precise temporal structure (specifically theta sequences) rather than only LFP features.

      b) The closed-loop manipulation enabled dissociation between deficits in theta sequences during a behavioural task and SWR-associated replay activity.

      c) As controls, the authors included rats with suboptimal viral transduction or optic-fibre placement, and, within subjects, both stimulation-on (stim-on) and stimulation-off (stim-off) trials. Notably, sequence disruption persisted into stim-off periods within the same session.

      Overall, this is a strong manuscript that will provide valuable insights to the field.

      After revision, the manuscript has been substantially strengthened. The authors did incorporate the vast majority of the reviewer's comments and have expanded the discussion of prior medial septal manipulations, clarified the rationale for their theta-sequence analyses, added analyses of SWR and replay in the rest/sleep box as well as provided additional methodological and histological validation.

      The new rest-box analysis is a great addition and directly addresses my request to distinguish aSWRs from events during longer off-track rest periods (rSWR). The data support the narrower conclusion that no large group difference was detected in rest-box ripple rate or duration.

      The new observation (in response to reviewer #2, point3.2) that on the W-track theta power does not fully recover during stimulation-off periods does change the interpretation of the results. It means that these epochs are then not a physiologically recovered control condition. Therefore, the persistent disruption of theta sequences during the middle block cannot, alone, demonstrate that disrupting sequences during the earliest experience produced a lasting plasticity-related effect. It could also reflect a lingering network effect of the stimulation that persists after laser delivery has stopped. In my opinion the discussion should present at least these two alternatives: the disruption of early experience-dependent plasticity, as well as the incomplete physiological recovery from the preceding stimulation. The linear-track recovery data is helpful, but it does not guarantee the same mechanisms/effects will be present on the novel Wmaze (versus the familiar linear track).

    3. Reviewer #2 (Public review):

      Summary:

      The authors of this study developed a closed-loop optogenetic stimulation system with high temporal precision in rats to examine the effect of medial septum (MS) stimulation on the disruption of hippocampal activity at both behavioral and compressed time scales. They found that this manipulation preserved hippocampus single-cell-level spatial coding but affected theta sequences and performance during a spatial alternation task. The performance deficits were observed during the more cognitively demanding component of the task and even persisted after the stimulation was turned off. However, the effects of this disruption were confined to locomotor periods and did not impact waking rest replay, even during the early phase of stimulation-on. Their conclusion is consistent with previous findings from the Pastalkova lab, where MS disruption (using different methods) affected theta sequences and task performance but spared replay (Wang et al., 2015; Wang et al., 2016). However, it differs from a recent study in which optogenetic disruption of EC inputs during running affected both theta sequences and replay (Liu et al., 2023).

      Strengths:

      The experiments were well designed and controlled, and the results were generally well presented.

      Comments on revised version.

      The authors of this study addressed all my concerns, some of them successfully. The stimulation disrupted theta oscillations, making quantification of theta sequences problematic. Within the constraints of their experimental design, the authors tried their best to address my concerns. Therefore, I am satisfied with the current version and express no further comments.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study employs a closed-loop, theta-phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats and reports that disrupting theta-timescale coordination impairs performance of challenging aspects of spatial behaviors, while sparing hippocampal replay and spatial coding in hippocampal place cells. The findings are expected to advance theoretical understanding of learning and memory operations and to provide practical implications for the application of similar optogenetic approaches. The experiments were viewed as technically rigorous, but the strength of evidence provided in the current version of the manuscript was viewed as incomplete, mostly due to limited analyses and the descriptions of some of the experimental protocols.

      We thank all reviewers for their overall assessment, thoughtful comments, and suggestions. We have now addressed each of the reviewers’ comments in detail and updated the manuscript on bioRxiv (URL: https://www.biorxiv.org/content/10.1101/2025.09.15.675587v2). In addition, we have shared the raw data, intermediate analysis files, and the complete repository to facilitate replication of the analysis and figures.

      Code repo: github.com/LorenFrankLab/ms_stim_analysis

      Data repo: dandiarchive.org/dandiset/001634

      Docker containers (see GitHub repo for use instructions):

      - Database: https://hub.docker.com/r/samuelbray32/spyglass-db-ms_stim_analysis

      - Python notebooks: https://hub.docker.com/r/samuelbray32/spyglass-hub-ms_stim_analysis

      (1) Novelty and contrast with earlier manipulations:

      We now explicitly contextualize our results with prior pharmacological (Wang et al., 2016; Wang et al., 2015; Koenig et al., 2011; Brandon et al., 2014), systemic (Robbe & Buzsaki 2009; Petersen and Buzsáki 2020), and behavioral (Drieu et al., 2018) manipulations that also assessed some of the physiological features we evaluated. This contrast helps us highlight both the insights and the discrepancies observed in the prior approaches. We also more clearly explain the novelty and importance of our specific approach for temporally and physiologically precise manipulation. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (2) Additional analysis on SWRs during rest:

      Since submitting the manuscript, we have conducted additional analysis on the rate and length of SWRs in the rest box (rSWRs) and found that the rate and length are also indistinguishable between targeted and control animals (effect of manipulation between control and targeted animals; rSWR rate: p=0.45; rSWR length: p=0.94, mixed-effects model). We also find evidence for sequential neural representations (“significant replay”) in the rest box when the encoding was performed in the behavioral arena. Example trajectories and full analysis per animal are shown in the new supplementary figure (Figure S6). These results are consistent with our observations on aSWR rate, length, and content in the behavioral arena. Additionally, based on the reviewer’s recommendation, we have evaluated the fraction of ripples with continuous trajectories during the rest box before the W-Track experience and in the subsequent sleep epochs after the first exposure. We find that with experience on the track, the proportion of continuous replays increases on average in both control and transfected animals, and both groups of animals show an overlapping range of continuous trajectory lengths.

      (3) Theta sequence measurement in the absence of theta:

      We now explicitly explain why our manipulation makes it more appropriate to measure sequential hippocampal representations during locomotion (i.e., theta sequences) without using theta oscillation or an epoch-averaged, relatively large sliding window as a reference. The key insight here is that our manipulation suppresses theta and thus makes it difficult or impossible to accurately identify theta phase. We explain that while theta-phase-based approaches were used in prior work; these prior analyses may have confounded the absence of hippocampal theta sequences during locomotion by the inability to detect theta oscillatory phase reliably. We show that our method of using clusterless Bayesian decoding, in which we estimate the decoded position at every 2ms timestep, is indeed able to capture endogenous hippocampal sequences even without imposing any requirements of aligning to theta oscillations, thus providing an unbiased estimate of the rhythmicity of hippocampal spatial representations.

      (4) Additional analysis on place cell stability and tuning:

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Joshi and colleagues demonstrates that the precise theta-phase timing of spikes is causal for CA1 hippocampal theta sequences during locomotion on a linear track and is necessary for learning the cognitively demanding outbound component of a hippocampus-dependent alternation task (W-maze), independently of replay during immobility. To reach these conclusions, the authors developed a theta-phase-specific, closed-loop manipulation that used optogenetic activation of medial septal parvalbumin (PV) interneurons at the ascending phase of theta during locomotion. This protocol preserved immobility periods, allowing a clean and elegant dissociation from SWR-associated replay.

      The manuscript is well written and was a pleasure to read. The work described is of high quality and introduces several notable advances to the field:

      (a) It extends prior studies that manipulated theta oscillations by examining precise temporal structure (specifically theta sequences) rather than only LFP features.

      (b) The closed-loop manipulation enabled dissociation between deficits in theta sequences during a behavioural task and SWR-associated replay activity.

      (c) As controls, the authors included rats with suboptimal viral transduction or optic-fibre placement, and, within subjects, both stimulation-on (stim-on) and stimulation-off (stim-off) trials. Notably, sequence disruption persisted into stim-off periods within the same session.

      Overall, this is a strong manuscript that will provide valuable insights to the field. I have only minor comments:

      (1) As the authors note, it is striking that both behavioural performance and spike patterns are altered during stim-off trials. They propose that "disruption of theta sequences during the initial experience in an environment is sufficient to have lasting effects," implying that rapid, experience-dependent plasticity is driven by sequential firing. Does this imply that if rats were previously trained on the task, subsequent stim-on and stim-off trials would yield different outcomes, with stim-off trials showing improved performance and intact theta sequences? For example, if the sequence of one-third stim-on, one-third stim-off, one-third stim-on were inverted to off-on-off, would theta sequences be expected to emerge, disappear, and potentially re-emerge? While I am not asking for additional experiments, I think the discussion could be extended in this aspect.

      Alternatively, could the number of stim-off trials (one third of the total) be insufficient to support learning/induce plasticity? In the controls, ~50-100 trials appear necessary to achieve high performance.

      We think it is likely that pretraining would result in a different outcome, although we did not test this possibility. We have modified the discussion to address this point:

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to recover. This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (2) In line with the point above, the authors characterise the behavioural changes induced by MS optogenetic stimulation specifically as a "learning deficit," as rats failed to improve across 300 trials in an initially novel environment (W-maze). While they present this as complementary to prior demonstrations of impaired performance on previously learned tasks (Zutshi et al., 2018; Quirk et al., 2021; Etter et al., 2023; Petersen et al., 2020), an alternative interpretation is a working-memory deficit. This would produce the same behavioural pattern, with reference memory (the less cognitively demanding trials) remaining intact despite stimulation and concomitant changes in theta sequences. This interpretation would also be consistent with work in certain disease models, where reduced synaptic plasticity and working-memory deficits co-occur with preserved place coding despite impaired theta sequences (e.g., Viana da Silva et al., 2024; Donahue et al., 2025).

      We agree that traditionally deficits in alternation tasks have been termed “working memory” but we also note that this may confuse some readers, as memory in these tasks does not engage persistent activity throughout delays in areas like the prefrontal cortex.

      (3) It was not immediately clear whether SWR-associated activity was derived from the interleaved ~15-min rest sessions in a rest box, or from periods of immobility or reward consumption in the maze (aSWR, as in Jadhav et al 2012). Regardless, it would be informative to compare aSWR events within the maze to rest-box SWRs that may occur during more prolonged slow-wave episodes (even if not full sleep). This contrasts with Liu et al. (2024), who analyzed replay during ~1.5-h sleep sessions.

      We thank the reviewer for this comment and suggestion. We will now explicitly mention in the manuscript that we have measured awake sharp wave ripples (aSWRs) on the track during immobility periods. In addition, in line with this and another reviewer’s recommendation, we have included analyses on the proportion of rest SWRs (rSWRs) between control and targeted animals in Supplementary Figure 6, replicating our findings during aSWRs. However, we note that the differences between Liu et al.’s (2024) study and ours. While they waited and analyzed replay during 1.5 hours of sleep sessions, in our study, and in the 1-day w-track learning protocol, sleep sessions are typically shorter (15-20 minutes). That said, control animals in our tasks with an intact hippocampus (previous studies) and intact theta sequences (our study control animals) can learn the task in one day, so a longer replay period is not necessary to learn the task.

      Reviewer #2 (Public review):

      Summary:

      The authors of this study developed a closed-loop optogenetic stimulation system with high temporal precision in rats to examine the effect of medial septum (MS) stimulation on the disruption of hippocampal activity at both behavioral and compressed time scales. They found that this manipulation preserved hippocampus single-cell-level spatial coding but affected theta sequences and performance during a spatial alternation task. The performance deficits were observed during the more cognitively demanding component of the task and even persisted after the stimulation was turned off. However, the effects of this disruption were confined to locomotor periods and did not impact waking rest replay, even during the early phase of stimulation-on. Their conclusion is consistent with previous findings from the Pastalkova lab, where MS disruption (using different methods) affected theta sequences and task performance but spared replay (Wang et al., 2015; Wang et al., 2016). However, it differs from a recent study in which optogenetic disruption of EC inputs during running affected both theta sequences and replay (Liu et al., 2023).

      Strengths:

      The experiments were well designed and controlled, and the results were generally well presented.

      Weaknesses:

      Major concerns are primarily technical but also conceptual. To further increase the impact of this study by contrasting findings from different disruptions, it is necessary to better align the analysis and detection methods.

      We thank the reviewer for their assessment and critical questions. We have addressed each of the comments below. As we note in our responses, our inclusion criteria were based on our analysis approach, in which we aimed to measure the impact of our manipulation where possible for each animal, and ideally at the level of every 20-minute run epoch. This is a strength of our experimental approach, and we will explicitly explain that in a next version of the manuscript.

      Major concerns:

      (1) To show that MS disruption does not affect spatial tuning, the authors computed the KL divergence of tuning curves between stimulation-on and stimulation-off conditions. I have two main questions about this analysis:

      (1.1) The authors seem to impose stringent inclusion criteria requiring a large number of spikes and a strong concentration of tuning curves. These criteria may have selected strongly spatially tuned cells, which are typically more stable and potentially less vulnerable to perturbations. Based on the Figure 2 caption, it seems that fewer than 10% of cells were included in the KL divergence analysis, which is lower than the usual proportion of place cells reported in the literature. What is the rationale for using such strict inclusion criteria? What happens to the cells that are not as strongly tuned but are still identified as significant place cells?

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      (1.2) The KL divergence was computed between stimulation-on and stimulation-off conditions within the same animal group. However, the authors also showed that MS stimulation had lasting effects on theta sequences and performance even during stimulation-off periods. Would that lasting effect also influence spatial tuning? Based on these questions, the authors should perform additional analyses that directly measure spatial tuning quality and compare results across control and experimental groups - for example, spatial information of spikes (Skaggs et al., 1996), tuning stability, field length, and decoding error during running.

      To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). Mean decoding error between animals depends on the recording quality. We have reported these values around stimulus times at the choice point in the previous version of the manuscript (Sup. Figure 3G). We also find that the distribution of place field coverage is not statistically different within and across animals (stimulation on versus stimulation off: p = 0.17, control versus targeted: p = 0.9, interaction of stimulation and targeting: p = 0.8, linear mixed effects model).

      (2) The authors compared their results with those from Liu et al. (2023) and proposed that the different outcomes could be explained by different sites of disruption. However, the detection and quantification methods for theta sequences and replay differ substantially between the two studies, emphasizing different aspects of the phenomenon. I am not suggesting that either method is superior, but providing additional analyses using aligned detection methods would better support the authors' interpretations and benefit the field by enabling clearer comparisons across studies. In the current analysis, the power spectrum of the decoded ahead/behind distance only indicates that there is a rhythmic pattern, without specifying the decoding features at different theta phases. Moreover, the continuous non-local representations during ripples could include stationary representations of a location or zigzag representations that do not exhibit a linear sequential trace. Given that, the authors should show averaged decoding results corrected by the animal's actual position within theta cycles and compute a quadrant ratio. For replay analysis, they could use a linear fit (as in Liu et al., 2023) and report the proportion of significant replay events.

      In Liu et al., 2023 study theta sequences were quantified by explicitly segmenting theta cycles into phase quadrants and evaluating the structure of the decoded representations within these phase-defined windows. This approach can be applied in studies that have a stable theta that can provide a temporal reference frame and where there is not enough spatial coverage in the spikes to identify the extent of ahead/behind representations. This analysis is hence not ideal to detect theta sequences in our data, as it is prone to errors due to our disruption of theta oscillatory activity itself. To account for this, we have used a Bayesian clusterless decoding approach in which we can measure the structure of hippocampal sequential representations even without imposing any restrictions on their temporal order. This method allows us to reliably capture theta sequences during locomotion in control animals (Fig. 4B) and their disruption in targeted animals (Fig. 4F).

      Since our experimental paradigm suppresses theta oscillations themselves, we used a clusterless decoding approach (as in Joshi et al., 2023) to obtain an unbiased estimate of rhythmicity for hippocampal spatial representations. Briefly, we estimated the peak of the posterior at every 2ms time step and computed the distance between that value and the actual position of the animal (decode-to-animal distance). We confirmed that, as expected, in control animals, we could detect “theta sequences” as in prior studies without explicitly requiring theta oscillatory cycle windows.

      We have now evaluated the distributions of SWRs that are labeled as continuous and have a trajectory displacement > 10 cm and consistently find continuous replays during both aSWRs and rSWRs (Fig. 5, Fig. S6). We also find that these distributions overlap between control and targeted animals (n=1216 targeted, n=3212 control, p=0.06).

      Author response image 1.

      (3) The finding that theta sequences and performance were impaired even during stimulation-off periods is particularly interesting and warrants deeper exploration. In the Discussion, the authors claim that this may arise from "the rapid plasticity engaged during early learning." However, this explanation does not fully account for the observation. Previous studies have shown that theta sequences can develop very rapidly (Feng et al., Foster lab, 2015; Zhou et al., Dragoi lab, 2025). If the authors hypothesize that rapid plasticity during early stimulation-on disrupts the theta sequence, then the plasticity window must also be short and terminate during the subsequent stimulation-off period. Otherwise, why can't animals redevelop theta sequences during stimulation-off? The authors should conduct additional analyses during the stimulation-off periods of the W-maze task. For example:

      (3.1) What is the spike-theta phase relationship? Do the phases return to normal or remain altered as during stimulation-on?

      We thank the reviewer for this question. We have now looked at theta power on the W Track and find that theta power does not fully recover on the W Track even during stimulation-off periods. We have included this in the results. See Figure S4.

      (3.2) Is there a significant place-field remapping from stimulation-on to stimulation-off? (Supplementary Figure 3F includes only a small subset of cells; what if population vector correlations are computed across all cells, or Bayesian decoding of stimulation-on spikes is performed using stimulation-off tuning curves?)

      We have addressed this question by computing the correlation between the peak of the place fields between stimulation-on and stimulation-off conditions and find that the distributions are largely overlapping between control and targeted animals on both the linear (Spearman correlation of place field peaks between stim-on and stim-off intervals, control vs targeted, p=0.68, t-test, n=6 control, n=4 targeted epochs) and wtrack (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.91, t-test, n=36 control, n=18 targeted epochs).

      (3.3) The authors should also discuss why the stimulation-off epochs were not sufficient to support learning, and if the stimulation-off place cell sequences could have supported replay.

      We do not find the aSWR-associated replay to be impacted as a result of our manipulation. We have not conducted a specific experiment to test the impact of longer stimulation-off periods on the formation of place cell sequences, but in response to this and another question from Reviewer 1, we have added the following speculation in the discussion.

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (4) Citations and/or discussion of key studies relevant to the current work are missing: Wang et al. in Pastalkova lab 2015-2016 studies for disruption of theta sequence (but not place cell sequence) disrupting learning but not replay, Drieu et al. in Zugaro lab 2018 study on disruption of theta sequence affecting sleep replay, Farooq and Dragoi 2019 for association between a lack of theta sequence and presence of waking rest replay during postnatal development, etc. The authors should discuss what the conceptually new findings in the current study are, given the findings of the previous literature above.

      We thank the reviewer for this question. We have substantially modified the introduction to include this prior work and highlight that our manipulation enabled us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (5) The assessment of theta sequence is not state-of-the-art:

      (5.1) Detecting the peak of cross-correlograms between neurons (CCG) relates to behavioral timescale CCG, not the theta sequence one; for the theta sequence, the closest to zero local peak should be used instead.

      Here we think we failed to explain our analyses clearly, as we did exactly that analysis. The cross-correlation peak in Fig. 2D is the peak within the theta timescale (+/- 100ms lag), not a slow behavior timescale (e.g. +/- 1s lag). In the revised version of the manuscript, we have improved the explanation so that this confusion does not arise.

      (5.2) How were other methods of detecting theta sequences performing on the stimulation-on/stimulation-off data: Bayesian decoding, firing sequences?

      In the absence of detectable theta oscillations, it is inappropriate to use the commonly used metric of detecting theta sequences. That method also assumes that the theta oscillatory cycle is the correct temporal “reference” for hippocampal theta sequences. To our knowledge, there is no direct evidence for this. Thus, in our manuscript, we have used two approaches to identify hippocampal spatial representations during stimulation-on and stimulation-off periods:

      (1) Clusterless Bayesian decoding approach

      (2) Pairwise correlations between neurons

      Analyses using these approaches provide an unbiased method to detect hippocampal spatial-temporal sequences during locomotion without using theta oscillations as a reference. Indeed, in control animals, we recover the endogenous theta timescale correlation and sequence structure during locomotion. Using the same approach in targeted animals reveals that even though we can measure hippocampal spatial sequential representations in targeted animals, the timing between them is altered.

      (5.3) How was phase precession during stimulation-on/stimulation-off?

      We cannot do this analysis for the theta manipulation condition since there is an unreliable phase estimate in the absence of theta on W Track. Based on the reviewers’ comments, we have now evaluated the theta phase precession on the linear track 10Hz stimulation condition. Phase precession could be observed even when evaluated against the entrained LFP. Further, at a population level, we observed that autocorrelograms followed the entrained 10Hz LFP in the stimulation-on condition compared to the stimulation-off condition (n=41 neurons, solid lines are medians, shaded areas 25/75 percentiles). We have added these results in Supplementary Figure 5.

      (6) It would be important to calculate additional variables in the replay part of the study to compare the quality of replay across the 2 groups:

      (6.1) Proportion of significant replay events out of the detected multiunit events.

      We assume the reviewer is recommending the analysis of significant replays as defined by performing a linear fit on the trajectory. We recognize that there is a fundamental diversity in the replay architecture, as has been shown using clustered and clusterless decoding approaches, and that assuming that only the replays with a linear fit are significant might bias us toward those trajectories. In aSWRs, we have shown that the replay events that are labeled “continuous” are equivalent in proportion between control and targeted animals. In the current version of the manuscript, for rSWRs, we similarly computed the proportion of replay events with a continuous trajectory that traverses at least 10cm on the w-track and found no differences between control and targeted groups. We also found the proportion of continuous rSWRs to increase in targeted animals, similar to control animals. We have included these results in the new Figure S6.

      (6.3) The average extent of trajectory depicted by the significant replay events in the targeted compared to the control, stimulation-on/stimulation-off.

      Here, we show the distributions of continuous replay trajectories (> 10 cm) between control and targeted animals, showing overlapping distributions (p=0.06). See Author response image 1.

      Reviewer #3 (Public review):

      Joshi et al. present an elegant and technically rigorous study examining how the temporal structure of hippocampal spiking during locomotion contributes to spatial learning. Using a closed-loop, theta phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats, the authors demonstrate that disrupting theta-timescale coordination impairs performance on the cognitively demanding component outbound trajectory of a spatial alternation task, while sparing hippocampal replay, place coding, and the simpler inbound learning. The work aims to dissociate the role of theta-associated temporal organization during navigation from sharp-wave ripple-associated replay during subsequent rest periods, providing a mechanistic link between theta sequences and learning. The findings have important implications for models of septo-hippocampal coordination and the functional segregation between online (theta) and offline (SWR) network states. That said, there are a few conceptual and methodological issues that need to be addressed.

      We thank the reviewer for the thoughtful comments and suggestions to strengthen our work. In a revised manuscript, we explicitly address prior publications where manipulations (either pharmacological or behavioral) have shown a dissociation between temporal sequence formation, place coding, and replay in the introduction and highlight the specific insights obtained from using our approach. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      One concern is the overall novelty of this work; the dissociation between online temporal sequence and offline replay events following memory deficits has previously been shown by Wang et al., 2016 elife. While the authors discuss Lui et al., 2023, which demonstrates MEC activation of inhibitory neurons at gamma frequencies during locomotion disrupts theta sequences, subsequent replay and learning (line 65-66), they do not reference Wang et al., 2016 who performed a very similar study with MS pharmacological inactivation, and report large decreases in theta power, attenuated theta frequencies together with behavioural deficits but SWR replay persisted. Given strong similarities in the manipulation and findings, this study should be discussed.

      We agree that this important study should be cited in the introduction, and have now included that in the revised introduction. Importantly in Wang et al., 2016 elife paper, they showed that while replay can exist after periods where theta sequences are disrupted using muscimol inactivation in the medial septum, the rate of replay was higher than in control, leaving open the possibility that the increased replays might contribute to consolidation of non-task memories. We also note that the manipulation was applied for hours, making it impossible to attribute a precise physiological basis for poor behavioral performance.

      Along the same lines, it should be noted that Brandon et al. (2014, Neuron) demonstrated that hippocampal place codes can still form in novel environments despite MS inactivation and loss of theta, indicating that spatial representations can emerge without intact septal drive. Referencing this study would strengthen the discussion of how temporal coordination, rather than spatial coding per se, underlies the learning deficits observed here.

      Thank you. We have referenced the study appropriately in the updated manuscript.

      Our findings, and previous dissociations between precise timescale and place field properties (Petersen and Buzsáki 2020; Liu et al. 2023; Wang, et al., 2016, Brandon et al., 2014) further suggest that different circuits with different time constants are responsible for processing spatial and temporal information in the hippocampal circuit. Spatial/contextual information may arrive from regions with slower timescales (such as the cortex), making them less susceptible to sub-second brief disruptions, while precisely timed inputs from the medial septum coordinate the tightly controlled timing offsets between hippocampal neurons. We hypothesize that learning requires the intersection of these two streams of information in the hippocampal network, and is impaired by the disorganization of the precise temporal templates in which internal plans can be matched to external inputs.

      The conclusion that disrupting "theta microstructure" impairs learning relies on the assumption that the observed behavioral deficits arise from altered temporal coding from within hippocampal CA1 only. However, optogenetic modulation of medial septal PV neurons influences multiple downstream regions (entorhinal cortex, retrosplenial cortex) via widespread GABAergic projections. While the authors do touch on this, their discussion should expand to include the network-level consequences of entorhinal grid-cell disruption and how this could affect temporal coding both online and offline.

      We agree and have expanded our previous discussion to include that possibility.

      Modified in discussion:

      “Understanding precisely why this temporal organization is critical will require more distributed measurements. Notably, MS targets include multiple cortical and subcortical targets (Joshi 2017), and our manipulation may have disrupted precise spike timing throughout these regions. Key amongst these regions include the pre- and para-subiculum, retrosplenial area and the entorhinal cortex, which also receive dense PV projection in addition to the CA3 and DG (Joshi et. al., 2017; Viney et. al., 2018; Salib et al., 2020). Disrupting the spatial code or spike-timing in these regions may contribute to the disruption of sequential activity we have observed. However, we do note that the first response of the stimulation to spiking activity in CA1 is consistent with a strong disinhibitory input to CA3, with spike latencies less than 20 milliseconds. Additionally, monitoring regions beyond the temporal cortex would be informative given the broad coordination between hippocampal theta and other systems (Joshi et al. 2023; Eichenbaum 2017; Buño and Velluti 1977; Berg, Whitmer, and Kleinfeld 2006; Ledberg and Robbe 2011).”

      The finding that replay content, rate, and duration are unchanged is critical to the paper's claim of dissociation. However, the analysis is restricted to immobility on the track. Given evidence for distinct awake vs. sleep replay, confirming that off-track rest and post-session sleep replays are similarly unaffected would confirm the conclusions of the paper. If these data are unavailable, the limitation should be acknowledged explicitly. Moreover, statistical power for detecting subtle differences in replay organization or spatial bias should be added to the supplement (n of events per animal, variability across sessions).

      We thank the reviewer for this suggestion. We have now explicitly evaluated replay properties during rest and indeed confirm that replay rate, length, and content in this manipulation are indistinguishable between targeted and control animals. We have now added statistical power to these claims by reporting the number of events per animal and variability across sessions. As we mentioned above, where possible, we have attempted to replicate each analysis per session (20min session), and all analyses are replicable for each targeted and control animal. Our internal controls are the power of our approach, as there can be significant animal-to-animal variability.

      The exact protocol for optogenetic stimulation is a bit confusing. For the task, the first and final third (66%) of trials were disrupted and were only stimulated when away from the reward well and only when the animal was moving. What proportion of time within "stimulated" trials remained unstimulated? Why were only 66% of trials stimulated?

      We developed a stimulation protocol that included an interval in which stimulation was not applied (stimulation-off) periods to assess the impact of the stimulation at the level of each animal and epoch. This analytical approach gives us the power to study the impact of the stimulation and our experimental approach to the spiking patterns and neural activity observed. We will modify the explanation in the methods. Based on the reviewer’s comment, we have now calculated the proportion of time within the epoch where the laser was on. The laser is ON for ~10% of the total time in the epoch. We implemented a closed-loop algorithm that had both the spatial location of the animal and the theta phase of a reference electrode as online inputs. The trigger was applied when three conditions were met: a spatial inclusion criterion, a speed criterion, and theta phase criteria.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1D & G: should include the frequency band of the filtered trace in the figure caption.

      We have now included the frequency band of the filtered trace (5-11Hz) in the figure caption.

      (2) The observation that hippocampal cells respond to stimulation within ~5-10 ms (Figure 2C), while theta power decreases significantly after 200 ms (Figure 1G), is interesting. Do the authors have any hypothesis explaining this discrepancy?

      Theta is likely best understood as the result of a complex feedback loop between the medial septum and various structures in the hippocampal formation. This would make it robust to individual perturbations. We now discuss this in the text.

      Modified in results:

      “Note, that while we can measure the response to hippocampal cells within 5-10ms, theta power decreases gradually over a period of 200 milliseconds. This is consistent with the view that theta oscillatory activity that can be measured in the hippocampus is a result of a multi-region feedback loop that involves various cortical and subcortical networks, a feature that may make it robust to individual perturbations.”

      (3) Figure 2A: The caption states that "in control rats, endogenous theta sequences are apparent." This is unclear. The purple region marks the ascending phase, but theta sequences are typically defined across peak-to-peak cycles. The shading may distract from this; consider adding dashed lines to indicate theta peaks.

      We thank the reviewers for this suggestion, and we have modified the visualization for this figure per the reviewers’ suggestion.

      (4) Figure 2B (top panel): How are the cells ordered? It appears they are sorted by peak firing time before stimulation (<0). This could misleadingly suggest a sequence before stimulation but not after. Some cells have peak firing in the 0-40 ms range-what is the intended interpretation of this ordering?

      We have modified the spiking by the spiking order after stimulation onset.

      (5) Line 239: Typo - "as for the linear track" should be "as for the W track."

      Modified.

      (6) Lines 279-282: The statement "This suggests a surprising level of preserved representational movement across frequencies ranging from 6 to 12 Hz" is unclear. What does "preserved representational movement" mean? A simpler explanation for the matching of power-spectrum peaks to stimulation frequency could be that stimulation increases firing rates, biasing decoding toward locations with higher mean firing, producing rhythmic fluctuations at the stimulation frequency under Poisson decoding assumptions.

      We have now added examples of the decode to animal distance in the different stimulation conditions to supplement this statement. We will also acknowledge that the change in firing rate might contribute to the observation.

      (7) Line 313: The phrase "while sparing learning on the interleaved trials where a less cognitively demanding choice was required" may be confusing. Although the authors refer to inbound runs, readers may interpret "interleaved trials" as stimulation-off trials.

      We agree and have modified this phrasing.

      (8) The claim that "representations of locations more distant from the animal were preserved during theta disruption" (line 350) is unclear-where is this shown?

      Fig S3G: Distribution of max decode to animal distance within theta cycles

      Consistent with the pairwise analyses (Supplementary Figure 3B-D), the ∼8 Hz peak in the power spectrum of the ahead/behind distance was significantly larger in control animals than in targeted animals across all conditions (Figure 4I; pooled comparison p’s < 10<sup>−4</sup>, hierarchical bootstrap p’s < 0.05). At the same time, the maximal extent of locations represented ahead and/or behind the animal near the choice point, where choices must be made on outbound and inbound trials, did not differ between control and targeted animals (Supplementary Figure 3G). Thus, our findings indicate that the precise timing of non-local representations during theta was disrupted, but the spatial extent was not.

      (9) The title of Figure S1 appears twice.

      Thank you. We have edited it.

      Reviewer #3 (Recommendations for the authors):

      (1) It should be noted that there are several referencing errors that should be addressed. Please check the following:

      (a) Lines 77-81 - a few incorrect references.

      (b) Line 155 - Disruption of hippocampal theta has consistently shown to preserve spatial properties of place cells: Brandon et al., 2014, Koenig et al., 2011.

      Thank you. We have corrected these references.

      (2) Figure 1C- Did the authors stain for colocalization with virus and PV in the MS?

      Yes, our viral construct has eYFP expressed together with channelrhodopsin, and this rat line and viral construct have been previously standardized (Yu et al., 2018, Lepperod et al., 2021). In addition, we conducted three standardization experiments and visually inspected the overlap between eYFP-positive cells and parvalbumin-expressing neurons (86/86 YFP-expressing neurons tested positive for PV). An example of the overlap is now included in Supplementary figure 7.

      (3) Figure 1 - Shows an impressive reduction of theta power. Can Figure 1G be extended to show theta recovery immediately following stimulation?

      Our stimulation protocol restricted the stimulation to periods that were within the spatial and speed inclusion criteria. The period immediately following the stimulation on every trial is reward delivery, during which the animal has already slowed down, and we do not expect high theta power. Based on the reviewer’s suggestion, we have inspected the first few trials after the stimulus is turned off in the linear track and have confirmed that theta power immediately recovers. We have added these additional figures in Sup. Figure S4H.

      Author response image 2.

      (4) Do authors have examples of the same trajectory where temporal coding is intact in baseline and disrupted during stimulation? Does an intact theta sequence ever develop in the target animals?

      Yes, target animals do exhibit intact spatial sequences during the stimulation-off periods on the linear track and w-track (example in a targeted animal during the stimulation-off condition on linear track below). As we show in Figure 4E and Supplementary Figure S4G, on the w-track, while spatial sequences exist, each cycle’s duration is not consistent across the behavioral experience. Thus, on average, power spectrum of the ahead-behind distance is not rhythmic at 8Hz.

      (5) Have the authors computed spike-phase relationships or shown phase-position to evaluate phase precession in individual cells in both the 10 Hz stim vs the phase-specific?

      As discussed above, due to unreliable phase estimates during theta suppression, we have chosen to base our analysis of sequential structure on temporal cross-correlation and decoding analysis. However, in response to the reviewer’s question, we have evaluated phase precession under 10Hz stimulation. Consistent with our overall results for the ahead-behind distance, we find that individual cells phase-precess with the newly entrained theta. Additionally, at a population level, we are able to visualize a clear shift in peak spiking frequency.

      However, we agree that these results do not completely rule out the contribution of additional spikes, and in a revised version of the manuscript, we have included that possibility.

      (6) Did authors perform other types of stimulations that either drove the dominant frequency out of theta range (gamma) or completely desynchronize the system (by stimulating along all different times of theta to perform a phase-specific "scramble")?

      We did not attempt a gamma or scrambling stimulation condition.

      (7) There is overall inconsistency in the formatting of references throughout the text.

      We apologize for these errors and have rectified them in the updated version.

    1. eLife Assessment

      This important study uses biochemical and single-molecule approaches to characterize how yeast Cdc13 assembles on single-stranded telomeric DNA and contributes to telomere protection. The evidence supporting the proposed model for Cdc13 assembly and its role at telomeres is solid. The findings provide insight into how Cdc13 contributes to the maintenance and protection of chromosome ends. This work will be of interest to researchers in telomere biology, DNA replication, biochemistry, and single-molecule biophysics.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how the Saccharomyces cerevisiae telomere-binding protein Cdc13 assembles on a 12-nucleotide single-stranded telomeric DNA substrate. Using complementary smFRET and CoSMoS measurements, together with dimerization- and DNA-binding-defective mutants, mass photometry, photobleaching analysis, and kinetic modeling, the authors assign state II to a DNA-bound Cdc13 monomer and state III to a stable complex containing two Cdc13 molecules. They propose that the stable DNA-bound dimer forms predominantly through sequential recruitment of two monomers, while direct binding of a preformed dimer represents a less frequent pathway. This work addresses an important mechanistic question in telomere biology because the pathway of Cdc13 assembly may influence telomere recognition, end protection, and recruitment of telomere-maintenance factors.

      Strengths:

      Overall, the manuscript is well written, and the combination of two complementary single-molecule approaches is a strength. The central model is interesting and potentially important.

      Weaknesses:

      Several issues require clarification or additional analysis.

      Major comments:

      (1) The mass photometry results are central to the mechanistic model and should be presented more prominently.

      The conclusion that Cdc13 binds DNA predominantly through sequential monomer recruitment depends strongly on the oligomeric state of Cdc13 at the concentrations used in the single-molecule experiments. The mass photometry results currently provide the principal direct evidence that Cdc13 is predominantly monomeric at 5 nM but exists as a mixture of monomers and dimers at 20 nM. These results should therefore be included in a main figure rather than only in the Supplementary Information. The authors should also provide a complete description of the mass photometry in the Methods section.

      Related to lines 191-192, the authors should discuss the estimated cellular or nuclear concentration and abundance of Cdc13 and compare these values with the experimental concentrations at which states II and III are populated. Because the relevant quantity may be the effective local concentration at a telomere rather than the average nuclear concentration, this distinction should also be acknowledged. Such a discussion is needed to establish under what physiological conditions sequential monomer loading versus binding of a preassembled dimer would be expected.

      (2) 75% labeling efficiency must be clearly defined and incorporated into both the stoichiometric and kinetic analyses.

      In line 251, the authors state that the labeling efficiency of DY-649P1-Cdc13 is 75%, but it is not clear how this value was measured. The authors should state whether 75% refers to the efficiency of the sortase reaction, the fraction of labeled molecules in the final purified preparation, or a value inferred from the plateau in Figure 3B. If it was inferred from the binding plateau, the plateau below 100% could also arise from inactive or inaccessible DNA molecules, incomplete colocalization detection, inactive protein, or an effect of the fluorophore on binding. An independent measurement, such as absorbance-based determination of the dye-to-protein ratio, quantitative gel analysis, or intact-mass analysis, would be preferable.

      Incomplete labeling has direct consequences for the interpretation of Figure 3. With a labeling probability of 0.75, a true Cdc13 dimer would contain zero, one, or two fluorophores. Thus, even among detectable dimers, 40% would appear as one-step photobleaching events. A one-step event therefore cannot automatically be equated with a monomer without correcting for labeling efficiency. The authors should quantitatively account for incomplete labeling when inferring the relative monomer and dimer populations from the photobleaching data.

      The same issue is even more important for the CoSMoS kinetic model. The authors should incorporate labeling efficiency into the observation model, or at minimum perform simulations or a sensitivity analysis demonstrating that the inferred transition rates and state assignments are robust to 75% labeling.

      (3) Apparent state I→III transitions do not by themselves demonstrate direct binding of a preformed Cdc13 dimer.

      At lines 326-330, the authors interpret state I→III transitions as direct binding of a solution dimer and conclude that sequential binding is approximately eightfold faster than direct dimer binding. This interpretation is not yet sufficiently established. An observed I→III transition demonstrates only that no state II intermediate was resolved; it does not distinguish true binding of a preformed dimer from sequential binding in which the state II lifetime is shorter than the temporal resolution of the experiment.

      This concern is particularly important because the CoSMoS experiments were conducted at only 0.3-1.25 nM Cdc13, whereas mass photometry indicates that Cdc13 is predominantly monomeric even at 5 nM. In addition, the smFRET signals were averaged over a sliding window of ten 50-ms frames, which could obscure short-lived intermediate states.

      The authors should show representative raw traces containing apparent I→III transitions in a supplementary figure and quantify the shortest state II dwell time that could be detected under the acquisition, smoothing, and HMM procedures used.

      (4) The interpretation of the WT-Cdc13/Cdc13^R635C mixture requires further clarification.

      For Figures 2G-H and lines 209-218, the reduced state III population in the mixture of 2.5 nM WT Cdc13 and 2.5 nM Cdc1^R635C is interpreted as evidence for a solution monomer-dimer equilibrium and formation of a nonfunctional WT-mutant heterodimer. However, at least two nonexclusive explanations should be considered:

      a) Formation of WT-mutant heterodimers in solution could reduce the concentration of free WT monomers and WT homodimers available to form state III.<br /> b) A WT-mutant heterodimer, or recruitment of Cdc13^R635C to a DNA-bound WT molecule through protein-protein interactions, could produce a DNA-bound complex that cannot adopt the state III conformation because only one subunit has an intact DNA-binding interface.

      The authors should discuss these possibilities explicitly and clarify expected FRET states.

      For direct visual comparison, Figure 2G should include the FRET histograms for 2.5 nM WT Cdc13 alone and 5 nM WT Cdc13 alone, in addition to the WT-mutant mixture. The concentrations of both the initially loaded WT Cdc13 and the WT or mutant protein added during the chase experiment in Figure 2I should also be stated in the main text and figure legend.

      (5) The physical basis of the different FRET values for states II and III should be explained earlier and more carefully.

      The assignment of state II and state III to one and two bound Cdc13 molecules is supported by the combined smFRET and CoSMoS results. However, the manuscript should explain earlier why the addition of a second Cdc13 molecule is expected to produce a further decrease in FRET. Because the fluorophores are attached to the DNA, the different FRET values imply a change in the distance, orientation, or local photophysical environment of the DNA-linked dyes when the second Cdc13 binds. CoSMoS establishes a change in protein stoichiometry, but it does not by itself establish that the DNA has undergone a particular conformational change.

      The discussion at lines 404-423 suggests that the second Cdc13 induces a rearrangement of the first Cdc13-DNA complex. This is a reasonable hypothesis, but it should be presented as an inference rather than as a demonstrated DNA conformational transition. References 44-46 describe different RPA binding modes and rearrangements of protein-DNA contacts; they do not directly demonstrate the specific DNA conformational change proposed here. The authors should either provide more direct support or revise the discussion accordingly. The distance estimates should also be described cautiously because they assume that dye orientation and photophysical properties are unchanged between states.

      The rationale for the internally positioned Cy3 constructs in lines 147-149 should also be explained more clearly. Why does it demonstrate that Cdc13 cannot bind duplex DNA?

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript presents an interesting and potentially important single-molecule study of Cdc13 assembly on telomeric ssDNA. The experimental observations are intriguing, particularly the identification of distinct FRET states associated with different Cdc13 occupancies. However, I have substantial concerns about whether the current data support the central mechanistic conclusion as strongly as the authors claim. In particular, the manuscript does not yet clearly distinguish between the observation that two Cdc13 molecules are associated with DNA and the stronger mechanistic claim that Cdc13 loads sequentially as monomers and subsequently dimerizes on DNA.

      Major concerns

      (1) The central mechanistic conclusion is stronger than the evidence

      The manuscript's principal model is that Cdc13 proceeds through the pathway monomer → DNA-bound monomer → recruitment of a second monomer → stable DNA-bound dimer. However, the experiments establish this sequence only indirectly.

      The authors show that FRET state II is associated with one Cdc13 molecule. FRET state III is associated with two Cdc13 molecules. Cdc13-DM produces state II but not state III. WT Cdc13 can undergo I→II→III transitions. Direct I→III transitions also occur.

      These observations are consistent with sequential binding, but they do not uniquely demonstrate that the second Cdc13 molecule first binds as a monomer and subsequently undergoes dimerization on DNA. In particular, the direct I→III events indicate that a preassembled dimer or another cooperative pathway can also contribute.

      I therefore recommend substantially tempering statements such as "Cdc13 initially loads onto telomeres as a monomer." A more defensible formulation would be: "The data support a kinetically favored sequential pathway involving an initial Cdc13 binding event followed by recruitment of a second Cdc13 molecule." The authors could still present sequential loading as the preferred model, but the language should clearly distinguish a kinetically supported pathway from a uniquely established molecular mechanism.

      (2) The assignment of FRET states II and III to exactly one and two Cdc13 molecules requires stronger validation

      The interpretation of states II and III as one- and two-Cdc13 states is central to the entire mechanistic model. Because labeling efficiency, incomplete labeling, photophysics, and heterogeneous molecular populations can all affect the observed distributions, the authors should provide a quantitative probabilistic model. Specifically, the authors should calculate the expected distribution of one- and two-labeled-Cdc13 species given the experimentally determined labeling efficiency and compare these expectations with the observed FRET-state populations. This analysis would provide an important independent validation of the proposed stoichiometric assignments.

      (3) The TG12 substrate raises an important stoichiometric and conformational question

      The manuscript argues that a single Cdc13 molecule binds approximately 11 nt, while two Cdc13 molecules can bind a 12-nt ssDNA substrate. This raises an immediate mechanistic question: if one Cdc13 occupies approximately 11 nt, how can two Cdc13 molecules simultaneously associate with only 12 nt of ssDNA? This issue should be addressed experimentally rather than only structurally or schematically. A particularly informative experiment would be to systematically vary ssDNA length. The probability and kinetics of state III formation could then be quantified as a function of substrate length. Such an experiment would determine whether formation of the two-Cdc13 state genuinely requires additional DNA and whether the two proteins occupy overlapping or distinct regions of the substrate.

      (4) The "salt-resistant" interpretation is overstated

      The manuscript repeatedly describes state III as "salt-resistant" and uses this observation to support the existence of a highly stable physiological complex. However, the reported experiment examines only a relatively modest range of NaCl concentrations (50, 75, and 100 mM).

      I recommend either expanding the salt-dependence analysis substantially or using more quantitative language. For example, the authors could report the fraction and lifetime of state III as a function of salt concentration and define explicitly what they mean by "salt-resistant." The current data support persistence under the tested conditions, but they do not by themselves establish exceptional physiological stability.

      (5) The mass-photometry experiment does not establish the physiological solution equilibrium

      The mass-photometry data are useful for demonstrating that Cdc13 can exist in monomeric and dimeric forms, but the current experiment does not establish the equilibrium between these species under physiologically relevant conditions.

      A concentration series would be valuable. The observed monomer/dimer populations should be fit to an explicit equilibrium model to obtain an apparent dimerization constant and assess how strongly the equilibrium depends on Cdc13 concentration. This would also help connect the solution behavior to the single-molecule observations and determine whether the observed DNA-bound dimer could plausibly arise from a pre-existing solution dimer.

      (6) The WT + R635C mixing experiment is overinterpreted

      The interpretation of the WT + R635C experiment is currently complicated and somewhat speculative. The experiment is potentially informative, but the conclusions appear stronger than what can be directly inferred from the data. The authors should clearly distinguish between the observations directly supported by the mixing experiment and the mechanistic interpretation proposed from them. In particular, the experiment does not necessarily establish the precise sequence of DNA binding and dimerization events.

      (7) The requirement for DNA-binding activity in both Cdc13 molecules is not fully established

      The manuscript concludes that both Cdc13 molecules must possess DNA-binding activity. However, the R635C experiment does not clearly distinguish between:<br /> two independently DNA-bound Cdc13 molecules; and one Cdc13 molecule directly bound to DNA plus a second molecule whose DNA-binding surface is required for allosteric stabilization of the dimer.

      This distinction is mechanistically important. Additional experiments using DNA-binding-defective mutants in defined heterodimeric configurations would help determine whether both molecules directly contact DNA or whether DNA binding by one molecule promotes recruitment/stabilization of the second through protein-protein interactions.

      (8) Stronger controls are needed for FRET-state assignment

      The FRET states are treated as discrete molecular states, but alternative explanations for heterogeneous FRET populations should be considered more explicitly.

      Important controls would include: concentration-dependent FRET measurements in the absence of DNA binding; fluorescence controls to determine whether the observed states could arise from dye-protein interactions; Cdc13 mutants with altered DNA-binding specificity; alternative dye positions; demonstration that the major FRET states are reproduced with independent labeling configurations. The duplex-positioned Cy3 controls, which show little FRET change, are useful. However, they do not completely exclude the possibility that protein-induced changes in DNA conformation contribute to the observed FRET states. Independent labeling geometries would substantially strengthen the assignment.

      (9) The physiological relevance of the 12-nt substrate requires better justification

      The authors use TG12 as their primary substrate and state that telomeres contain approximately 12-14 nt of ssDNA during most of the cell cycle. This rationale requires greater biological context. Telomere length and the extent of the exposed G-rich strand are dynamic and heterogeneous, and Cdc13 has established functions throughout telomere replication. The authors should explain more carefully why TG12 is biologically representative and how the proposed mechanism is expected to behave on substantially longer telomeric substrates. The TG25 experiment is useful, but at present it functions primarily as a stoichiometric observation. A systematic substrate-length analysis, as suggested above, would turn this observation into a mechanistic test.

      (10) The relationship to existing structural studies requires deeper discussion

      The manuscript presents sequential loading as a novel mechanism, but existing structural and biochemical studies of Cdc13 dimerization and Cdc13-DNA architecture are essential for interpreting these observations. The authors should explicitly reconcile their proposed model with the existing structural literature. In particular, they should address:

      Does the known Cdc13 dimerization interface permit simultaneous DNA binding by both subunits?

      Is the dimerization interface compatible with the proposed DNA-bound state II?

      Could DNA binding alter the dimerization interface?

      Are the two Cdc13 molecules predicted to bind overlapping or distinct portions of the telomeric sequence?

      Can the structural models accommodate the apparent stoichiometry on a TG12 substrate?

      Without this reconciliation, the proposed sequential-loading mechanism remains somewhat disconnected from the established structural framework.

      Other important issues:

      (11) The Kd comparisons are confusing

      The manuscript should include a table summarizing the different Kd values and explicitly explaining why they differ. In particular, describing 3.7 nM as the Kd for the first monomeric binding step while reporting an apparent Kd of 1.2 nM requires careful kinetic and statistical justification. The authors should distinguish clearly among microscopic Kd values, apparent Kd values, and parameters inferred from kinetic models.

      (12) The direct-dimer pathway deserves greater attention

      The occurrence of direct I→III transitions is mechanistically important and should not be treated primarily as an exception to the sequential pathway. A more balanced conclusion would be:

      "Both pathways contribute to formation of the final Cdc13-DNA complex, with the sequential pathway being kinetically favored under the experimental conditions." This interpretation appears better aligned with the data and would still constitute a strong mechanistic conclusion.

      (13) The "kinetic proofreading" interpretation is currently speculative

      The Discussion proposes that the first Cdc13 monomer provides a kinetic proofreading step. This is an interesting hypothesis, but it is not directly demonstrated by the current experiments.

      I recommend changing this to language such as "a kinetic proofreading-like mechanism may be possible" unless the authors can provide direct evidence that the first binding event selectively promotes productive complex formation or rejects nonproductive substrates.

      (14) Biological-function claims should be clearly separated from the in vitro findings

      The manuscript frequently connects the stable state III complex with telomere protection, telomere length regulation, CST formation, Est1 recruitment, Pol α recruitment, and prevention of DNA degradation. None of these functions are directly tested in the present study.

      The experiments establish a biochemical/single-molecule mechanism in vitro. They do not establish that state III is the functional protective species in vivo.

      This distinction should therefore be maintained throughout the Abstract, Discussion, Key Points, and concluding statements. The authors can appropriately discuss these possibilities as<br /> implications or hypotheses, but should avoid presenting them as demonstrated functions of state III.

    1. eLife Assessment

      This valuable proof-of-concept study tests and validates a novel magnetic resonance-based method - proton-observed proton-edited MRS - for tracking glucose metabolism in the brain beyond simple substrate uptake, without the need for specialist hardware typically restricted to research settings. The approach offers a promising way to assess excitation/inhibition balance by detecting labelled glucose metabolites. The pre-clinical data in mouse models provide convincing support for the authors' hypothesis. The evidence in the human brain is incomplete: the results would require further optimisation of the experimental set-up and validation in a larger group of volunteers. This work will be of broad interest to researchers studying psychiatric, neurodevelopmental, and related brain disorders.

    2. Reviewer #1 (Public review):

      Excitation/inhibition (E/I) balance between excitatory (glutamate) and inhibitory (GABA) neurotransmission is being increasingly studied using magnetic resonance spectroscopy (MRS), for example in autism spectrum disorder, schizophrenia and attention deficit/hyperactivity disorder. These are typically measured using standard single-voxel MRS methods (eg PRESS, sLASER) to measure glutamate/glutamine or "Glx" and spectral editing methods (eg MEGAPRESS, MEGA-sLASER) techniques to measure GABA. Such methods only give a measure of the total MR-visible metabolite concentration in the voxel. That is, they don't distinguish between glutamate/GABA involved in neurotransmission or in other metabolic processes.

      Cherix et al present a method for measuring glucose metabolism, with the potential to be used on a standard clinical MRI scanner. This proof-of-concept study focused on measuring proton signals from glucose metabolites, including lactate, glutamate, GABA & Glx. The method works by administering 13C universally labelled glucose (where all six carbon atoms are substituted with 13C). When the glucose is metabolised, 13C label is incorporated into specific positions within its metabolites. Protons attached to 13C don't produce a signal in 1H-MRS in a subsequent MEGA-sLASER scan, leading to a drop in the signal as the labelled metabolite concentration builds up. At the same time, "satellite resonances" appear for protons coupled to 13C, which increase as the labelled metabolite concentration builds up. Metabolite concentrations were inferred using a simple dynamic model.

      The main strength of the method is that it enables dynamic metabolic information that would typically only be available with multi-nuclear MRS capability (13C or 2H) to be achievable using standard preclinical or high field (>= 7T) human MRI systems, with widely available spectral-editing acquisitions.

      The results in the mouse spectra seem very convincing for lactate and GABA/Glx. For the human scans, changes in lactate weren't detectable, which is not surprising given how little lactate appears in the normal brain. In the discussion, the authors argue the method can potentially be used in a standard, 3T clinical scanner. It may be too soon to conclude that, as it's not yet clear there would be sufficient SNR in spectra at that field strength. Additionally, the heteronuclear coupling constants are quite high. The authors recognise that this may complicate detection of satellite resonances due to signal dephasing. Another potential complication at 3T (or 2.9T) is the potential for the satellite resonances to come close to the GABA peak at 3 ppm. More accurate measurements of coupling constants will allow that to be determined. Another potential limitation of the method is macromolecule contamination of the 3 ppm GABA peak. That may be overcome by using macromolecule-nulled MEGA-editing, though frequency navigators may be necessary to overcome the increased sensitivity to frequency drift.

      The authors achieved their aims of showing that imaging glucose metabolites was possible using standard proton-only MRI systems, without the need for additional multinuclear coils, transmitters and receivers. The evidence if very compelling for the mouse scans but only incomplete for the human scans.

      The ability to quantify metabolites involved in E/I balance has the potential to revolutionise studies into disorders where changes in E/I balance are implicated. This is especially the case for preclinical models. Such studies may be less feasible in clinical studies due to the high cost of universally 13C-labelled glucose, but this proof-of-concept is a promising start.

    3. Reviewer #2 (Public review):

      Summary:

      The main aim of the presented manuscript was to test and validate proton observed proton edited 13C MRS and track the 13C label from uniformly labeled U-13C-glucose into glutamine, glutamate, GABA and lactate in rodent and human brain in vivo.

      Strengths:

      In contrast to already established methods of 13C and 2H MRSI this method applies only proton RF and thus can potentially be implemented on a standard clinical scanner.

      Weaknesses:

      The validation of the method in a human setting is rudimentary. First, even at ultra-high field strength of 7T, the authors did not reach sufficient SNR to detect and quantify GABA with good CV, and further, the experiment needs optimisation to reach metabolic and fractional enrichment steady state and/or for additional conditions to show the possibility of lactate detection.

    4. Reviewer #3 (Public review):

      Summary:

      In their paper, the authors propose using dynamic MEGA-sLASER acquisitions to track the incorporation of ¹³C from uniformly ¹³C-labeled glucose into glutamate(+glutamine) and GABA pools. This approach enables the direct and simultaneous investigation of excitatory and inhibitory neurotransmission and metabolism without requiring a ¹³C radiofrequency coil. While the authors' goals and efforts are commendable, the work has several significant limitations.

      Strengths:

      Use of excellent hardware (¹H cryoprobe, ultra-high-field MRI scanners); Ambitious objectives.

      Weaknesses:

      MRS data analysis requires improvement; Small cohort sizes; No demonstration that the proposed acquisition scheme is superior in terms of robustness, accuracy, or performance.

    1. eLife Assessment

      There is a significant need for improved understanding of the neural circuit mechanisms for learning and memory dysfunction in Alzheimers disease. In this valuable study, Zheng and colleagues compare neural coding between wild-type and App knock-in rats experiencing different contexts. Solid evidence of a difference emerges when "spatial" and "temporal" neural coding dynamics are considered separately. However, the lack of a link between neural coding and behaviorally measured learning, and minor issues with the analyses, are limiting factors of the evidence and the significance of the work.

    2. Reviewer #1 (Public review):

      Summary:

      Remapping is clearly degraded in amyloid models, but people and animals with a lot of pathology often hold onto more function than their spatial maps would predict. The authors' idea is that CA1 carries two things at once: an explicit code where rate maps onto position, and an implicit one in the temporal relationships between cells, and that AD hits the first much harder. They recorded CA1 with tetrodes in App(NL-G-F) and WT rats running an A-B-B-A sequence of open field sessions, repeated daily for six days, with rest in between. They then compared rate map measures against a cofiring measure (pairwise Kendall's tau) and looked at SWR reactivation during rest.

      Strengths:

      (1) The design is right for the question. Alternating back to the familiar arena separates "can the network register a new context" from "can it get back to the old one," and the finding that App rats look OK on the first A to B transition but fall apart on the return is the most striking thing here. The confused cell result, 24% vs 4.5%, is easy to read and hard to dismiss.

      (2) Six consecutive days is also worth something. Most work on this is cross-sectional, and looking at how the coding changes with accumulated experience is the right way to ask about plasticity.

      (3) The PIR analysis in Figure 5 is the bit I liked best. Subtracting each cell's position-predicted rate before computing tau is a reasonable check that the cofiring effects aren't place fields in disguise, and it's more than most papers making this argument do. The theta index control is a good instinct too, though see below on how it's analysed.

      Weaknesses:

      (1) No behaviour. This is the main problem and everything else is secondary. The title, abstract, intro and discussion all turn on preserved learning and memory, and the rats were never tested on anything. They foraged for popcorn in an open field. No discrimination measure, no novelty preference, no probe, nothing. So "learning" ends up being defined as "decoder accuracy went up across days," which makes the central claim circular. Either add a behavioural readout in these animals or take the learning language out of the title and abstract and say what was actually measured, which is experience-dependent change in neural coding.

      (2) Four animals per group is fine for this kind of work, but the statistics don't respect it. Degrees of freedom in the thousands and tens of thousands (t(8616), t(36293), F(1,2674)) treat cells and cell pairs as independent, which they aren't. The mixed models in Figure 1 are the right approach, and I couldn't see why they were dropped everywhere else. The theta result is the clearest casualty: a null across 36,293 cell pairs from four rats isn't evidence that theta coordination is preserved; it's an untested question with n=4.

      (3) Also, the reported df do not always match the stated n. The Methods say 4 per genotype, but several animal-level tests give t(4), which implies 3. Figure 4C gives t(32), and Figure 5C gives t(31) for what look like per-day measures. I couldn't work out what the sampling unit was in each case.

      (3) Some statistics can't be right. I noticed three, without looking hard. For example. Figure 1C, rate overlap: t(471) = 2.3 with p = 2.0e-7. Fig 1G: t(163) = 5.3 with p = 0.48. Fig 3F: F(1,483) = 36.5 with partial eta squared = 0.7, when the almost identical test in the previous sentence gives 0.07. These are likely all typos, but there are enough that the whole set needs going through.

      (4) The dissociation isn't tested with matched methods. Explicit coding gets rate map correlations, PVC, rate overlap, and field size. Implicit coding gets an SVM across six days. The claim is that one improves with experience and the other doesn't, but they're never put through the same analysis. The authors should run the identical decoder on rate vectors and on tau vectors, same cross-validation, day by day, and show the slopes diverging. That would be a real dissociation. Figure 1G does run a rate decoder but only pooled, not across days. As it stands, the difference in learning slopes could partly be the two analyses having different sensitivity.

      (5) The PIR residual may not be as clean as it looks. PIR is observed rate minus rate predicted by the cell's own spatial map. If the spatial map is a worse model of firing in App rats, which is the paper's own claim, then less gets subtracted and more is left in the residual. So a group difference in residual tau structure could be partly downstream of the group difference in place coding quality rather than something independent. This is worth some kind of check, e.g. matching cells on spatial information, or at least reporting how much variance the spatial model explains in each group.

      (6) Reactivation consistency. This carries a lot of the interpretation, and it's the thinnest evidence in the paper. r = 0.2, p = 0.04 in App against r = -0.2, p = 0.06 in WT. That's a difference in significance, not a tested difference between groups, and with both p-values sitting on either side of 0.05, I wouldn't build a mechanism on it. Needs a group x day interaction. Separately, mean pairwise correlation across SWR population vectors depends on how many cells are active, how many events there are (panels show 189 to 491) and how sparse the firing is. SWR rate, duration, participating cells and firing rates would let the reader judge whether "more consistent reactivation" means what's claimed.

      (7) The hyperexcitability to excessive replay to Hebbian consolidation story on pp 21 to 22 runs about a page on the back of one marginal correlation. This section should be cut down and flagged as speculation.

      (8) Missing controls. Things I expected and didn't find: histology confirming tetrode placement in CA1, any pathology verification in this cohort rather than a citation to Pang 2022, and A1-A2 spatial correlation shown next to B2-A2. That last one matters. If the App representation of A just drifts across the day, that's a different story from a specific failure to reinstate A, and the confused cell analysis as built can't tell them apart. Also, with the threshold set at the 95th percentile of the A1B1 baseline, the WT value of 4.5% is basically the chance floor by construction, so the number that carries information is the App one.

      (9) The issue of males only should be mentioned.

    3. Reviewer #2 (Public review):

      This study by Wang et al. longitudinally tracks hippocampal CA1 population dynamics in Alzheimer model rats during repeated exposure to different environments. The authors dissociate two levels of neural coding. "Explicit" spatial coding, assessed by rate maps and population vector correlations, is severely impaired in AD rats and does not improve with experience. In contrast, "implicit" temporal cofiring structure, quantified by pairwise Kendall's tau and population cofiring correlations, becomes progressively more context-specific over days, mirroring behavioral learning. This preserved temporal coding is not merely a byproduct of spatial overlap, as position-independent rate analysis confirms that learning-dependent discrimination arises from intrinsic temporal dynamics rather than from residual spatial tuning. Moreover, offline sharp-wave ripple reactivation shows increasing consistency across days specifically in AD rats. These findings reveal a dissociation in the AD hippocampus and propose that temporally structured population dynamics, rather than spatially selective firing, may support residual cognitive function.

      Overall, the findings are novel and thought-provoking; the analyses are comprehensive and well-controlled, and the proposed re-registration framework offers a compelling new lens for understanding cognitive resilience in Alzheimer's disease, with clear potential to guide future neuromodulation interventions.

      I have only a few minor comments.

      (1) Justification of terminology ("explicit" vs. "implicit").

      The manuscript should clearly define the two terms early in the Introduction, ideally in a dedicated paragraph. In my opinion, the current use of "explicit" and "implicit" is not intuitive and may even be misleading.

      (2) Does the dissociation reflect differential impairment between rest and running states in AD?

      The authors could elaborate on this.

      (3) Effect size in Figure 1D - the difference does not appear very large.

      The statistical significance in Figure 1D is accompanied by relatively modest effect sizes. The authors may tone down the claim about "failure to distinguish" in the manuscript ("failed to distinguish different contexts during the second transition" on Page 8).

      (4) Definition of "confused cells" - why not a fixed correlation threshold?

      The authors define "confused cells" as those whose B2‑A2 spatial correlation exceeds the 95th percentile of the A1‑B1 baseline distribution within the same animal. This seems to be unnecessary. A fixed correlation threshold (e.g., r > 0.5 or r > 0.6) would have a direct biological interpretation: it would identify cells that maintain similar firing fields across two putatively distinct environments, i.e., cells that truly fail to remap. In contrast, a relative percentile threshold defines "confusion" not against an absolute standard of similarity, but against the degree of remapping observed during the first A‑B transition.

      (5) SVM decoding on spike‑train temporal structure - missing justification.

      The authors state that "the temporal structure of spike trains contained distinct contextual information" (P9) and then directly apply an SVM decoder to the data, but the logical bridge is missing. Why is a decoder necessary here, and is it useful for such a task?

      (6) Figure 6 - learning occurs during reactivation but does not transfer to the online state (theta state), even after days of training.

      What does this mean in terms of different phases of memory? Is the consolidation phase affected more? The authors may provide more discussion along these lines.

    4. Reviewer #3 (Public review):

      Summary:

      Determining the ways in which Alzheimer's disease (AD) impacts the neural instantiations of memory is a fundamental aim of neuroscience and likely to be critical to understanding and treating disease progression. Previous work has highlighted how the hippocampal spatial code - typically context-specific - fails to discriminate between different environments in rodent models of AD. Here, Wang et al. leverage new analyses in a rat model of AD to test whether this deficit in 'remapping' reflects an impairment in intrinsic temporal coding. The authors find that the intrinsic temporal code of AD rats does come to effectively discriminate between environments (albeit more weakly than their wild-type counterparts), despite persistent impairments in remapping.

      Strengths:

      One of the major strengths of this work is its focus on distinguishing between two different types of neural impairments in AD. While previous work has highlighted impairments in remapping, these impairments could be due to: (a) impairments in intrinsic neural computations, or (b) impairments in the way these intrinsic neural computations are anchored to the world. The author's evidence supports the latter, with important implications for the nature of these impairments.

      Weaknesses:

      A handful of weaknesses could be addressed to strengthen this work. Firstly, a number of key measures of place code quality and behavioral quality between groups are omitted. Given that the author's interpretations rely on comparisons between groups and often use decoding analyses for central conclusions, indicating whether the recordings are comparable between groups in terms of behaviors and cell counts would be helpful for the reader. Even better, matching cell counts between groups for decoding analyses could ensure that outcomes are not driven by this potential confound.

      Another weakness that is worth addressing involves the pivotal comparison between time-averaged remapping (Figure 4) and intrinsic temporal code discrimination (Figure 3). For temporal code discrimination, the analysis relies on comparisons between pairs of epochs (e.g. from A1B1 to A2B2), while for remapping, the comparison averages together epochs (A1+A2 and B1+B2). As a result, these comparisons are characterizing two slightly different things. Given that this is a key comparison for this paper (cofiring evolves to discriminate contexts while the spatial code does not), I think it would be important to demonstrate that remapping between pairs of epochs also stagnates to more closely mirror the cofiring analysis.

      A final weakness is the sharp wave ripple (SWR) ensemble analysis. Here, the authors extract population vectors during SWR (SWR-PVs) during rest in both WT and AD rats. Next, they compare the extent to which SWR-PVs on AVERAGE resemble other SWR-PVs for that epoch. The authors report that this measure increases with experience in AD rats but not wild-type (WT) rats. While the authors interpret this to mean that SWR-PVs come to be more reliable with experience in AD rats, it could alternatively mean that SWR-PVs become less diverse or 'muddier' with experience, while SWR-PVs in WT rats continue to represent diverse trajectories. To make this analysis compelling, a better measure might be something that quantifies the diversity or content of SWR-PVs and something that quantifies similarity between each SWR-PV and the most similar other SWR-PVs.

    1. eLife Assessment

      This study provides valuable findings on computational measures of learning in human subjects with current depression, individuals remitted from depression, individuals at familial risk, and healthy controls using reinforcement learning and risky decision-making tasks. The evidence is solid but could benefit from clear interpretation of model parameters and psychiatric symptoms. The paper will be of interest to scientists interested in learning, reward processing, value-based decision-making, and psychopathology.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates state-, trait-, and recovery-related computational phenotypes in Major Depressive Disorder (MDD) using an unmedicated, four-group cross-sectional design comprising currently depressed participants, remitted individuals, first-degree relatives at familial risk, and healthy controls. Utilizing a volatile four-armed bandit task and an explicit risk gambling task alongside hierarchical Bayesian modeling, the authors report that remitted participants uniquely display a lower punishment learning rate relative to all other groups. The authors conclude that recovery from MDD does not represent a simple normalization to healthy baseline levels, but rather involves a protective computational recalibration that dampens reactivity to negative outcomes to sustain remission.

      Strengths:

      (1) Highly Valuable Clinical Sample: Evaluates a rare, unmedicated sample across four distinct clinical stages (MDD, REM, REL, CTR), providing an exceptionally controlled framework to disentangle state, trait, and recovery markers.

      (2) Combines dynamic reinforcement learning under environmental volatility with prospect-theoretic decision-making under explicit risk to capture multiple dimensions of value-based choice.

      (3) Challenges the conventional assumption of clinical "normalization," offering a compelling hypothesis that psychiatric recovery may depend on active, compensatory recalibrations of cognitive parameters.

      (4) Utilizes hierarchical Bayesian parameter estimation and leverages Bayesian Model Averaging (BMA) to mitigate single-model selection bias.

      Weaknesses:

      (1) The Discussion characterizes MDD and familial-risk groups as exhibiting "noisier" choice behavior in the gambling task, which directly contradicts the reported higher inverse temperature values that mathematically denote more deterministic choices.

      (2) Framing a reduced punishment learning rate as a "recovery mechanism" overinterprets single-timepoint data, which cannot differentiate an acquired post-episode adaptation from a pre-existing resilience trait.

      (3) The theoretical claim that lower punishment learning is protective in remission directly conflicts with the authors' dimensional findings, where lower punishment learning correlates with worse subclinical apathy and anhedonia in non-depressed participants.

      (4) Model-agnostic choice repetition yielded no significant group effects, contrasting sharply with the robust group differences in model-derived parameters and necessitating posterior predictive checks.

      (5) Fails to provide parameter recovery analyses to demonstrate that punishment learning rates can be reliably disentangled from lapse rates and outcome sensitivities across a 200-trial task structure.

      (6) Selects a lower-ranked model under LOOIC without sufficient quantitative justification, and lacks sensitivity analyses to confirm that group-specific hierarchical priors did not skew estimates given unequal group sizes.

      (7) Relies on several marginal p-values bordering across multiple parameters and symptom correlations without establishing a clear family-wise error or FDR correction strategy.

    3. Reviewer #2 (Public review):

      This manuscript reports on reinforcement learning in participants with current depression, remitted depression (without current depression), in people without depression but with first-degree relatives with depression, as well as healthy controls. Participants completed two common tasks measuring reinforcement learning and risk aversion, and their behavior was fit to computational models assessing processes on these tasks.

      Participants with remitted depression showed a lower punishment learning rate and more value-concordant decisions on the reinforcement learning task. Relatives of depressed participants, as well as people who are currently depressed, had higher inverse temperature, indicating more value-driven choices. Within the non-depressed participants, the punishment learning rate was negatively associated with symptoms of anhedonia and apathy.

      I have reviewed this paper at a previous journal. This revised version is responsive to most of my concerns, particularly in terms of placing the manuscript more in the context of other related literature and providing more details on methods. There are some remaining concerns about sample size and the appropriateness of some of the methods (e.g., interpreting participant-level parameter estimates from hierarchical models estimated using BMA), but the latter has been adequately addressed with sensitivity analyses.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript describes an interesting study (preceded by a pilot study) that combined computational modeling with two tasks to try to tease apart state, trait, and vulnerability-related effects of depression. Specifically, the authors administered a bandit task and a gambling task to healthy controls, currently depressed adults, formerly depressed adults, and adults at familial risk of depression, and then used a suite of computational models to identify group differences in key model parameters. Key results included the detection of lower punishment learning rates (during the bandit task) in the remitted depressed group, which the authors hypothesize may be a compensatory mechanism to counteract the over-reaction to negative feedback that often characterizes depression (and indeed, punishment learning rates were elevated in currently depressed adults). Among non-depressed adults, lower punishment learning rates were associated with more apathy and anhedonia. The remitted depressed group also showed lower lapse rates in the bandit task. In the gambling task, the inverse temperature parameter was elevated in currently depressed adults and in adults at high familial risk for depression; exactly how to interpret this last result seems unclear. This is a revised manuscript, and from what I can see it appears that the authors were responsive to earlier comments.

      Strengths (and summary of weaknesses):

      I think the study has several noteworthy strengths. Testing four groups is a strength, as there is great interest in teasing apart risk factors of depression vs. "scars" of the illness, and the use of unmedicated individuals removes a common confound. The work is hypothesis-driven, the modeling is sophisticated, and the paper is well-written. Moreover, as the authors note, modeling allowed the authors to identify group differences that were not evident in raw behavioral analyses.

      However, I think the manuscript could be further improved because some aspects of the methodology are a bit confusing and/or do not seem optimal. I list these issues in the comments below, but to briefly summarize: (a) I did not understand why the authors estimated separate reward and punishment values for each bandit; (b) the rationale for Bayesian Model Averaging could be strengthened; (c) examining relationships between model parameters and symptom scores only in the control group is suboptimal given the goal to better understand depression; and (d) the group differences in inverse temperature seem like they may depend on potential outliers. I think addressing these concerns would improve the paper and ensure that the study has a strong impact on the field.

      Details regarding weaknesses/concerns:

      (1) I found a basic aspect of the 4-arm bandit modeling confusing - namely, the use of separate reward and punishment value estimates (see equations 1 and 2 in the supplement). The participants are choosing among the bandits (presumably) based on value estimates for each bandit. Clearly, delivery of rewards and punishments affect those estimates, but it is not clear to me why or how there are separate value terms for rewards and punishments; typically, rewards and punishments influence one overall value estimate per bandit. Can the authors clarify? (I see that similar models were used in references 29 and 30, but additional clarification for readers of the current manuscript would be helpful)

      (2) Bayesian Model Averaging (BMA) is new to me and may be new to many readers, and it would be helpful to provide a stronger rationale for the approach. The manuscript argues that model selection amounts to a "winner-takes-all" approach that may introduce bias; maybe, but typically the goal is to figure out which mechanism(s) best explain behavior, and so winner-takes-all is often appropriate. My limited understanding is that BMA is often used when prediction-rather than mechanistic understanding-is the goal. Why is BMA the right choice here?

      (3) The authors performed a confirmatory factor analysis on questionnaire data from healthy controls out of concern that extreme scores in the other groups might bias the factor structure; consequently, they can only relate model parameters to their latent factors in the healthy controls. This does not seem optimal. While the negative relationships between punishment learning rate and both anhedonia and apathy in controls are interesting, the controls are not struggling with anhedonia or apathy. It would be valuable to know if similar relationships obtain in the other groups, particularly because this would speak to the study's goal of distinguishing between risk for depression and state/trait aspects of depression.

      (4) Figure 5 gives the impression that group differences in inverse temperature may depend on potential outliers in the relative and MDD groups. Is that correct?

      (5) The paper notes a group difference in IQ as estimated from the WTAR, and looking at Table 1 it appears that the group difference is driven by a lower WTAR score in the healthy volunteers (HV) from the pilot study. First things first: the score in the HV group is too low to be a standardized WTAR score. Like IQ scores, standardized WTAR scores typically have a mean around 100; scores for the four Study 2 groups look alright, but the Study 1 mean score of 40 is much too low. Can the authors clarify?

    1. eLife Assessment

      This important study provides evidence that Clock and fat body proteasome subunits contribute to dietary restriction-mediated lifespan extension in Drosophila. The evidence supporting these conclusions is solid, yet whether dietary restriction-induced daily rhythmicity of proteasomal genes is CLK-dependent remains incompletely demonstrated, and alternative explanations involving feeding differences and non-circadian functions of CLK cannot be, with the current dataset, excluded. The work will definitely be of interest to researchers studying biological rhythms, nutrition, and aging.

    2. Reviewer #1 (Public review):

      The studies by Hwangbo et al. diligently attempt to account for many of the typically neglected dietary and non-dietary factors.

      Strengths:

      • Work addresses many potential artifacts of dietary (e.g., dehydration stress, macronutrient ratios, and protein source) and non-dietary (e.g., leaky expression of S106-GAL4) manipulations-important factors that are too often overlooked.

      • Balanced and complementary behavioral, molecular, and bioinformatic experiments

      • Show necessity of proteostatic subunits in the fat body for DR-mediated longevity. The findings in the current manuscript lay the ground for future studies that test sufficiency of fat body prosβ3 and rpn7, or necessity of other proteostatic genes in other tissues.

      Comments on revised version:

      The revised manuscript is substantially improved and addresses many of the prior concerns. I have only a few minor recommendations and remaining issues:

      Clarify the interpretation of the Con‑Ex feeding data. The authors describe the ~70% higher intake on 1SY in Clk^Jrk as modest, and note a ~40% higher mean intake on 5SY that is not statistically significant. However, lack of significance can reflect limited power, and these differences are potentially biologically meaningful, given that relatively small changes in nutrient ingestion can substantially affect lifespan. If the average effects are real, the Clk^Jrk flies would be ingesting an effective diet closer to ~1.7SY and ~7SY relative to controls. A shift of the diet-lifespan response curve in Clk^Jrk therefore cannot be fully excluded, particularly given the absence of intermediate diets between 1SY and 5SY and the observation that Clk^Jrk is sometimes shorter‑ and sometimes longer‑lived than controls across different trials and diets.

      Although the core finding is strengthened by using several diet formulations, most additional experiments continue to rely on whole‑food dilution, even as the field is moving toward more defined DR regimens (e.g., yeast‑only or yeast‑extract-based protocols). There remains considerable variability and, in some cases, a lack of clear DR‑mediated lifespan extension in control cohorts (for example, in some GeneSwitch experiments using whole‑food dilution). It would be helpful if the authors briefly commented on this variability and justified their continued use of whole‑food dilution in these experiments.

      Please add a clear Methods description of the feeding assay (Con‑Ex), including fly age, assay duration, dye or tracer conditions, sample processing, quantification, and statistical analysis.

      Please ensure that the survival data shown in Figure 5 and associated supplements are explicitly linked to Cox proportional hazards analyses in the text or figure legends, with clear indication of the models used (e.g., gene, diet, and gene×diet interaction terms). The Methods state that diet is used as a continuous variable; given the non‑linear (U‑shaped) lifespan-diet reaction norm (reduced survival at both 1SY and higher yeast), it would be important to clarify whether 1SY was excluded from these Cox models, or alternatively, to model diet categorically, restrict the continuous analysis to 5-20SY, or apply an appropriate non‑linear transformation (e.g., splines). As written, it is not clear how the Cox model accommodates the non‑linear diet response.

    3. Reviewer #2 (Public review):

      Summary:

      Dietary restriction (DR) increases lifespan, an effect that has been consistently observed in several organisms, but we still lack a clear mechanism to explain this phenomenon. In this work, Hwangbo et al. revisited the role of the circadian clock in DR-mediated lifespan effects. They found that the increase in lifespan produced by DR is missing on a clock mutant, a clock dependency that is also observed at the level of nutrient-dependent egg laying. By conducting RNA-seq with an impressive temporal resolution, they showed that DR triggers an increment in the number of cycling genes expressed in the fat body, the fly functional analog of the mammalian liver. Interestingly, from these genes, a group of them are de novo daily expressed genes, meaning that their expression was not rhythmic under the control diet but appear rhythmically expressed under DR. Among those, genes encoding proteasome subunits are enriched. The authors finally showed that adult-specific knockdown of these genes in the fat body prevents the increase in lifespan under DR, further supporting a role of the proteasome in this process. Overall, the conclusions are mostly supported by the evidence presented, and the authors' discussion nicely frame their results with other research in the field.

      Strengths:

      - Many studies have limited their observations of DR on lifespan to a few dietary conditions which makes the reach of some previous conclusions somewhat limited. The dilution strategy that the authors used in this work provides a strong indication that the effect of DR on lifespan relies on clock expression regardless of the conditions used. Furthermore, the inclusion of the egg-laying assay is a good addition to support this hypothesis.

      - Because the strength of the rhythmicity statistics relies heavily on the number of data points collected, the temporal resolution used for the RNA-seq experiments (every 2 hrs per 48hrs) is remarkable. This allows exquisite dissection of the phase of rhythmic genes in different conditions. The dataset produced in this work might be of use to other groups interested in weighting the role of other represented gene clusters in DR.

      Weaknesses:

      I see only minor flaws in this work, that if addressed, might strengthen the authors' conclusions, particularly:

      - The results of the lifespan assays are quite variable and in some instances contradictory (Fig. S8) across trials, possibly because there are other unaccounted variables we still do not understand. The fecundity assay, in contrast, seems to be a better readout (Fig. 2). Confirming at least the two genes picked for the study (Fig. 5) would be good support for the claim that the proteasome mediates the effects of DR.

      - According to the model, the acute effect of DR on gene expression is related to CLOCK protein function. However, I am not sure how this link was established. It is tempting to assume that CLOCK upstream is the reason for having an increase in rhythmic genes under DR, but the experiments did not test this. The tests conducted either assessed the role of clk or the effect of an impaired proteasome on DR-dependent extension of lifespan. Thus, it is difficult to assert the authors' claims on the link between CLK and the changes in cycling genes and to the proteasome upon DR.

      Comments on revised version:

      In this new version, Hwangbo and colleagues add new data to support a role of the clock in the effects of DR on lifespan. While adding this new data helps to alleviate some of the concerns previously raised, I think there are still some gaps. Below are my main concerns:

      ClkJrk transcriptomic data: Adding this data supports a role of the clock in daily rhythmicity of genes in the fat body, likely due to a circadian role. However, it does not show that de novo rhythmicity of proteasomal genes is clock-related since there is no ClkJrk DR dataset. Thus, there is still a possibility this is a pleiotropic effect. While redoing an entire RNA-seq dataset might not be feasible, a possible way to support the circadian claim would be to use proxy genes observed in the Ctrl vs DR conditions and compare them by qPCR in ClkJrk Ctrl vs DR.

      Feeding data: Are the flies reared in DR conditions, or do they just start the DR at the beginning of the experiment? If so, is it possible the flies will show a different feeding pattern after consecutive days of DR affecting overall (5-10 days) food consumption?

      Clk expression is important for non-circadian roles in the ovaries (Wang et al., Cell Mol Life Sci, 2025). Therefore, it is possible that the fecundity effect is at the low level of the ovary/egg development instead of integration and processing of DR. This might be a confounding effect when interpreting data in Fig 2 in ClkJrk as solely the effect of DR.

      Line 215: Considering the discussion above, I'd rather change "circadian-dependent change" to "daily"<br /> Considering that the ClkJrk RNA-seq transcriptomic was generated, presumably, at a different date/time than the original transcriptomic data from Control vs DR, comparative metrics (seq depth, mapping rate, etc.) between these are needed.

      How is the feeding analysis conducted? I believe the method information was not updated.

    4. Author response:

      The following is the authors’ response to the original reviews

      Correction: In the process of revising the preprint, we discovered that for the 15SY dataset that a single time point (ZT2) out of the 12 timepoint series was inadvertently combined with temporally adjacent time samples (ZT20, 22, 24). We corrected the accompanying GEO submission (Series GSE145509). With this update, we repeated the rhythm analysis with an updated RAIN algorithm as the original Boot-eJTK could not be run as it was outdated with dependent packages no longer maintained or supported, and some are no longer available through standard package managers. The analysis with the corrected ZT2 sample and did not find any significant changes in the major claims of the papers. One minor change is that we no longer observe significant DR-dependent increases in proteasome gene levels. Nonetheless, we still find DR-dependent cycling of proteasome genes and proteasome module network connectivity consistent with the DR-sensitivity of the proteasome pathway. Figure 4 has been updated to reflect this change. After correcting this issue, we revised the manuscript in response to the reviewer comments.

      eLife Assessment

      This study describes important findings on how a core component of the circadian clock impacts the effect of dietary restriction (DR) on longevity and fecundity in Drosophila, which lead the authors to postulate rhythmic control of proteostasis in the fat body as a critical aspect of DR effects. The evidence presented is still incomplete, not fully supporting the conclusions of the study, as alternative hypotheses/explanations have not yet been systematically explored. The work will nevertheless be of substantial interest to researchers working in circadian and cell biology, metabolism, and aging, with an interesting hypothesis to be explored further.

      We sincerely thank eLife for considering our manuscript and express our gratitude to all three reviewers for their time and constructive comments. While acknowledging that there are alternative hypotheses for some of our findings, which make our evidence incomplete, we appreciate that our manuscript was recognized as important and of substantial interest. We have revised the manuscript to acknowledge that the major findings under light-dark conditions could be attributable to being driven by light rather than the circadian clock. We add new data demonstrating that the far majority of cycling genes in LD in control flies are disrupted in Clk<sup>Jrk</sup> consistent with circadian regulation (see also below). Future investigations are needed to test the alternative hypotheses and explanations. Nevertheless, we believe that the manuscript still represents a meaningful advancement in understanding how molecular circadian clocks in peripheral tissues interact with diet to influence systemic lifespan and aging.

      Public Reviews:

      Reviewer #1 (Public Review):

      The studies by Hwangbo et al. diligently attempt to account for many of the typically neglected dietary and non-dietary factors.

      Strengths:

      - Work addresses many potential artifacts of dietary (e.g., dehydration stress, macronutrient ratios, and protein source) and non-dietary (e.g., leaky expression of S106-GAL4) manipulations-important factors that are too often overlooked.

      - Balanced and complementary behavioral, molecular, and bioinformatic experiments

      - Show necessity of proteostatic subunits in the fat body for DR-mediated longevity. The findings in the current manuscript lay the ground for future studies that test sufficiency of fat body prosβ3 and rpn7, or necessity of other proteostatic genes in other tissues.

      Weaknesses:

      - Could the lack of DR response in clock mutants across dietary concentrations be simply because the clock mutants are better at compensatory feeding adjustments to dietary dilutions? If this were the case, there are two major implications to the authors' conclusions:

      a) The Clk mutants are differently responding to dietary dilutions, not to dietary restriction, per se.

      b) Nutritional intake was unaffected by the dietary manipulations. If the changes in fat body proteostasis and lifespan were due to nourishment, it would be expected that the physiology and lifespan do not change.

      Accurate measurements of food consumption and the resulting protein intake could potentially clarify this critical question.

      We thank the reviewer for their positive feedback and also appreciate their raising the important issue of whether Clk^Jrk mutants may be more effective at compensatory feeding. Xu et al. (2008) reported that overall food consumption in Clk^Jrk flies was indistinguishable from control flies.

      We also directly assessed food intake in ~1 week old flies on three diets (1% SY, 5% SY, 15% SY) over two days (48 hours) using the Con-Ex method (Shell et al. 2018). We observed either no significant changes or relatively modest changes in food consumption between iso31 controls and Clk^Jrk flies that are limited compared to the large differences in caloric content between the diets. While we cannot rule out changes in feeding patterns throughout the lifespan, we believe that minimal differential compensatory feeding in the Clk^Jrk mutants are not sufficient to be the primary cause of the lifespan differences observed across diets. These data are added as new Figure 1-figure supplement 3.

      Reviewer #1 (Recommendations For The Authors):

      Hard to find information:

      - Type of yeast used. Should be addressed in the Methods section at least, instead of having to dig through several paragraphs into the Results section.

      Missing information:

      - Agar type and concentration

      - Mifepristone diets: pipetted on top or mixed into food?

      The relevant information has been updated in the Materials and Methods

      Wrong information:

      - Line #159-160: "Fig. 1 and Fig. S3" should be "Fig. 1 and Fig. S2" and then "(Fig. S3)" added to the end of the sentence.

      This information has been corrected in the revision

      Presentation:

      - Paragraphs are very long.

      - Interactions should be denoted by ×, not * or x.

      - Remove markers from mortality graphs. Having bulky markers reduces perceived differences between curves.

      These changes have been updated in the revised version.

      Reviewer #2 (Public Review):

      Dietary restriction (DR) increases lifespan, an effect that has been consistently observed in several organisms, but we still lack a clear mechanism to explain this phenomenon. In this work, Hwangbo et al. revisited the role of the circadian clock in DR-mediated lifespan effects. They found that the increase in lifespan produced by DR is missing on a clock mutant, a clock dependency that is also observed at the level of nutrient-dependent egg laying. By conducting RNA-seq with an impressive temporal resolution, they showed that DR triggers an increment in the number of cycling genes expressed in the fat body, the fly functional analog of the mammalian liver. Interestingly, from these genes, a group of them are de novo daily expressed genes, meaning that their expression was not rhythmic under the control diet but appear rhythmically expressed under DR. Among those, genes encoding proteasome subunits are enriched. The authors finally showed that adult-specific knockdown of these genes in the fat body prevents the increase in lifespan under DR, further supporting a role of the proteasome in this process. Overall, the conclusions are mostly supported by the evidence presented, and the authors' discussion nicely frame their results with other research in the field.

      Strengths:

      - Many studies have limited their observations of DR on lifespan to a few dietary conditions which makes the reach of some previous conclusions somewhat limited. The dilution strategy that the authors used in this work provides a strong indication that the effect of DR on lifespan relies on clock expression regardless of the conditions used. Furthermore, the inclusion of the egg-laying assay is a good addition to support this hypothesis.

      - Because the strength of the rhythmicity statistics relies heavily on the number of data points collected, the temporal resolution used for the RNA-seq experiments (every 2 hrs per 48hrs) is remarkable. This allows exquisite dissection of the phase of rhythmic genes in different conditions. The dataset produced in this work might be of use to other groups interested in weighting the role of other represented gene clusters in DR.

      We are grateful for the reviewer’s positive feedback regarding the robust experimental design in the manuscript.

      Weaknesses:

      I see only minor flaws in this work, that if addressed, might strengthen the authors' conclusions, particularly:

      - The results of the lifespan assays are quite variable and in some instances contradictory (Fig. S8) across trials, possibly because there are other unaccounted variables we still do not understand. The fecundity assay, in contrast, seems to be a better readout (Fig. 2). Confirming at least the two genes picked for the study (Fig. 5) would be good support for the claim that the proteasome mediates the effects of DR.

      We appreciate the reviewer’s comment of seeing “only minor flaws”. We reiterate that we focused on those results which were replicated across trials, providing confidence in the overall conclusions. Nonetheless, we agree that exploring the proteasome role on the DR effect on fecundity would be intriguing and may complement the lifespan data. We now acknowledge this point in our discussion.

      - According to the model, the acute effect of DR on gene expression is related to CLOCK protein function. However, I am not sure how this link was established. It is tempting to assume that CLOCK upstream is the reason for having an increase in rhythmic genes under DR, but the experiments did not test this. The tests conducted either assessed the role of clk or the effect of an impaired proteasome on DR-dependent extension of lifespan. Thus, it is difficult to assert the authors' claims on the link between CLK and the changes in cycling genes and to the proteasome upon DR.

      We have updated our text and model figure suggesting a direct CLK role in the Discussion.

      Reviewer #2 (Recommendations For The Authors):

      As mentioned before, the experiments, in particular the RNA-seq datasets are excellent. Additionally, the discussion provides a good overview of other relevant papers on DR, and the conclusions are mostly supported by the data. Here I provide a couple of suggestions that I believe might improve this work:…

      - Although the model is simple and understandable (Fig. 6), the inclusion of an overall summary or explanation in the figure legend would be appreciated, especially for readers that are not familiar with the terminology.

      - It might be the formatting while parsing the files but some of the in-text citations are between curly brackets (e.g., lines 80, 92, 93).

      - By definition, and unlike Canton-S or Oregon-R strains, w1118 flies are not wild-type but a genetic control. I believe that reference to this on the figures and text may need correction.

      We updated and/or our corrected each of these in the revised version. For wild-type, we more explicitly define this as wild-type for the relevant genetic locus.

      Reviewer #3 (Public Review):

      In this study, Hwangbo and co-workers investigate the extent to which the well-established life extending effects of DR rely on the molecular circadian clock and how the landscape of clock-controlled gene expression changes in the face of DR within the fat body of the fly, a tissue that performs the functions associate with both the liver and adipose tissue of mammals. The authors evidence that DR extends lifespan in a manner that depends on only one of the two major limbs of the fly's molecular circadian clock, namely the positive limb, that DR produces major changes in the identities of cycling clock output genes, and that genes related to the proteosome represent a major component of DR-induced transcript cycling. Though interesting, these conclusions are not strongly supported by the data and there are two major reasons for this. First, the authors rely on only one loss of function genotype each for the loss of positive and negative limb clock gene function. Second, though they wish to address the "circadian transcriptome" under normal and DR conditions, the authors conduct all their work under strong Light/Dark cycles, making it impossible to address circadian phenomena. These shortcomings are problematic in the extreme, as they leave open obvious alternative explanations for the results and fail to directly determine if the rhythmic expression, they observe are clock controlled or merely driven by the light/dark cycles, which themselves produce major effects on activity, feeding, etc., that may be responsible for differentially driving rhythmic transcripts under normal and DR conditions in the fat bodies.

      Major Weakness One: The use of only genotype each for the loss of positive (Clk^JRK) and negative (Per^01) limb of the circadian represents a major challenge for a central conclusion of the study. Phenotypes caused by the loss of a single clock gene may be due to the loss of circadian timekeeping, or they may represent a pleiotropic effect of the loss of function mutant being used. There are multiple precedents for pleiotropic (non-circadian) effects of clock gene mutants. It is, therefore, possible that the differences in the extent of DR mediated life extension between Clk^JRK and Per^01 may not represent a difference between breaking the positive and negative limbs of the clock but may simply reflect a pleiotropic effect of the dominant negative Clk^JRK. This possibility is acknowledged by the authors (lines 343-344). This could be addressed quite easily by extending the analysis to other loss of function mutants, for example, tim01 for the negative limb and cyc01 for the positive. Given the central focus here on the "circadian transcriptome," leaving open this alternative explanation for Clk's role in DR induced life extension represents a major weakness of the study. Furthermore, given the fact that Clk^JRK appears to be short lived on most of the media tested in the study, is it really surprising or informative that they would display lower life extension under DR?

      We confirmed that the large majority of LD oscillating genes in wild-type controls are disrupted in ClkJrk consistent with circadian clock regulation (Figure 3-figure supplement 1). As noted, we formally acknowledged that the circadian clock mutant alleles used here, and in fact any circadian clock alleles, can have pleiotropic, i.e., non-circadian, clock effects. This would only be partially mitigated by adding more (but also potentially pleiotropic) clock mutant alleles. Very challenging circadian resonance experiments (see Xu et al, 2019) are the gold standard for resolving circadian clock v. non-clock effects which are beyond the scope of this study which we now add to our discussion.

      We also note that foxo mutants are both short-lived and exhibit a robust lifespan extension to dietary restriction and thus the ClkJrk mutant is distinct in this regard. We have added this point to the Discussion.

      Major Weakness Two: The authors have not established that any of cycling transcripts they have detected in the fat body under normal and DR conditions are driven by the circadian clock. This is because: 1.) they have conducted their transcriptomic analysis on cells taken from flies entrained to light dark cycles, which can themselves drive daily changes in expression levels and 2.) they have not shown that the cycling measured on normal diet or DR conditions depends on a functional circadian clock. The "significant reorganization of the circadian transcriptome" is presented as a major conclusion of this study, but the authors have not addressed circadian control of transcription at all here, either by an examination of transcription under free-running conditions and/or in loss of function clock mutants.

      In addition, there is a logical gap in this study. The authors have shown that DR produces less life extension in Clk^JRK mutants than Per^01 or wild-type controls. They then show that DR produces changes in the rhythmic transcriptome when flies are place on DR. The central model presented in Fig. 6 shows/concludes that CLK drives increases in proteome-related transcript rhythms under DR. This conclusion could have been directly tested by asking if the changes in rhythmic gene expression induced by DR are gone the loss of function Clk mutants, or if the transcriptomic landscapes fail to differ between feeding conditions in these mutants.

      In conclusion, the study falls far short of directly testing the ideas it puts forth, greatly limiting its impact and interest.

      As noted above, we also examined the diurnal transcriptome in ClkJrk (at 4 hour resolution) and found that of the 290 genes that were detectably rhythmic in wild-type just 13 were rhythmic in ClkJrk consistent with the notion that oscillations depend on Clk (Figure 3-figure supplement 1). We now add this analysis to the manuscript. Nonetheless, we cannot exclude a role for light and thus have opted to use “diurnal” in place of “circadian” where appropriate for observed rhythms under LD conditions.

      Reviewer #3 (Recommendations For The Authors):

      Line 140 "showed an almost identical response" was a little hard to understand at first. Consider clarifying.

      This has been rephrased for clarity

      The authors claim that Clk mutants are "much longer lived" than wild-type controls on two of the relatively low calorie diets. Figure S3C certainly argues otherwise, and it's not clear how the data in 1C and S2C and warrant the use of "much" here.

      This wording has been rephrased and corrected in the revised version. The low-calorie diet shown in Figure 1-figure supplement 3 contains a higher sucrose concentration (5%) than those used in Figure 1 and Figure 1-figure supplement 2 (1%). This observation suggests that sucrose may play an independent role in the survival of ClkJrk mutants under malnutrition conditions.

      The authors should provide the rationale for the use of a dominant negative form of Clk for their experiments. Would the available amorphic allele be a better choice?

      As ClkJrk is the first described Clk allele and it is probably the most well characterized. As a dominant negative version which is still capable of dimerizing and binding DNA it is less susceptible to compensation by redundant bHLH transcription factors as has been observed for between mouse Clock and NPAS2 (Debruyne et al, 2006).

      It is not clear why the authors have chosen to examine transcriptomes so soon after transfer to DR. Why not wait longer. The authors provide context that changes are already taking place at the early time-point used, but would waiting a bit provide a more robust indication of how DR is changing the fat body?

      We noted in the manuscript that the effects of DR on survival are evident relatively soon (~2d) after a diet shift. We were interested in identifying those changes in daily transcription that would be occurring during that early time span and potentially be a cause rather than an effect of survival changes.

    1. eLife Assessment

      This valuable study provides a detailed three-dimensional characterization of primary cilia organization in the postnatal mouse growth plate, revealing reproducible spatial patterns in ciliation, ciliary length, and orientation. The evidence supporting these descriptive findings is solid, based on high-quality quantitative imaging and complementary genetic, mechanical, and transcriptomic approaches. However, evidence for the broader mechanistic conclusions is incomplete, particularly regarding whether ciliary orientation is uncoupled from basal-body and cell orientation and whether its stability under altered mechanical loading reflects a cell-intrinsic program. The study provides a helpful foundation for understanding primary cilia organization and mechanobiology in the growing skeleton, while the mechanisms underlying these observations remain to be established.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript provides fundamental insight into ciliary biology, specifically, how ciliary axoneme orientation is governed by microenvironmental bending or intrinsic cytoskeletal steering rather than strictly basal body docking coordinates.

      Strengths:

      There are three major strengths in this manuscript. First, combining high-resolution imaging with deep-tissue sectioning yields impressive lateral resolution, enabling robust separation of dual centrioles within the crowded chondrocyte extracellular matrix. Second, the authors established an automated pipeline that evaluates thousands of individual cells across multiple anatomical regions and differentiation zones, lending strong statistical weight to positional and volumetric measurements. Lastly, they demonstrate that ciliation peaks in the peripheral resting zone and ciliary length peaks in hypertrophic cells, providing a compelling cellular explanation for why Ift88 deletion impacts peripheral growth plate geometry and hypertrophic expansion.

      Weaknesses:

      There are three major weaknesses in this manuscript. First, the paper lacks explicit descriptions of data mentioned in the Abstract and Methods, including the RNA-seq differential expression, WGCNA modules, and immobilization/ciliary alignment data. Second, while the Methods section mentions correcting for "Z-blur" / point-spread function distortion in 3D spherical coordinate calculations, further detail is required on how orientations are disambiguated from optical sectioning depth artifacts. Lastly, the RNA-seq analysis demonstrates that ambulatory unloading alters hedgehog and primary cilia gene signatures, yet axoneme orientation itself remains static. The narrative requires a clearer mechanistic synthesis regarding how mechanical loading modulates ciliary signaling if physical alignment of the cilia is refractory to mechanical force in chondrocytes.

    3. Reviewer #2 (Public review):

      Summary:

      The aim of this work was to characterise in detail the cellular organisation of the growth plate in the growing mouse limb, the effect of mechanical loading due to physical activity, and the role of primary cilia in mediating this effect as putative mechanosensors. Primary cilia are generally thought to function as mechanosensors in a range of different organs and tissues, e.g. in the kidney. Exposure to mechanical loading is a normal part of post-natal limb growth, and the central hypothesis underlying this work was that primary cilia will play an important role in transducing the effects of mechanical loading into cellular responses during this process. To investigate this, the authors used a mouse model encoding fluorescent markers for primary cilia, and also allowing conditional knockout of IFT88, a protein that plays a key role in primary cilium biogenesis. The authors compared the effects of mechanical loading by comparing tissue from mice with normal or surgically immobilised limbs. To investigate the role of primary cilia, the authors performed the same experiments with mice treated with tamoxifen to induce IFT88 knockout.

      Strengths:

      A major strength of this work is that it studied primary cilia in a fully in vivo system. The authors employed cutting-edge imaging approaches and image analysis pipelines to rigorously study cellular organisation, centriole positioning, cilium length, and orientation across thousands of cells in limbs from 18 animals. The imaging results presented are of an exceptionally high-quality. The analysis of imaging data is extremely quantitative and robust and utilised appropriate statistical analyses, which were clearly stated throughout. The imaging data were complemented with transcriptomic data and analyses, which provided an orthogonal dimension for understanding cellular responses. An interesting and unexpected outcome from this is that the primary cilia are oriented in the same direction, and inclined at a roughly 45{degree sign} angle to the mediolateral and proximal-distal axes of the limb. The authors speculate as to how this might arise and the role it may play in sensing.

      Weaknesses:

      Overall, the work reported in this study was of a very high quality, and I could not find any significant shortcomings. However, I did feel that the paper was not very well written in many places, which made it difficult to read.

      Overall, I think that the authors did achieve their aims in this study. It will be interesting to unpick the cellular mechanisms that lead to alignment of the cilia, the role of this alignment in mechanosensing, and the molecular mechanisms by which the cilia sense mechanical strain. Thus, this work provides a fertile ground for future studies, which will have important consequences for the study of primary cilia in vivo.

    4. Reviewer #3 (Public review):

      Summary:

      This study aims to characterize the three-dimensional organization of primary cilia in the postnatal mouse growth plate and to determine how ciliary prevalence, length, position, and orientation vary across anatomical regions and stages of chondrocyte differentiation. The authors combine volumetric imaging and automated image analysis with conditional disruption of IFT88, limb immobilization, and bulk transcriptomics. The work provides a valuable anatomical dataset and identifies several interesting spatial patterns, including increased ciliation in the lateral resting zone, longer cilia in hypertrophic chondrocytes, and, most notably, a non-random orientation of ciliary projections despite a much broader distribution of centriole positions.

      The descriptive evidence is largely solid, and the imaging dataset should be useful to researchers studying primary cilia, skeletal development, and tissue mechanobiology. However, the evidence is incomplete for several of the central mechanistic conclusions. In particular, the current analyses do not establish that ciliary orientation is uncoupled from basal-body position, and they leave unresolved how ciliary orientation relates to cell orientation. The immobilization experiment also supports a narrower conclusion than the proposed cell-intrinsic orientation program, while the final model linking ciliary angle to multidirectional signal integration remains speculative. Overall, the authors succeed in identifying an interesting and reproducible anatomical pattern, but the mechanism underlying that pattern remains largely open.

      Strengths:

      A major strength is the scale and anatomical context of the imaging. Quantifying thousands of cilia and tens of thousands of centrioles in three dimensions while preserving information about growth-plate zone and position across the limb is technically demanding. This allows the authors to identify regional differences that would be lost in dissociated cells or bulk tissue measurements.

      The study also provides several potentially useful observations. Ciliation is higher in the lateral periphery, particularly in the resting zone, cilia are longer in hypertrophic chondrocytes, and ciliary projections show a reproducible non-random orientation. The latter is the most interesting result of the study and provides a useful foundation for asking how organelle orientation is established within a developing tissue.

      I also appreciated that the authors place these observations in several biological contexts rather than stopping at a descriptive atlas. IFT88 disruption alters ciliation and growth-plate cell organization, while immobilization changes growth-plate dimensions, cell morphology and orientation, and the transcriptome even though the measured ciliary properties remain comparatively stable. These perturbations give the anatomical observations useful biological context.

      Weaknesses

      The main concern is that the claim that ciliary orientation is uncoupled from basal-body position is not directly demonstrated. The manuscript shows that centriole positions are broadly distributed and, separately, that ciliary projections have a preferred tissue-level orientation. Different population-level distributions, however, do not establish independence within individual cells. For example, basal bodies could be broadly distributed while their position determines which of two opposite directions along a common tissue axis the cilium adopts. This possibility is particularly relevant because Centrin-2 labels both centrioles, whereas only one serves as the basal body. The current data therefore support a difference between the population distributions of position and orientation, but not yet the stronger claim that the two are uncoupled.

      A related gap is the relationship between cell orientation and ciliary orientation. The manuscript measures both, and cell orientation changes with growth-plate region, IFT88 deletion, and immobilization, while ciliary orientation appears relatively stable. Yet the two measurements are never directly related within the same cells. It is therefore unclear whether cilia adopt a reproducible angle relative to the major axis of their own cell, whether this relationship changes between the center and periphery, or whether changes in cell organization can occur independently of local ciliary alignment. This seems important for interpreting the tissue-level orientation pattern.

      The statistical treatment of orientation also deserves caution. These are circular or spherical measurements, yet much of the analysis relies on linear distributions and Kolmogorov-Smirnov tests. This is particularly problematic around the 0{degree sign}/360{degree sign} boundary, where values near 350{degree sign} and 30{degree sign} are geometrically close but appear separated in a linear representation. The manuscript also does not clearly distinguish between a preferred axis, where opposite directions are equivalent, and a preferred polarity, where one direction is favored. In addition, thousands of cilia are nested within a much smaller number of mice, so the apparent statistical power should not be driven primarily by pooled object counts. Given that regional differences are a central theme of the paper, it would also be useful to know more clearly whether the preferred axis or the strength of the orientation bias differs between the middle and lateral growth plate at the animal level. I also could not identify a formal comparison of the centriole or ciliary distributions with an appropriate uniform circular or spherical null. The reported tests mainly compare zones and regions, so the claims that centriole position is non-preferential and ciliary orientation is non-random are not yet statistically established in the form presented.

      The IFT88 conditional knockout is not carried through to the principal orientation question. The authors examine cell size, cell-axis organization, ciliation, and cilium length, but do not report whether the remaining cilia retain the preferred orientation or whether centriole positioning changes. Given the central role of this genetic perturbation in the manuscript, this leaves the genetic and orientation arms of the study somewhat disconnected.

      Finally, the immobilization experiment and the mechanistic interpretation should be separated more carefully. Only four animals were analyzed in the offloaded and contralateral conditions, and medial and lateral regions were averaged because of the small sample size. The experiment shows that the established ciliary orientation remains relatively stable over a two-week postnatal interval despite clear changes elsewhere in the tissue. It does not exclude a role for mechanical forces earlier in establishing the axis, nor does a nonsignificant difference with four animals demonstrate equivalence. Similarly, the proposed cell-intrinsic orientation program and the final "single-axis blindness" model are interesting hypotheses, but the study does not yet identify what establishes the axis or how the observed ciliary angle would alter sensitivity to a force or biochemical gradient. Those ideas are worth discussing, but they should remain clearly separated from the observations directly supported by the data.

    5. Author response:

      In response to the valuable reviewers’ comments and suggested changes, we are finalising changes to the manuscript, in order to resubmit a revised version, and a document with full author responses, that reflects all the review comments.

      These changes include adding points of clarity, improving accuracy on wording of key messages, adding additional interpretation of the data, including additional data and analyses that reflect open questions raised, and more discussion concerning these unanswered questions, which are subjects of future work.

      In response to specific points we were asked to provisionally address (actions in italics):

      Reviewer #1. We are pleased the reviewer sees the insight these data bring. We indeed think it likely that cilia axoneme orientation is governed by the immediate microenvironment and/or changes to the cytoskeleton and are actively looking to explore this.

      To address the 3 areas of concern:

      (1) Our re-writing of the abstract and the results concerning transcriptomic data seeks to overcome weaknesses in descriptions of these data.

      (2) Further detail is being added on how z-distortion is corrected for, so accuracy is the same in all axis and orientation measurements are robust.

      (3) We will add to the discussion to add our thoughts as to why ciliary and cilia signalling genes are regulated by immobilisation, but that immobilisation does not apparently affect cilia structure.

      Reviewer #2. Thank you for such broadly positive comments, we are pleased the scale and depth of the quantitative analyses comes across, but will make sure that revisions throughout improve the quality of the writing describing these. We are actively exploring means to test ideas for how cilia axoneme become orientated in this way and what the function is. These preliminary ideas will be reflected in the discussion.

      Reviewer #3. Thank you for such a detailed and thoughtful review. They will ensure the data are presented to their very best and we will address the concerns raised. This descriptive study had a hypothesis, and made discoveries which surprised us, we have tried, as you say, to put this in some context of the role of cilia and the role of mechanical forces in GP biology. Most notably we are considering that uncoupled is not the correct term here. To address areas of concern;

      (1) We agree ‘uncoupled’ is not the correct word here. We cannot find a pattern of correlation between centriole position and orientation. However, the two can’t be ‘uncoupled’ and there is no proven independence on a single-cell level (that cilia position and cilia orientation are not in any way mechanistically linked). We will make changes and add more details on what we have considered in this area.

      (2) Similarly, we have not correlated in each cell, cellular orientation and ciliary orientation, though we have made attempts and not yet found a relationship. However, again, this is not the same as one being absent. We might have expected ciliary orientation to change as cell orientation does (through zones or with pertubations) as we have seen in vitro but this remains to be fully explored and is one subject of follow-up work. We are considering column populations and per animal considerations of the data.

      (3) We did take a cautious approach to statistics and specialist advice, but advice was not to overcomplicate things when there are 2 main messages related to centriole position and cilia orientation. Firstly, centriolar position appears random or without preference thus distribution of position on cell, is homogenous. Second, ciliary orientation angle is not a homogenous distribution, as would be expected if random with this number of measurements. We do, and will add comments to this effect, have to mindful of large dataset, but do not think this means we are looking at false discoveries due to number of comparisons. We are considering how better to reflect this. To address these important points we will add a section to the methods and discussion and will endeavour to change results to this end as it is a central point of the manuscript.

      (4) We will add new data concerning IFT88cKO and centriole position and orientation.

      (5) We will add a critique of the immobilisation experiments to ensure the relatively diminished power is clear. We agree force may have set things up initially, we will ensure this is discussed and we will ensure our proposal for why cilia orientation is this way is framed as a hypothesis. We have preliminary data, but this is the subject of an entire new project thus not yet supported by robust experimental evidence so is speculative at this stage and we will ensure this is clear.

    1. eLife Assessment

      This Review Article addresses a significant and timely topic that concerns the role of cryo-EM in transforming RNA structural biology from static, individual conformations toward the reconstruction of dynamic conformational ensembles and energy landscapes. The authors describe eight case studies spanning different RNA classes and argue that conformational heterogeneity is a source of mechanistic information, rather than a limitation.

    2. Reviewer #1 (Public review):

      Summary:

      The review addresses an important and timely topic that concerns the role of cryo-EM in transforming RNA structural biology from a "static" discipline to one increasingly concerned with conformational ensembles and molecular dynamics. The scope is well within the eLife standards, and the style and general architecture do fit eLife.

      Strengths:

      The review is extremely well written, well-conceived and clear. The main strengths are in the breadth of coverage, the clear theme, and the inclusion of practical examples that explain in detail the construct design, sample preparation, vitrification, and data analysis. The manuscript will certainly be impactful and valuable, especially for readers who are not specialists in cryo-EM, as it provides an accessible overview of recent advances across a wide range of RNA systems.

      Weaknesses:

      My only reservation is that, currently, the review reads too much like a list of examples. The authors should make an effort, and I am sure they are well up to it, to try to synthesise the message, provide more critical insights and amalgamate the text better, to really reach a wider audience.

      If revised along the lines detailed below, I am sure that the review will become an authoritative and influential resource for the RNA structural biology community.

    3. Reviewer #2 (Public review):

      Summary:

      In this review, the authors set out to synthesize how cryo-EM is reshaping RNA structural biology, moving the field from the determination of static, individual conformations toward the reconstruction of dynamic conformational ensembles and energy landscapes. Through eight case studies spanning ribozymes, riboswitches, viral RNAs, and synthetic RNA assemblies, they aim to show how cryo-EM has revealed mechanisms of RNA motion, including folding, ligand-dependent switching, and cooperative assembly, and to provide a practical account of the experimental and computational challenges (construct design, sample preparation, vitrification, data analysis) that are specific to dynamic RNA targets.

      Strengths:

      The manuscript succeeds in bringing together a wide and genuinely current range of case studies illustrating the field's shift toward dynamics-focused cryo-EM, several published within the last one to two years. The dedicated "Challenges" section, which walks through construct engineering, buffer and vitrification optimization, grid screening, and heterogeneity-resolving computational approaches, is a particularly useful and practical contribution; it goes beyond simply cataloguing structures and gives readers new to the area a genuine methodological roadmap. The figures are detailed and well matched to the quantitative claims made in the text (helix rotations, distances, RMSDs), which strengthens the paper's value as a reference resource.

      Weaknesses:

      The coverage of two areas in particular, the SL5 viral RNA element and RNA quaternary/multimeric assemblies, would benefit from incorporating additional recent primary literature that is directly relevant but currently omitted. This does not undermine the manuscript's core narrative, but it means the review is presently less complete than it could be as a field synthesis, particularly for readers using it to identify the full body of recent work on these specific RNA classes.

      More concerning is that one specific structural claim, describing conformer heterogeneity in the cobalamin riboswitch (Case study 3, holo dimer 4), appears to invert the finding reported in its own source paper (Ding, Deme et al., 2023). As written, the manuscript states that P2 and the distal half of P6 are structured in dimer 4, whereas the source paper reports that these are the regions that could not be modeled. This is worth flagging prominently because it is a factual claim about a specific structure, not an interpretive point, and readers relying on this review as a secondary source could come away with an inverted understanding of that structure's flexibility.

      The manuscript also contains several minor internal inconsistencies. None of these individually threatens the paper's core arguments, but together they suggest the manuscript would benefit from a careful proofreading and reference-list audit pass.

      Overall, the authors largely achieve their stated aim. The case studies convincingly illustrate that cryo-EM can now resolve discrete and continuous conformational states of RNA at near-atomic resolution, and the Challenges section substantiates the claim that construct design, sample preparation, and computational innovations have been jointly responsible for this progress. The gaps in coverage of SL5 and multimeric RNA literature, and the inverted claim in Case study 3, are the main respects in which the manuscript falls short of being a fully comprehensive and accurate synthesis at this stage, but these are correctable issues rather than flaws in the overall argument or framework.

      This review is likely to be a useful entry point and practical reference for researchers moving into RNA cryo-EM, particularly given the level of methodological detail in the Challenges section. Its impact would be strengthened by two additions. First, a short discussion of recent cryo-EM advances in tRNA would round out the manuscript's coverage of classical RNA structural targets alongside the ribozyme, riboswitch, and viral RNA case studies already included. Second, the raiA non-coding RNA, currently mentioned only briefly, has in the last two years become a genuine model system for cryo-EM-based ncRNA structure determination, progressing from a single novel-fold discovery to a comparative structural framework spanning multiple raiA subtypes and candidate protein partners. Expanding this into a full case study would let the manuscript showcase, in a single worked example, exactly the kind of field-level progression (from novel fold to scaffold-based strategy and further to comparative structural framework) that the review's own framing describes as the trajectory of the field as a whole.

    4. Reviewer #3 (Public review):

      This review describes how cryo-EM is moving RNA structural biology from the determination of static structures toward the characterization of conformational ensembles. Eight case studies spanning self-splicing introns, riboswitches, viral RNA elements and synthetic assemblies show how cryo-EM has captured folding intermediates, hinge-mediated domain motions and ligand-dependent switching, and the authors pair these examples with practical guidance on construct design, sample preparation, vitrification and heterogeneity analysis. The argument that conformational heterogeneity is a source of mechanistic information, rather than a limitation, is well supported, and Table 1 will be a useful reference for laboratories entering the field. The manuscript is timely and of broad interest, and I recommend publication after minor revision.

      (1) Scaffold-based structure determination (page 14, lines 5-9). The section on chimeric RNAs presents the scaffold strategies as a single group and cites Haack et al. (2025) and Langeberg & Kieft (2023) in one parenthetical, so individual results are not attributed to their sources. The distinction between these studies is substantive. Earlier scaffolds based on the Tetrahymena group I intron (Langeberg & Kieft, 2023) or on RNA origami (Sampedro Vallina et al., Nucleic Acids Res. 51:4613-4624, 2023) resolved the appended RNAs at approximately 4.4-5 Å. Haack et al. (2025) reported the first RNA scaffold to yield a high-resolution structure of the target itself, resolving the ligand-binding pocket of the thiamine pyrophosphate (TPP) riboswitch at 2.5 Å and extending nucleotide-level cryo-EM analysis to small RNAs that had previously been intractable. I ask the authors to attribute each result to its source and to state explicitly that Haack et al. (2025) achieved the first high-resolution structure of a scaffolded target RNA. Because the same study captured the ligand-free TPP riboswitch in an open, Y-shaped conformation, a direct example of the ligand-dependent switching that is central to this review, the TPP riboswitch should also appear in the riboswitch section (page 8, lines 15-36), together with the fluoride riboswitch of Langeberg & Kieft (2023).

      (2) Page 6, lines 8-14. Self-splicing introns are described collectively as evolutionary ancestors of the spliceosome that reside in pre-mRNA transcripts. The proposed ancestral relationship to spliceosomal introns and snRNAs applies to group II introns. Group I introns, including the Tetrahymena intron, which interrupts a pre-rRNA, initiate splicing with an exogenous guanosine and are not considered spliceosomal precursors.

    1. eLife Assessment

      This study provides an important insight into how the medial and lateral entorhinal cortices interact through distinct excitatory and inhibitory pathways. Using anatomical tracing, optogenetics, and electrophysiology, the authors show that glutamatergic medial entorhinal neurons provide broad excitatory input to lateral entorhinal, while long-range SST+ interneurons deliver selective inhibition to layer I. These findings reveal a novel layer- and cell-type-specific organization of medial to lateral entorhinal connectivity with implications for spatial and episodic memory. The work is convincing, but validation of injection specificity and viral spread is needed to fully confirm the anatomical interpretations; with these clarifications, this will be a significant contribution to understanding entorhinal-hippocampal circuit organization.

    2. Reviewer #1 (Public review):

      The study addresses the organisation of synaptic connections from medial to lateral entorhinal cortex. Classic anatomical work has suggested these connections exist but very little is known about their identity or functional impact. The manuscript argues that these projections are mediated by glutamatergic neurons, providing excitatory input from MEC to all layers of LEC, and by SST+ve interneurons sending inhibitory projections to L1 of LEC. This appears the most likely interpretation of the data. Potential concerns about confounds due to spread of virus/tracer from the injection site are addressed in the supplemental figures. My view is that the weight of evidence favours the authors' interpretation although the evidence isn't quite compelling.

      Knowing the configuration of projections from MEC to LEC is important for thinking about circuit mechanisms for spatial cognition and episodic memory. This study adds to an emerging view that MEC and LEC can interact directly, indicating that the cell-type level organisation of these interactions is asymmetric and identifying an intriguing long range inhibitory pathway.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Nilssen et al. presents a comprehensive study of the circuitry linking the medial and lateral entorhinal cortices (MEC and LEC). Using a combination of anatomical tracing, optogenetics, and in vitro electrophysiology, the authors convincingly demonstrate that the MEC sends both glutamatergic and long-range inhibitory SST+ GABAergic projections to the LEC, with distinct laminar and cell-type-specific targeting. Notably, they reveal that SST+ inhibitory projections selectively suppress the activity of layer IIa neurons, whereas excitatory inputs preferentially engage neurons in layers IIb and III, thereby differentially modulating hippocampal-projecting populations.

      Strengths:

      The experiments are carefully executed, the results are compelling, and the conclusions are well supported by the data. This work will be of broad interest to researchers studying memory circuits, cortical inhibition, and the organization of long-range connectivity.

      Weaknesses:

      Although the in vivo relevance of these connections remains to be determined, this is an important and timely contribution to our understanding of entorhinal-hippocampal interactions.

      Comments on revised version.

      The authors have addressed my comments satisfactorily, and I am satisfied with the changes made in the revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study addresses the organisation of synaptic connections from the medial to the lateral entorhinal cortex. Classic anatomical work has suggested these connections exist, but very little is known about their identity or functional impact. The manuscript argues that these projections are mediated by glutamatergic neurons, providing excitatory input from MEC to all layers of LEC, and by SST+ve interneurons sending inhibitory projections to L1 of LEC. This appears to be the most likely interpretation of the data, although in my opinion, more could be done to rule out the possible impact of the spread of the virus/tracer from the injection site.

      While this concern might seem overly picky, the importance of this level of detail is nicely shown by the authors' previous work clarifying connectivity from postrhinal to entorhinal cortices through careful analysis of similar types of data (Doan et al. 2019). If additional analyses/data can address the concern here, then I think this will be an important set of fundamental results that will influence thinking about circuit mechanisms for spatial cognition and episodic memory. In particular, it will nicely add to an emerging view that MEC and LEC can interact directly, showing that the organisation of these interactions is asymmetric and identifying a potentially interesting long-range inhibitory pathway.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Nilssen et al. presents a comprehensive study of the circuitry linking the medial and lateral entorhinal cortices (MEC and LEC). Using a combination of anatomical tracing, optogenetics, and in vitro electrophysiology, the authors convincingly demonstrate that the MEC sends both glutamatergic and long-range inhibitory SST+ GABAergic projections to the LEC, with distinct laminar and cell-type-specific targeting. Notably, they reveal that SST+ inhibitory projections selectively suppress the activity of layer IIa neurons, whereas excitatory inputs preferentially engage neurons in layers IIb and III, thereby differentially modulating hippocampal-projecting populations.

      Strengths:

      The experiments are carefully executed, the results are compelling, and the conclusions are well supported by the data. This work will be of broad interest to researchers studying memory circuits, cortical inhibition, and the organization of long-range connectivity.

      Weaknesses:

      Although the in vivo relevance of these connections remains to be determined, this is an important and timely contribution to our understanding of entorhinal-hippocampal interactions.

      The request for validation of injection specificity and viral spread, as detailed in the comments and suggestions of the two reviewers has been provided in the revised version. We added supplementary figures 1,2 and 6 as well as an extra insert into the old supplementary figure 5, now supplementary figure 9.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Interpretation of the retrograde labelling experiment in Figure 1A-C is challenging, as the spread of the tracer at the injection site is not shown. It's important to see the full dorsal-ventral extent of the injection site in order to establish that the labelling of neurons in MEC results from projections to LEC and not adjacent areas (including MEC).

      As mentioned in our initial reply, we are fully aware of the risks associated with an incomplete assessment of injection sites and viral spread, so we provide a new Supplementary Fig. 1 showing 6 dorsoventral levels of the case shown in Fig. 1A. The injection site in case of FG often shows a core of damaged tissue with a halo of substantial unspecific fluorescence. Outside of the injection side, one only sees retrogradely labeled somata (often recognizable by FG signal clustered in lysosomes) and dendritic elements. As indicated in the legend, we report some tracer leakage along the needle track in temporal and perirhinal cortex, areas that receive only sparse MEC projections, but there is no apparent spread of the injection into MEC.

      Numbers are also quite low. E.g., Figure 1A is N=1/2 for retrograde labelling experiments.

      The reviewer would be correct if the experiments were meant to analyze MEC projections to LEC in full anatomical detail. This was not our intention (several papers addressed this pathway in detail), we merely aimed to distinguish between glutamatergic and potential GABAergic contributions to this pathway and to establish optimal coordinates in slices to prepare for the electrophysiological recording experiments. Two animals suffice for this purpose and using more would be against the aim of reducing the use of experimental animals as much as possible.

      The rationale here for the use of AAV2-CAG-tdTomato as a retrograde tracer is unclear. My understanding is that this is more effective as an anterograde tracer. Some clarification and validation would be important.

      The reviewer is correct that AAV2 is generally considered an effective anterograde tracer, but tracing the connectivity of entorhinal cortex with AAVs has been proven to be notoriously difficult, in particular retrograde tracing of inputs to layer II. In a neighboring lab in the centre, headed by Edvard and May-Britt Moser, it was established that AAV2 types show very efficient retrograde transport and that is why we decided to use the virus. Also, in our hands the virus showed excellent retrograde transport that served our purpose

      (2) Interpretation of the anterograde experiments in Figure 1D-E would also benefit from showing evidence that the injection sites are restricted to MEC. It should be straightforward to make a supplemental figure showing labelling at all dorsoventral levels.

      More careful analysis of the axon labelling in the dentate gyrus could also help make a case for the selectivity of the injection site for eGFP. In this case, only the intermediate portion of the molecular layer of the DG should be labelled. In the image shown, the labelled band is quite wide, but it's hard to tell if this reflects the plane of section or is because it also includes labelling in the outer molecular layer (which would be indicative of LEC expression).

      We thank the reviewer for these two suggestions, and we have prepared a new Supplementary Fig. 2 in line with this.

      Numbers are also on the low side for these experiments.

      See our response above

      (3) For optogenetic experiments in Figure 2, the selectivity of targeting of AAV to MEC is assessed through the specificity of labelling in the DG. This is great, but it's important to show that this specificity is maintained at all dorsoventral levels.

      Higher resolution images of labelling in LEC could also be helpful. It's hard to tell from the images in 2A if labelling is axonal or is in the soma adjacent to the nuclear NeuN signal (which would indicate a lack of selectivity for MEC).

      We thank the reviewer for these two suggestions and provide a new Supplementary Fig. 6, showing both the details of AAV1 being present only in neuropil in MEC not in somata as well as the specific labeling in the middle molecular layer of DG in detail. Including all dorsoventral levels would not provide additional information in view of the very well-established topographical organization of the entorhinal to dentate projection, reaching approximately 20 -25 % of the full long axis of DG (Van Groen et al., 2003)

      In addition, we have again carefully screened all tissue from the electrophysiological experiments for possible leakage of virus from MEC to LEC. We decided to exclude recordings from one mouse, which had labelling in MEC that was close to the border with LEC. Neuron counts have therefore been adjusted (pages 7-9) and the example recording showing responses to TTX/4-AP exposure in Figure 2B has been exchanged.

      (4) The analysis of excitatory and inhibitory opto-responses in Figure 2 is nice. It may be helpful to report quantification of the rise and decay kinetics of the synaptic currents. They appear much slower for the inhibitory input, which may be functionally important.

      This would indeed be nice to add, but it would not significantly impact or change the main message of our study. Since the lab of the senior author (MPW) has been discontinued and the resources for conducting these analyses are not readily available anymore, we have found it difficult to comply with the reviewer’s request

      (5) More direct evidence for SST axons projecting from MEC to LEC would strengthen the conclusions made. E.g., in experiments where the SST neurons are labelled, is it possible to follow the axons? Do they project as expected from the MEC to the LEC?

      In our view the tracing data provide convincing evidence in support of a direct projection from MEC to LEC by SST neurons, as shown in horizontal brain sections where SST axons labelled in MEC of an SST<sup>Cre</sup> mouse projects within Layer I from the site of origin in MEC to Layer I of MEC (Supplementary Figure 3). Similar visualizations were not possible to obtain in our electrophysiological experiments where semicoronal slices were used. This cutting angle has been shown to be optimal to preserve most of the axon and the dendritic tree of LEC neurons (Tahvildari and Alonso, 2005; Canto and Witter 2012), but does not maintain the projection from MEC to LEC.

      Minor Points:

      (1) "These layers are heavily innervated by medial entorhinal axons (Figure 1F...". I don't see a 1F.

      This has been corrected; should have been Figure 1E.

      (2) Methods should report series resistance values for patch-clamp experiments (range and mean).

      Fully agree and this information has now been added on page 22 of the manuscript:

      Under Voltage clamp: ‘Recordings with series resistance ≤ 25 MΩ were accepted, with an average of 16.2 MΩ for voltage clamp recorded neurons (range, 4.9 – 24.9 MΩ).’

      Under Current clamp: ‘All recordings (series resistance: 18.9 MΩ, 5.0 – 66.7 MΩ; mean, range) were included for analysis.’

      Reviewer #2 (Recommendations for the authors):

      (1) Please specify in the figure or, alternatively, in the figure legend which virus was used in each group shown in Figures 2H and 2I. This is somewhat confusing, since Figure 2E illustrates a specific combination of viruses and mouse lines that only corresponds to part of Figure 2H. While this information is provided in the text, including it directly in the figure would help the reader.

      We thank the reviewer for this excellent suggestion, and we have implemented this in the new version of figure 2.

      (2) In Figure 4, regarding the inputs from PIR, cLEC, and PER to LEC, the inhibitory components recruited by each input were not examined as thoroughly as for the MEC inputs. In fact, some inhibitory interneurons were double-labeled in the GAD67 mice (Figure 4B), which could also influence the responses of LEC neurons, especially for PER inputs. Recordings in Figure 4C appear to have been obtained near the reversal potential for inhibition, which may have prevented the observation of inhibitory effects. The authors could discuss this point in the Results.

      The reviewer is correct and this issue is now addressed in the relevant section in the results (page 11):

      ‘It should be noted, however, that it is possible that inhibitory effects could have been masked in some recordings, due to the resting membrane potential in our recordings being close to the theoretical chloride equilibrium potential. This could be particularly relevant for the inputs from PER, an area where we found LEC-projecting GABAergic neurons (Fig. 4B) and which is known to provide long-distance inhibition to LEC (Pinto et al., 2006; Apergis- Schoute et al., 2007).’

      (2) A diagram summarizing the known connections among MEC, LEC, and the hippocampal formation, highlighting the relevant cell types, layers, and the new connections identified in this study, would be a valuable addition, perhaps as a supplementary figure.

      We appreciate the suggestion, though find a full summary of known connectivity a bit overdone. Instead, we included a new figure 6 that summarizes the main new findings of the paper in the context of LEC projections to the hippocampal formation.

      (3) Although the main focus is on MEC-LEC connectivity, the experiments examining interactions with other cortical areas and converging inputs would benefit from a discussion of how MEC-driven inhibition of LEC might influence those inputs and shape the resulting output to the hippocampus. Including a short paragraph addressing this in the Discussion section would strengthen the manuscript.

      Excellent suggestion although we did speculate briefly in the result section on the possible effect. We have added a short paragraph in the discussion (page 15), reiterating the part in the results (page 13, last paragraph of results). We also briefly discussed the potential functional relevance of the suppression of the pathway from layer IIa to DG-CA3/CA2 versus the facilitation of activity in the pathway from layers IIb/III to CA1 and subiculum (last section of the discussion).

      (4) Lastly, it would be interesting to know what the main source of activation is for the SST long-range LEC projecting neurons. Are these neurons recruited in a feedback manner by the activity of MEC excitatory cells? I realize this question is beyond the scope of the present study, but if the authors have any data or insights related to this point, including a brief discussion would be valuable.

      This is an interesting thought, and we have included a new supplementary figure (supplementary Figure 5) showing data from experiments mapping monosynaptic inputs to MEC SST neurons using rabies virus. Although it was not possible to target only MEC SST neurons that project to LEC, the data show which are the main extrinsic inputs to the population of MEC SST neurons, most likely including those that project to LEC.

    1. eLife Assessment

      This is an important study of critical period plasticity, focused on temperature manipulations, and how different parts of the Drosophila larval motor circuit adapt or maladapt. The work convincingly demonstrates that components of the motor network respond in distinct ways to heat shocks, and the combination of functional, structural, and electrophysiological approaches makes the study of significant interest. The work points to central interneurons as primary drivers of maladaptive changes, while motoneurons and neuromuscular junctions show compensatory or homeostatic adjustments. The authors have thoughtfully and substantially addressed the concerns raised in the original reviews through new temporal-window experiments, additional receptor and behavioral analyses, clearer discussion of the study's limitations, and an integrative model.

    2. Reviewer #1 (Public review):

      Summary:

      The authors examine the impact of heat stress during an embryonic CP in Drosophila, focusing on the larval locomotor network. They show that elevated temperature increases neuronal activity and, when applied during the CP, results in long-term instability of the network which manifests in prolonged seizure recovery times. At the neuromuscular junction, substantial structural changes occur, including terminal overgrowth and altered receptor composition, yet synaptic transmission remains preserved due to homeostatic regulation. Motoneurons display reduced excitability but receive increased synaptic input from premotor interneurons. These findings suggest that maladaptive instability originates within the central circuitry rather than at the neuromuscular junction, where changes seem to be homeostatically compensated. The study concludes that different network components exhibit distinct and hierarchical responses to CP perturbations, with premotor interneurons setting the tone for downstream adjustments in motoneurons.

      Strengths:

      The work takes advantage of the unique accessibility of the Drosophila system. A major strength of the study is the integration of structural, physiological, and behavioral analyses, which allows the authors to draw a comprehensive picture of how CP perturbations shape the locomotor network. The choice of an ecologically relevant stimulus (heat stress) is particularly convincing, as it links experimental manipulations more closely to natural environmental conditions. The experiments are carefully designed, and the results are robust and consistent with previous findings in the field, while also extending them in new directions. Importantly the work clarifies how temperature perturbations within distinct developmental time windows affect different properties of motor circuit formation.

      Weaknesses:

      A small limitation of the study is that it remains difficult to integrate maladaptive (seizure recovery) and adaptive/homeostatic phenotypes within a single mechanistic framework, leaving some space for interpretation.

      Comments on revised version.

      I think the authors did a great job at revising the manuscript and they addressed all my comments.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a thoughtful and well-executed study of critical period plasticity in the Drosophila larval motor circuit. The authors examined how transient heat, 32C, during embryonic stage, altered network properties, showing that premotor interneurons A27h increase excitatory drive onto motoneurons, which respond with a reduction in excitability. At the NMJ, synaptic terminals expand and GluRIIA distribution shifts, yet synaptic transmission remains largely unaffected. Despite these local compensations, the treated larvae display slower crawling and prolonged recovery from seizures, indicating that the network is functionally compromised.

      Strengths:

      (1) One of the major strengths of this study is the elegant dissection of a defined circuit, tracking changes from premotor interneurons through motoneurons to the NMJ. The multimodal approach provides a comprehensive view of how connected elements respond to CP perturbations.

      (2) An interesting finding is that NMJ morphology changes dramatically without corresponding deficits in synaptic transmission, challenging the common assumption that larger boutons necessarily indicate stronger synapses.

      (3) Another intriguing result is that even with two layers of homeostatic compensation, locomotor behavior is still impaired, highlighting the limits of compensation and underscoring the critical role of CP timing.

      (4) Beyond these scientific insights, the study benefits from a well-defined, tractable system and simple experimental manipulations, which together make the results highly interpretable and reproducible.

      Comments on revised version.

      The authors have carefully considered my comments and recommendations and have made substantial efforts to improve the clarity and validity of the study. Although additional electrophysiology experiments using shorter heat stress windows were not feasible, the authors performed additional analyses of postsynaptic GluRs and provided a clearer discussion of the study's limitations. Overall, this is a strong and well-written paper that establishes a valuable foundational framework for addressing interesting and important questions about adaptive responses in developing neural circuits.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This is an important study of critical period plasticity, focused on temperature manipulations, and how different parts of the Drosophila larval motor circuit adapt or maladapt. The work convincingly demonstrates that components of the motor network respond in distinct ways to the heat shock, and the combination of functional, structural, and electrophysiological approaches makes the study of significant interest. The work points to central interneurons as primary drivers of maladaptive changes, while motoneurons and neuromuscular junctions show compensatory or homeostatic adjustments. The study is methodologically rigorous, contributing important insights into critical period biology using a tractable invertebrate model.

      We thank the reviewers for their thoughtful critique and suggestions. We agree with these and, where possible, we have attempted to address these, improving this study. As outlined below, we have revised the manuscript substantively and included additional data and figures.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors examine the impact of heat stress during an embryonic CP in Drosophila, focusing on the larval locomotor network. They show that elevated temperature increases neuronal activity and, when applied during the CP, results in long-term instability of the network, which manifests in prolonged seizure recovery times. At the neuromuscular junction, substantial structural changes occur, including terminal overgrowth and altered receptor composition, yet synaptic transmission remains preserved due to homeostatic regulation. Motoneurons display reduced excitability but receive increased synaptic input from premotor interneurons. These findings suggest that maladaptive instability originates within the central circuitry rather than at the neuromuscular junction, where changes seem to be homeostatically compensated. The study concludes that different network components exhibit distinct and hierarchical responses to CP perturbations, with premotor interneurons setting the tone for downstream adjustments in motoneurons.

      Strengths:

      The work takes advantage of the unique accessibility of the Drosophila system. A major strength of the study is the integration of structural, physiological, and behavioral analyses, which allows the authors to draw a comprehensive picture of how CP perturbations shape the locomotor network. The choice of an ecologically relevant stimulus (heat stress) is particularly convincing, as it links experimental manipulations more closely to natural environmental conditions. The experiments are carefully designed, and the results are robust and consistent with previous findings in the field, while also extending them in new directions.

      Weaknesses:

      The study leaves some uncertainty regarding the experimental design and interpretation. The change from short to prolonged heat shock manipulations raises the possibility that the effects observed may not be confined to the critical period alone - this could be experimentally addressed or simply rephrased in the text.

      We agree that clarity about the experimental paradigm is important and have addressed this as suggested, within text, figures and figure legends: the duration of embryo exposure to 32˚C heat stress is now unambiguously stated and each figure has a graphical illustrations of the heat stress paradigm. For example, experiments represented in Figures 1, 3 (new data) and 8 (new data) used short, defined periods of a few hours of heat stress, aimed to identify specific windows of development that are sensitive to 32˚C heat stress. These also show that behavioural changes result from heat stress experienced during the specific 2-hour window that defines the critical period of the developing central locomotor circuitry, namely from 17-19 hours after egg laying - previously identified by Giachello & Baines (2015). Longer exposure to 32˚C heat stress during embryogenesis result in the same phenotypes when this 2-hour window is included, causing the same level of reduced larval crawling speed and lowered network stability, which manifests in increased seizure recovery times. This is as one might expect from a critical period of nervous system development.

      Where the neuromuscular junction is concerned, where we identified embryonic heat stress causing phenotypes that are evident at late larval stages, a more complex model has emerged. Following suggestions from both reviewers to explore shorter heat stress exposures during embryogenesis, additional experiments (see Figure 3) we identified what might be a critical period for the body wall muscles. This is an earlier window of development, within 13-16 hours after egg laying, which is sensitive to heat stress in terms of the levels of the GluRIIA glutamate receptor subunit that will be expressed in the late larva. This developmental period is characterised by muscles acquiring their electrical properties (Broadie & Bate, 1993), i.e. comparable to the central locomotor network transitioning through its critical period at the time that it becomes active. The neuromuscular junction is composed of both presynaptic motoneurons and postsynaptic muscles, and therefore this composite structure is subject to multiple, sequential critical periods. The characterisation of changes to neuromuscular junction synaptic physiology was carried out using heat stress throughout most of embryogenesis (Figure 4). We think this appropriate from the perspectives of having included all relevant critical periods (muscle and CNS) to explore how this composite structure responds to environmental heat stress; also based on our observations that for each critical period phenotypes are defined by the experience during the critical period and not exacerbated by prolonged heat stress either side.

      In addition, the maladaptive (seizure recovery) and adaptive/homeostatic phenotypes are not always clearly distinguished or highlighted, which makes it harder to appreciate how the different levels of the network plasticity fit together into a single mechanistic framework.

      Following the suggestion, we have tried to clarify the mechanistic framework in a new figure that aims to summarise the model in Figure 9.

      The question of whether phenotypes that result from an embryonic heat stress manipulation are adaptive or maladaptive is difficult to resolve. This is partly due to the nature of critical periods, since perturbations during these developmental windows can cause significant, long-lasting maladaptations that are challenging to reconcile from a perspective of adaptive plasticity. Secondly, in light of the nature of this animal, which has evolved a particularly rapid development and large brood sizes, any deviation from the evolved optimum developmental temperature of 25˚C could constitute a reduction in fitness. We interpret the phenotypes we see along those lines: network instability that results from critical-period perturbations is a manifestation of a sub-optimally tuned network, as is a reduction in larval crawling speed.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a thoughtful and well-executed study of critical period plasticity in the Drosophila larval motor circuit. The authors examined how transient heat, 32 {degree sign}C, during the embryonic stage, altered network properties, showing that premotor interneurons A27h increase excitatory drive onto motoneurons, which respond with a reduction in excitability. At the NMJ, synaptic terminals expand and GluRIIA distribution shifts, yet synaptic transmission remains largely unaffected. Despite these local compensations, the treated larvae display slower crawling and prolonged recovery from seizures, indicating that the network is functionally compromised.

      Strengths:

      (1) One of the major strengths of this study is the elegant dissection of a defined circuit, tracking changes from premotor interneurons through motoneurons to the NMJ. The multimodal approach provides a comprehensive view of how connected elements respond to CP perturbations.

      (2) An interesting finding is that NMJ morphology changes dramatically without corresponding deficits in synaptic transmission, challenging the common assumption that larger boutons necessarily indicate stronger synapses.

      (3) Another intriguing result is that even with two layers of homeostatic compensation, locomotor behavior is still impaired, highlighting the limits of compensation and underscoring the critical role of CP timing.

      (4) Beyond these scientific insights, the study benefits from a well-defined, tractable system and simple experimental manipulations, which together make the results highly interpretable and reproducible.

      Weaknesses:

      There are a few areas where the manuscript could be strengthened.

      (1) Although A27h premotor neurons are well characterized, the claim that they are the causal driver of downstream changes would be strengthened by additional experiments or a clearer discussion of the temporal hierarchy.

      We have tried to clarify the model of the temporal hierarchy (new Figure 9). This is a model and as such will hopefully help us collectively to think about this system and underlying processes, while also inviting this perspective to be challenged. The model we propose is compatible with our observations, namely that the premotor circuitry might change in response to a critical period heat stress (e.g. increasing their synaptic drive onto motoneurons), followed by homeostatic adjustment by the postsynaptic motoneurons (e.g. by reduction of their excitability), thus serving to maintain overall normal motoneuron firing patterns (see Figure 6).

      However, synaptic communication is commonly regulated in both antero- and retrograde directions. Therefore, while compatible with the observations we have made, bi-directional information flow could also be instructive during the CNS critical period.

      (2) While 32 {degree sign}C heat stress is presented as ecologically relevant, it produces maladaptive behavioral outcomes, raising questions about the ecological and mechanistic interpretation of the model. In particular, most experiments, with the exception of Figure 1, used prolonged (24h) heat treatments, which could introduce developmental effects beyond the CP itself. Comparing shorter and longer heat exposures would help clarify the specificity of the CP response.

      We agree. For a detailed response on this point, please response to Reviewer #1 above.

      (3) While there are schematics for experimental procedures, a circuit diagram tracing information flow and indicating where structural and functional changes occur would help readers better understand the findings.

      We have created Figure 9 as a working model.

      (4) Finally, the main paradox of the study, that robust homeostatic compensations occur yet behavior remains impaired, could be explored in more depth in the Discussion.

      We have tried to address this in the discussion.

      Reviewer #3 (Public review):

      Summary:

      During development, neural circuits undergo brief windows of heightened neuronal plasticity (e.g., critical periods) that are thought to set the lifelong functional properties of underlying circuits. These authors, in addition to others within the Drosophila community, previously characterized a critical period in late fly embryonic development, during which alterations to neuronal activity impact late-stage larval crawling behavior. In the current study, the authors use an ethologically-relevant activation paradigm (increased temperature) to boost motor activity during embryogenesis, followed by a series of electrophysiology and imaging-based experiments to explore how 3 distinct levels of the circuit remodel in response to increases in embryonic motor activity. Specifically, they find that each level of the circuit responds differently, with increased excitatory drive from excitatory pre-motor neurons, reduced excitability in motor neurons, and no physiological changes at the NMJ despite dramatic morphological differences. Together, these data suggest that early life experience in the motor neuron drives compensatory changes at each level of the circuit to stabilize overall network output.

      Strengths:

      The study was well-written, and the data presented were clear and an important contribution to the field.

      Weaknesses:

      The sample sizes and what they referred to throughout the distinct studies were unclear. In the legends, the authors should clearly state for each experiment N=X, and if N refers to an NMJ, for example, instead of an individual animal, they should state N=X NMJs per N=X animals. This will help readers better understand the statistical impact of the study.

      This is a good point. For the majority, each data point is derived from a unique specimen, unless explicitly stated otherwise, for NMJ size on muscle DA1 (Figure 3) and for larval crawling data, where each larva was measured up to three time, once per unique 5-minute crawling interval.

      Recommendations for the authors:

      Reviewing Editor Comments:

      In addition to revising the text and making interpretive changes as suggested by the reviewers, we invite you to consider the following:

      (1) Either rephrase the conclusions on the role of the 2h CP and discuss the effects of temperature during embryonic development. Alternatively, to validate the idea of a longer CP window, directly compare the results of a few key experiments using the 2h and 24h heat treatment.

      We have addressed this within the text, as suggested. In the text and figure legends, we have made clear distinctions between exposure to 32˚C heat stress during most of embryogenesis vs a specific developmental window of a few hours. In figures, we have provided diagrams that graphically illustrate the period of heat stress exposure.

      Our ability to experimentally test differences between precise vs broader heat stress periods during embryonic development have been constrained due to the departure of scientists, who were able to carry out electrophysiological recordings (as also explained below). As the next best alternative, we focus on imaging and behavioural analyses. These demonstrated that the developing body wall muscles are sensitive to heat stress during an earlier phase, from 13-16 hours after egg laying, when the body wall muscles become electrically active. It precedes the critical period of the central locomotor circuitry (17-19 hours after egg laying), when neurons in the CNS become electrically and synaptically active (16 hours after egg laying). We think this an exciting additional insight, demonstrating sequential critical periods as different parts of the locomotor network become active: first the body wall muscles, followed by the central circuitry.

      (2) Clarifying the homeostatic responses and shedding light on how they engage with the maladaptive changes described would benefit the study. Furthermore, adding more information about the anatomical and structural changes and how they relate to the intrinsic and synaptic changes would also benefit the study.

      We have tried to address this within the text and with a summary diagram (Figure 8), as suggested.

      Reviewer #1 (Recommendations for the authors):

      (1) It remains unclear whether the authors want to conclude that reduced network stability is not due to changes at the motoneuron level, but rather at the premotor level. Although this idea is mentioned in the results and discussion, it does not appear in the abstract or introduction, which leaves the different findings disconnected. Clarifying and highlighting this conclusion throughout the manuscript would strengthen the narrative.

      We have added additional experiments and changed the manuscript to address this point. These showed that there are distinct phases of embryonic development during which heat stress causes changes to NMJ structure vs to larval behaviour (seizure recovery times/network stability and crawling speed) - additional data in suppl. Fig. 2 and Fig. 7). In the Drosophila embryo, the body wall muscles develop and acquire their electrical properties before central neurons do, and these phases correlate with sensitivity to heat stress.

      (2) In the results section related to Figure 1, the logical link between the CP protocol and the functional assessment of the locomotor network at different temperatures is not sufficiently explained. It is not clear what this assessment is meant to test or demonstrate. A more explicit statement of the rationale and correction of what seems to be a typographical error in the final sentence of the paragraph would help to clarify the authors' intent.

      We have tried to rectify this by changes in the manuscript and to Figure 1, to make the sequence of panels more intuitive.

      (3) In the second results section, the experimental strategy shifts from using a short 2-hour heat shock to a 24-hour manipulation. The reasoning - that short manipulations in different windows yield no phenotype - is understandable, but a 24-hour perturbation may have broader consequences beyond the CP, simply by virtue of its longer duration. Moreover, 24h is roughly the duration of embryonic development at 25C. When at 32C, embryos should develop faster; therefore, is the 24h heat shock extending to L1?

      Yes, the 24 hour heat stress extends into the first few hours of the L1 larval stage.

      In order to validate the use of a longer window, the author should show how it affects the developmental time. Moreover, one should test that a few important observations remain the same with 2h and 24h heat perturbation. Alternatively, one cannot conclude that the phenotype is due to the rather narrow previously defined CP rather than to other effects associated with the overall embryonic developmental time and coordination. This would not make the results less interesting, but it would be important to assess whether the effects can be solely attributed to the 2h CP.

      We have compared the impact of heat stress experience during the majority of embryogenesis, including the CP that had been defined for the central locomotor network (17-19 hours after egg laying) with shorter heat stress manipulations during consecutive phases of embryogenesis until larval hatching. As outlined above in response to point (1) by the Reviewing Editor, reduced stability of the central network and associated reduction in larval crawling occurs when heat stress is experienced during the CP of the central locomotor network (17-19 hours after egg laying). Prolonged heat stress experience for 24 hours leads to indistinguishable outcomes, as long as this 2-hour CP window is included (see Fig. 1 and Fig. 8).

      However, this suggestion by Reviewer #1 led us to identify a second CP for the body wall muscles (see Fig. 3). NMJ overgrowth and changes to the postsynaptic glutamate receptor composition result from earlier heat stress experiences, and those are comparable to the effects caused by 24-hour heat stress exposure when this earlier muscle CP is included.

      Therefore, NMJ development is affected by consecutive CPs, an earlier one linked to body wall muscle development, followed by a later one that impacts the presynaptic motoneurons and their upstream circuitry. Nevertheless, the larval NMJ and behavioural phenotypes that we have identified appear to result from sensitivity to heat stress during these respective CP windows, with no clear evidence of cumulative effects on these phenotypes resulting from longer heat stress exposure during embryogenesis.

      (4) In session 3, the authors note that GluRIIA reductions were most pronounced in proximal regions of the NMJ. However, this is not explicitly quantified in the figures or methods. Including such quantification, or clarifying where it can be found, would make this observation more convincing.

      We have analysed anti-GluRIIA signal intensities in proximal vs distal boutons, comparing different ROI selection processes (e.g. thresholding to a full NMJ/anti-HRP mask and to an anti-GluRIIB mask, which is more selective to postsynaptic sites). Analysis of multiple data sets did not show statistical significance, but instead confirmed that comparable reductions in anti-GluRIIA signal manifest in both proximal and distal boutons, following an embryonic 32C heat stress, relative to controls. We have therefore removed relevant speculative statements.

      (5) In session 4, the authors conclude that motoneurons undergo a decrease in excitability to adjust to greater premotor drive. Is there anatomical evidence for this, such as an increase in input synapses?

      We previously quantified change in excitatory presynaptic synaptic contact number onto aCC motoneuron dendrites in third instar larvae following an embryonic pharmacological activity manipulation: overexcitation of the developing network following introduction of PTX via feeding to gravid females*. No significant structural changes were seen. Although this is a different manipulation of the developing network, all our data to date suggest that heat stress manipulations during the embryonic critical period signal via the same pathways, at least in part due to temperature increases leading to activity increases. Because such a quantification is technically challenging and extremely time-consuming due to the low level of marking individual motoneurons, we did not think it informative or in scope for this project.

      *See Figure 5 in this publication: Hunter I, Coulson B, Pettini T, Davies JJ, Parkin J, Landgraf M, Baines RA. Balance of activity during a critical period tunes a developing network. Elife. 2024 Jan 9;12:RP91599. doi: 10.7554/eLife.91599. PMID: 38193543; PMCID: PMC10945558.

      The interpretation of the optogenetic experiments would also benefit from clarification. If motoneurons are less excitable yet receive more drive, one might expect no net change, rather than the differences observed. Alternatively, could the excitability of the premotor neuron itself have changed, either intrinsically or in relation to Chronos expression? Measuring premotor activity directly during optogenetic activation could help to resolve this ambiguity.

      These are good suggestions. Yes, we think that motoneuron excitability has changed as a result of heat stress - see paper submitted in parallel and published since: Sobrido-Cameán et al., 2025, PLoS Biology. Unfortunately, the team members, who could have carried out this type of analysis had moved on by submission of the manuscript. Therefore, we were no able to experimentally pursue these questions further.

      (6) In session 6, the authors report slower propagation of premotor activity waves after CP heat stress, but the logic of the experiment is not sufficiently explained. How does this finding relate to the enhanced premotor drive described earlier? Only timing is quantified; information about amplitude and wave dynamics would strengthen the interpretation. These results could also be discussed in relation to Figure 1B, where acute heat stress increased motoneuron activity. One interesting possibility could be that CP manipulations might adaptively prepare the larva to function at different temperatures. Experiments testing wave propagation at 32 {degree sign}C (Figure 6) or, conversely, motoneuron activity after CP manipulations (Figure 1B paradigm) would provide valuable evidence for such an adaptive role.

      As per above, unfortunately, the team member who could have carried out this type of analysis had moved on by submission of the manuscript. However, we have tried to address the question of whether there is an adaptive element to the adjustments that result from embryonic heat stress experience. Specifically, we carried out behavioural tests on how larvae respond with changes in crawling speed to acute changes in ambient temperature (new Fig. 7).

      We interpret our findings as follows: that heat stress during embryonic development leads to sub-optimal outcomes with regard to network stability as well as default and maximum crawling speed. The precise causes for this will be difficult to unpick. Behaviourally, when challenging larvae with an acute change in ambient temperature, we saw that slow crawling larvae do respond comparatively normally to a relative increase in ambient temperature by speeding up (in effect an escape response). This demonstrates that embryonic heat stress causes a change in the default crawling speed, while principally maintaining behavioural responses to changes in ambient temperature. It appears that animals that had experienced heat stress during embryonic development, by adjusting their default speed downward, maintain a dynamic response range into a higher temperature range than controls (35C vs 29C, respectively). Potentially, this could be an adaptive outcome to living at higher temperatures, though such an interpretation would require a body of work. 

      Nevertheless, every aspect we have assayed suggests that heat stress experience during the CP leads to sub-optimal outcomes: of network instability, slower default and slower maximal crawling speeds.

      Minor points:

      (1) Abstract: "has suboptimal outcomes;" should be corrected to "has suboptimal outcomes,".

      Corrected

      (2) Abstract: "we find that transient embryonic..." would improve readability with a capitalized "We".

      We are unsure about the sentence this refers to. If this sentence, then we suggest that this could remain as was.

      "Within the central nervous system, we find transient embryonic CP perturbation leads to increased synaptic drive from premotor interneurons to motoneurons..."

      The intention here is to differentiate between changes within the CNS vs at the NMJ.

      (3) Abstract: The sentence "Present the larva ... as an experimental model system..." overstates novelty, as the system has already been established in prior work. Instead, this study could highlight temperature manipulation as an ecologically relevant way to probe CPs.

      We have adopted this suggestion.

      Reviewer #2 (Recommendations for the authors):

      I would recommend:

      (1) Perform additional experimental support and dissuasions for the causal role of the premotor neuron and the network disability.

      Unfortunately, it has not been possible to carry out additional e-phys experimental work due to key people having moved on and now unable to carry out such experiments, and no replacements in sight to do so. Instead, we have focused on other work that we could do, namely to test the effect of different heat stress windows during embryonic development on GluRIIA vs GluRIIB expression at postsynaptic sites. This shows that indeed the effect seen following a 24-hour heat stress is replicated by a much shorter window of heat stress. For NMJs GluRII composition the critical period is different from the critical period of the CNS. This replicates the different developmental timings of maturation: the body wall muscles express ion channels and attain their electrical properties several hours before central neurons. We have provided these additional data as a new supplementary figure to Fig. 2. We have changed the main text to note the caveat of longer heat stress manipulations potentially leading to additional or more exacerbated phenotypes.

      (2) I would suggest that the authors expand the discussion on why two layers of homeostatic adjustment fail to preserve behavior. Is this simply a limit of plasticity?

      This is a difficult aspect to address well, beyond the purely speculative and potentially confusing. We would like to suggest that to do so requires a basic understanding on what pressures neurons/networks respond to (heat-caused over-activation and/or metabolic); and from that perspective to gauge how they adjust to those pressures, i.e. what the adjustments are trying to "achieve". This in turn should inform on whether such adjustments are homeostatic or anti-homeostatic in nature, and whether there are limits beyond which we consider a system "breaking".

      (3) 32C heat was described as "ecologically relevant". However, it produces maladaptive outcomes. The author should consider reframing it as a stressor that reveals CP sensitivity, rather than an adaptive signal.

      This is a good suggestion, and we have implemented changes accordingly.

      (4) It is important to distinguish the effects of transient (2h) vs prolonged heat exposure to confirm the manipulation targeted CP specifically.

      We have tried to make these distinctions clearer within text and figure legends. As per response to (1) above, we generated and analysed additional data to show that changes in GluRIIA are induced during a defined shorter developmental time window, not exacerbated by prolonged heat stress exposure during embryogenesis.

      (5) I would definitely recommend adding a diagram tracking information flow and showing where structural and functional changes occur.

      This is a helpful suggestion, which we have tried to implement with a new figure (Fig. 8).

      Overall, this is a strong and well-written paper that produced some unexpected results, and added a solid model circuit to study CP plasticity at the circuit level.

      Reviewer #3 (Recommendations for the authors):

      I identified one typo: "activity manipulations during the embryonic CP are artificial, We asked to what extent" The "W" of "We" should be lower case.

      Now corrected.

    1. eLife Assessment

      This paper represents a valuable contribution to our understanding of how local field potential (LFP) oscillations and beta band coordination between the hippocampus and prefrontal cortex of rats may relate to learning. Through a set of appropriate analyses, including oscillation-oscillation coupling, oscillation-spike modulation, and behavioral dependant measurements, the study presents convincing evidence for uncoupled beta activity between the two regions. This work will be of interest to researchers working on interaction of brain regions during spatial learning in rodents.

    2. Reviewer #1 (Public review):

      Wang, Zhou et al. investigated coordination between prefrontal cortex (PFC), and hippocampus (Hp), during reward delivery via analyzing beta oscillation. Beta oscillations are associated with various cognitive functions but their role in coordinating brain networks during learning is still not thoroughly studied. Authors focused on the changes in power, peak frequencies and coherence of beta oscillations in two regions when rats learn a spatial task thru days. Contradicting with authors hypothesis, beta oscillations in those two regions during reward delivery were not coupled in spectral or temporal aspects. They were, however, able to show reverse changes in beta oscillations in PFC and Hp as the animal's performance got better. Authors were also able to show a small subset of cell population in PFC that are modulated by both beta oscillations in PFC and sharp wave ripples in Hp. A similarly modulated cell population was not observed in Hp. These results are valuable in pointing out distinct periods during a spatial task when two regions modulate their activity independent from each other.

      Authors made a detailed analysis of the data to support their conclusions. Few more points of discussion would clarify the results of the paper.

      (1) One of the big conclusions of the paper is how the beta burst power is changing after learning the task (Figure 3). Authors have also showed in Figure 6-1, how the SWR power and rate are changing thru the training days. Did they observe a change in coordination of Beta bursts and SWR between the days, which would also reflect how experience changes the coordination?

      (2) Authors have shown in detail the opposite relationship between Beta phase locking and SWR modulation in Hippocampus in Figure 7I. This might require a different analysis, but is it possible to make a discussion on predicting a cell firing in a SWR after it fires in a beta burst.

      Other than these two points, authors have addressed previous comments and made a convincing analysis of their data.

    3. Reviewer #2 (Public review):

      Using electrophysiological recordings in freely moving rats during a spatial navigation task, this study investigated the role of beta oscillations in the hippocampal-prefrontal network. Through a set of appropriate analyses-including oscillation-oscillation coupling, oscillation-spike modulation, and behavioral dependant measurements-the study presents convincing evidence for uncoupled beta activity between the two regions. These findings offer important insights into the network mechanisms of spatial navigation and may have significant implications for related neurological disorders.

      Comments on revised version.

      The authors have carried out additional analyses and made corresponding revisions to the manuscript in response to the earlier review comments, which have made the conclusions more convincing. However, please note that the last question about coexistence of beta and SWR has not been fully answered. Please supplement your response to address this point completely.

    4. Reviewer #3 (Public review):

      This paper explored the role of beta rhythms in the context of spatial learning and mPFC-hippocampal dynamics. The authors characterized mPFC and hippocampal beta oscillations, examining how their coordination and their spectral profiles related to learning and prefrontal neuronal firing. Rats performed two tasks, a Y-maze and F-maze, with the F-maze task being more cognitively demanding. Across learning, prefrontal beta oscillation power increased while beta frequency decreased. In contrast, hippocampal beta power and beta frequency decreased. This was particularly for the well-performed and well-learned Y-maze paradigm. The authors identified the timing of beta oscillations, revealing an interesting shift in beta burst timing relative to reward entry as learning progressed. They also discovered an interesting population of prefrontal neurons that were tuned to both prefrontal beta and hippocampal sharp-wave ripple events, revealing a spectrum of SWR-excited and SWR-inhibited neurons that were differentially phase locked to prefrontal beta rhythms.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Wang, Zhou et al. investigated coordination between the prefrontal cortex (PFC) and the hippocampus (Hp), during reward delivery, by analyzing beta oscillations. Beta oscillations are associated with various cognitive functions, but their role in coordinating brain networks during learning is still not thoroughly understood. The authors focused on the changes in power, peak frequencies, and coherence of beta oscillations in two regions when rats learn a spatial task over days. Inconsistent with the authors' hypothesis, beta oscillations in those two regions during reward delivery were not coupled in spectral or temporal aspects. They were, however, able to show reverse changes in beta oscillations in PFC and Hp as the animal's performance got better. The authors were also able to show a small subset of cell populations in PFC that are modulated by both beta oscillations in PFC and sharp wave ripples in Hp. A similarly modulated cell population was not observed in Hp. These results are valuable in pointing out distinct periods during a spatial task when two regions modulate their activity independently from each other.

      The authors included a detailed analysis of the data to support their conclusions. However, some clarifications would help their presentation, as well as help readers to have a clear understanding.

      (1) The crucial time point of the analysis is the goal entry. However, it needs a better explanation in the methods or in figures of what a goal entry in their behavioral task means.

      We appreciate Reviewer 1 pointing out this shortcoming and will clarify the description in the revised manuscript. Each goal is located at the end of the arm, and is equipped with a reward delivery unit. The unit has an infrared sensor. The rat breaks the infrared beam when it enters the goal. Figures 1 and 2 have been updated to clearly indicate the time of goal entry. The main text and methods have been updated with the explanation.

      (2) Regarding Figure 2, the authors have mentioned in the methods that PFC tetrodes have targeted both hemispheres. It might be trivial, but a supplementary graph or a paragraph about differences or similarities between contralateral and ipsilateral tetrodes to Hp might help readers.

      We appreciate this suggestion, which has led to an interesting finding. The coherence and burst coordination were similar for ipsi- and contralateral PFC and hippocampus. Interestingly, we found PFC beta activity was more coherent within each PFC hemisphere compared with across hemispheres. This was observed for coherence and burst time. This suggests there is hemispheric localization of beta oscillations. These results are shown in Fig. 2-1.

      (3) The authors have looked at changes in burst properties over days of training. For the coincidence of beta bursts between PFC and Hp, is there a change in the coincidence of bursts depending on the day or performance of the animal?

      This is now reported in Fig. 3-3. After quantifying the proportion of independent and coincident bursts as function of experiment day or performance, we found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (4) Regarding the changes in performance through days as well as variance of the beta burst frequency variance (Figures 3C and 4C); was there a change in the number of the beta bursts as animals learn the task, which might affect variance indirectly?

      The difference in the burst count across days did not explain the results. We performed a permutation test (Fig. 4-4), where we randomly shuffled the day identity to control for the count difference across days. The change in variance remains significant.

      (5) In the behavioral task, within a session, animals needed to alternate between two wells, but the central arm (1) was in the same location. Did the authors alternate the location of well number 1 between days to different arms? It is possible that having well number 1 in the same location through days might have an effect on beta bursts, as they would get more rewards in well number 1?

      The central arm remained the same across days since we needed the animals to learn the alternation task. In our experience, the animal needs a few days to learn the alternation rule when we switch the central arm location. For this experiment, we were interested in the initial learning process, and we kept the central arm constant. Switching the central arm location is a great suggestion for a follow-up experiment where we can understand the effects of reward contingency change on beta bursts.

      (6) The animals did not increase their performance in the F maze as much as they increased it in the Y maze. It would be more helpful to see a comparison between mazes in Figure 5 in terms of beta burst timing. It seems like in Y maze, unrewarded trials have earlier beta bursts in Y maze compared to F maze. Also, is there a difference in beta burst frequencies of rewarded and unrewarded trials?

      We performed the analysis and found burst timing was similar between the two mazes (Fig. 4-2). Bursts on rewarded trials occurred later than those on unrewarded trials (Fig. 4). Interestingly, PFC bursts on rewarded trials were lower in frequency compared with unrewarded trials. CA1 bursts during rewarded and unrewarded trials had similar frequencies (Fig. 4-3).

      (7) For individual cell analysis, the authors recorded from Hp and the behavioral task involved spatial learning. It would be helpful to readers if authors mention about place field properties of the cells they have recorded from. It is known that reward cells firing near reward locations have a higher rate to participate in a sharp wave ripple. Factoring in the place field properties of the cells into the analysis might give a clearer picture of the lack of modulation of HP cells by beta and sharp wave ripples.

      As recommended, we quantified the mean speed, mean distance to goal locations, and spatial information for CA1 cells (Fig. 7 J-L). We found SWR reactivated CA1 cells had higher speed and were spiking further away from the goals compared with non-reactivated CA1 cells. This is consistent with prior work that shows SWR-associated reactivation in dorsal CA1 can correspond to trajectories taken as the animal moves towards goals. In intermediate CA1, the content of reactivations is biased toward place representations closer to goals (Jin et al., 2024). CA1 cells with or without phase locking to beta oscillations had similar spatial firing properties.

      Reviewer #1 (Recommendations for the authors):

      (1) Please make a figure representing what the goal entry means in Figure 1.

      We have updated Fig. 1 to clearly show the definition of goal entry. We also edited the main text and methods to better explain the definition of goal entry.

      (2) For Figure 1-1, please either change the contrast of the pictures, or define the lesioned areas, as it is a bit difficult to see the lesioned parts, especially in PFC.

      We have increased the contrast for Fig. 1-1 for the histology to better show the lesions.

      (3) For Figure 1-2, is it possible to show a beta burst from Hp?

      Yes, we provided two examples of beta bursts from the hippocampus alongside burst examples from PFC in Fig. 1-2.

      (4) Is it possible to make a supplementary table showing the number of tetrodes recorded from each animal per day, plus the number of isolated single cells?

      Yes, we have included tetrode counts in Table 1-2 and cell counts in Tables 5-2 to 5-4.

      Reviewer #2 (Public review):

      (1) When presenting the power spectra for the representative example (Figure 1), it would be appropriate to display a broader frequency band-including delta, theta, and gamma (up to ~100 Hz), rather than only the beta band.

      We agree the extended frequency range provides a better overview of the spectral characteristics during the goal period. We have now included example spectrograms up to 100 Hz to show the spectral content for a wider range of frequencies (Fig. 1-2). Further, we have included additional analyses to compare the spectral characteristics between periods when the animal was moving on the maze or immobile at the goal, for frequencies up to 100 Hz (Fig. 1C-H, Fig. 1-3 and 1-4). We used both Welch’s periodogram (Fig. 1-3) and continuous wavelet transform (Fig. 1-4) to demonstrate our findings on beta oscillations in both regions are robust and consistent.

      What was the rat's locomotor state (e.g., running speed) after entering the reward location, during which the LFPs were recorded?

      Because goal entry is defined as the time the animals break the infrared beam at the goal (response to Reviewer 1), the rat would have come to a stop. We have added the time-aligned speed profile to the spectra and raw data examples in the manuscript (Fig. 1B, Fig. 1-4, Fig. 2A, and Fig. 6A). In addition, we added the quantification of the animal’s speed at the time of beta bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1D).

      If the rats stopped at the goal but still consumed the reward (i.e., exhibited very low running speed), theta rhythms might still occasionally occur, and sharp-wave ripples (SWRs) could be observed during rest.

      We typically find low theta power in the hippocampus after the animal reaches the goal location and as it consumes reward. Reviewer 2 is correct about occasional theta power at the goal. To compare differences in LFP characteristics between maze running and goal locations, we added additional analyses in Fig. 1-3 and 1-4. We did find SWRs during goal periods (Fig. 6) and we quantified SWR properties in an additional analysis in Fig. 6-1.

      Do beta bursts also occur during navigation prior to goal entry? It would be beneficial to display these rhythmic activities continuously across both the navigation and goal entry phases.

      We did not find consistent beta bursts in PFC or CA1 on approach to goal entry. We generated an additional goal entry-aligned spectrogram, and quantification (Fig. 1-4) to show that beta oscillations in both regions increased after goal entry. This was also supported by the Welch’s periodogram method (Fig. 1C-H, Fig. 1-3). Beta oscillations in the hippocampus during locomotion or exploration have been reported (Ahmed & Mehta, 2012; Berke et al., 2008; França et al., 2014; França et al., 2021; Iwasaki et al., 2021; Lansink et al., 2016; Rangel et al., 2015).

      Additionally, given that the hippocampal theta rhythm is typically around 7-8 Hz, while a peak at approximately 15-16 Hz is visible in the power spectra in Figure 1C, the authors should clarify whether the 22 Hz beta activity represents a genuine oscillation rather than a harmonic of the theta rhythm.

      We performed further spectral analysis comparing times when the animal is moving on the maze with times when the animal is immobile at the goal (Fig. 1-3 and 1-4). The results point to the beta frequency oscillations in both regions are unlikely to be harmonics of theta. The beta frequency bands in the spectrogram are independent of the theta band. We were initially concerned about the possibility that the 22 Hz power in CA1 may be a harmonic rather than a standalone oscillation band. If these are harmonics of theta, we should expect to find coincident theta at the time of bursts in the beta frequency. In Fig. 1B, Fig. 1-5, and Fig. 2A, we show examples of the raw LFP traces from CA1. Here, the detected bursts are not accompanied by visible theta-frequency activity. For PFC, we do not always see persistent theta-frequency oscillations like CA1. In PFC, we found beta bursts were frequent and visually identifiable when examining the LFP. We provided examples of the PFC LFP (Fig. 1B, Fig. 1-5, and Fig. 2A). In these cases, we see clear beta frequency oscillations lasting several cycles and these are not accompanied by any visible oscillations in the theta frequency in the LFP trace.

      (2) The authors claim that beta activity is independent between CA1 and PFC, based on the low coherence between these regions. However, it is challenging to discern beta-specific coherence in CA1; instead, coherence appears elevated across a broader frequency band (Figure 2 and Figure 2-1D). An alternative explanation could be that the uncoupled beta between CA1 and PFC results from low local beta coherence within CA1 itself.

      This is a legitimate concern, and we used three methods to characterize coherence and coordination between the two regions. First, we calculated coherence for tetrode pairs for times when the animal was at goals (Fig. 2B), which provides a general estimation of coherence across frequencies but lack any temporal resolution. Second, we calculated burst-aligned coherence (Fig. 2-2), which provides temporal resolution relative to the burst, but the multi-taper method is constrained by the time-frequency resolution trade-off. Third, we quantified the timing between the burst peaks (Fig. 2D), which described the timing differences but the peaks for the bursts may not be symmetric. Each method has its own caveats, but we drew our conclusion from the combination of results from these three analyses, which pointed to similar conclusions.

      Reviewer 2 is correct in pointing out the uniformly high coherence within CA1 across the frequency range we examined. When we inspected the raw LFP across multiple tetrodes in CA1, they were similar to each other (Fig. 2A). This likely reflects the uniformity in the LFP across recording sites in CA1, which is what we saw with coherence values across the frequency range (Fig. 2B). We found that CA1 coherence between tetrode pairs within CA1 was statistically higher than tetrode pairs in PFC across the frequency range (Fig. 2B and C), thus our results are unlikely to be explained by low beta coherence within CA1 itself. The burst-aligned coherence using a multi-taper method also supports this. The coherence values within CA1 at the time of CA1 bursts were ~0.8-0.9.

      (3) In Figure 2-1E-F, visual inspection of the box plots reveals minimal differences between PFC-Ind and PFC-Coin/CA1-Coin conditions, despite reported statistical significance. It may be necessary to verify whether the significance arises from a large sample size.

      We will include the sample sizes in Table 2-1. We repeated the analyses based on average values per day for each animal (Fig. 2-2 E-F). The pattern of significance remains consistent.

      (4) In Figure 3 and Figure 4, although differences in power and frequency appear to change significantly across days, these changes are not easily discernible by visual inspection. It is worth considering whether these variations are related to increased task familiarity over days, potentially accompanied by higher running speeds.

      We agree with Reviewer 2 that familiarity increases across days, and the animal is likely running faster. The analysis for Fig. 3 (previously Fig. 3 and 4) includes only data from periods when the animal was at the goal and was not moving. We added a supplemental figure (Fig. 3-1), which shows that the speed of the animal was below 0.5 cm/s at the time of the analyzed bursts. We used linear mixed-effects models to quantify the relationship between power, frequency and day or behavioral quintile, which accounts for repeated measurements across animals.

      (5) The stronger spiking modulation by local beta oscillations shown in Figure 6 could also be interpreted in the context of uncoupled beta between CA1 and PFC. In this analysis, only spikes occurring during beta bursts should be included, rather than all spikes within a trial. The authors should verify the dataset used and consider including a representative example illustrating beta modulation of single-unit spiking.

      We agree with Reviewer 2 that the stronger modulation to local beta is another piece of evidence indicating uncoupled beta between the two regions. We appreciate this suggestion and have revised Fig. 5 (previously Fig. 6) to include examples illustrating beta modulation for single units. These are spike-phase raster and histograms. We want to clarify that in the revised manuscript, the spikes were only from periods when the animal was at the goal location (5 s after entry) and did not include the running period between goals. Although beta power fluctuates in bursts, our data show there are ongoing beta oscillations throughout this period (Fig. 1-3 and Fig. 1-4), which prompted us to examine the entire period.

      (6) As observed in Figure 7D, CA1 beta bursts continue to occur even after 2.5 seconds following goal entry, when SWRs begin to emerge. Do these oscillations alternate over time, or do they coexist with some form of cross-frequency coupling?

      This is a very helpful suggestion and led to some interesting findings. We performed two additional analyses: 1) burst/SWR cross-correlation on a shorter timescale and 2) SWR-aligned spectrogram for the beta frequency range. PFC beta burst timing and power appear to be anti-correlated with SWRs detected in CA1 (Fig. 6G and K). In contrast CA1 beta bursts were more likely to occur with SWRs (Fig. 6H and L). These results suggest there is temporal coordination between ongoing beta oscillations in PFC and SWRs in the hippocampus during waking. Cortical beta oscillations are reduced during hippocampal SWRs, perhaps to support the transient switch in global cortical states that accompanies SWRs.

      To examine potential cross-frequency coupling between SWRs and beta oscillations, we computed the mean SWR band power (150-250).

      Reviewer #3 (Public review):

      Summary:

      This paper explored the role of beta rhythms in the context of spatial learning and mPFC-hippocampal dynamics. The authors characterized mPFC and hippocampal beta oscillations, examining how their coordination and their spectral profiles related to learning and prefrontal neuronal firing. Rats performed two tasks, a Y-maze and an F-maze, with the F-maze task being more cognitively demanding. Across learning, prefrontal beta oscillation power increased while beta frequency decreased. In contrast, hippocampal beta power and beta frequency decreased. This was particularly the case for the well-performed and well-learned Y-maze paradigm. The authors identified the timing of beta oscillations, revealing an interesting shift in beta burst timing relative to reward entry as learning progressed. They also discovered an interesting population of prefrontal neurons that were tuned to both prefrontal beta and hippocampal sharp-wave ripple events, revealing a spectrum of SWR-excited and SWR-inhibited neurons that were differentially phase locked to prefrontal beta rhythms.

      In sum, the authors set out to examine how beta rhythms and their coordination were related to learning and goal occupancy. The authors identified a set of learning and goal-related correlates at the level of LFP and spike-LFP interactions, but did not report on spike-behavioral correlates.

      Strengths:

      Pairing dual recordings of medial prefrontal cortex (mPFC) and CA1 with learning of spatial memory tasks is a strength of this paper. The authors also discovered an interesting population of prefrontal neurons modulated by both beta and CA1 sharpwave ripple (SWR) events, showing a relationship between SWR-excited and SWR-inhibited neurons and beta oscillation phase.

      Weaknesses:

      Moreover, there is little detail provided about sample sizes and how data sampling is being performed (e.g., rats, sessions, or trials), raising generalizability concerns.

      We appreciate Reviewer 3’s thoughtful suggestions for making our claims convincing. We have included information about sample sizes in the revised manuscript.

      The authors report on a task where rats were performing sub-optimally (F-maze), weakening claims.

      Our experiment was designed to create a scenario in which one task was learned (Y-maze) and another was not (F-maze). This contrast allows us to determine differences in neural correlates of learning versus familiarity. The design produced a learned and not learned task with similar levels of familiarity over 5 days, within the same animal.

      Likewise, it is questionable as to whether mPFC and hippocampus are dually required to perform a no-delay Y-maze task at day 5, where rats are performing near 100%.

      We agree with Reviewer 3 that the mPFC and hippocampus may not be required when the animal reaches stable performance on day 5 (Deceuninck & Kloosterman, 2024). The data we collected spans the full range of early learning (day 1) to proficiency (day 5). We wanted to understand the dynamics of beta across these learning stages, which have not been reported previously.

      Recent studies suggest mPFC and hippocampus are likely to be needed, in some capacity, for learning continuous spatial alternation tasks on a range of maze geometries. Lesions, inactivation or waking activity perturbation of hippocampus or hippocampus and mPFC on the W maze alternation task slowed learning (Jadhav et al., 2012; Kim & Frank, 2009; Maharjan et al., 2018). More recently, optogenetic silencing of mPFC after sharp wave ripples on the Y-maze alternation affected performance when the center arm was switched (den Bakker et al., 2023). The Y and F-mazes in our study both share the continuous alternation rule, where the animal needed to avoid visiting a previously visited location on the outbound choice relative to the center, and always return to the center location.

      Further, the performance characteristics on the outbound and inbound components of our Y task are similar to the W task. We have analyzed the “inbound” and “outbound” performance of the animals on the Y-maze alternation task, and they are similar to the W maze alternation task. The “inbound” or reference location component is learned quickly whereas the “outbound”, alternation component is learned slowly.

      There would be little reason to suspect strong oscillatory coupling when task performance is poor and/or independent of mPFC-HPC communication (Jones and Wilson, 2005) potentially weakening conclusions about independent beta rhythms.

      Although many studies have examined the oscillatory coupling properties at the theta frequency between mPFC-HPC (Hyman et al., 2005; Jones & Wilson, 2005; Siapas et al., 2005), our understanding of beta frequency coordination between the two regions is less established, especially at goal locations. Our work suggests beta frequency coordination at goal locations does not share properties with those of theta frequency coupling between mPFC and HPC, which occurs primarily during movement on the maze. Our first novel finding is that beta oscillations occur at goal locations in mPFC and HPC; our second is that the beta frequency dynamics in these regions are surprisingly distinct. We are not aware of prior work describing these properties at goal locations in spatial navigation tasks, especially their temporal coordination.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusions from this article would be made much stronger if the authors (1) record from rats performing a task known to be dependent on mPFC-HPC communication (e.g. a spatial working memory task) or (2) record from the F-maze in well-trained rats (75-80% performance is common), or (3) show that Y-maze task performance is dependent on mPFC-hippocampal communication (see Maharjan et al., 2018, which used a W-track). It is possible that learning the Y-maze depends on mPFC-hippocampal communication, but that with asymptotic performance, this changes. This would put beta oscillation coupling findings into a nuanced perspective.

      We appreciate the recommendation. The objective of this current manuscript is to present previously unknown properties of beta oscillations in hippocampal-prefrontal cortical networks. We agree that further investigation is required to fully dissect the functional contribution of beta dynamics in these networks. We are in the process of doing that.

      The rule on the Y-maze in our experiment is identical to the continuous alternation rule on the W maze in Maharjan et al., 2018 and Kim and Frank, 2009 (Author response image 1), which showed the PFC and hippocampus are required for normal learning, respectively. The Y and W mazes share the same topology; there is one junction connecting three arms. The W maze has two 90-degree-angle turns which are not choice points. Thus, both maze tasks have one choice point and involve learning an alternation rule. Further the learning properties share similarities. For the Y-maze, the inbound portion (return to center) (Author response image 1) was quicker to learn than the outbound portion (alternation) (Author response image 1). The same pattern is observed in the W maze learning task (Maharjan et al., 2018, Fig. 3A and D, Kim and Frank, Fig. 4C). We agree an experiment is needed to show that Y-maze task learning is dependent on the function of hippocampal-prefrontal cortical networks. The shared topology, rule definition, and behavior profile suggest that the Y and W maze tasks engage similar learning processes.

      Author response image 1.

      W and Y maze alternation tasks share the same rule. Schematics illustrate a comparison of W and Y maze alternation task rules. The W maze has been inverted for visual comparison with the Y maze. Performance grouped by in- or outbound trials. Inbound trials originate from the side goals (2 or 3). Outbound trials originate from the center goal (1). Performance on the inbound trials was higher than that on the outbound trials, which is comparable to previously published alternation tasks on W-shaped mazes.

      (2) Typically, when analyzing LFP profiles, experimenters include running velocity/speed. It should be ruled out whether spectral changes are confounded in any way by speed or time spent in the goal zones.

      We appreciate this suggestion for better conveying our definition of goal period. This was raised by other reviewers. We have now included the speed profiles for the goal-entry-aligned spectrograms (Fig. 1B, 1-4, 2A, and 6A), as well as the speed quantification at the times of bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1 D). We strictly define the goal period after entry into the sensor on the reward delivery device at the end of the arms. This is to ensure the animal is immobile for the period to avoid confounds related to speed. We also ensured the time periods are comparable between the trials we analyzed.

      (3) The authors should describe in each analysis how data are sampled, and for extracellular electrophysiology experiments, cell counts from each rat reported. A table would be ideal. For example, it is unknown if entrainment analysis is performed on data collected from one rat or from all rats.

      We have now added the missing information (Table 1-2, 5-1 to 5-3). We performed analyses using data from all animals.

      (4) There was no profiling of mPFC neurons in terms of their behavioral correlates. The authors should strongly consider examining how individual neurons encode task variables (e.g., trial correctness, reward location...) and can do so using a generalized linear model. Adding an analysis of behavioral correlates could nicely tie into the beta-SWR analyses. For example, are SWR-beta rhythm-modulated neurons also behaviorally modulated?

      This is a very helpful suggestion. We have added analyses on the behavioral correlates of PFC and CA1 neurons. We quantified three metrics: 1) whether PFC and CA1 spiking activity can distinguish goal location based on firing rate or phase preference, 2) firing distance to goal and 3) spatial information (Fig. 5 and 7). We also performed the same analysis based on SWR and beta modulation status. We found PFC cells that were both SWR- and beta-modulated showed the strongest task firing relationship (Fig. 7).

      (5) Figure 2:

      Are these data analyzed from well-trained rats? What is your N (rats/sessions/trials/epochs)?

      These results in Fig. 2 are from all days. We added the breakdown in Table 2-1.

      (6) Figure 3:

      (a) The authors show that on the Y-maze, performance, beta oscillations power, and beta oscillation frequency change over days. However, for the F-maze, performance improves but appears to taper off at 60% and PFC beta frequencies do not change with learning. Do you have rats performing this task well above chance (e.g., 75-80%?), and if so, do beta oscillation frequencies in the mPFC gradually change?

      We did not observe rats performing above chance on the F-maze over 5 days. They do not appear to learn the alternation rule on this task. This is why we used the F-maze as the “non-learner” control. The F-maze task design for this study appears to be difficult to learn and would likely require a much longer training period for the performance to exceed chance. We agree an important follow-up question is whether the effects reported here are generally observed across learning in different tasks. This is a future direction we are actively pursuing.

      One existing data point may partially address the relevance of beta power change with learning. We analyzed the power as a function of performance quintile on the Y-maze (learned) or F-maze (not learned). The power changed with performance quintile on the Y-maze (learned) but not the F-maze (not learned) (Fig. 3D), pointing to the change in power being associated with learning status.

      (b) Does the proportion of beta bursts change with learning? What about the proportion of coherent events?

      We performed this analysis and the results are shown in Fig. 3-3. We found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (c) Is the reduction in beta oscillation frequency in the mPFC related to running behavior? This should be ruled out. Same for the hippocampus.

      The reduction is not due to running because we only included bursts when the animal was at the goal and immobile. We added a figure to show speed at the time of bursts (Fig. 3-1).

      (d) Why don't you also show coherence as a function of learning?

      This is a great suggestion. Coherence as a function of day or performance is now shown in Fig. 3-2. Overall, the trends were weak suggesting there was no strong change in coherence over days or as a function of learning. Although some of the linear mixed effects models were statistically significant, the marginal R<sup>2</sup> values (R<sup>2</sup><sub>m</sub>) were very low, indicating the effects of performance or day on coherence were small. Beta frequency coherence within each brain region either remained the same or slightly decreased across days (Fig. 3-2 D and F) or with performance (Fig. 3-2 J). For coherence across regions, we only found a weak but significant increase for the well-performed Y-maze task across days (Fig. 3-2 B).

      (e) The statistics in the caption are great, but maybe consider using a table and a supplemental figure

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (7) Figure 4

      (a) This figure is better suited as supplemental to Figure 3.

      We have combined Fig. 4 with Fig. 3.

      (b) Statistics might be better suited in a table and a supplemental figure.

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (8) Figure 5

      (a) The shift of beta rhythms being linked to learning is interesting, albeit confusing, given that error trials were accompanied by even more 'precise' beta rhythm timing. Is it possible that beta rhythms are accompanied by an error signal? Or maybe instead related to running behavior?

      We verified this observation was not due to speed since we only included periods when the animal was at the goal and immobile. We added a new supplemental figure (Fig. 4-1) to report this analysis.

      (b) Changes in beta rhythm timing in the F-maze could be entirely explained by the data in the Y-maze, given that F-maze performance was, in general, very low. To test if beta rhythm shifting is a consequence of learning, the authors should show performance on the F-maze without the Y-maze, and vice versa.

      As suggested, we also analyzed frequency variance data from each maze separately and found learning-related changes were more prominent on the Y-maze (learned) than on the F-maze (not learned) (Fig. 4-4 C). The F-maze data serves as a valuable control within the same animal for the learning-related effects we are reporting.

      We agree that the Y-maze data is consistent with performance-related changes to beta timing. The addition of the F-maze provides a within-animal comparison for a task in which the animal performed less well. We agree an additional control group would provide further evidence to support learning-related changes in beta timing. However, the F-maze will likely take much longer to learn, therefore the additional days of exposure on the task will be a different confound. Thus, we wanted to ensure that we compared a “learned” with a “not learned” task within the same animal, with a comparable exposure period.

      (9) Figure 6

      (a) How many cells are analyzed here? Are they all pyramidal neurons? The authors should add the number of cells analyzed from each rat. A table would suffice.

      We added tables (Table 5-2 to 5-4) to show this information. For our analysis, we included all units.

      (b) The authors should include statistics about recorded units. Peak-to-trough timing and interspike interval are commonly used. Even demonstration of action potentials from individual cells is valuable for proof of concept.

      We have added more information on the cell type classification (Table 5-2) and how we classified the cells (Methods) and also added example waveforms in Fig. 5 (previously Fig. 6).

      (c) Why did the authors stop comparing the F-maze and Y-maze for entrainment analysis?

      The data for each maze was comparable. Per the reviewer’s suggestion, we added the entrainment analysis for each maze separately (Fig. 5-1).

      (10) Figure 7

      (a) The authors should minimally include instantaneous velocity to show that during putative ripple events, the animal is in fact, quiescent.

      We added the speed at the time of ripples in Fig. 6-1 D.

      (b) Figure 7C: The authors should show how time spent in the reward zones varies over days. Is SWR power simply changing due to behavioral occupancy differences (e.g., a statistical sampling problem)? An analysis of the SWR rate over days would be valuable.

      We agree and have added additional SWR properties across days in Fig. 6-1. This includes SWR power (Fig. 6-1 A), rate (Fig. 6-1 B), duration (Fig. 6-1 C) and speed (Fig. 6-1 D). To ensure we are comparing the equivalent time spent at goals, the SWR properties were calculated from the 10 s after goal entry (Fig. 6-1) and we only included goal visits that lasted at least 10 s. This controls for the potential confounds from occupancy differences across days.

      (c) Why did the authors stop comparing the F-maze and Y-maze for SWR analysis?

      The SWR properties were similar in both tasks. We have now added separate analyses for each maze in Fig. 6-1 and 6-2.

      (11) Figure 8

      (a) The authors discovered that mPFC neurons more strongly entrained to SWRs compared to their own beta rhythms, potentially indicating coordination of mPFC and hippocampus (lines 275:276). What about mPFC spiking to CA1 beta?

      We quantified mPFC spiking entrainment to CA1 beta in Fig. 5 O and P. We found mPFC cells were more strongly entrained to the local beta within mPFC compared with CA1 beta.

      (b) Lines 277:278: How did the authors arise at 8.6% being the expected proportion of neurons entrained by both SWRs and beta rhythms?

      We have added a better explanation of how we arrived at the expected proportion. The calculation is based on the joint probability between the proportion of neurons modulated by SWRs (30%) and the proportion of neurons with phase locking to beta rhythms (22%). The expected proportion (6.6%) is calculated by multiplying the two values (30% ´ 22%) under the assumption that the two classes are independent. Deviations from the expected proportions are determined using a Fisher’s exact test.

      We note that in the revised manuscript we reanalyzed the data with more stringent criteria. We now include SWRs with durations greater than 50 ms, rather than 15 ms, and we quantify beta spike phase-locking using only the spikes within the first 5 s after goal entry, for goal visits lasting at least 5 s.

      With these more conservative criteria, the proportion of SWR- and beta-modulated cells (8%) is no longer significantly different from the expected proportion (6.6%, Fisher’s exact test p=0.09). Although this differs from our original significance test, the trend, and importantly, the distinct task correlates for this population (Fig. 7 C-D) still hold. We have updated the revised manuscript with these findings.

      Our original result was: “The subset of PFC cells that are modulated by both SWR and beta (11%) is greater than the expected proportion (8.6%) under the assumption that SWR and beta can modulate the population independently (Fisher exact test, p=0.021), although the size of the difference is small.”

      (c) The discovery that neurons modulated by both beta and SWRs are unique from those simply modulated by beta is really interesting. The authors discovered an interesting relationship among dually entrained mPFC neurons whereby SWR-excited mPFC neurons were entrained to the peak of beta, whereas SWR-inhibited mPFC neurons were entrained to the trough of beta. Then, in lines285:286, the authors write:

      (12) "This relationship was not observed for PFC cells that are modulated by beta but not modulated by SWRs (Fig. 8C, right, Fig. 8-1A)."

      Why would the authors expect this relationship to exist when the mPFC neurons were not SWR modulated? Were the authors referring to something else?

      We should have phrased this more clearly. We edited the text to better convey the expected SWR and beta modulation patterns for the control populations. We wanted to ensure the relationship between SWR modulation direction and beta phase preference was not observed for the cells that were not SWR-modulated.

      Furthermore, what about mPFC neurons that are SWR modulated but not beta modulated? For completeness, the authors should examine these.

      We agree this is an important comparison and have now included this population in the new Fig. 7. As expected, we do not find any relationship between SWR modulation direction and spike preference to beta for the SWR-modulated and not beta-modulated population.

      (13) Lines 294:296: The authors should elaborate on how they obtained an expected proportion of CA1 beta modulated neurons.

      We have now added a description for calculating the expected number of beta-modulated CA1 neurons. These are the neurons that have a Rayleigh test p-value less than 0.05.

      (14) What are these neurons doing to predict task information? Do they at all? Do they differ from beta-only neurons or non-phase-locked neurons?

      This is a great suggestion, and we have added analyses to show the task correlates for the cells. We found brain region-specific differences for task correlates that depended on how the cell was modulated by SWRs or local beta oscillations. This is shown in the updated Fig. 7 and the accompanying supplemental Figs. 7-1 to 7-2.

      For PFC cells, SWR and beta modulation status defined a subpopulation with a strong task structure correlation. There was a positive correlation between the direction of SWR modulation and the spiking distance relative to goals. SWR-excited cells were more active further away from goals, whereas the SWR-inhibited cells were active closer to goals. This correlation was not found for SWR-modulated PFC cells that were not beta-modulated. For PFC cells, the direction of SWR modulation is known to be correlated with movement speed, consistent with the hypothesis that movement-active cells become reactivated during SWRs, and immobility-active cells become suppressed (Jadhav et al., 2016; Yu et al., 2017). We found the same relationship, the direction of SWR modulation was positively correlated with mean spiking speed. Our results show beta modulation marks a subpopulation of PFC cells with stronger task-structure correlates.

      For CA1 cells, we found the expected relationship between SWR modulation and task-structure correlates. SWR-excited CA1 cells spiked further away from the goal and when the animal was moving. This is consistent with the reactivation of trajectory-related spatial firing patterns during movement on the maze. This pattern was observed irrespective of beta modulation status.

      Ahmed, O. J., & Mehta, M. R. (2012). Running speed alters the frequency of hippocampal gamma oscillations. J Neurosci, 32(21), 7373-7383. https://doi.org/10.1523/JNEUROSCI.5110-11.2012

      Berke, J. D., Hetrick, V., Breck, J., & Greene, R. W. (2008). Transient 23-30 Hz oscillations in mouse hippocampus during exploration of novel environments. Hippocampus, 18(5), 519-529. https://doi.org/10.1002/hipo.20435

      Deceuninck, L., & Kloosterman, F. (2024). Disruption of awake sharp-wave ripples does not affect memorization of locations in repeated-acquisition spatial memory tasks. Elife, 13. https://doi.org/10.7554/eLife.84004

      den Bakker, H., Van Dijck, M., Sun, J. J., & Kloosterman, F. (2023). Sharp-wave-ripple associated activity in the medial prefrontal cortex supports spatial rule switching. Cell Rep, 42(8), 112959. https://doi.org/10.1016/j.celrep.2023.112959

      França, A. S., do Nascimento, G. C., Lopes-dos-Santos, V., Muratori, L., Ribeiro, S., Lobão-Soares, B., & Tort, A. B. (2014). Beta2 oscillations (23-30 Hz) in the mouse hippocampus during novel object recognition. Eur J Neurosci, 40(11), 3693-3703. https://doi.org/10.1111/ejn.12739

      França, A. S. C., Borgesius, N. Z., Souza, B. C., & Cohen, M. X. (2021). Beta2 Oscillations in Hippocampal-Cortical Circuits During Novelty Detection. Front Syst Neurosci, 15, 617388. https://doi.org/10.3389/fnsys.2021.617388

      Hyman, J. M., Zilli, E. A., Paley, A. M., & Hasselmo, M. E. (2005). Medial prefrontal cortex cells show dynamic modulation with the hippocampal theta rhythm dependent on behavior. Hippocampus, 15(6), 739-749. https://doi.org/10.1002/hipo.20106

      Iwasaki, S., Sasaki, T., & Ikegaya, Y. (2021). Hippocampal beta oscillations predict mouse object-location associative memory performance. Hippocampus, 31(5), 503-511. https://doi.org/10.1002/hipo.23311

      Jadhav, S. P., Kemere, C., German, P. W., & Frank, L. M. (2012). Awake hippocampal sharp-wave ripples support spatial memory. Science (New York, N.Y.), 336(6087), 1454-1458. https://doi.org/10.1126/science.1217230

      Jadhav, S. P., Rothschild, G., Roumis, D. K., & Frank, L. M. (2016). Coordinated Excitation and Inhibition of Prefrontal Ensembles during Awake Hippocampal Sharp-Wave Ripple Events. Neuron, 90(1), 113-127. https://doi.org/10.1016/j.neuron.2016.02.010

      Jin, S. W., Ha, H. S., & Lee, I. (2024). Selective reactivation of value- and place-dependent information during sharp-wave ripples in the intermediate and dorsal hippocampus. Sci Adv, 10(32), eadn0416. https://doi.org/10.1126/sciadv.adn0416

      Jones, M. W., & Wilson, M. A. (2005). Theta Rhythms Coordinate Hippocampal–Prefrontal Interactions in a Spatial Memory Task. PLoS Biology, 3(12). https://doi.org/10.1371/journal.pbio.0030402

      Kim, S. M., & Frank, L. M. (2009). Hippocampal Lesions Impair Rapid Learning of a Continuous Spatial Alternation Task. PLoS ONE, 4(5). https://doi.org/10.1371/journal.pone.0005494

      Lansink, C. S., Meijer, G. T., Lankelma, J. V., Vinck, M. A., Jackson, J. C., & Pennartz, C. M. (2016). Reward Expectancy Strengthens CA1 Theta and Beta Band Synchronization and Hippocampal-Ventral Striatal Coupling. J Neurosci, 36(41), 10598-10610. https://doi.org/10.1523/JNEUROSCI.0682-16.2016

      Maharjan, D. M., Dai, Y. Y., Glantz, E. H., & Jadhav, S. P. (2018). Disruption of dorsal hippocampal-prefrontal interactions using chemogenetic inactivation impairs spatial learning. Neurobiol Learn Mem, 155, 351-360. https://doi.org/10.1016/j.nlm.2018.08.023

      Rangel, L. M., Chiba, A. A., & Quinn, L. K. (2015). Theta and beta oscillatory dynamics in the dentate gyrus reveal a shift in network processing state during cue encounters. Front Syst Neurosci, 9, 96. https://doi.org/10.3389/fnsys.2015.00096

      Siapas, A. G., Lubenov, E. V., & Wilson, M. A. (2005). Prefrontal Phase Locking to Hippocampal Theta Oscillations. Neuron, 46(1), 141-151. https://doi.org/10.1016/j.neuron.2005.02.028

      Yu, J. Y., Kay, K., Liu, D. F., Grossrubatscher, I., Loback, A., Sosa, M.,…Frank, L. M. (2017). Distinct hippocampal-cortical memory representations for experiences associated with movement versus immobility. Elife, 6, e27621. https://doi.org/10.7554/eLife.27621

    1. eLife Assessment

      This study provides convincing evidence that the genetics of local adaptation in Arabidopsis is shaped by fluctuations in the environment and interactions with genotype and location. This is an important dataset contributing to the developing understanding of non-linear selection in plants and beyond.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review both by revising their original presentation where necessary and by providing responses to the reviewers in cases where they disagreed that revision of the original presentation would be necessary.]

      Summary:

      As a general phenomenon, adaptation of populations to their respective local conditions is well-documented, though not universally. In particular, local adaptation has been amply demonstrated in Arabidopsis thaliana, the focal species of this research, which is naturally highly selfing. Here, the authors report assays designed to evaluate the spatial scale of fitness variation among source populations and sites, as well as temporal variability in fitness expression. Further, they endeavor to identify traits and genomic regions that contribute to the demonstrated variation in fitness.

      Strengths:

      With many (200) inbred accessions drawn from throughout Sweden, the study offers an unusually fine sampling of genetic variation within this much-studied species, and through assays in multiple sites and years, it amply demonstrates the context-dependence of fitness expression. It supports the general phenomenon of local adaptation, with multiple nuances. Other examples exist, but it is of value to have further cases illustrating not only the context-dependence of fitness expression but also the sometimes idiosyncratic nature of fitness variation. I commend the authors on their cautionary language in relation to inferences about the roles of particular genomic regions (e.g.l.140-144; l.227)

    3. Reviewer #2 (Public review):

      Summary:

      The goal of this study was to find evidence for local adaptation in survival and fecundity of the model plant Arabidopsis thaliana. The authors grew a large set of Swedish Arabidopsis accessions at four common garden sites in northern and southern Sweden. Accessions were grown from seed in trays, which were laid on the ground at each site in late summer, screened for survival in fall and the following spring, and fecundity was determined from rosette size and seed production in spring. Experiments were complemented by 'selection experiments', in which seeds of the same accessions were sown in plots, and after two years of growth, plants were sampled to determine fitness from genotype frequencies, providing a more comprehensive evaluation of lifetime fitness than can be gleaned from fecundity alone.

      As the main result, southern accessions had higher mortality in northern sites in one of two years, but also suffered more slug damage in southern sites in one year, indicating a potential link between frost tolerance and herbivore resistance. Fecundity of accession was highest when growing close to the 'home' environment, but while accessions from one sand dune population in southern Sweden had among the lowest fecundities overall, they consistently had the highest fitness in the selection experiment. Accessions from this population had large seed size and rapid root growth, which might be related to establishment success when arriving in a new, partially occupied habitat. However, neither trait could fully explain the very high fitness of this population, suggesting the presence of other, unmeasured traits.

      Overall, the authors could provide clear evidence of local adaptation in different traits for some of their experiments, but they also highlight high temporal and spatial variability that makes prediction of microevolutionary change so challenging.

      Strengths:

      A major strength of this study is the highly comprehensive evaluation of different fitness-related traits of Arabidopsis under natural conditions. The evaluation of survival and fecundity in common garden experiments across four sites and two years provides an estimate of variability and consistency of results. The addition of the 'selection experiment' provides an extended view on plant fitness that is both original and interesting, in particular highlighting potential limitations of 'fitness-proxies' such as seed production that don't take into account seedling establishment and competitive exclusion.

      Throughout the study, the authors have gone to impressive depths in exploring their data, and particularly the discovery of 'native volunteers' in selection experiment plots and their statistical treatment is very elegant and has resulted in compelling conclusions. Also, while the authors are careful in the interpretation of their GWAS results, they nonetheless highlight a few interesting gene candidates that may be underlying the observed plant adaptations, and which likely will stimulate further research.

      Overall, this study will likely make an important contribution to the field of evolutionary biology, and it is another very strong example of how the extensive molecular tools in Arabidopsis can be leveraged to address fundamental questions in evolution and ecology, to an extent that is not (yet) possible in other plant systems.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript presents a large common garden experiment across Sweden using solely local germplasm. Additionally, there is a collection of selection experiments that begin investigating the factors shaping fecundity in these populations. This provides an impressive amount of data and analysis investigating the underlying factors involved. Together, this helps support the data showing that fluctuations and interactions are key components determining Arabidopsis fitness and are more broadly applicable across plant and non-plant species.

      Strengths:

      The field trials are well conducted with extensive effort and sampling. Similarly while the genetic analysis is complex it is well conducted and reflects the complexity of dealing with population structure that may be intricately linked to adaptive structure. This has no real solution and the option of presenting results with and without correction is likely the only appropriate option.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      As a general phenomenon, adaptation of populations to their respective local conditions is well-documented, though not universally. In particular, local adaptation has been amply demonstrated in Arabidopsis thaliana, the focal species of this research, which is naturally highly selfing. Here, the authors report assays designed to evaluate the spatial scale of fitness variation among source populations and sites, as well as temporal variability in fitness expression. Further, they endeavor to identify traits and genomic regions that contribute to the demonstrated variation in fitness.

      Strengths:

      With many (200) inbred accessions drawn from throughout Sweden, the study offers an unusually fine sampling of genetic variation within this much-studied species, and through assays in multiple sites and years, it amply demonstrates the context-dependence of fitness expression. It supports the general phenomenon of local adaptation, with multiple nuances. Other examples exist, but it is of value to have further cases illustrating not only the context-dependence of fitness expression but also the sometimes idiosyncratic nature of fitness variation. I commend the authors on their cautionary language in relation to inferences about the roles of particular genomic regions (e.g.l.140-144; l.227)

      Weaknesses:

      To my mind, the manuscript is written primarily for the Arabidopsis community. This community is certainly large, but there are many evolutionary biologists who could appreciate this work but are not invited to do so. The authors could address the broader evolution community by acknowledging more of the relevant work of others (I've noted a few references in my comments to the authors). At least as important, the authors could make clearer the fact that A. thaliana is (almost) strictly selfing and how this feature of its biology both enables such a study and also limits inferences from it. Further, it seems to me that though I could be wrong, readers would appreciate a more direct, less discursive style of writing, and one that makes the broader import of the focal questions clearer.

      We agree that connecting the paper better to the broader field is desirable, and have tried to do this in the revision, including adding a brief description of A. thaliana that mentions the fact that it is highly selfing. That the availability of inbred lines influences the study design is fairly obvious. As for the inferences, we would argue that our main point—“know your organism”—is universal. For example, our selection experiments and common garden experiments produced diametrically opposite results in terms of fitness, and we attribute this to the importance of seedling establishment. It is hardly surprising that seedling establishment matters for an annual plant that favors disturbed habitats, but this is probably true for both outcrossers and selfers (although it is probable that selfing itself is an adaptation to such habitats), and could well be true for some long-lived trees as well. The Devil is in the details.

      As a reader, I would value seeing estimates of the overall fitness of the accessions in the different conditions, i.e., by combining the survival and fecundity results of the common garden experiments.

      Combining estimates would be possible in the common garden experiments, and would bring us somewhat closer to total fitness estimates, although as noted by another reviewer (and also emphasized by us), the time scale of our experiment is not sufficient to evaluate the trade-off between survival and fecundity. Furthermore, we would still be missing the establishment component of fitness, which we found to be extremely important. Therefore little would be gained by combining the estimates, while at the same time losing resolution to disentangle the fitness components. We thus decided to focus on the individual fitness components and leave a qualitative consideration of their joint effect for the Discussion.

      Reviewer #2 (Public review):

      Summary:

      The goal of this study was to find evidence for local adaptation in survival and fecundity of the model plant Arabidopsis thaliana. The authors grew a large set of Swedish Arabidopsis accessions at four common garden sites in northern and southern Sweden. Accessions were grown from seed in trays, which were laid on the ground at each site in late summer, screened for survival in fall and the following spring, and fecundity was determined from rosette size and seed production in spring. Experiments were complemented by 'selection experiments', in which seeds of the same accessions were sown in plots, and after two years of growth, plants were sampled to determine fitness from genotype frequencies, providing a more comprehensive evaluation of lifetime fitness than can be gleaned from fecundity alone.

      To clarify, fecundity was determined from total plant area using photos of the mature stems, not the rosettes or direct counting of seeds. That said, it is true that our fecundity estimate was well correlated with rosette area. Furthermore, we validate our fecundity estimates by showing they were highly correlated with seed production estimated by measuring and counting siliques on a separate set of plants grown under common garden conditions in one of our sites (Brachi et al. 2022).

      As the main result, southern accessions had higher mortality in northern sites in one of two years, but also suffered more slug damage in southern sites in one year, indicating a potential link between frost tolerance and herbivore resistance. Fecundity of accession was highest when growing close to the 'home' environment, but while accessions from one sand dune population in southern Sweden had among the lowest fecundities overall, they consistently had the highest fitness in the selection experiment. Accessions from this population had large seed size and rapid root growth, which might be related to establishment success when arriving in a new, partially occupied habitat. However, neither trait could fully explain the very high fitness of this population, suggesting the presence of other, unmeasured traits.

      Another clarification: the small set of “beach” accessions that performed well in the selection experiments came from several beaches on the Baltic coast of Skåne, and we also had one such accession from the island of Gotland, 400 km away as the bird flies. Furthermore, we sampled measured seed size in several hundred additional accessions and demonstrated that large seeds are indeed characteristic of this habitat.

      Overall, the authors could provide clear evidence of local adaptation in different traits for some of their experiments, but they also highlight high temporal and spatial variability that makes prediction of microevolutionary change so challenging.

      Strengths:

      A major strength of this study is the highly comprehensive evaluation of different fitness-related traits of Arabidopsis under natural conditions. The evaluation of survival and fecundity in common garden experiments across four sites and two years provides an estimate of variability and consistency of results. The addition of the 'selection experiment' provides an extended view on plant fitness that is both original and interesting, in particular highlighting potential limitations of 'fitness-proxies' such as seed production that don't take into account seedling establishment and competitive exclusion.

      Throughout the study, the authors have gone to impressive depths in exploring their data, and particularly the discovery of 'native volunteers' in selection experiment plots and their statistical treatment is very elegant and has resulted in compelling conclusions. Also, while the authors are careful in the interpretation of their GWAS results, they nonetheless highlight a few interesting gene candidates that may be underlying the observed plant adaptations, and which likely will stimulate further research.

      Overall, the authors provide a rich new resource that is relevant and interesting both in the context of general evolutionary theory as well as more specifically for molecular biology.

      Weaknesses:

      While the repetition of the common garden experiments over two years is certainly better than no repetition (hence its mention also under 'strengths'), the very high variability found between the two years highlights the need for more extensive temporal replication. In this context, two temporal replicates are the bare minimum, and more repeats in time would be necessary to draw any kind of conclusion about the role of 'high mortality' and 'low mortality' years for the microevolution of Arabidopsis. It also seems that the authors missed an opportunity to explore potentially causal variation among years, as they did not attempt to relate winter mortality to actual climatic variables, even though they discuss winter harshness as a potential predictor.

      We agree that two years is insufficient to understand how variation in selective pressures compound over time to generate micro-evolutionary change. The eight-year data in Oakley et al. (2023), which we discuss in the paper, support this. Our results are nonetheless sufficient to demonstrate the idiosyncratic nature of selection. In the revision, we further emphasize that far longer time series would be needed for definitive conclusions on long-term micro-evolutionary change.

      Our short time series is one reason why we do not try to correlate with climate data, as this would amount to doing statistics with four data points (mostly two groups of accession N vs S, with mostly homogenous climates within groups, and two years). Another reason is that we simply have no idea how the kind of meteorological data that are publicly available would affect plant performance on a local scale. Hence we make no claims..

      The low temporal variation also makes the accidental slug herbivory appear somewhat random. Potted plants are notoriously susceptible to slug herbivory, and while it is certainly nice that slug damage predominantly affected one group of accessions, it nonetheless raises the question whether this reflects a 'real' selection pressure that plants commonly face in their respective local environments.

      Characterizing the plants as “potted” is not accurate. A more serious objection is having what is effectively an A. thaliana monoculture. But indeed we have no idea whether slugs exert a significant selective pressure on A. thaliana in Sweden, and we make no claims to that effect. The evidence for selection on glucosinolates by generalist herbivores such as slugs is fairly strong, but the precise agent is not known, and probably varies over time and space. Our results merely demonstrate one possibility.

      The addition of the 'selection experiment' is certainly original and provides valuable additional insights, but again, it seems a bit questionable which natural process really has affected this outcome. While the genetic and statistical analysis of this experiment seems to be state-of-the-art, the experimental design is rather rudimentary compared to more standard selection experiments. Specifically, the authors added seeds from greenhouse-grown mothers to experimental plots and only sampled plants two years later. This means that, potentially, the first very big bottleneck was germination under natural conditions, which may have already excluded many of the accessions before they had a chance to grow. While this certainly is one type of selection, it is not exactly the type of selection that a 2-year selection experiment is set up to measure. Either initially establishing the selection experiment from plants instead of seeds, or genotyping the population over several generations, would have substantially strengthened the conclusions that could be drawn from this experiment.

      We 100% agree that more data would have been beneficial, and hence we do not make any claims about the nature of selection. The selection experiment was an experiment per se, and we were lucky that very large fitness differences turned out to exist. As for initial selection on dormancy “set” incorrectly in greenhouse-grown seeds, we agree that this is a possibility, but we do not think it is likely to explain the data for several reasons. First, why would these differences favor a small set of beach accessions over every other accession in four very different field sites? Second, existing dormancy estimates do not predict fitness in our selection experiments. Third, the same seed batches germinated uniformly in the common-garden experiments with minimal stratification. Fourth, a much simpler explanation—seed size—exists. We clarify this in the revision while retaining our original message that further experiments are needed (and are underway).

      Also, the complete lack of information on population density is a bit problematic. It is not clear if there were other (non-Arabidopsis) plants present in the plots, how many Arabidopsis plants were established, if numbers changed over the year, etc. Given all of these limitations, calling this a 'selection experiment' is in fact somewhat misleading.

      Seeds were introduced into sites that appeared appropriate for A. thaliana, leaving the background community intact. We provided information on sowing density; the density of plants (A. thaliana and other species) that we obtained during the course of the experiments varied considerably between sites, much like in natural populations, although we lack systematic measurements. We provide more information (including photos) in the revision.

      Despite these weaknesses, the authors could achieve their main goals, and despite the somewhat minimal temporal replication, they were lucky to sample two fairly distinct years that provided them with interesting variation, which they could partially explain using the variation among their accessions. Overall, this study will likely make an important contribution to the field of evolutionary biology, and it is another very strong example of how the extensive molecular tools in Arabidopsis can be leveraged to address fundamental questions in evolution and ecology, to an extent that is not (yet) possible in other plant systems.

      Reviewer #3 (Public review):

      Summary:

      The manuscript presents a large common garden experiment across Sweden using solely local germplasm. Additionally, there is a collection of selection experiments that begin investigating the factors shaping fecundity in these populations. This provides an impressive amount of data and analysis investigating the underlying factors involved. Together, this helps support the data showing that fluctuations and interactions are key components determining Arabidopsis fitness and are more broadly applicable across plant and non-plant species.

      Strengths:

      The field trials are well conducted with extensive effort and sampling. Similarly while the genetic analysis is complex it is well conducted and reflects the complexity of dealing with population structure that may be intricately linked to adaptive structure. This has no real solution and the option of presenting results with and without correction is likely the only appropriate option.

      Weaknesses:

      A significant finding from this study was that fecundity is shaped more by yearly fluctuations and their interaction with genotype than it is by the main effect of location or genotype. Another significant finding is that the strength of selection can be quite strong, with nearly 5x ranges across accessions. It should be noted that there are a number of other studies using Arabidopsis in the wild with multiple years and locations that found similar observations beyond the Oakley citation. In general, the context of how these findings relate to existing knowledge in Arabidopsis is a bit underdeveloped.

      We have tried to remedy this in the revision (see also comments by Reviewer #1).

      The effects of the populations across the locations seem to rely on individual tests and PC analysis. It would seem to be possible to incorporate these tests more directly in the linear modeling analysis, and it isn't quite clear why this wasn't conducted.

      We respond to this question below, in Recommendations for the Authors (first item from you).

      I'm a bit puzzled by the discussion on how to find causative loci. This seems to focus solely on GWAS as the solution, with a goal to sequence vast individuals. But the loci that the manuscript discussed were found by a combination of structured mapping populations followed by molecular validation that then informed the GWAS. As such, I'm unsure if the proposed future approach of more sequencing is the best when a more balanced approach integrating diverse methods and population types will be more useful.

      We are puzzled by this comment in return. Our statement about more sequencing (penultimate sentence of discussion) was referring to achieving a better understanding of the history of migration and selection rather than identifying causative loci.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers provide consistent and complementary recommendations that you should consider in a revision. In particular, you should try to better embed your results into the broader literature, both from Arabidopsis and other plant species. Also, try to make your text more accessible to readers outside the special topic.

      Agreed!

      Reviewer #1 (Recommendations for the authors):

      (1) l.545: Here, the text states, importantly, that the accessions were randomized in the greenhouse. However, I have found no mention of this in the referenced Brachi et al. paper. I'm concerned that this be reported accurately and also that, if the accessions were not, in fact, randomized in the greenhouse, the potential for environmentally induced maternal influences to be confounded with genetically based differences be acknowledged, esp. in the case of the seed size difference found for B accessions.

      Accessions were continuously randomized in the greenhouse although we stopped moving them when they flowered to reduce the risk of contamination. All the seeds for all the accessions were produced in the same greenhouse, in the same conditions, in one single planting. We state this clearly in the revised paper.

      As for the seed size variation, note that the measurements we use in the paper were not taken on the seed we used in the experiments, so there is no reason they should be correlated with fitness if fitness was influenced by maternal effects. Moreover, we have plenty of data demonstrating that seed size variation is mostly genetic, and we added a supplementary figure showing that the measurements we used are strongly correlated with those of other experiments, including one done in the field. Finally, we already included a figure demonstrating that beach populations generally have much larger seeds.

      (2) l.593: BLUPs are estimated with statistical uncertainty. I realize that it would be far from straightforward to take into account their sampling variances in the GWAS, but it should be acknowledged that ignoring the uncertainty of the BLUPs has an unknown impact on the findings from the GWAS.

      All phenotypic measures come with error, and we had more replication than most GWAS studies (including essentially all human GWAS). The effect of such error is to reduce heritability and decrease the power of GWAS.

      (3) l.48: As an earlier important reference demonstrating this point: Antonovics, Clay, Schmitt, 1987 Oecologia.

      Indeed: added. Thanks!

      (4) l.49: studied rapid evolution in a natural population[s] - delete 's'

      Done.

      (5) l.502: lead -> led

      Corrected.

      (6) l.53: The present study sought to gain insight into local adaptation in A. thaliana.; Vague, as is the rest of the intro.

      Agreed. As noted above, the intro has been extensively reworked.

      (7) l.533: What is alpha?

      Alpha is the regularization parameter for the snmf algorithm. We added this clarification.

      (8) l.553: Were the accessions also randomized in the field planting?

      Yes, and we now state this.

      (9) l.597: As I've noted in my public comments, the authors have done well to couch their genomic inferences with caveats. In line with such caution, I urge the authors to consider alternate phrasing to 'genetically determined' -> genetically influenced. Also l. 166: 'control of' -> influence on (among other instances).

      No, “genetically determined” was correct on line 597, as this describes the model assumptions. But note that the liability-threshold model is by no means genetically deterministic: it is widely used in epidemiology to model genetic predisposition to diseases caused by the environment. For example, whether you get lung cancer or not primarily depends on luck and on your exposure to pollution (in particular smoking), but there are genes that influence how susceptible you are (to pollution; there are no genes that influence luck). There were some words missing in the sentence; hopefully things are clearer now.

      Re l. 166, "control of” was changed to “effect on”, which is more accurate. Note, however, that there is nothing genetically deterministic here: the context is a variance-partitioning, and the results show very clearly that genetic factors play a minor role relative to environmental ones—and that most of the variance remains unexplained. (Luck?)

      (10) l.595-602: Not clear.

      As noted, some words were missing. Hopefully it is clearer now.

      (11) l.744: Could have been confounded by cryptic native -- missing or extra word? Also next line.

      Corrected.

      (12) l.755-7: More direct phrasing would help readers understand this point.

      Indeed. Sentence straightened out.

      (11) l.129: 'The peak appears to involve a haplotype over 30 kb in length, which is consistent with a history of strong selection on this locus'. I don't understand the logic here. In what way should the length of the haplotype relate to the strength of selection? I would think this relates quite directly to the high degree of inbreeding.

      Strong selection causes rapid allele-frequency change, leaving less time for recombination to break up associated haplotypes, causing increased linkage disequilibrium. This is standard population genetics (e.g., Maynard Smith and Haigh 1974). Inbreeding also causes increased haplotype sharing, but genome-wide. Telling them apart is difficult, hence we used “consistent”. However, Reviewer 3 points to the existence of a segregating inversion in this region, which is an even more likely explanation, and we focus on this in the revision.

      (12) l.132: 'This suggests that the variation for slug damage seen in Figure 4 may be partly mediated by glucosinolate production.' I find the logic unclear here, as well.

      “This” referred to several observations, which is poor English. We rearranged the sentences to make our meaning clearer.

      (13) l.173: 'the accession-effect on fecundity' -> variation among accessions wrt fecundity.

      We rewrote this clunky sentence.

      (14) l.175: PCA unclear. Is this the same PCA described at l. 638? That is the only mention of PCA that I find in the methods, but I don't see that it connects here.

      No this doesn’t refer to the same analysis. In line 175 we talk about a vanilla, textbook PCA: confronted with 200 fecundity estimates in 8 experiments, we used PCR to look for patterns across the 8 dimensions, and found that 3 dimensions captured most of the variation (as shown in the heatmap). We have added a few sentences about this in the methods. Line 638 describes the way images of plants were treated to estimate variation in rosette color.

      (15) l.188: 'reveals what is causing them': Consider rephrasing to avoid language of causality, consistent with care taken elsewhere.

      Well, the context here is very different: we are effectively doing a post hoc analysis, and there is nothing wrong with saying that “the significant value in the chi-square test is caused by an excess of…”, for example. The patterns we see in the PCA are caused by the kinds of patterns we go on to discuss—it captures these patterns, among other things. We changed “reveals” to “helps us understand” to be less biblical.

      (16) l.219: downstream of -> conditional on.

      Yes.

      (17) l.235: Although native plants WERE growing nearby.

      Yes.

      (18) l.264: nearby: make this more explicit, i.e. within x km.

      We changed the sentence to “Although native plants were growing within less than a hundred meters in most cases,...”

      (19) l.285: none of our MAJOR conclusions depend on this.

      We disagree. No conclusion in the paper could be confounded by potential natives.

      (20) l.315: albeit it not as large as B accessions: delete 'it'.

      Corrected.

      (21) l.355: genetic basis of fitness -> genetic contribution to variation in fitness.

      No, this is pretty much exactly Lewontin’s usage—the sentence (and section) is about changes in allele frequencies. The alternative suggested is less precise as we have no fitness estimates, and cannot say anything about variance.

      (22) l.380-2: It seems to me that the authors could articulate a more compelling case.

      Well, what we wrote is the truth. Of course all of this work also fits into a broader context, but those were the specific goals—and we achieved them. There are many papers that ask big questions, but actually answer much more limited ones.

      (23) l.401: latitude vs. latitude???

      Oops. Changed to “north vs. south”, which is hopefully less obscure!

      (24) l.402: 'textbook local adaption,' -- the authors should acknowledge that rigid/simplistic thinking about LA has been recognized as a caricature for quite some time.

      We do. Using “textbook” is meant to convey this: textbooks tend to be full of cartoonish simplifications, and most chapters on local adaptation will have a figure (cartoon or based on real data) showing reaction norms as two crossing lines.

      (25) l.406: could maintain variation AMONG POPULATIONS [right?]

      Could be either, but “among populations” is better in this context.

      (26) l.419: Is there a possibility of a source envt maternal effect? See my comment on l.545.

      As stated in response to the previous comment: In principle yes, but if so, then this environmental maternal effect happens to be strongly correlated with seed size, a trait we know to be genetically controlled and which is also extremely likely to influence seedling establishment. Experiments to confirm these results are underway—meanwhile we would be happy to accept bets against!

      (27) l.434: 'very stable environments dating back to the last glaciation.' REF?

      Ha! That would be Wikipedia references as the precise location of the “Littorina sea” and existence of obvious ancient beachlines now inland have entertained Swedish and Danish school children for generations. But the details are not important: the point is that beaches are vast (by A. thaliana standards) disturbed habitats maintained by the sea, and while they change, they do so on a geological time scale. Of course not all beaches harbor A. thaliana—the beaches of Hanö Bay are geologically unusual for Sweden—but this is beyond the scope of this publication. The sentence has been changed to something less specific making the relevant points.

      (28) l.436: 'It is likely that the existence of S1 and S2 accessions is far more uncertain': unclear wording.

      Yup. We rewrote the paragraph.

      (29) l.446-7: 2nd person is jarring.

      Reworded.

      (30) ll.452-474: This is a welcome acknowledgement of the limits of molecular approaches to elucidating selection. It would be good scholarship to acknowledge earlier authors making such a case. I could suggest Rockman 2012. Evolution; Travisano and Shaw 2013. Evolution; Hoban et al. 2016. Am.Nat., and there are others.

      True; added; thanks!

      (31) l.490: 'might have to run a gauntlet of linkages to genes directly involved in local adaptation': This teleological wording is especially jarring.

      (32) l.494: 'resistance allele to sweep': In concluding this manuscript, it does not seem appropriate to use as an example a single locus case. More broadly, I question the value/effectiveness of this paragraph as a conclusion to this manuscript.

      We agree. The paragraph, gauntlets and all, has been replaced by discussion of why dissecting adaptive traits is hard.

      Reviewer #2 (Recommendations for the authors):

      (1) At the end of the introduction, the authors briefly outline the geographic scale of their study and mention both common garden and experimental evolution plots, but at this point in the manuscript, it is not clear how these differ, and as the methods only come at the very end, it would be better if some more details are provided here. Specifically, it could be mentioned explicitly here that common garden experiments consisted of placing greenhouse-grown plants in pots on the ground (but not burying/planting them), whereas selection experiments were established by sprinkling seeds in 1m 2 plots. This doesn't really become clear anywhere except in the methods, but this is fairly important for the interpretation of results.

      We agree, and have expanded the description of the experiments in the legend to Figure 1, which also links to supplementary photos of the sites. We also note that the differences between the common-garden and the selection experiments should not be exaggerated. In particular, we did not use “greenhouse-grown plants”, but seedlings that had been allowed to establish outdoors in sheltered conditions, and we did not put “pots on the ground” but buried trays with holes in the bottom so that the plants could root in local soil. The main differences between the experiments are guaranteed establishment and lack of competition (from conspecific and other plants). This has also been clarified in the figure legend and in Methods.

      (2) The authors discuss a difference in mortality between the two years of their common garden experiment and suggest that harsher winters could have been the cause of mortality in the north. However, they do not provide meteorological data to support this. Was the 2011-2012 winter harsher than the 2012-2013 winter, and did NM experience harsher conditions?

      The problem is that we do not know what constitutes harshness from the point of view of the plant. Winters differ in many ways: temperature, duration, snow cover, etc. We have changed the relevant paragraph to make clear that our observations are consistent with those of Oakley et al., and that some aspect of winter weather is a plausible explanation.

      (3) The authors find an indication for the involvement of the AOP cluster in overwinter survival, which, among other things controls the accumulation of hydroxyl/alkenyl/methylsulfinyl glucosinolates. Is the chemotype of the accessions in this study known?

      Indeed they are! We had downloaded the data from Katz et al (2021), but the analysis did not make it into the first version of this paper. Thanks for the encouragement. We added a figure showing that their chemotypes are strongly associated with slug damage, supporting a causal relationship.

      (4) L68: 'we added one field site in eastern Skane' - which site does this refer to? It seems odd to mention this before the common garden/selection experiments are mentioned. It should be clear that this is specifically referring to the latter.

      Yes, this was out of place. The site is discussed later, when it becomes relevant.

      (5) Figure 1: It would be useful if common garden and experimental evolution sites used different colors. In contrast, the use of colors for seasons in part C is unnecessary and distracting, as the red and blue colors for fall and winter are very similar to the colors for S1 and B.

      Fixed.

      (6) L90: The authors discuss mortality, but figures show survival. This seems an unnecessary complication for the reader, and the same unit should be used for discussion and presentation.

      Survival is one minus mortality. We trust the readers to be able to do this conversion.

      (7) L99: 'the converse was not true' - it is not immediately clear what this refers to. Referring to the non-linear relationship in panel 3a would make this clearer.

      It refers to the result that the accessions with high mortality in NM did not necessarily have high mortality in NA. We think this is clear from the sentence in question.

      (8) L108: 'S2 accession were also strongly affected' - this can be gleaned from the figures, but it is not immediately obvious as they are relatively complex. Could mean mortality/survival rates be provided to facilitate this?

      We could but, throughout the paper, we have attempted to improve readability by not interrupting the text with numbers or repeating details that are presented in the figures.

      (9) L117: 'however, an indirect association is likely' - what is meant by this?

      We meant that it seems more likely that some accessions are more sensitive to stress regardless of source. This has been clarified.

      (10) Figure 6: What are the two horizontal bars? What are dashed lines? Provide appropriate labels and a figure legend.

      The top is a zoom-in of the bottom and the dashed lines outline the zoomed-in region. Clarified in legend.

      (11) L174: Define 'BLUPS' here.

      Done.

      (12) L198: 'S1 and S2 accessions generally had higher fecundity in 2012-2013 than in 2011-2012' - this is not obvious from Figure 9. Especially for S1 (dark blue presumably), there appears to be no difference visible between years.

      Changed text to note that differences are sometimes small, but that the stated pattern is seen in 13 out of 16 comparisons (which has p = 0.01 using a sign test).

      (13) L401: 'latitude vs. latitude' - what is meant by this? Or is this a mistake?

      Northern vs. southern. This has been clarified. Also noted by Reviewer 1.

      (14) L436: 'the existence of S1 and S2 accessions is much more uncertain' - this is a bit of an odd expression. Could it be rephrased?

      This has been clarified and the paragraph rewritten. Also noted by Reviewer 1.

      (15) L673: Were plots tilled before sowing, or cleared of vegetation? If not, what other vegetation was present at sowing?

      No clearing of vegetation was done except at the SR site, which was a weedy agricultural field. This and other information has been added. No vegetation surveys were made, but we added some more photos to give an idea of what the sites were like.

      Reviewer #3 (Recommendations for the authors):

      (1) In the section on lines 156-190 the large dataset is analyzed by accession in the linear model presented in Figure 6. Then, in Figure 8, the proposed populations appear to be tested individually in each experimental unit, leading to the probability argument in line 193. Is there a reason not to simply have accession nested within population in the model shown in Figure 6 as a way to directly test the between vs within population level components influencing the model? If the accessions are different, then it would be viable to treat this as random rather than fixed. Similarly, the field sites could be parsed into subsets as well.

      Excellent question. We thought a lot about this. Fig. 6 presents a standard ANOVA that simply shows that the overall pattern is what we hoped for: very large effects of site, year, and accession, plus substantial interaction effects. This motivates the exploratory analysis presented in Figs. 7-9, where we focus on how the relative performance of accessions within experiments depended on the fixed effects (year and site) and whether this matched their (arbitrarily but independently defined) group designation—as would be expected under local adaptation.

      We could explicitly add “group” to the model, as suggested, but this would give it a reality we do not think it deserves. Site, year, and accession are very much real, whereas group is the outcome of a somewhat arbitrary clustering of genotypes (i.e., accessions), analogous to race in human genetics, but without the sociological factors that sometimes warrant including race in a model not only because of direct genetic effects. Of course we could try to partition genotype and phenotype into within- and between-population variation using the classical quantitative genetics framework (i.e. Fst/Qst), but then we should have an a priori definition of “population”, which we don’t have. In addition to this fundamental objection, limited experimentation suggests that fitting a far more complex, nested mixed-effects model to our data is difficult in practice.

      Thus we prefer our original approach. We have changed the writing to clarify our logic.

      (2) Similarly, I'm not quite sure that the PC analysis is helping as the section is somewhat difficult to read given that the same impressions from Figure 7 are more explicitly shown in Figures 8 and 9.

      Well, the PCA (Fig. 7) is what led us to the analyses in Figs 8-9, where we interpret the PCs in terms of group behavior. We have tried to clarify our logic in writing.

      (3) Line 87-88 - An honest question, what is considered substantial divergence? The Fst values range from 0.05 to 0.27, suggesting that the divergences range across a spectrum and not all are substantial.

      We removed the sentence. That the divergence is substantial enough to make GWAS difficult becomes clear later (although the extent to which this is due to selection rather than marker divergence is not clear).

      (4) Line 129-130 - The 30kb AOP haplotype identified is likely the inversion associated with the major phenotypic variation identified in Sweden within Katz et al 2021. As such, the local LD structure may represent blocked recombination as much as linked selection.

      Of course; we missed that. Thanks!

      (5) Line 131-132 - I'm not quite sure what the evidence is that the MAM locus is more important.

      The AOP and MAM loci are epistatic, and both have been found to have influence on fitness in Arabidopsis and Brassica ssp in the field, along with sequence signatures of selection for both loci in multiple Brassicaceae. I'm unsure if this statement is supported.

      Indeed. This was a lazy and misleading reference to the fact that the MAM peak in Katz et al explains more of the variance than the AOP peak. It has been consigned to the dustbin of history. Instead we have added a figure showing that the glucosinolate profiles presented in that paper are highly correlated with slug damage in our study. Based on these results, I believe AOP explains more of the variation, but we leave pursuing this for those directly working on these pathways. The correspondence between our studies is certainly a nice confirmation of both.

      (6) Line 131-132 - It should also be noted that the MAM locus is variable in Sweden, albeit having multiple independent haplotypes that convergently create the same phenotype. There is an indication from Gloss 2022 that these haplotypes may create small-effect phenotypic variation.

      Correct again. What we meant to say was that the major polymorphism that is responsible for the highly significant MAM peak does not appear to segregate in Sweden, hence it is not surprising that we do not find an association at this locus either. We now say this. Needless to say, there could still be multiple variants at this locus that we do not have the power to detect.

    1. eLife Assessment

      This important study demonstrates that pyroglutamylation of peptide ligands alters G protein-coupled receptor (GPCR) activation in a subtype-specific manner and proposes a novel mechanism whereby the posttranslational modification steers the distal conformation of the peptide within the orthosteric binding site. The strength of the evidence supporting receptor-specific functional effects is convincing, but the strength of the evidence supporting the proposed structural mechanism is incomplete. This study will be of interest to pharmacologists and cell biologists interested in posttranslational modifications (PTMs) and activation mechanisms of GPCRs.

    2. Reviewer #1 (Public review):

      Summary:

      The study identifies two previously uncharacterized endogenous receptors for Aplysia PRXamide peptides and shows that N-terminal pyroglutamylation can produce opposite effects on receptor activation. The authors further propose that this modification acts indirectly by altering peptide conformation and that receptor pocket properties determine the direction of its effect. The findings are potentially significant for understanding how peptide modifications influence receptor selectivity, but the evidence supporting the proposed molecular mechanism is not yet sufficiently strong.

      Strengths:

      The identification and functional characterization of the two receptors are valuable. The contrasting effects of pyroglutamylation, together with peptide and receptor mutagenesis and computational analyses, provide an interesting framework for investigating peptide-receptor selectivity.

      Weaknesses:

      Most of the statements about novelty and generality are stronger than warranted by the current data. The central mechanistic model relies heavily on computational predictions and indirect functional measurements, without direct structural or receptor-proximal evidence. In addition, the human receptor data provide only partial support for the proposed mechanism, particularly because the pyroglutamylated and non-pyroglutamylated peptides are not significantly different at human receptor 2. Thus, the data support differential effects of pyroglutamylation but do not yet fully establish the proposed distal conformational mechanism.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Chang et al. presents the discovery and deorphanization of two bona fide PRXamide GPCRs, and their peptide agonists, in Aplysia californica (sea hare, a type of sea slug). Focusing on MMG2-DPb, the most potent peptide agonist, the authors demonstrate that pyroglutamination on the extracellularly facing N-terminus of the peptide reduces its agonist potency towards ApPRXa-R1 and increases potency towards ApPRXa-R2. Model-guided pocket and peptide mutagenesis demonstrate multiple differential dependencies and sensitivities of ApPRXa-R1 and ApPRXa-R2, supporting distinct mechanisms of peptide binding. The authors then show that a pyroglutaminated version of the dog Neuromedin U peptide has an increased potency towards human NMUR1, relative to the free-Gln version or the non-Gln-containing human peptide, whereas the potency towards NMUR1 is slightly reduced. These examples delineate N-terminal pyroglutamination of a peptide agonist as a receptor subtype selectivity controlling mechanism even when the peptide binds the receptor C-terminus.

      Strengths:

      (1) Identification of a class of GPCR-signaling peptides in a poorly characterized organism (Aplysia californica, a type of a sea slug).

      (2) Deorphanization of two Aplysia californica GPCRs and the establishment of an Aplysia analog of the human Neuromedin U signaling system.

      (3) A rigorous structure-function study of the Aplysia receptors and peptides.

      (4) Identification of pyroglutamination as a mechanism controlling subtype selectivity in Aplysia PRXamide receptor system.

      (5) Very clear and compelling graphics.

      (6) Overall, the biochemical and pharmacological part of the study is strong, exciting, and well-presented.

      Weaknesses:

      (1) My biggest problem is with the emphasis on the 'non-contacting role' of the N-terminal pGlu: "purely on positional grounds, N-terminal pQ/Q is unlikely to form direct contacts with receptor residues", "Despite the absence of direct pQ/Q-receptor interactions...", "indirect PTM distal steering", etc. The absence of direct contacts between pGlu and the receptor residues is likely an artifact of limited precision in modeling and reflects insufficiently advanced modeling tools. The authors used Swiss-Model, Robetta, and HPEPDOCK: why not go for the state-of-the-art AI modeling software like Boltz, Chai, or AlphaFold? A quick AlphaFold3 generates a model where Gln and pGlu form perfect contacts with the receptor N-term and ECL2 (in ApPRXa-R2) or N-term and ECL3 (in ApPRXa-R1). The N-termini of both receptors are quite long, form well-defined tertiary structures, and fold onto the receptor extracellular loops; in the case of ApPRXa-R1, the N-terminal domain is stabilized by an intra-domain disulfide bond C15(NT)-C133(NT) and stapled to ECL2 via another disulfide bond C16(NT)-C322(ECL2), suggesting that these interactions are real. The 3D models generated by the authors are not available for review but based on figures, I don't see these N-terminal domains modeled at all, which would obviously affect the docked position of the peptide N-terminus and its contacts with the receptor. For example, in an AlphaFold model, ApPRXa-R1 F161, E317, R340 are in direct vicinity of the peptide's Gln / pGlu and are quite distinct in ApPRXa-R2 - direct contacts with these or nearby residues may easily explain the differential preferences for free Gln vs pGlu but none of them were tested via mutagenesis.

      (2) Some conclusions and interpretations in the manuscript are overstated and lack precision. For example: "...a single N-terminal lactam modification converts a minimal chemical difference into bidirectional GPCR subtype outputs" makes the reader think about something as dramatic as an inversion of efficacy, an agonist becoming an antagonist, or at least the introduction of signaling bias, whereas in reality, the authors established that pyroglutamination reduces the potency of the peptide at one receptor but increases it at the other receptor. The use of PTMs and proteolytic processing is a broadly utilized mechanism for controlling GPCR peptide selectivity. The effects of pyroglutamination on the potency of MMG2-DPb towards its two Aplysia receptors very well align with this mechanism and should be described as such. Other examples of sentences that are similarly misleading: "... the same ligand, ..., produces opposite functional outcomes: pQ suppresses ApPRXa-R1 activation while enhancing ApPRXa-R2 activation" (in reality, both receptors are activated and the functional outcomes are the same, just achieved with different potency), "A single N-terminal lactam enables bidirectional GPCR subtype tuning...", "This bidirectional, subtype-specific effect of pQ is unprecedented in neuropeptide signaling..." (not bidirectional and not as unprecedented as the authors make it sound), etc.

      (3) Overall, what the authors present as a new "indirect PTM distal steering" mechanism is really not supported by the data. The described systems fit the classical "lock-key" and "induced fit" paradigms rather than challenging them. I recommend the authors revisit their modeling approaches and model-guided interpretations of the biochemical/pharmacological data. Importantly, I do not think such revision would undermine the key strengths of the paper (listed above in Strengths). Based on experimental data alone, this is a very strong and interesting study and with more rigorous modeling and adequate interpretations, it can be made exceptional.

    1. eLife Assessment

      This important paper studies how adiposity affects brain development in the ABCD study. Major findings include: (1) central adiposity indices (BRI, WHtR) show stronger cross-sectional associations with cognition than BMI, partially mediated by frontotemporal cortical morphology; (2) the rate of adiposity accrual, rather than static adiposity at either timepoint, predicts follow-up cognition and altered (attenuated) cortical thinning; and (3) among overweight/obese adolescents, central fat reduction is accompanied by accelerated cortical thinning and catch-up in inhibitory control. The strength of evidence is solid, with the main limitation related to only looking at the cerebral cortex rather than examining the whole brain. The subgroup observation that fat reduction may be accompanied by normalization of both cortical trajectories and inhibitory control is potentially clinically meaningful, suggesting reversibility rather than fixed deficit, and would be of broad interest if it withstands more rigorous analysis.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors used the ABCD study to understand how changes in body fat composition affect brain development during the teenage years. They found that the key determinants of changes in the brain were changes in fat accumulation rather than just baseline measures. Less importantly, but still interesting, they reaffirmed that more complex measures, such as body roundness index, better captured the impact of obesity than body mass index.

      Strengths:

      This study uses an excellent open access dataset, asks an important set of questions, and is executed with diligence for proper methodology. The finding that trajectories in adiposity matter more than baseline differences is important both for understanding how the brain adapts to body composition changes and has potential impact for public health interventions.

      Weaknesses:

      While I am overall impressed with the study, there are a few weaknesses I would like to see the authors address:

      (1) There is no reason to limit analyses to the cerebral cortex. Body composition changes are as likely to affect the subcortex or cerebellum as the neocortex. In some cases, like the hypothalamus, potentially even more likely.

      (2) Confounders should be examined in more detail. How are adiposity changes mediated by (for example) socio-economic status?

      (3) I'd like to see more rationale for excluding 483 kids with extreme adiposity indicators. Is it that they are untrustworthy entries? Otherwise, they could be particularly informative.

      (4) Are there any blood measures of glucose/insulin available?

      (5) Some of the figures could benefit from having raw data included rather than just showing the fitted trend lines.

      (6) Are there any alternate explanations for baseline differences? Particular examples that came to mind are the influence that maternal adiposity has on offspring brain development.

    3. Reviewer #2 (Public review):

      This manuscript addresses an important public health question in adolescents using a large longitudinal ABCD cohort. It is generally well organized and presents a compelling story linking central adiposity, brain development, and cognition. The authors use baseline and 4-year follow-up data from the ABCD Study (N = 8,519 baseline; N = 1,873 longitudinal, the use of 4-year follow-up only should be justified) to examine how adiposity relates to cognitive performance and cortical structure in early adolescence. Major findings include: (1) central adiposity indices (BRI, WHtR) show stronger cross-sectional associations with cognition than BMI, partially mediated by frontotemporal cortical morphology; (2) the rate of adiposity accrual, rather than static adiposity at either timepoint, predicts follow-up cognition and altered (attenuated) cortical thinning; and (3) among overweight/obese adolescents, central fat reduction is accompanied by accelerated cortical thinning and catch-up in inhibitory control.

      The use of large longitudinal data from ABCD is considered a strength as most prior neuroimaging work is cross-sectional. The focus on adiposity change velocity within adolescence is considered novel over single-timepoint designs, and the comparison of central-fat indices (BRI/WHtR) against BMI in a neurodevelopmental context is important given growing interest in these markers in cardiometabolic research. The subgroup observation that fat reduction may be accompanied by normalization of both cortical trajectories and inhibitory control is potentially clinically meaningful, suggesting reversibility rather than fixed deficit, and would be of broad interest if it withstands more rigorous analysis.

      However, several major issues were identified. For example, the strength of the causal language is not supported by the design, the headline comparison between adiposity indices rests on coefficients that are described as standardized but evidently are not, and several analytic and reporting issues (attrition and selection, family clustering, regression to the mean in the subgroup analysis, internal inconsistencies in reported p values) must be resolved before the conclusions can be considered established.

      Major Concerns

      (1) ABCD provides more follow-ups and should be included. Also, there is a severe longitudinal attrition (8519 vs 1873) that needs to be addressed. No comparison of completers versus non-completers is provided. The cohort description of having 78.5% White seems to be wrong, suggesting either a coding error in the race variable or strong selection introduced by the exclusions. The exclusion of 483 participants for "extreme adiposity indicator values" is especially concerning in a study of obesity and should be justified with explicit criteria in the main text, with sensitivity analyses retaining these participants where possible.

      (2) Adiposity is heavily influenced by socio-economic status, which itself contributes heavily to cognition and mental health as well. Many ABCD studies have reported this. Another important factor contributing to adiposity is sleep, which itself can contribute significantly to cognition and mental health. Many ABCD-based sleep studies have been published and can be reviewed.

      (3) The manuscript repeatedly uses causal language that the observational design cannot support (e.g., "longitudinal fat accumulation ... drives cortical alteration," "fat reduction activated adaptive neural change," "neuroprotective management"). The cited literature (Likhitweerawong et al., 2022) is explicitly bidirectional: poor inhibitory control may promote weight gain rather than the reverse. The baseline mediation analysis is particularly problematic because exposure, mediator, and outcome were measured concurrently, so the temporal ordering required for mediation is assumed rather than established. Causal claims should be tempered throughout. Similarly, mediation analyses should be interpreted cautiously as statistical rather than causal.

      (4) zBMI, WC, BRI, and WHtR are highly intercorrelated, as are follow-up adiposity and delta-adiposity (Table 3). Entering them simultaneously invites unstable estimates and sign flips. Please report pairwise correlations and variance inflation factors, and demonstrate that the key conclusions are robust.

      (5) Subgroup analysis: regression to the mean, group labels, and internal inconsistencies. (a) Stratifying on delta-BRI >= 0 enriches the "decreasing" group for high baseline values, so the observed cognitive "catch-up" and accelerated thinning may partly reflect regression to the mean; the percentile-based sensitivity analysis mitigates but does not resolve this. Group comparisons should adjust for baseline BRI and baseline cognition. (b) zBMI >= 1 corresponds to overweight by WHO criteria, not obesity; labelling this group "obese adolescents" throughout (including the Abstract) is inaccurate. (c) The Abstract claims a Flanker slope difference of p < 0.05, but the Results report p = 0.058 versus controls (described, incorrectly, as significant) and p = 0.021 only versus the increasing-BRI group; these statements must be reconciled. (d) It is unclear whether subgroup p values were FDR-corrected and how t-tests/Wilcoxon tests produced "adjusted" p values; if covariate adjustment was intended, regression models are required.

      Tables 2-4 mention that beta values are standardized, yet the magnitudes are clearly scale-dependent (e.g., WHtR beta = -7.5 vs. zBMI beta = -0.42 in Table 2; WHtR delta-beta = -780 in Table 3). This undermines the central claim that BRI/WHtR "outperform" BMI, since predictors cannot be compared on unstandardized coefficients. Please either truly standardize all predictors and outcomes, or compare indices formally (e.g., delta-R2, AIC/BIC, or non-nested model comparison tests), and report 95% confidence intervals for all estimates.

      (6) Underspecified mixed models and family clustering. The random-effects structure of the LME models is never stated. ABCD includes twins and siblings, so family nesting must be modelled (and the treatment of site - fixed covariate versus random effect - clarified). As written, the analyses may be anti-conservative. Please provide full model specifications and confidence intervals.

      (7) Interpretation of accelerated cortical thinning is overly strong. The manuscript equates faster thinning with better maturation and slower thinning with delay. This reading is contested: apparent cortical thinning also reflects myelination-related signal changes, and thickness-cognition associations are age- and region-dependent. Moreover, the present null finding for baseline adiposity contradicts Kaltenhauser et al. (2023), who reported that baseline adiposity attenuates thinning in ABCD; the discrepancy is cited but never reconciled and deserves direct discussion (differences in sample, covariates, or modelling) Unclear units and implausible mediation estimates. Table 1 gives delta-WC = 0.264, interpretable only if delta is per month (age in months), yet the text describes "annualized" change. Units for all delta variables must be stated explicitly and used consistently. Mediation effects reported as wide ranges (e.g., ACME = -29.95 to -0.18; "ACME = -0.1.17" is a typographical error) are uninterpretable and, at face value, implausibly large relative to the cognitive score scale; please report per-indicator ACMEs with confidence intervals and proportion mediated. Not sure how the ranges such as "beta = -0.29 to -0.00" cited as significant effects are informative.

      (8) Pubertal confounding. Only baseline pubertal category is adjusted for, but pubertal tempo over the follow-up window confounds both adiposity change and cortical development. This should be adjusted for change in pubertal stage or acknowledged as a substantive limitation. Puberty interactions and sex-specific analysis should be performed too.

      (9) Effect sizes should be discussed.

      (10) Inclusion of both baseline adiposity and adiposity change in longitudinal models should be justified.

      (11) Would there be nonlinear age/puberty effects?

  2. Sep 2026
    1. eLife Assessment

      This valuable study investigates the action of a floral homeotic transcription factor in different cell types. The results provide solid evidence that PhDEF shows more extensive binding and regulatory effects in the epidermis than in the mesophyll, supporting the conclusion that cell-layer context influences homeotic gene function. While there are some limitations, they do not undermine the principal observation of layer-dependent PhDEF activity, but the interpretation of the data is complicated by uncertainties in cell-type assignment and composition in the scRNA-seq data, the cell-type-enriched rather than strictly cell-type-specific nature of the ChIP-seq data, and differences in developmental stage between the single-cell and ChIP-seq experiments. The work will be of interest to developmental biologists studying transcriptional regulation and cell identity.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      Summary:

      The authors previously generated two cell-layer-specific mutants of petunia for the petal identity gene PhDEF. In this study, they profiled differential gene expression in those mutants through single-cell RNA sequencing (scRNA-seq). They found that more genes are highly and specifically expressed in the epidermal cell layer than in mesophyll cells. In addition, they identified cell-layer-specific and -aspecific PhDEF target genes. Using the extensive single-cell transcriptome and layer-specific target identification, the authors concluded that different cell identities affect homeotic regulator PhDEF, thereby influencing transcriptional regulation.

      Major comments from first round of review:

      This presented work provides comprehensive evidence, that pre-existing cell layer identity (epidermis and mesophyll) modulate transcriptional output of homeotic transcription factor, PhDEF.

      However, a disconnection between PhDEF bindings to genome and transcriptional output undermines the robustness of their conclusion although some binding loci were shown to be correlated with DEG. This may indicate the chromatin state, the existence of interacting partners, and the non-productive binding of PhDEF, suggesting that PhDEF binding alone is not sufficient to predict transcriptional outcomes and additional regulatory mechanisms that shape gene expression in addition to the layer-specific regulatory mechanisms. This disconnection may also be due to the developmental timing. Indeed, it appears authors used different flower stages for ChIP-seq and scRNA-sequencing. In fully differentiated organs, PhDEF binding itself may be no longer transcriptionally productive, and differential gene expression results primarily from the pre-established cell identity rather than directly from the homeotic regulation of PhDEF. Therefore, the main question the authors asked-how homeotic identity works with cell-layer identity and how the homeotic gene, PhDEF, acts in mature organs-was not clearly explained by this study.

      In Figure 2, the use of the term "target" is potentially misleading. It sounds like direct target genes (direct binding and differential expression) for PhDEF, but it refers only to DEGs.

      Lines 496-497: When the authors state, "~ demonstrates for the first time that the regulatory function of homeotic factor is influenced by cell layer identity," it sounds overstated, as prior studies have shown that pre-existing tissue or cell identity can shape transcriptional activity and developmental output.

      Significance:

      This study is well-designed and technically sound. They utilize single-cell transcriptomics and ChIP-seq by using genetically well-defined genetic materials and layer-specific PhDEF deletion mutants. The analysis showed where PhDEF binds to genomic loci and which genes are differentially expressed in petal epidermis and mesophyll, providing evidence of cell-layer-specific function of homeotic gene in mature organs. Although certain mechanistic aspects were not elucidated, the data from the extensive genome-wide study contributed to drawing their conclusions.

      Advances: This research goes beyond classical models of floral organ identity by showing that homeotic gene function is not uniform in the same floral organ. It represents a conceptual advance in our understanding by integrating cell layer identity into the framework of homeotic gene regulation.

      Audience: This study will be of broad interest to scientists who study transcription networks, cell and organ identity in the context of plant development.

      My field of expertise: Transcriptional regulation by transcription factor, epigenetic regulation of gene expression, plant development.

      Comments on latest version:

      Thank you for sharing the assessment. I am happy with the proposed eLife assessment. I do not have any further amendments to my review.

    3. Reviewer #2 (Public review):

      Summary:

      This study from Cavallini-Speisser et al. cleverly leverages a tissue layer-specific mutant, single cell and bulk RNA-sequencing, and ChIP-sequencing to decipher tissue layer-specific regulation of petal development in petunia by the PhDEF transcription factor. The authors find common and unique targets of PhDEF between the epidermis and mesophyll and conclude that the activity of transcription factors like PhDEF are influenced by the pre-existing environment in the cell they are expressed in. Understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology that can be elusive outside of highly tractable model systems. As such, I think this study is of high value and has strong potential to expand our understanding of how developmental specificity is mediated by commonly employed transcriptional regulators. However, I think there are some issues with possible over-interpretation and some places where documentation of experimental design and data quality control are lacking. I elaborate on these concerns below.

      Major Comments From Review at Review Commons:

      (1) Line 155: Assigning mesophyll clusters by default without any positive marker genes strikes me as problematic, especially as much of the analysis rests on comparing the transcriptomes of epidermis and mesophyll cells. Can the authors perhaps leverage published scRNA-seq datasets to find potential mesophyll markers, even homologs from other species, to improve confidence in the cluster assignment?

      (3) Line 228: Through the description and interpretation of the ChIP-seq dataset, the authors use the fact that peaks are more abundant and bigger in the epidermis to conclude that binding of PhDEF is "stronger" in the epidermis. This implies a difference in physical interaction between the TF and the DNA that I don't think can be concluded from the data presented. This could be confounded by biology; if expression of PhDEF is more heterogeneous in the mesophyll than in the epidermis, the peaks from that tissue will be averaged out and appear smaller when in fact the binding is the same strength. This could also be a technical artifact if the ChIP was less efficient in one sample versus another. This conclusion requires reinterpretation. The authors have the power to address this at least partially with the scRNA-seq by measuring PhDEF heterogeneity. I believe assessing ChIP efficiency would have required a spike in control, but perhaps there is a computational way to address this. It is important to discuss these confounding factors in the text.

      (3) Line 376: The authors risk overinterpreting a lack of differential gene expression detection in their analysis of PhDEF binding profiles. This can be affected by how deeply a library was sequenced or how many cells were analyzed per cell type. A gene might not be found to be DE if low depth or few cells resulted in noise or dropout. Lack of detection does not mean lack of differential regulation so the biological relevance of this portion of the analysis should be interpreted with caution.

      (4) Line 476: The authors state there is a mismatch in developmental timing between the RNAseq and ChIP datasets. Why is this? This is mentioned briefly in the Discussion, but has the potential to be majorly confounding to the joint interpretation of the ChIP and RNAseq datasets. This experimental design choice should be justified more thoroughly and a consideration of the limitations it brings to data interpretation should be more prominent in the text.

      Significance:

      Strengths: The authors employ a unique and powerful mutant system to explore a fundamental developmental biology question. In addition, the datasets generated will likely be useful to other researchers working in petunia or flower development.

      Limitations: While the mutant system employed here is a creative way to get at tissue-layer specific transcription factor activity, the ChIP samples still include heterogeneous cell types, which may impact the findings presented here.

      Advance: This study uses a unique system to test the function of a transcription factor in distinct cell types. As stated above, understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology.

      Audience: I believe this work will be of primary interest to the plant development, single cell, and chromatin biology communities. These are specialized, basic research communities.

      Reviewer Expertise: I am a plant developmental biologist who works with multiple modes of cell-type-specific NGS datasets including bulk and single cell RNAseq and ChIPseq among others.

      Comments on latest version:

      I find the authors' response to my review to be comprehensive and I think the revision looks solid. The eLife assessment works for me.

    4. Reviewer #3 (Public review):

      Summary:

      This study addresses a fundamental but underexplored aspect of homeotic gene function: how regulators of cell identity act during late stages of organ development. The authors take advantage of layer-specific mutants of the MADS-box gene PhDEF in Petunia hybrida to dissect the roles of this floral identity regulator in the epidermis and mesophyll. By combining single-cell RNA sequencing with ChIP-seq analyses in wild-type and mutant chimeric petals, the work demonstrates that, although PhDEF is expressed at comparable levels in both petal layers, it binds to and regulates a substantially larger and more layer-enriched set of genes in the epidermis than in the mesophyll. The identification of both layer-specific and shared PhDEF binding sites supports a model in which pre-existing layer identity modulates the regulatory output of homeotic transcription factors.

      Major comments from review at Review Commons:

      Dissecting PhDEF binding preferences in the epidermis versus the mesophyll using the star and wico mutants is a clever and powerful approach. However, conclusions involving cell identity should be drawn with caution for two reasons. First, cell identities appear to be altered in the mutants: scRNA-seq data suggest that even cells assigned to the same cluster can be molecularly distinct across genotypes. For example, the transcriptomes of wild-type epidermis and mesophyll are highly similar (Pearson correlation R = 0.94), yet both show lower correlation with epidermal cells from star or wico petals. These results raise questions such as are the cells identified as epidermal cells really strictly epidermal in the mutants? Do you need to take into consideration cell composition of the mutants when you do differential expression analysis? Second, PhDEF is expressed in both epidermal and mesophyll clusters in all genotypes, albeit at different levels and in fewer cells in the mutants. As a result, the ChIP-seq profiles should be interpreted as cell type-enriched rather than cell type-specific. And there are a few things that need clarification:

      a) Protoplasting and tissue dissection analyses suggest that mesophyll cells constitute more than 80% of the cells in wild-type petals, whereas the scRNA-seq data indicate a substantially lower proportion. Could this discrepancy reflect technical biases in cell recovery or capture efficiency, or issues related to cell identity assignment during clustering and annotation? Notably, the scRNA-seq data from star and wico petals show mesophyll proportions close to 80%. Is this difference due to an increased abundance of mesophyll cells in the mutants, or could it instead reflect differences in transcriptomic separability? In wild-type petals, the epidermal and mesophyll transcriptomes are highly correlated and express similar numbers of genes, with epidermal cells distinguished mainly by higher expression of a subset of genes. This raises the possibility that mesophyll cells in the wild type occupy a more plastic or less differentiated transcriptional state and may therefore be misclassified as epidermal cells, whereas disruption of regulatory mechanisms in the mutants enhances transcriptional divergence and alters cell clustering outcomes.

      b) Cluster 0 appears to show internal heterogeneity, as the expression patterns of KCS3 and LLE2 are largely mutually exclusive. Do these patterns reflect the presence of distinct epidermal cell types within the limb that are currently grouped into a single cluster?

      c) Clusters 0 and 7 both exhibit high expression of pigmentation genes, while cluster 7 additionally shows strong enrichment for cell division genes. Are cell cycle genes the primary features distinguishing these two clusters? If cell cycle effects are regressed out, would cluster 7 merge with cluster 0, potentially yielding a more continuous cell state trajectory and helping to resolve the pattern noted in point (b)?

      d) Pearson correlation is largely driven by highly expressed genes and may therefore be insensitive to changes in cell identity markers. The conclusions that the less clear separation of mesophyll and epidermal cells in star is due to altered cell identity would be more convincing if supported by independent validation, such as in situ hybridization or reporter analyses, to directly visualize molecular alterations in the relevant cell types when comparing wild-type and mutant tissues.

      Significance:

      Overall, this work significantly advances our understanding of late homeotic gene function. It establishes compelling evidence for how developmental context constrains transcription factor activity and offers broadly relevant insights for studies of organ patterning. The combination of genetic mosaics with single-cell and chromatin-level analyses represents a powerful and generalizable strategy that will be of interest to both plant developmental biologists and researchers studying transcriptional regulation.

      My expertise: single cell genomics, chromatin biology and plant development.

      Comments on latest version:

      The authors have addressed my concerns and questions. And I agree with your assessment of "valuable" for significance and "solid" for strength of evidence.

    5. Author response:

      General Statements

      We wish to bring to the attention of the editor and reviewers that the title of the manuscript has been modified to more accurately reflect the conclusions of our study, i.e. that PhDEF has a major binding and regulatory role in the petal epidermis.

      Point-by-point description of the revisions

      We thank the three reviewers for their critical reading of our manuscript. We also appreciate that the three reviewers highlighted the conceptual advances and the broad findings that our study brings to the community. They have identified key limitations of our analyses, and we have either provided explanations for our choice, or performed new analyses to circumvent biases. We believe that the manuscript is greatly improved and that our main conclusion, that PhDEF has a major binding and regulatory action in the petal epidermis, is strongly supported by our data.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      The authors previously generated two cell-layer-specific mutants of petunia for the petal identity gene PhDEF. In this study, they profiled differential gene expression in those mutants through single-cell RNA sequencing (scRNA-seq). They found that more genes are highly and specifically expressed in the epidermal cell layer than in mesophyll cells. In addition, they identified cell-layer-specific and -aspecific PhDEF target genes. Using the extensive single-cell transcriptome and layer-specific target identification, the authors concluded that different cell identities affect homeotic regulator PhDEF, thereby influencing transcriptional regulation.

      Major comments:

      This presented work provides comprehensive evidence, that pre-existing cell layer identity (epidermis and mesophyll) modulate transcriptional output of homeotic transcription factor, PhDEF.

      However, a disconnection between PhDEF bindings to genome and transcriptional output undermines the robustness of their conclusion although some binding loci were shown to be correlated with DEG. This may indicate the chromatin state, the existence of interacting partners, and the non-productive binding of PhDEF, suggesting that PhDEF binding alone is not sufficient to predict transcriptional outcomes and additional regulatory mechanisms that shape gene expression in addition to the layer-specific regulatory mechanisms. This disconnection may also be due to the developmental timing. Indeed, it appears authors used different flower stages for ChIP-seq and scRNA-sequencing. In fully differentiated organs, PhDEF binding itself may be no longer transcriptionally productive, and differential gene expression results primarily from the pre-established cell identity rather than directly from the homeotic regulation of PhDEF. Therefore, the main question the authors asked-how homeotic identity works with cell-layer identity and how the homeotic gene, PhDEF, acts in mature organs-was not clearly explained by this study.

      We thank the reviewer for raising this important issue. We agree that the difference in developmental timing between the scRNA-Seq and ChIP-Seq experiments might contribute to the disconnection that this reviewer pointed out.

      First, we want to explain that the reason for performing scRNA-Seq on fully differentiated petals was purely technical, as we were initially aiming to obtain protoplasts at stage 8 (stage used for the ChIP-Seq) but could never retrieve enough of them for proper encapsulation in the 10X Chromium chips. We have now clearly explained this in the manuscript (lines 104-107).

      A hypergeometric test shows that our ChIP-Seq and scRNA-Seq datasets overlap more than by chance (p = 0.000137); however, we agree that differences in developmental stages possibly bring a confounding effect to our conclusions. Therefore, we have decided to add to our manuscript the intersection between ChIP-Seq and bulk RNA-Seq data performed on WT, star and wico flowers at stage 8, that we published previously (Chopy et al., 2024). In that case, both datasets have been obtained with the same exact genetic material and at the same exact developmental stage.

      This intersection confirms that very similar binding profiles are observed for PhDEF target genes, whether they are differentially expressed in the epidermis (star only), in the mesophyll (wico only) or in both layers (star and wico) (Figure 4A). However, star-specific DEGs were more often bound by PhDEF by epidermis+shared binding sites, and wico-specific DEGs more often with mesophyll-specific binding sites, suggesting a weak but significant association between binding and regulatory profiles. Performing similar tests for individual binding categories for DEGs identified by scRNA-Seq also revealed that epidermis-specific DEGs displayed more epidermis+shared binding sites than expected. Since the association between epidermal DEGs and shared+epidermal binding sites is found both in the intersection with RNA-Seq and scRNA-Seq data, we have now stated that “layer-specific binding and transcriptional regulation are partially linked, at least in the epidermis” (line 354).

      In Figure 2, the use of the term "target" is potentially misleading. It sounds like direct target genes (direct binding and differential expression) for PhDEF, but it refers only to DEGs.

      Indeed, the term "target" was referring to both indirect and direct targets of PhDEF. To avoid any possible confusion, we have replaced it by differentially expressed genes (DEGs) throughout the manuscript, when appropriate.

      Lines 496-497: When the authors state, "~ demonstrates for the first time that the regulatory function of homeotic factor is influenced by cell layer identity," it sounds overstated, as prior studies have shown that pre-existing tissue or cell identity can shape transcriptional activity and developmental output.

      We agree that previous studies have shown that cell identity influences transcriptional activity in general. While this might not have been specifically assessed in the context of different cell layers, we have rewritten this sentence accordingly.

      Minor comments:

      In the UMAP presentation, as depicted in Figures 2C, S3, and S5, the cells with zero expression can be colored in light gray (or an inverted color scheme). The purple hue masks the gene expressions of other cells, making it difficult to see the yellow or light green colored cells.

      We have modified all UMAPs depicting gene expression as suggested, in Figures 2C and S5, and replaced UMAPs with DotPlots in Figure S3.

      Reviewer #1 (Significance):

      General assessment: This study is well-designed and technically sound. They utilize single-cell transcriptomics and ChIP-seq by using genetically well-defined genetic materials and layer-specific PhDEF deletion mutants. The analysis showed where PhDEF binds to genomic loci and which genes are differentially expressed in petal epidermis and mesophyll, providing evidence of cell-layer-specific function of homeotic gene in mature organs. Although certain mechanistic aspects were not elucidated, the data from the extensive genome-wide study contributed to drawing their conclusions.

      Advances: This research goes beyond classical models of floral organ identity by showing that homeotic gene function is not uniform in the same floral organ. It represents a conceptual advance in our understanding by integrating cell layer identity into the framework of homeotic gene regulation.

      Audience: This study will be of broad interest to scientists who study transcription networks, cell and organ identity in the context of plant development.

      My field of expertise: Transcriptional regulation by transcription factor, epigenetic regulation of gene expression, plant development

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      This study from Cavallini-Speisser et al. cleverly leverages a tissue layer-specific mutant, single cell and bulk RNA-sequencing, and ChIP-sequencing to decipher tissue layer-specific regulation of petal development in petunia by the PhDEF transcription factor. The authors find common and unique targets of PhDEF between the epidermis and mesophyll and conclude that the activity of transcription factors like PhDEF are influenced by the pre-existing environment in the cell they are expressed in. Understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology that can be elusive outside of highly tractable model systems. As such, I think this study is of high value and has strong potential to expand our understanding of how developmental specificity is mediated by commonly employed transcriptional regulators. However, I think there are some issues with possible over-interpretation and some places where documentation of experimental design and data quality control are lacking. I elaborate on these concerns below.

      Major Comments:

      Line 155: Assigning mesophyll clusters by default without any positive marker genes strikes me as problematic, especially as much of the analysis rests on comparing the transcriptomes of epidermis and mesophyll cells. Can the authors perhaps leverage published scRNA-seq datasets to find potential mesophyll markers, even homologs from other species, to improve confidence in the cluster assignment?

      We thank the reviewer for raising this important point. Our statement of defining mesophyll identity by default was not entirely true (and we have now removed it), since it is supported by GO-enriched terms for cluster markers. For cluster "mesophyll 3", there is a strong enrichment for photosynthesis-related genes; and for cluster "mesophyll 2", there is a strong enrichment for water transport-related genes, both functions being fulfilled by the petal mesophyll. For cluster "mesophyll 1", the most enriched GO term is "glutathione metabolic process" that rather points to stress response. These three clusters also do not express any of the epidermal-identity genes, in contrast to the clusters that we assigned as epidermal. Cluster markers from the mesophyll display the lowest enrichment of all clusters (the best cluster markers have a log2FC between 2.8 and 4.9, in contrast to a log2FC between 8.1 and 9.9 for all other clusters), which is in line with our finding that the mesophyll expresses less specific genes than the epidermis, and suggests a basal identity with a transcriptomic signature that is less clear than in the epidermis. Therefore, it is not entirely trivial to find positive marker genes for the mesophyll with a strong specificity.

      In order to be more transparent about the expression patterns of the genes we selected to assign cluster identity, we modified Figure S3 to include DotPlots of selected photosynthesis-related genes, histone genes, S-phase genes and vasculature genes in the same figure, to compare with DotPlots of epidermal genes and pigmentation genes from Figure 1D, to allow for an informed comparison. Our conclusions remain the same as previously: epidermal clusters are defined based on the specific expression of epidermal genes and/or pigmentation genes. We notice, however, that the cluster "upper limb epidermis" strongly expresses photosynthesis genes, which is likely why it is close in the UMAP space to the "mesophyll 3" cluster. We have no explanation for that, but the extremely high expression of pigmentation genes in this cluster, however, identifies it as epidermal without a doubt. We have also added in Figure S3 the DotPlots of expression levels of homologs of 15 epidermis-enriched and 9 mesophyll-enriched genes from tobacco petal scRNA-Seq published by Kang et al. (doi: 10.1111/nph.17992). This shows that our definition of epidermal and mesophyll clusters and the one from Kang et al. generally overlap.

      Line 228: Through the description and interpretation of the ChIP-seq dataset, the authors use the fact that peaks are more abundant and bigger in the epidermis to conclude that binding of PhDEF is "stronger" in the epidermis. This implies a difference in physical interaction between the TF and the DNA that I don't think can be concluded from the data presented. This could be confounded by biology; if expression of PhDEF is more heterogeneous in the mesophyll than in the epidermis, the peaks from that tissue will be averaged out and appear smaller when in fact the binding is the same strength. This could also be a technical artifact if the ChIP was less efficient in one sample versus another. This conclusion requires reinterpretation. The authors have the power to address this at least partially with the scRNA-seq by measuring PhDEF heterogeneity. I believe assessing ChIP efficiency would have required a spike in control, but perhaps there is a computational way to address this. It is important to discuss these confounding factors in the text.

      This is indeed another important point. We agree that the word "stronger", to describe PhDEF binding in the epidermis, was not appropriate and we have replaced it with "more frequent" which is a more factual interpretation of our results. We also agree that even this interpretation depends on potential ChIP artifacts that we have now evaluated.

      We have used our WT scRNA-Seq data to explore the heterogeneity of PhDEF expression in the mesophyll and the epidermis, as suggested. The barplot in Figure S10B represents the number of cells (y-axis) with given PhDEF RNA counts (x-axis) in the clusters that we assigned as epidermis (left) and mesophyll (right). We have performed this analysis after removing cells that do not express PhDEF at all, which represents 47.7% and 47.4% of epidermal and mesophyll cells, respectively, hence very similar proportions. The distributions of expression of PhDEF in the epidermis and in the mesophyll are within the same ranges, with a slightly higher expression of PhDEF in the mesophyll than in the epidermis. The coefficients of variation (cv) of the two distributions are similar and slightly higher in the epidermis (cv = 0.34 in the mesophyll and cv = 0.36 in the epidermis, p = 0.00106 with Feltz and Miller’s asymptotic test). Therefore, it appears that the expression of PhDEF is actually higher and slightly less variable in the mesophyll than in the epidermis, meaning that it should not result in averaging out the peaks detected. This relies on the assumption that PhDEF protein levels are directly correlated with PhDEF RNA levels, which we have not explored in this study and remains a limitation. We have included this analysis lines 275-278 and Figure S10B.

      Regarding ChIP efficiency, we had run different tests prior to sequencing: first, we have tested different amounts of chromatin, keeping the quantity of antibody unchanged, and tested ChIP enrichment by qPCR on a set of two positive (PhDEF and Pos2) and one negative (Neg1) control binding sites. Second, after selecting the best chromatin quantity, we have performed 4 independent ChIP replicates for each genotype (split between two assays named ChIP-1 and ChIP-2) and measured ChIP efficiency by qPCR. This is depicted in Figure S10A, with the replicates chosen for sequencing highlighted with an orange star.

      We have now explained in greater detail in the Methods our preliminary tests. There is indeed variation of enrichment between replicates, and particularly between assays here (ChIP-1 vs. ChIP-2), which is inherent to the ChIP experiment. It might be particularly prominent in our case due to our custom antibody directed against PhDEF, in contrast to commercial antibodies that are commonly used in ChIP experiments with tagged transgenic lines. However, we see consistently lower enrichment for star as compared to wico and WT, in line with the more frequent epidermal binding of PhDEF. We have followed the ENCODE guidelines for our analysis pipeline, in particular applying the IDR. We have now added other mapping statistics in Table S5 including the FRiP (fraction of reads in peaks) metric, that is in the range of expected values but is lower for star (around 0.5%) than for WT and wico (0.7-1.2 %), again consistent with the more frequent epidermal binding of PhDEF.

      Line 376: The authors risk overinterpreting a lack of differential gene expression detection in their analysis of PhDEF binding profiles. This can be affected by how deeply a library was sequenced or how many cells were analyzed per cell type. A gene might not be found to be DE if low depth or few cells resulted in noise or dropout. Lack of detection does not mean lack of differential regulation so the biological relevance of this portion of the analysis should be interpreted with caution.

      We have added to this new version of the manuscript an intersection between ChIP-Seq and bulk RNA-Seq in WT, star and wico, as bulk RNA-Seq is much more sensitive than scRNA-Seq in detecting lowly expressed genes. We have also added the sentence that bulk RNA-Seq "better captures lowly expressed genes" than scRNA-Seq, line 333. This intersection revealed an association between epidermal DEGs (star-specific DEGs) and the presence of epidermal+shared binding sites for PhDEF. We have modified our conclusions accordingly.

      Line 476: The authors state there is a mismatch in developmental timing between the RNAseq and ChIP datasets. Why is this? This is mentioned briefly in the Discussion, but has the potential to be majorly confounding to the joint interpretation of the ChIP and RNAseq datasets. This experimental design choice should be justified more thoroughly and a consideration of the limitations it brings to data interpretation should be more prominent in the text.

      This concern has also been raised by the first reviewer, and we have now added to our study an intersection between ChIP-Seq and bulk RNA-Seq performed at the same stage. We have also explained the technical reasons for performing scRNA-Seq at a mature stage only (lines 104-107). Indeed, the intersection between bulk RNA-Seq and ChIP-Seq performed at the same stage revealed a significant association between epidermal DEGs (star-specific DEGs) and the presence of epidermal+shared binding sites for PhDEF. Testing for individiual binding categories, we could also detect an enrichment of epidermal+shared binding sites for epidermal-specific DEGs identified by scRNA-Seq. Therefore, we have now stated that “layer-specific binding and transcriptional regulation are partially linked, at least in the epidermis” (line 354).

      Minor Comments:

      Line 229: The authors compare correlations between pseudo-bulked transcriptomes and argue that in the star mutant the epidermis adopts a mesophyll-like identity. The correlation between epidermis and mesophyll in star is 0.97 and the correlations were 0.94 and 0.93 in the other genotypes tested. What is the meaningful cutoff for saying the transcriptomes are similar or not? Is 0.97 so much higher than 0.94 that this conclusion is supported?

      The comparison of pseudo-bulk transcriptomes is a very global and exploratory approach. Given the high number of genes underlying these pseudo-bulk datasets, any difference in the correlation between them is statistically significant, which is not very informative. We agree with this reviewer that the interpretation of these correlation coefficients is somewhat arbitrary. We have simplified this part of the manuscript and have mostly focused on comparing star and wico pseudo-bulk transcriptomes to the WT ones, but not to each other's, which aligns well with our main message of a specific epidermal identity, easily shifting to a mesophyll-identity when PhDEF is missing or not entirely functional. We have also removed Figure 2E to a supplementary figure to give less emphasis to this analysis.

      Line 264: What are the "manually chosen thresholds for differential expression"? Can the authors explain and justify this? There is very little detail on this in the materials and methods and this raises some concerns regarding how a threshold was chosen.

      Seurat gives a default threshold of 0.25 for log2FC, which we found to be very permissive. In order to capture the most informative targets of PhDEF, but still to capture a meaningful number of targets, we empirically decided to increase this threshold to 0.75. On the WT scRNA-Seq dataset, we observed that the layer-specificity factor was also capturing meaningful differences in layer-specific expression (see Figure 1E). We chose a cut-off at 10% since it was the lowest to give a significant difference in the numbers of epidermis- vs. mesophyll-enriched genes in the WT petal. We have added these explanations in the methods.

      Figure 3G: Could this plot be annotated with the classification of peak layer specificity? It is a little difficult for me as the reader to keep up with all the categories in the text, and showing them in the figure might make that easier to follow.

      We have now annotated the Venn diagram in Figure 3G with the classifications of binding profiles.

      Figure 4A: A comparison of only two cell categories should not use scaled expression, as this can over-emphasize small differences in gene expression. Can this be replaced with a dot plot that uses unscaled expression values?

      We thank the reviewer for noticing this issue, we have built a DotPlot with unscaled values and replaced it in Figure 4A, which does not change our conclusions. We have also used unscaled values in Figure S14.

      Figure 4C: Could this be represented more legibly with stacked, space-filled bar charts? As is, this is difficult to read, and might be impossible for someone who is color blind. In addition, the authors claim this analysis shows similar proportions across all categories, but I wonder if that would hold true if they performed an over-representation analysis normalized to the categories shown in "all genes expressed". This could allow them to statistically test whether there are real differences in representation among the categories.

      We have now used a different and color-blind-friendly palette for pie charts of PhDEF binding profiles. Following reviewers' comments, we have analyzed the intersection of bulk RNA-Seq with ChIP-Seq (both performed at the same stage), and performed Chi2 goodness-of-fit tests that indeed support some association between binding profile and regulation profile, although this remains limited.

      Supplemental Figure 2: It's great that the authors include these metrics, but it would be ideal to also include the plots of standard QC metrics for scRNA-seq such as those found here to give a better sense of per cell quality: https://satijalab.org/seurat/articles/pbmc3k_tutorial

      In addition, it is important to include QC metrics for ChIPseq, which I did not find in the supplement. Metrics such as FRiP are important for interpreting ChIP library quality.

      We have now added the standard QC plots (Feature number per cell and RNA counts per cell) to Figure S2, as suggested. Mitochondrial and ribosomal genes are not annotated in the Petunia axillaris nuclear genome that we used, therefore we could not compute mitochondrial or ribosomal gene counts. We have removed cells expressing less than 200 genes, but did not apply any upper thresholds as there were no obvious outliers.

      FastQC reports for scRNA-Seq and ChIP-Seq have been deposited at https://entrepot.recherche.data.gouv.fr/dataverse/PhDEF_Flower_layer, as indicated in the Methods.

      We have also added ChIP metrics in Table S5, including the number of reads, duplicated reads, mapped reads and computed the FriP score. This score ranges between 0.5 % and 1.2 %, which is satisfactory and above the minimum recommended score  by ENCODE of 0.3 %.

      Line 654: What model was used for DESeq2?

      We have used default parameters for DESeq2, ie a negative binomial GLM fitting and Wald significance tests. We have added this information in the Methods.

      Line 752: Can the authors justify why peaks were called separately on input and ChIP samples rather than allowing MACS2 to call peaks in the ChIP sample over input background? That differs from the standard MACS2 pipeline and no explanation for this is provided in the text.

      Our analysis pipeline has indeed been customized, in particular to detect peaks in our positive control PhDEF, for which a binding site of PhDEF in the promoter has been demonstrated experimentally by others in many different species. This binding has a strong experimental support, and we expected PhDEF to bind to its own promoter in the two cell layers. Our ChIP-Seq results show a posteriori that this particular peak is far from being the strongest one over the genome; therefore, we believe it represents a good control to detect binding enrichment for average targets. We have first tried the standard MACS2 pipeline that calculates the enrichment of IP over Input, but we could only detect PhDEF binding for one WT and one wico replicate, although the peaks were visually clear in the two replicates. Our input sample being generally noisy, we explored how separate peak calling between IP and Input would behave (as already done in e.g. Durand et al., 2023, doi: 10.1093/plcell/koad025). We also explored how thresholds for FDR in MACS2, and IDR thresholds for reproducibility between IP samples, would influence peak detection in IP and Input.

      We found that calling peaks on the IP with a relaxed FDR (0.1), then applying the IDR at 0.1, allowed the capture of PhDEF binding to its own promoter in the two wico replicates (but still not in WT, because the peaks were lost after applying the IDR threshold). No peak was detected in the input with these settings, however for other genes we observed spurious peak detection in the input, therefore we decided to lower the FDR thresholds for input peaks to 0.05. Our choices have been made in an attempt to increase specificity at the risk of losing sensibility, and we probably lose true binding events. Considering that this ChIP has been performed on the endogenous PhDEF protein directly, and in chimeric flowers that only express PhDEF in half of the tissue, adapting the ChIP analysis pipeline was a necessary step.

      We have now added a few lines in the Methods (lines 696-706) to explain our rationale.

      Reviewer #2 (Significance):

      General Assessment:

      Strengths: The authors employ a unique and powerful mutant system to explore a fundamental developmental biology question. In addition, the datasets generated will likely be useful to other researchers working in petunia or flower development.

      Limitations: While the mutant system employed here is a creative way to get at tissue-layer specific transcription factor activity, the ChIP samples still include heterogeneous cell types, which may impact the findings presented here.

      Advance: This study uses a unique system to test the function of a transcription factor in distinct cell types. As stated above, understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology.

      Audience: I believe this work will be of primary interest to the plant development, single cell, and chromatin biology communities. These are specialized, basic research communities.

      Reviewer Expertise: I am a plant developmental biologist who works with multiple modes of cell-type-specific NGS datasets including bulk and single cell RNAseq and ChIPseq among others.

      Reviewer #3 (Evidence, reproducibility and clarity):

      Summary:

      This study addresses a fundamental but underexplored aspect of homeotic gene function: how regulators of cell identity act during late stages of organ development. The authors take advantage of layer-specific mutants of the MADS-box gene PhDEF in Petunia hybrida to dissect the roles of this floral identity regulator in the epidermis and mesophyll. By combining single-cell RNA sequencing with ChIP-seq analyses in wild-type and mutant chimeric petals, the work demonstrates that, although PhDEF is expressed at comparable levels in both petal layers, it binds to and regulates a substantially larger and more layer-enriched set of genes in the epidermis than in the mesophyll. The identification of both layer-specific and shared PhDEF binding sites supports a model in which pre-existing layer identity modulates the regulatory output of homeotic transcription factors.

      Major comments:

      Dissecting PhDEF binding preferences in the epidermis versus the mesophyll using the star and wico mutants is a clever and powerful approach. However, conclusions involving cell identity should be drawn with caution for two reasons. First, cell identities appear to be altered in the mutants: scRNA-seq data suggest that even cells assigned to the same cluster can be molecularly distinct across genotypes. For example, the transcriptomes of wild-type epidermis and mesophyll are highly similar (Pearson correlation R = 0.94), yet both show lower correlation with epidermal cells from star or wico petals. These results raise questions such as are the cells identified as epidermal cells really strictly epidermal in the mutants? Do you need to take into consideration cell composition of the mutants when you do differential expression analysis?Second, PhDEF is expressed in both epidermal and mesophyll clusters in all genotypes, albeit at different levels and in fewer cells in the mutants. As a result, the ChIP-seq profiles should be interpreted as cell type-enriched rather than cell type-specific.

      We fully agree with this reviewer's comments, and alterations in cell identity in the star and wico mutants is indeed a main pitfall for our analysis. The phdef mutation alters the transcriptomic signatures of cells physically located in the epidermis or in the mesophyll, which can result in their artificial clustering with cells located elsewhere in the petal. It might be particularly true for the epidermis, since we see strong depletion in epidermal clusters in the star flowers (and even in the wico flowers), whereas the mesophyll is not much affected in wico flowers. As a result, it is likely that we lose many epidermal cells in star flowers that end up labeled as mesophyll cells, resulting in the under-estimation of DEGs in this layer. We had explored other possible ways to define epidermal or mesophyll cells in our dataset, based for instance on PhDEF or PhGLO1 expression, but since for the majority of cells PhDEF expression is simply not captured, we would have wrongly assigned phdef mutant identity to WT cells. In spite of this limitation, we find more DEGs in the epidermis than in the mesophyll, showing that even an underestimation of DEGs in the epidermis does not affect our main conclusion, which is that PhDEF has a major regulatory action in the epidermis. We have explicitly written this limitation in our manuscript, lines 234-236.

      We agree that ChIP-Seq profiles are rather cell type-enriched than cell type-specific, which we have stated explicitly in the sentence line 308, saying that "differential binding between layers is quantitative". However, for simplicity we prefer to retain the term "specific" since our conclusions are based on the definition of peaks that can either be present or absent in a given genotype, and hence in a given cell layer.

      And there are a few things that need clarification:

      a) Protoplasting and tissue dissection analyses suggest that mesophyll cells constitute more than 80% of the cells in wild-type petals, whereas the scRNA-seq data indicate a substantially lower proportion. Could this discrepancy reflect technical biases in cell recovery or capture efficiency, or issues related to cell identity assignment during clustering and annotation? Notably, the scRNA-seq data from star and wico petals show mesophyll proportions close to 80%. Is this difference due to an increased abundance of mesophyll cells in the mutants, or could it instead reflect differences in transcriptomic separability? In wild-type petals, the epidermal and mesophyll transcriptomes are highly correlated and express similar numbers of genes, with epidermal cells distinguished mainly by higher expression of a subset of genes. This raises the possibility that mesophyll cells in the wild type occupy a more plastic or less differentiated transcriptional state and may therefore be misclassified as epidermal cells, whereas disruption of regulatory mechanisms in the mutants enhances transcriptional divergence and alters cell clustering outcomes.

      Our protoplasting and tissue sections show that the mesophyll should represent 80% of cells in WT tissue, while we estimate it at 70% in our WT scRNA-Seq data based on our cluster assignment (mesophyll = 60% + vasculature = 10%, that we separated from the mesophyll cells but is actually embedded within this tissue). This is not a very strong difference, and considering the multiple steps that protoplasting and cell capture entail, we considered that this was reasonably close to the expected proportions.

      In the star and wico flowers, we assign mesophyll identity to a greater proportion of cells, but as explained above, we believe that cells with altered epidermal identity are easily clustered as mesophyll cells, since they lose their specific transcriptomic signature. This is actually in line with one of the main messages of our article, that petal epidermis transcriptional identity is highly specific.

      b) Cluster 0 appears to show internal heterogeneity, as the expression patterns of KCS3 and LLE2 are largely mutually exclusive. Do these patterns reflect the presence of distinct epidermal cell types within the limb that are currently grouped into a single cluster?

      Indeed, there appears to be some internal heterogeneity within cluster 0. Since our main focus was to compare epidermal and mesophyll clusters, we did not explore further the heterogeneity within epidermal clusters and kept a coarse resolution.

      c) Clusters 0 and 7 both exhibit high expression of pigmentation genes, while cluster 7 additionally shows strong enrichment for cell division genes. Are cell cycle genes the primary features distinguishing these two clusters? If cell cycle effects are regressed out, would cluster 7 merge with cluster 0, potentially yielding a more continuous cell state trajectory and helping to resolve the pattern noted in point (b)?

      To answer one of Reviewer 1's comments, we have added additional DotPlots to better describe our clusters, in Figure S3. Clusters 0 (limb epidermis), 6 (upper limb epidermis) and 7 (replicating cells) exhibit high expression of pigmentation genes (in particular cluster 6), as depicted in Figure 1D. Cluster 7, consisting of only 30 cells, is the only cluster expressing histone genes and S-phase genes. It is possible that these few cells are cluster-6 cells that are replicating, but considering the very low number of cells involved, we have not explored any further their identity and decided to remove them, as it would only marginally affect the conclusions of our analyses. It is indeed surprising that cluster 6 (upper limb epidermis) is quite distinct in the UMAP space to the other epidermal clusters 0 (limb epidermis) and 4 (upper tube epidermis), and we have not observed similar situations in other scRNA-Seq studies. We speculate that this is due to the joint expression of anthocyanin-related and photosynthesis-related genes, which convey a very strong transcriptomic signature to these cells that distinguish them from other epidermal cells.

      d) Pearson correlation is largely driven by highly expressed genes and may therefore be insensitive to changes in cell identity markers. The conclusions that the less clear separation of mesophyll and epidermal cells in star is due to altered cell identity would be more convincing if supported by independent validation, such as in situ hybridization or reporter analyses, to directly visualize molecular alterations in the relevant cell types when comparing wild-type and mutant tissues.

      Following another reviewer's comments, we have now given less emphasis on the Pearson correlation analysis, that was to some extent subjective. Therefore, our conclusion that mesophyll and epidermal cells in star are less separated than in WT has been removed.

      Minor comments:

      (1) For color-coded figure legends (e.g., Fig. 1C), please also include the cluster numbers. This would facilitate interpretation, particularly for readers with reduced color sensitivity.

      We have now used colour-blind-friendly palettes and we have added cluster numbers in Figure 1C.

      (2) For figures containing abbreviations (e.g., Fig. 1D, st. / ca.), please explicitly define all abbreviations in the figure legend.

      We have defined all abbreviations in the figure legends.

      (3) For all UMAP figures, and for figures involving comparisons across clusters, please use consistent color schemes for the same clusters throughout the manuscript.

      We have modified color schemes across the manuscript for colour-blind-friendly palettes, consistently used throughout the manuscript.

      (4) In Fig. S6, the plot showing all genes does not exactly match Fig. 1C, although it appears to represent the same data. Please use the same version of packages, seed values and parameters for all UMAP plots to avoid such discrepancies.

      Figure S6 is the result of integrating with Harmony the WT dataset (as shown in Figure 1C) with WT datasets after removing genes differentially expressed by the protoplasting process. Therefore these UMAPs are a result of integrating different datasets than in Figure 1C, which changes the shape of the UMAPs but cannot be controlled by the seed values, to our knowledge.

      Reviewer #3 (Significance):

      Overall, this work significantly advances our understanding of late homeotic gene function. It establishes compelling evidence for how developmental context constrains transcription factor activity and offers broadly relevant insights for studies of organ patterning. The combination of genetic mosaics with single-cell and chromatin-level analyses represents a powerful and generalizable strategy that will be of interest to both plant developmental biologists and researchers studying transcriptional regulation.

      My expertise: single cell genomics, chromatin biology and plant development.

    1. eLife Assessment

      This important study uses the Cntnap2 mouse model of autism to investigate how reduced activity in the dorsal CA1 region of the hippocampus contributes to impaired temporal binding and memory flexibility. The authors combined trace conditioning and radial maze assays with fibre photometry, optogenetic rescue, and brain-wide c-Fos mapping to provide convincing evidence that dorsal CA1 hypoactivity limits the retention of temporally separated associations and biases learning towards less flexible strategies. The work advances understanding of hippocampal contributions to cognitive alterations associated with autism and will be of broad interest to researchers studying memory, hippocampal function, and neurodevelopmental disorders.

    2. Reviewer #2 (Public review):

      The authors investigate the contribution of dorsal CA1 hippocampal dysfunction to cognitive impairments in the Cntnap2 knockout mouse model of autism spectrum disorder. Using two complementary behavioral paradigms, trace fear conditioning and a relational/declarative memory radial maze task, together with fiber photometry, optogenetic manipulation, and cFos mapping, they examine whether altered CA1 function contributes to deficits in temporal binding and memory flexibility.

      A major strength of the study is the combination of behavioral, recording, and causal manipulation approaches. The trace fear conditioning experiments show that Cntnap2 knockout mice retain associations across shorter temporal intervals but fail when the temporal gap is increased to 40 s. Fiber photometry reveals reduced dorsal CA1 activity under these conditions, and, importantly, optogenetic activation of dorsal CA1 pyramidal neurons during the trace interval rescues subsequent memory performance. This provides compelling evidence for a causal contribution of dorsal CA1 activity to the temporal binding deficit observed in this model.

      The radial maze experiments extend these findings to a more complex form of relational/declarative memory. Cntnap2 knockout mice are able to acquire the task but show impaired flexibility when previously learned spatial relations must be recombined. Their behavior is also characterized by greater lateralization, consistent with increased reliance on an egocentric rather than an allocentric spatial strategy. The revised manuscript now explains more clearly how this paradigm distinguishes relational/declarative from procedural learning strategies and how the 20-s inter-trial interval introduces a temporal binding requirement. This clarification substantially improves the conceptual link between the two behavioral paradigms.

      The accompanying cFos analyses further show reduced recruitment of hippocampal regions and increased engagement of striatal and prefrontal regions in Cntnap2 knockout mice. These findings are consistent with altered recruitment of memory systems accompanying the behavioral strategy shift. However, unlike the optogenetic experiments in the trace fear conditioning paradigm, the cFos measurements remain correlational. They therefore support, but do not by themselves establish, a causal reorganization from hippocampal-dependent declarative memory toward striatum-dependent procedural learning. The authors have appropriately moderated this interpretation in much of the revised manuscript, although some statements, particularly in the Abstract, could still be phrased more cautiously.

      The revised manuscript also addresses the potential influence of the well-described hyperactivity of Cntnap2 knockout mice. The distinction between hyperactivity and impulsive-like behavior is now more explicitly discussed. In the trace fear conditioning experiments, comparable resting and freezing behavior under control conditions argues against locomotor activity accounting for the memory phenotype. In the radial maze, shorter decision latencies combined with longer post-choice running times are more consistent with reduced deliberation than with a simple increase in locomotor activity. Importantly, the authors acknowledge that the task was not specifically designed to measure impulsivity.

      Overall, the revision has strengthened the manuscript and addressed my main concerns. The principal conclusion-that reduced dorsal CA1 activity contributes causally to impaired temporal binding in Cntnap2 knockout mice-is well supported by the converging behavioral, photometric, and optogenetic evidence. The broader proposal that hippocampal dysfunction is accompanied by greater reliance on egocentric/striatal learning strategies is also supported by the behavioral and cFos data, provided that it is interpreted as an association rather than a demonstrated causal reorganization of memory systems.

      The study makes a valuable contribution by linking hippocampal circuit dysfunction to specific components of declarative memory in a widely used model of autism. More broadly, the combination of temporal binding and memory flexibility paradigms provides a useful framework for investigating how hippocampal dysfunction may alter the organization and flexible expression of memory in neurodevelopmental disorders.

    3. Reviewer #3 (Public review):

      Summary:

      The manuscript evaluated behavioral phenotypes in the Cntnap2 knockout mouse using two behavioral paradigms: trace fear conditioning and a radial maze task. The trace fear conditioning training is normal, but memory generalization is impaired. The inflexibility is suggested to be related to low activity in dCA1 neurons, which can be rescued by ChR2. The radial maze task data suggested a similar conclusion. Brain-wide cFos mapping indicated impairments in the Cntnap2 knockout mouse. The brain-wide cFos mapping does not show direct correlations with Cntnap2, limiting the interpretation of these data in the context of this paper.

      Strengths:

      The behavior data are solid.

      Comment on revised version.

      The authors have addressed all my concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The uniqueness of this paper is the study of the formation of temporal binding-dependent memories in the cntnap2 mouse, a long-standing mouse model of autism that has been used to test therapeutic modalities.

      Strengths:

      I liked the combination of optical recordings and interventions and the backup of primary observations with control experiments.

      Weaknesses:

      (1) Fiber photometry recordings are too coarse to give salient clues to the underlying mechanism.

      We acknowledge that fiber photometry provides population-level measurements and does not resolve the activity of individual neurons or synaptic mechanisms. Our aim was to identify alterations in the activity of defined neuronal populations during behaviour rather than to delineate the underlying cellular mechanisms. We agree that future studies employing higher-resolution approaches, such as two-photon calcium imaging, in vivo electrophysiology, or single-cell recordings, would provide important mechanistic insights into the circuit changes underlying the observed activity patterns.

      (2) Are perturbed pyramidal cells causally responsible for the altered trace? What can be concluded about the possible role of inhibitory interneurons as potential drivers? The observations focus on abnormal regional activity as observed with fiber photometry and manipulated by optogenetics. The authors should state clearly the limits of their conclusions.

      Our optogenetic experiments demonstrate a causal role for the targeted neuronal population in modulating the observed activity and behavioural phenotype. However, we do not conclude that this population is solely responsible for generating the altered fiber photometry signal, nor do we infer that it represents the exclusive driver of the underlying circuit dysfunction. Rather, our findings demonstrate that manipulating this population is sufficient to alter circuit activity, while acknowledging that the recorded signals likely reflect interactions between multiple neuronal populations. We have revised the manuscript to make these distinctions clearer.

      (3) I found the "trace" nomenclature confusing. "....in which mice are required to memorize the association between a tone (Conditioned Stimulus) and a mild electric foot-shock (Unconditioned Stimulus), separated by a time interval called Trace (Sellami et al., 2017)." It seems that the conceptual model invokes the creation of an [eligibility] trace, characterized by its progressive disappearance over time. It may be a convention in the field or a matter of language, but it seems perverse to use "trace" to label the time interval rather than the entity that is decaying. If this is an accepted convention going back to Howard Eichenbaum, the authors should cite the paper that first introduced the convention.

      We thank the reviewer for this comment. This is indeed an established convention in the classical/Pavlovian conditioning literature rather than terminology specific to our study or to Sellami et al. (2017). The term "trace conditioning" was coined by Pavlov (1927), who used "trace interval" to designate the empty period separating CS offset from US onset, precisely because — as the reviewer intuits — successful conditioning across this gap requires the organism to maintain a memory trace of the CS. In other words, the interval is named for the cognitive/neural entity (the decaying CS trace) that must be sustained across it, not because the interval itself is thought to be a physical or decaying object. So "trace interval" is shorthand for "the interval across which a trace must be maintained," analogous to how "delay conditioning" refers to a paradigm named for a temporal property of the procedure rather than the mechanism per se. We have added a citation to Pavlov (1927) at first use of the term, and clarified the phrasing to make the etymology explicit, as suggested.

      (4) I would advocate for the addition of some discussion points for the authors to consider.

      (a) Is the retention of activity in CA1 related to phenomena at the cellular or subcellular level in CA1 pyramidal cells? I'm thinking of dendritic, delayed, and stochastic CaMKII activation (DDSC) as defined by Yasuda's group or short-term and associative plasticity of calcium dynamics (STAPCD) as delineated by Caya-Bissonette and Beique.

      We thank the reviewer for this insightful suggestion. Our study was designed to investigate network-level dynamics, and the approaches used do not allow us to determine whether the retained activity in CA1 arises from intracellular mechanisms, such as dendritic calcium dynamics or CaMKII-dependent signalling (including mechanisms such as DDSC or STAPCD), recurrent circuit interactions, or a combination of both. We therefore cannot directly assess the contribution of these cellular and subcellular processes. We have now added a paragraph to the Discussion acknowledging that persistent dendritic calcium signalling and CaMKII-dependent plasticity are plausible contributors to sustained CA1 activity and represent an important avenue for future investigation.

      (b) Was the optogenetic intervention ever administered in a delayed fashion, capitalizing on the temporal advantages of optogenetics to probe dynamics?

      This is indeed an important control. While we did not perform this intervention in the current study, this control was part of our seminal study demonstrating the causal role of the dCA1 in temporal binding (Sellami et al., PNAS, 2017, Fig. 1E doi: https://doi.org/10.1073/pnas.161965711). We showed that aged mice had reduced temporal binding capacity, associated with decreased dCA1 activity. We successfully rescued temporal binding capacity in aged mice through ChR2-induced activation of dCA1 pyramidal neurons during the trace interval, but not outside of the trace interval. Given the striking similarity between the temporal binding deficits observed in aged mice and those reported here in Cntnap2 KO mice, we did not repeat this previously established control in the present study. To acknowledge this point, we have also added a statement to the Discussion noting that confirming the temporal specificity of dCA1 optogenetic manipulation in the Cntnap2 KO model will be an important direction for future studies.

      (c) Is the newfound reliance on corticostriatal pathways something more than compensation at the behavioral level? Could it be driven in part by the ASD-related genetic changes?

      We agree that the increased reliance on corticostriatal pathways could reflect both compensatory recruitment at the behavioral level and a direct consequence of ASD-related genetic alterations affecting circuit development and function. Our current data demonstrate a shift in circuit engagement but do not allow us to distinguish whether this represents an adaptive compensation or a primary consequence of Cntnap2 deficiency. However, given that ASD-associated mutations can alter the development, connectivity, and plasticity of corticostriatal circuits, it is possible that the observed changes reflect intrinsic circuit reorganization rather than solely a compensatory behavioral strategy. We have revised the Discussion to acknowledge this possibility and to clarify that altered corticostriatal recruitment may represent a direct consequence of the genetic disruption that subsequently shapes behavioral strategies.

      Reviewer #2 (Public review):

      The authors investigate the contribution of dorsal CA1 hippocampal dysfunction to cognitive impairments in the Cntnap2 knockout mouse model of autism spectrum disorder. Building on previous evidence implicating the hippocampus in episodic and relational memory processes, they combine trace fear conditioning, fiber photometry, optogenetic manipulation, a relational/declarative memory radial maze task, and cFos mapping to test whether altered CA1 function contributes to deficits in temporal binding and memory flexibility.

      The study has several important strengths. First, the work addresses a relatively understudied aspect of autism-related cognition, namely hippocampal-dependent memory processes, whereas much of the literature has focused on social behavior, cortical circuits, or striatal dysfunction. Second, the authors employ multiple complementary approaches that converge on a coherent mechanistic hypothesis. The behavioral data demonstrate a reduced ability of Cntnap2 knockout mice to retain associations across long temporal gaps. Fiber photometry recordings reveal reduced dorsal CA1 activity during conditions that challenge temporal binding, and optogenetic activation of dorsal CA1 neurons during the trace interval is sufficient to rescue memory performance. Together, these findings provide strong support for a causal contribution of dorsal CA1 activity to temporal binding deficits in this model.

      The second major strength of the manuscript is the extension of these findings to a more complex hippocampus-dependent memory task. The radial maze experiments indicate that Cntnap2 knockout mice show impaired memory flexibility and a greater reliance on egocentric learning strategies. The accompanying cFos analyses suggest altered recruitment of hippocampal and striatal networks during learning, providing a systems-level framework that may explain the observed behavioral phenotype.

      Overall, the main conclusions regarding impaired temporal binding and reduced dorsal CA1 engagement are well supported by the data. The optogenetic rescue experiments are particularly compelling because they move beyond correlation and directly test causality. The manuscript therefore makes a meaningful contribution to our understanding of how hippocampal dysfunction may contribute to cognitive abnormalities associated with autism.

      Weaknesses:

      Some conclusions are necessarily more inferential than others. In particular, the interpretation that the observed behavioral phenotype reflects a broader shift from hippocampal-dependent declarative memory toward striatum-dependent procedural learning is supported primarily by cFos activity patterns and behavioral strategy measures. While the data are consistent with this interpretation, they do not directly demonstrate a causal reorganization of memory systems. Similarly, although the findings identify a mechanism in the Cntnap2 model, caution is warranted when extrapolating these conclusions to autism spectrum disorder more broadly; but I believe this caution is addressed in the discussion.

      We agree that our data do not directly demonstrate a causal reorganization of memory systems with cFOS activity patterns. Our intention was to propose that the behavioural strategy together with the brain-wide cFos activation patterns are consistent with a shift in the relative engagement of hippocampal- and striatal-dependent networks. We have revised the manuscript to moderate our interpretation throughout, replacing causal language with wording that reflects an association between the observed behavioural changes and altered recruitment of these memory-related circuits.

      Despite these limitations, the study presents a coherent and well-executed body of work that provides novel mechanistic insight into hippocampal contributions to cognitive dysfunction in a widely used autism model. The findings should be of considerable interest to researchers studying hippocampal function, memory systems, and neurodevelopmental disorders.

      Reviewer #3 (Public review):

      Summary:

      The manuscript evaluated behavioral phenotypes in the Cntnap2 knockout mouse using two behavioral paradigms: trace fear conditioning and a radial maze task. The trace fear conditioning training is normal, but memory generalization is impaired. The inflexibility is suggested to be related to low activity in dCA1 neurons, which can be rescued by ChR2. The radial maze task data suggested a similar conclusion. Brain-wide cFos mapping indicated impairments in the Cntnap2 knockout mouse. The brain-wide cFos mapping does not show direct correlations with Cntnap2, limiting the interpretation of these data in the context of this paper.

      We agree that brain-wide cFos mapping does not directly identify the molecular or cellular mechanisms by which Cntnap2 deficiency alters circuit function. Rather, we use cFos as a functional readout of network recruitment during behaviour. Our interpretation is therefore limited to identifying differences in activity patterns associated with the behavioural phenotype, rather than establishing direct mechanistic links to Cntnap2 function. We have clarified this point in the revised manuscript and moderated the text accordingly.

      Strengths:

      The behavior data are solid.

      Weaknesses:

      The underlying mechanism is not fully investigated.

      Major points:

      (1) The authors should thoroughly check their manuscript as there are many typos in the current version that affect the readability.

      (2) In trace fear conditioning, the tone test impairment can be rescued by ChR2. Have the authors tried rescue experiments with Cntnap2? Rescue experiments in the radial maze task are also essential, either with ChR2 or Cntnap2.

      We agree that rescue experiments in the radial maze would provide additional mechanistic insight by testing whether restoring dCA1 activity optogenetically in Cntnap2 KO mice is sufficient to rescue declarative memory flexibility. However, we have already established the causal role of dCA1 in temporal binding and relational/declarative memory in the radial maze task in our earlier work (Sellami et al., PNAS, 2017, Figure 2E). Indeed, optogenetic inhibition of dCA1 in the radial maze task, specifically during the inter-trial interval, prevented flexible relational/declarative memory formation. In the present study, we confirmed the causal contribution of dCA1 activity to temporal binding in Cntnap2 KO mice in the trace fear conditioning paradigm (Figure 1J-L). Therefore, repeating the optogenetic experiment in the radial maze would provide only limited additional information, while requiring dedicated experimental cohorts and additional validation. Moreover, as the team is based in 2 different institutions and countries (France and Australia), we unfortunately do not have the capacity to perform these experiments.

      Regarding a rescue with the Cntnap2 protein, here, we used the Cntnap2 KO as a well-accepted model of autism spectrum disorder (ASD), rather than understanding the role of this protein in temporal binding capacity and memory formation. The present study was hence designed to determine how dCA1 activity relates to temporal binding and memory performance in ASD, building on the causal evidence established in our previous work. We therefore consider Cntnap2-specific rescue experiments an important follow-up to further test the sufficiency of dCA1 activation, rather than an essential experiment for establishing the conclusions of the present study.

      We have highlighted these points as an important direction for future investigations and revised the manuscript to better distinguish our findings from this additional mechanistic question.

      (3) The quality of the cFos example image in Figure 3 is too low. The authors should also provide example images for the other brain regions in the supplementary data, if possible.

      The low quality of the cFos image in Figure 3 resulted from a formatting issue during figure uploading. This has now been corrected, and a higher-resolution image has been included in the revised manuscript. We have also added representative images for the other brain regions analyzed to the Supplementary Information, as requested.

      (4) The causal link between the brain-wide cFos mapping and the Cntnap2 knockout is weak. How to explain the increase of cFos cell densities in some brain regions, but the decrease in others?

      We agree that cFos mapping cannot explain the mechanisms underlying the regional increases and decreases in activity. We interpret these bidirectional changes as reflecting differential recruitment of distributed brain networks rather than direct effects of Cntnap2 deficiency on individual brain regions, and have clarified this in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Improve clarity of the Figure layout: Figure 1, panel I legend. I think the identification of experimental groups is in the wrong place. It really applies to panels K, N, and L, not to panel I. This is important because the genetic stimulation in panels K, N, and L is the key experiment in this whole figure.

      We agree that the placement of the experimental group identification in the original figure legend could be confusing. We have revised the figure layout and legend so that the description of the experimental groups is now associated with panels K and L, where the optogenetic stimulation experiments are presented, thereby improving the clarity of the figure.

      (2) Fix Figure references. "This indicates impaired retention of trace fear memory and temporal binding ability, relative to WT mice (Figure 2I)....: [there is no Fig. 2I, so I'm guessing you mean Fig. 1I; subsequent Figure references also seem to confuse Fig. 1 and 2].

      We thank the reviewer for identifying these errors. We have carefully checked and corrected all figure citations throughout the manuscript, including the reference to Figure 1I-L in this section, to ensure that each citation now refers to the appropriate panel.

      (3) Missing words? In both T20 and T40 conditions, regardless of genotype, [the response to] tone frequency significantly decreased.

      We thank the reviewer for spotting this omission. The sentence has been corrected to read: "In both T20 and T40 conditions, regardless of genotype, the response to tone frequency significantly decreased."

      (4) Jargon that should be declared at an appropriately early point. R/DM = relational/declarative memory.

      We have now introduced the term as relational/declarative memory (R/DM) at its first appearance in the manuscript and have ensured that the abbreviation is used consistently thereafter.

      (5) Confusing statement? "In the 40s trace-conditioned group, while the frequency of the calcium transients did not differ between WT and Cntnap2 KO mice throughout conditioning (Figure S1G), their amplitude was significantly reduced both during the presentation of the second tone and consistently across the three trace intervals in Cntnap2 KO mice (Figure 1I)."

      We have revised this sentence to improve its clarity and explicitly distinguish between the frequency and amplitude of calcium transients. The revised text now makes clear that, although the frequency of calcium transients did not differ between genotypes throughout conditioning, their amplitude was significantly reduced in Cntnap2 KO mice during the second tone presentation and across the three trace intervals.

      Reviewer #2 (Recommendations for the authors):

      (1) The manuscript would benefit from a clearer distinction between conclusions directly supported by the data and broader interpretations. In particular, statements suggesting a shift from declarative to procedural memory systems could be presented more cautiously, as the cFos analyses provide indirect rather than causal evidence for such reorganization.

      We thank the reviewer for this constructive comment. We have revised the manuscript throughout to more clearly distinguish conclusions that are directly supported by our data from broader interpretations. In particular, statements referring to a shift from declarative to procedural memory systems have been moderated.

      (2) Additional clarification of the behavioral interpretation of the radial maze task would be valuable for readers who are less familiar with this paradigm, particularly regarding its relationship to relational/declarative memory and its distinction from procedural learning strategies. In addition, the results' interpretation in this paradigm is unclear to me. What is the significance of the 20-second delay in this task (why not use 10 sec or 60 sec)? I believe the paradigm tests spatial rather than temporal distant items (in contrast with trace fear conditioning that refers to temporally distant events)? Please explain further the link between the two behavioral paradigms.

      We thank the reviewer for these comments.

      Regarding the relationship between relational/declarative and procedural learning: The R/DM task was designed by our group (Marighetto and colleagues; Mingaud et al., 2007; Sellami et al., 2017, 2018) specifically to dissociate two learning/memory systems that can support the same behavioral output (correct arm choice) but that differ fundamentally in their underlying representations and, critically, in their flexibility.

      During acquisition, an animal can solve each of the three arm-pair discriminations either by forming a flexible, relational representation of the whole spatial configuration (i.e. an allocentric/hippocampus-dependent "cognitive map" strategy, in which the reward's location is encoded relative to distal cues and to the other pairs) or by learning a set of rigid, response-based rules (i.e. an egocentric/striatum-dependent "turn left/turn right" procedural strategy tied to each specific pair).

      Both strategies can produce equivalent accuracy during initial acquisition, which is why acquisition performance alone cannot distinguish them. The flexibility probe (recombining pairs A and B into a novel pair AB, without moving the reward) is the critical dissociation: only an animal that encoded the reward's location relationally, within a broader spatial map, can generalize correctly to this untrained configuration; an animal that relied on a rigid stimulus– response rule for each pair individually would be unable to solve the recombined pair above chance. Performance on pair AB is therefore the read-out of hippocampus-dependent relational/declarative memory, while the degree of left–right lateralization during acquisition (quantified by a lateralization index) is a converging behavioral signature of reliance on the egocentric/procedural system.

      We have clarified this distinction in the Methods and Results to facilitate interpretation of the paradigm.

      Regarding the 20-s delay and its relationship to trace fear conditioning: The reviewer is correct that the R/DM task and TFC differ in the information being integrated: the R/DM task involves spatially distinct discrimination episodes, whereas TFC involves temporally separated stimuli presented in the same context. However, both tasks require information separated by an interval to be integrated into a unified memory representation, a process referred to here as “temporal binding” (Sellami et al., 2017). The 20-s inter-trial interval in the R/DM task was based on our previous work, in which this interval was shown to be sensitive to age-related deficits in temporal binding and relational memory (Sellami et al., 2017). Similarly, our TFC experiments established a longer temporal-binding capacity in young mice (up to 40 s) compared with aged mice (20 s). Thus, although the two paradigms involve different types of information, they share the requirement to maintain and integrate information across a temporal gap. The different intervals used in the two paradigms reflect their distinct task structures and demands.

      (3) Although the authors address potential locomotor confounds, a brief discussion of how hyperactivity and impulsivity might influence behavioral performance in the different paradigms would strengthen the interpretation of the results.

      We thank the reviewer for raising this point. We agree that distinguishing hyperactivity from impulsivity strengthens the interpretation of our behavioral phenotypes, and we have added a discussion paragraph making this distinction explicit. Briefly: in the TFC paradigm, our existing open-field data (Figure S1D) show that while Cntnap2 KO mice travel more distance and move faster, their resting time is unaffected relative to WT; so, the reduced freezing/temporal binding phenotype at the 40 s trace cannot be attributed to a general locomotor confound, since freezing (immobility) was actually comparable between genotypes during acquisition. In the R/DM radial maze task, however, we agree the phenotype is better read as impulsivity than hyperactivity per se: KO mice show shorter decision latencies specifically at the choice point (Figure 2G), with no corresponding genotype effect on their post-choice running speed for correct trials (Figure 2H), indicating that the effect is on deliberation before response rather than on general motor output. We have revised the Discussion to make this distinction, and its implications for interpreting the egocentric/procedural bias, explicit.

      (4) Minor presentation issues:

      - Ensure that all statistical tests, sample sizes, and post hoc comparisons are reported consistently throughout the manuscript and figure legends.

      We have carefully reviewed the statistical reporting throughout the manuscript and figure legends to ensure consistency. We have clarified sample sizes, statistical tests, and post hoc comparisons where required and in the methods.

      - Carefully proofread the manuscript for minor typographical and grammatical errors. For instance, a) several figures numbers indicated in the main text do not correspond to the figures the authors refers to (page 5 Figure 2I, page 6 Figure 2J and 2K, page 7 "data in fig S2"... and so on); b) rephrase (not clear) page 9: "Differences in the relative contribution of individual substructures between trained WT and Cntnap2 KO mice involved dCA1, dCA2, and dCA3, and favored the DMS, IL, and ACC."

      We thank the reviewer for highlighting these issues. We have carefully proofread the manuscript and corrected all figure reference errors and typographical inconsistencies throughout the text. We have also revised the unclear sentence on page 9 to improve clarity and accuracy.

      - Page 9, an additional sentence is needed to explain the significance of the bibliographical reference: "Atypicalities in declarative memory have been reported in ASD, with specific impairments in relating items (Minor et al., 2023)."

      We have revised this section to clarify the significance of the cited study and its relevance to our findings. Specifically, we have expanded the sentence to highlight that impairments in relational processing in ASD may reflect difficulties in integrating individual experiences into coherent memory representations, a process that relies on hippocampal function.

      Reviewer #3 (Recommendations for the authors):

      Major Points:

      As stated in the public review, the authors need to thoroughly check their manuscript before submission. There are too many typos in the current version that affect the readability. For example, but not limited to:

      (1) Page 5, "Figure 2I" should be "Figure 1I"; "Figure 2J" should be "Figure 1J"; "Figure 2K" should be "Figure 1K".

      (2) The figure legends of Figure 1J-L are missing.

      (3) Page 8, "In contrast, Cntnap2 KO mice showed no learning-induced dCA1 activation together with hypoactivation of dCA3 in both naïve and trained groups compared to WT mice (Figure S2A)", should it be dCA2 instead of dCA1? Please double-check the manuscript.

      We thank the reviewer for highlighting these issues. We have thoroughly checked the manuscript and corrected all typos, figure reference errors, and fixed figure legend information. Regarding the point (3), we did mean to describe the lack of training-induced activation of dCA1, but failed to reference back to Figure 3, likely causing the confusion. This has now been clarified. The manuscript has also been carefully proofread to improve clarity and readability.

      Minor points:

      There are two duplicate rows (two wt) in their raw datasheet, Fig3BCD & S2 and Fig3EF. Please correct them in case of further problems.

      WT Naive 178.736 103.821 250.703 185.229 437.575 50.535 266.254 306.187 306.187 WT Naive 178.736 103.821 250.703 185.229 437.575 50.535 266.254 306.187 306.187

      WT Naive 46.87 37.94 15.19 8.57 4.98 12.02 8.88 20.98 2.42 12.77 14.68 14.68 WT Naive 46.87 37.94 15.19 8.57 4.98 12.02 8.88 20.98 2.42 12.77 14.68 14.68

      We thank the reviewer for noticing this mistake. These duplicated rows have been corrected.

    1. eLife Assessment

      This study provides valuable evidence that glycogen phosphorylase is unlikely to represent an effective insecticidal target in Plutella xylostella and that diflubenzuron does not directly inhibit this enzyme. The combination of biochemical, molecular and physiological approaches provides convincing data to support these conclusions, although the proposed metabolic compensation mechanism is supported primarily by indirect evidence and would benefit from direct demonstration.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      In this study, the authors investigate whether glycogen phosphorylase represents a molecular target of benzoylphenylurea insecticides and evaluate the physiological consequences of suppressing glycogen phosphorylase activity in the diamondback moth Plutella xylostella. The authors combine recombinant protein biochemistry, enzyme inhibition assays, RNA interference, structural modelling, metabolite profiling, gene expression analyses, and physiological measurements to determine whether diflubenzuron directly inhibits glycogen phosphorylase and whether suppression of this enzyme is sufficient to impair insect development. Based on these experiments, the authors conclude that diflubenzuron does not directly inhibit glycogen phosphorylase and that insects tolerate substantial suppression of this enzyme through compensatory metabolic responses.

      Strengths:

      This study addresses an important question in insect toxicology by systematically evaluating glycogen phosphorylase as a potential insecticidal target. The authors combine complementary biochemical, molecular, physiological, and structural approaches, including recombinant enzyme characterization, inhibitor assays, RNA interference, metabolite profiling, structural modelling, and measurements of fitness-related traits. This integrative approach provides a comprehensive evaluation of the biological consequences of glycogen phosphorylase suppression. In particular, the biochemical evidence that diflubenzuron does not inhibit glycogen phosphorylase, together with the observation that strong suppression of glycogen phosphorylase produces only transient physiological effects without measurable impacts on development or reproduction, provides strong support for the conclusion that glycogen phosphorylase is unlikely to represent an effective standalone insecticidal target.

      Weaknesses:

      The main limitation concerns the proposed mechanism underlying metabolic compensation. The observed increases in gluconeogenic gene expression, changes in metabolite abundance, and reductions in total protein are consistent with activation of compensatory metabolism, but are insufficient to directly demonstrate increased gluconeogenic flux or establish that amino acid-derived carbon is incorporated into newly synthesized glucose. Similarly, although the analyses of glycogen-associated enzymes strengthen the discussion of alternative metabolic pathways, changes in gene expression alone do not demonstrate that these pathways contribute to glycogen utilization in vivo.

      Some mechanistic interpretations therefore extend beyond the data presented. For example, decreases in total protein are interpreted as evidence of protein catabolism fuelling gluconeogenesis, yet they do not directly demonstrate amino acid mobilization or incorporation into glucose. Likewise, increased expression of gluconeogenic genes is interpreted as evidence of increased pathway activity, although transcriptional changes do not necessarily reflect metabolic flux. Finally, the absence of major developmental defects following glycogen phosphorylase suppression is attributed primarily to metabolic compensation, but an alternative explanation is not fully considered. Such explanation could be that glycogen phosphorylase is not rate-limiting for glucose homeostasis under the nutrient-rich experimental conditions, where dietary carbohydrates are continuously available. Consequently, the proposed compensatory mechanism remains plausible and well supported by indirect evidence, but several aspects would benefit from more cautious interpretation.

      Overall, the authors successfully achieve their primary objective of evaluating glycogen phosphorylase as a candidate insecticidal target. The study provides useful biochemical and physiological evidence that this enzyme is unlikely to represent an effective target for insecticide development in P. xylostella, while highlighting the importance of metabolic plasticity when assessing metabolic targets. The experimental approaches and datasets presented here should be valuable to researchers studying insect metabolism, insecticide mode of action, and target validation, although the precise mechanisms underlying the proposed metabolic compensation remain an important subject for future investigation.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This study addresses an important question in insect toxicology by systematically evaluating glycogen phosphorylase as a potential insecticidal target. The authors combine complementary biochemical, molecular, physiological, and structural approaches, including recombinant enzyme characterization, inhibitor assays, RNA interference, metabolite profiling, structural modelling, and measurements of fitness-related traits. This integrative approach provides a comprehensive evaluation of the biological consequences of glycogen phosphorylase suppression. In particular, the biochemical evidence that diflubenzuron does not inhibit glycogen phosphorylase, together with the observation that strong suppression of glycogen phosphorylase produces only transient physiological effects without measurable impacts on development or reproduction, provides strong support for the conclusion that glycogen phosphorylase is unlikely to represent an effective standalone insecticidal target.

      We thank the reviewer for the positive assessment of our study and for recognizing the value of our integrative approach, including recombinant enzyme characterization, RNAi, metabolite profiling, structural modelling, and fitness measurements. We are also grateful for the acknowledgement that our biochemical evidence and phenotypic observations provide strong support for the conclusion that GP is unlikely to be an effective standalone insecticidal target.

      Weaknesses:

      (1) The proposed metabolic compensation mechanism is supported by indirect evidence.

      We agree with the reviewer that our data demonstrate correlation with, rather than direct proof of, increased gluconeogenic flux. We have revised the manuscript throughout to moderate our interpretations and to clearly frame the compensation model as a plausible interpretation supported by multiple lines of indirect evidence, rather than an established mechanism. Specific revisions are detailed below.

      Specific revisions:

      (1) Abstract (Lines 35–39):

      “We provide evidence that insects compensate through a multi-layered metabolic response: upregulation of gluconeogenic enzymes (PEPCK, G-6-Pase), selective upregulation of glycogen branching enzyme (GBE) but not α-amylase, and changes in protein content suggestive of catabolic substrate mobilization”

      (2) Abstract (Lines 41–42):

      “Indicating that the metabolic compensation response is ultimately effective in sustaining development”

      (3) Abstract (Lines 42–44):

      “These findings suggest that GP is functionally non-essential under these conditions, likely through a combination of gluconeogenic compensation and the availability of alternative carbon sources”

      (4) Introduction (Lines 88–90):

      “Our findings reveal that PxGP is functionally non-essential for larval development under the tested conditions, reflecting a previously uncharacterized gluconeogenic compensation mechanism”

      (5) Results – Gene expression (Line 265):

      “Gene expression analysis is consistent with gluconeogenic activation”

      (6) Results – Protein content (Line 278):

      “Changes in protein content suggestive of catabolic substrate mobilization”

      (7) Results – Protein decline (Line 283):

      “This protein decline is consistent with substrate mobilization”

      (8) Results – GBE expression (Line 389):

      “PxGP knockdown selectively upregulates glycogen branching enzyme expression”

      (9) Results – GBE differential response (Lines 400–403):

      “This differential expression pattern—selective GBE upregulation with unchanged α-amylase expression—indicates that the compensatory response to GP suppression involves targeted remodeling of glycogen structure rather than a generalized upregulation of all glycogen-degrading enzymes.”

      (10) Results – Fitness assessment (Lines 411–412):

      “To determine whether the metabolic changes observed following PxGP knockdown are associated with measurable physiological consequences”

      (11) Results – Fitness interpretation (Lines 426–430):

      “GP suppression likely triggers protein catabolism to supply amino acids for gluconeogenesis, contributing to transient weight loss. In the presence of continuous dietary carbohydrate supply, the compensatory response appears sufficiently effective to restore metabolic homeostasis before developmentally critical transitions”

      (12) Discussion – Glycogen accumulation paradox (Line 520-521):

      “a paradox that could be explained by the activation of GP-independent glycogen catabolism via alternative enzymes like...”

      (13) Results – Trehalose and G6P (Lines 299–303, 315–317):

      “Trehalose levels remained stable at 48-72 h but increased substantially by 96 h (7.44-fold elevation, P < 0.05) (Figure 9F). This increase coincided with the significant upregulation of gluconeogenic genes (e.g., PEPCK and G-6-Pase), consistent with the idea that gluconeogenesis contributes to trehalose synthesis and ensures adequate carbohydrate reserves for the upcoming pupation”

      And: “Together, the coordinated behavior of G6P and trehalose is consistent with gluconeogenesis-derived glucose being converted into storage and transport carbohydrates”

      (14) Results – Integrated interpretation (Lines 330–332):

      “Collectively, these metabolite and gene expression data reveal a biphasic metabolic adaptation that provides a coherent explanation for why substantial GP suppression causes no mortality or developmental defects”

      (15) Discussion – Definitive proof (Lines 527–531):

      “At 96 h, trehalose and G6P levels in dsGP-treated larvae were maintained at markedly higher levels than in dsGFP controls, which had declined by this time point, coinciding with 3–4‑fold upregulation of PEPCK and G‑6‑Pase. These changes strongly support, albeit indirectly, de novo glucose synthesis as the primary rescue mechanism.”

      (16) Discussion – Protein catabolism confirmation (Lines 505):

      “provides independent biochemical support for protein catabolism”

      (2) Some mechanistic interpretations extend beyond the data presented.

      We accept this criticism and have systematically revised the manuscript to distinguish more carefully between observed transcriptional/metabolic changes and functional interpretations. We now consistently frame these as correlative evidence consistent with—but not proving—the proposed mechanisms.

      Specific revisions:

      The revisions listed under Weakness 1 above—particularly those modifying language around protein decline (Lines 278, 283), gene expression (Lines 265, 389), and trehalose/G6P changes (Lines 299–303, 314–317)—also directly address this weakness by removing the implication that transcriptional or protein-level changes constitute functional proof of pathway activation. Additionally, we have made the following revisions:

      (17) Additional discussion of GBE/α-amylase limitation (Line 405-408)

      We have added explicit acknowledgement that transcriptional changes alone do not demonstrate functional pathway activation:

      “However, as these observations are limited to the transcript level, further enzymatic activity assays or glycogen structure analyses would be required to determine whether this transcriptional change translates into functional alterations in glycogen mobilization.”

      (18) Fitness section – clarification (Lines 426–430):

      We have revised the interpretation of transient larval weight loss as described above, introducing “likely” and clarifying that the data are consistent with the model rather than conclusively demonstrating it.

      (3) An alternative explanation for the limited phenotype is not fully considered.

      We thank the reviewer for this important point. We agree that the continuous availability of dietary carbohydrates under our experimental conditions represents a plausible alternative explanation for the lack of phenotype. We have added explicit discussion of this alternative interpretation in the Abstract, Introduction, and Discussion sections.

      Specific revisions:

      (19) Abstract (Lines 42–44):

      As shown above, we have revised the abstract to include: “These findings suggest that GP is functionally non-essential under these conditions, likely through a combination of gluconeogenic compensation and the availability of alternative carbon sources.”

      (20) Introduction (Lines 88–90):

      As shown above, we have revised the introduction to include: “under the tested conditions” and to acknowledge that the observed phenotype reflects a combination of factors.

      (21) Discussion – Biphasic response (Lines 531–539):

      We have revised the discussion of the biphasic response to moderate the causal language and acknowledge alternative interpretations:

      “This biphasic response—initial metabolic stress followed by delayed (48-72 h) and robust compensation—creates a temporal buffer by 96 h and is consistent with the absence of overt phenotypes despite substantial GP suppression. This temporal coordination raises the possibility of a developmentally programmed response, potentially involving hormonal regulation, that anticipates energy demands at critical transitions. Critically, this pattern suggests a potential vulnerability window (~72 h) during which combined inhibition of glycogenolysis and gluconeogenesis might overcome compensation, although this remains speculative and would require experimental validation.”

      (22) Discussion – Alternative explanation paragraph (Lines 540–547):

      We have added a following paragraph explicitly addressing the alternative explanation:

      “We also acknowledge that the absence of a severe phenotype may reflect an additional, non‑mutually exclusive explanation: under our experimental conditions, with continuous dietary carbohydrate availability, GP activity may not be rate‑limiting for maintaining glucose homeostasis. The observed transient larval weight loss could thus reflect both active metabolic compensation and the inherently low demand on GP for glucose supply in a feeding larva. Distinguishing between active compensation and the non‑limiting nature of GP will require future studies under nutrient‑restricted conditions or during fasting intervals, where GP's role is likely to become more critical.”

      Additional revisions to moderate language throughout the manuscript

      Beyond the specific revisions addressing each weakness, we have also made the following broader language modifications to ensure consistent and cautious interpretation throughout:

      (23) Changes from “ablation” to “suppression” (Lines 180, 195–197, 357):

      “Ablation” to “suppression” throughout.

      (24) Conclusions – Fundamental principle (Lines 619):

      “our findings uncover a fundamental principle of...”

      (25) Conclusions – Strategic path (Lines 621–623):

      “This metabolic plasticity renders GP non-viable as a standalone insecticidal target but illuminates a potential strategic path forward”

      (26) Discussion – Trehalase inhibitors (Lines 556–558):

      “For example, trehalase inhibitors such as Validamycin A are potent insecticides [44, 45], suggesting that trehalase inhibition may not be subject to the same degree of metabolic compensation”

      (27) Discussion – Metabolic pincer (Line 561):

      “a 'metabolic pincer' attack that might overcome adaptive compensation”

      (28) MM/GBSA and structural predictions (Lines 27–32, 450, 455–457):

      Abstract (Lines 27–32)

      “Molecular docking and MM/GBSA analysis predict that this selectivity reflects differential side-chain engagement: GPI is predicted to occupy the allosteric site at the dimer interface via contacts with seven residues spanning both subunits (ΔG = −34.63 kcal/mol), whereas DFB's difluorobenzoyl moiety is predicted to remain solvent-exposed without productive protein contacts (ΔG = −29.29 kcal/mol).”

      Line 450

      “MM/GBSA analysis indicated”

      Lines 455–457

      “These structural predictions are consistent with established structure–activity relationships of acyl urea compounds [15], in which target selectivity is governed by the side-chain substitution pattern rather than by the shared acyl urea core”

      (29) Scope of conclusion (Lines 471–472):

      “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for diflubenzuron”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major recommendations

      (1) Moderate the interpretation of the proposed compensatory mechanism throughout the manuscript.

      We fully agree with the reviewer's assessment. Throughout the manuscript, we have systematically moderated the language used to describe the compensatory mechanism, ensuring that our interpretations are framed as plausible inferences supported by multiple lines of indirect evidence rather than as established conclusions. The specific revisions are detailed in items (1)–(7), items (10)–(16) of our response to Weakness 1 above, and items (24)–(25) of our response to Weakness 3 above.

      Additional language modifications addressing the reviewer's concern about assertive terminology are detailed in our response to Minor Recommendation 1 below.

      Line 495

      “Our investigation shows that this tolerance reflects a robust”

      Line 498

      “As insects elicited a compensatory gluconeogenic pathway.”

      Line 518

      “Direct metabolite quantification reflected this compensation”

      (2) Clarify the evidence supporting alternative glycogen degradation pathways.

      We thank the reviewer for this valuable suggestion. We have revised the GBE and α-amylase sections to more clearly distinguish between transcriptional changes and functional pathway activation. The revised wording now explicitly acknowledges that our observations are limited to the transcript level and that functional confirmation would require additional experiments. The specific revisions are detailed in items (8)–(9) and item (12) of our response to Weakness 1 above. We have also added a sentence acknowledging that these observations are limited to the transcript level (item 17) of our response to Weakness 2 above.

      (3) Discuss alternative explanations for the limited physiological phenotype.

      We thank the reviewer for this important point. We agree that the continuous availability of dietary carbohydrates under our experimental conditions represents a plausible alternative explanation for the lack of phenotype. We have incorporated this alternative interpretation into the Abstract, Introduction, and Discussion sections, as detailed in items (3)–(4) of our response to Weakness 1 above, and items (21)–(22) of our response to Weakness 3 above.

      We have also revised the subsequent conclusion (Lines 553–554) to align with this alternative interpretation, changing “GP fails as a standalone target due to compensation via gluconeogenesis” to “GP appears to fail as a standalone target, at least in part because of compensation via gluconeogenesis.” This ensures consistency with the preceding discussion of GP’s potentially non‑rate‑limiting role.

      Minor recommendations

      (1) Review the manuscript for statements using terms such as "demonstrates," "confirms," "proves," or "establishes."

      We have systematically reviewed the entire manuscript and replaced overly assertive terms (e.g., “demonstrates,” “confirms,” “proves,” “establishes”) with more cautious language (e.g., “suggests,” “is consistent with,” “provides evidence that,” “indicates”) wherever they refer to the proposed metabolic mechanism. Specific revisions are listed under Major Recommendation 1 above. In addition:

      MM/GBSA and structural predictions (Lines 27–32, 450, 455–457): Abstract (Lines 27–32): Changed “Molecular docking and MM/GBSA analysis reveal that...” to “Molecular docking and MM/GBSA analysis predict that...”

      Results (Line 450): Changed “MM/GBSA analysis confirmed” to “MM/GBSA analysis indicated.”

      Results (Lines 455–457): Changed “These structural data confirm that...” to “These structural predictions are consistent with...”

      Conclusion scope (Lines 471–472): Changed “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for BPUs” to “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for diflubenzuron.”

      Changes from “ablation” to “suppression” (Lines 180, 195–197, 357): Changed “ablation” to “suppression” throughout the manuscript where referring to GP knockdown.

      (2) Ensure that the distinction between changes in transcript abundance, enzyme activity, and metabolic flux is maintained consistently throughout the Results and Discussion.

      We have carefully reviewed the Results and Discussion sections to ensure that we consistently distinguish between transcript abundance (measured by RT-qPCR), enzyme activity (measured by activity assays), and inferred metabolic flux (not directly measured). The revisions listed under Major Recommendations 1 and 2 above directly address this point. Key changes include:

      Consistently using “transcriptional upregulation” or “expression” when referring to qPCR data, rather than “activation” or “pathway activity.”

      Explicitly acknowledging that transcriptional changes do not necessarily reflect metabolic flux.

      Using “is consistent with” rather than “demonstrates” when linking gene expression changes to functional outcomes.

      Additional clarifications and corrections to data presentation

      During the preparation of the revised manuscript, we also made several corrections and clarifications to the data presentation:

      RNAi knockdown efficiency (Line 364–367): We have added a clarification that the knockdown efficiency in the cohort used for enzyme activity and fitness assays (54.66% at 48 h) was lower than that achieved in the dose–response experiment (87.59% at 48 h), reflecting batch-to-batch variation between independently injected cohorts. We have also noted that the enzyme-activity data should be interpreted against the transcript reduction measured in this same cohort.

      Total protein data comparability (Lines 374–380): We have added a note clarifying that the total-protein data shown in Figure 9D and Figure 10–figure supplement 2 derive from independent experimental cohorts processed at different homogenization ratios; absolute protein concentrations are therefore not directly comparable between the two panels. Both datasets nonetheless show a transient decline of approximately 30% in total protein in dsGP-treated larvae within the first 72 h.

      Trehalose and G6P data interpretation (Lines 304–309, 314–315, 345–353): We have revised the description of trehalose and G6P levels at 96 h to clarify that the large fold-differences primarily reflect a pronounced decline in dsGFP control values at this time point, rather than a net increase in absolute metabolite content in dsGP-treated larvae. The revised wording now indicates that GP-suppressed larvae maintain their trehalose and G6P pools at a stage when control larvae are actively depleting them.

      Gene expression recovery kinetics (Lines 251–252): We have corrected the description of PxTre and PxHex expression at 96 h to state that they “returned to, and modestly exceeded, control levels” rather than “returned to near baseline levels.”

      Cell line description update (Lines 644–654): In response to editorial requirements, we have updated the cell line description to include supplier authentication, mycoplasma testing status, and confirmation that the cells were used experimentally within one year of purchase.

      Figure corrections:

      Figure 2 legend: revised to state “maximum inhibition did not exceed 55.07 ± 7.05%” to match the main text.

      Figure 4 legend: corrected normalization description from “control set to 1.0 at each time point” to “calibrated to the 24 h control sample (set to 1.0).”

      Figure 6 legend: corrected reference sample from “L1 set to 1.0” to “Egg set to 1.0.”

      Figure 8D: corrected y-axis label from “Adult emergency (%)” to “Adult emergence (%).”

      Figure 9C: corrected in-figure label from “G-6-P” to “G-6-Pase.”

      Terminology correction: We have standardized the use of “similarity” (rather than “identity”) when referring to sequence comparisons (Lines 190, 1075).

      PROSITE motif name correction (Lines 181–184): We have corrected “phosphatase-pyridoxal phosphate linkage site” to “phosphorylase pyridoxal-phosphate attachment site.”

      Figure 11 statistical description (Lines 1138–1142): We have revised the figure legend to specify that data in (B) and (C) are mean ± SEM of three biological replicates, while data in (D–F) are shown as box plots with sample sizes indicated, and the statistical comparison is dsGP vs. dsGFP (independent samples t‑test). These revisions do not affect any results or conclusions.

      Gene expression peak description (Line 301): Changed “coincided temporally with peak expression of gluconeogenic gene” to “coincided with the significant upregulation of gluconeogenic genes” to avoid implying a single defined peak.

      Minor textual corrections

      We have also corrected the following textual issues:

      Sentence structure (Lines 433–435): We have revised the sentence structure to correct a comma splice and clarify the logical relationship. The original sentence “To provide structural insight into the observed selectivity, GPI potently inhibits PxGP (IC<sub>50</sub> = 2.96 nM) while DFB does not, we performed...” has been revised to “To provide structural insight into the observed selectivity—in which GPI potently inhibits PxGP (IC<sub>50</sub> = 2.96 nM) while DFB does not—we performed...”

      Additional minor corrections: We have corrected a small number of typographical errors and stylistic inconsistencies throughout the manuscript.

      All of these corrections are limited to data presentation, figure labeling, and textual clarity. They do not alter any experimental results, quantitative conclusions, or the overall interpretation of the study.

    1. eLife Assessment

      The paper presents a valuable finding that the human brain and models that incorporate sentence structures can capture sentence-level semantics beyond word meaning, while large language models behave differently. The evidence supporting the authors' claims is solid, though the stimuli are highly controlled and some analyses could be more thorough. This work will be of interest to researchers in language neuroscience and those developing language models.

    2. Reviewer #3 (Public review):

      Summary:

      Large Language Models have revolutionized Artificial Intelligence and can now match or surpass human language abilities on many tasks. This has fueled interest in cognitive neuroscience in exposing representational similarities between Language Models and brain recordings of language comprehension. The current study breaks from this mold by: (1) Systematically identifying sentence structures for which brain and Large Language Model representations diverge. (2) Using models structured by semantic roles to help account for divergences. As such the study may now fuel interest in characterizing how Large Language Models and brain representations differ, which may prompt new more brain like language models.

      Strengths:

      (1) This study challenges a literature trend that has touted similarities between Transformer models and human cognition based on representational correlations with brain activity. This challenge is substantiated by identifying sentences for which brain and model representations of sentences diverge.

      (2) This study conducts a rigorous pre-registered analysis of a comprehensive selection of the state-of-the-art Large Language Models, on a controlled sentence comprehension fMRI dataset. The analysis is conducted within a Representation Similarity framework to support similarity comparisons between graph structures and brain activity without needing to vectorize graphs. Transformer language models are predicted and shown to diverge from brain representations on subsets of sentences with similar word-level content but different sentence structures.

      (3) The study introduces a 7T fMRI sentence comprehension dataset and accompanying human sentence similarity ratings which may be a fruitful resource for developing more human-like language models. Unlike other model-based sentence datasets, the relation between grammatical structure and word-level content is controlled, and subsets of sentences for which models and brains diverge are identified.

      Weaknesses:

      (1) The interpretation of findings is nuanced. Although Transformers underperform as brain models on the critical subsets of controlled sentences, a Transformer outperforms all other models when evaluated on the union of all sentences when both word-level content and structure vary. Transformers also yield equivalent or better models of human behavioral data. Thus, although Transformers have demonstrable flaws as human models which are pinpointed here, when tested on all sentences (some) Transformers are more human-like than the other models considered.

      (2) There may be confounds between the critical sentence structure manipulations and visual processing. This is inconvenient because activation in brain regions that process semantics tends to partially correlate with low-level representations of sentence surface features encoded in visual cortex. Although the study commendably controls for confounds associated with sentence length, correlations with key sentence structure models are most salient in visual cortex and diminish in other brain networks when V1-V4 activation is controlled for.

      (3) Sentence similarity computations are emphasized as the basis for unifying comparative analyses of graph structures and vector data. A strength of this approach is that correlation is not always the ideal similarity metric. However, a weakness is that similarity computations are not unified across models. This has practical consequences because different similarity metrics applied to the same model produce positive or negative correlations with brain data.

      Comments on revised version.

      Thanks again for the responses. In particular, the new sentence-length control is helpful for interpreting the anomalous outcomes from the DIEM analysis, and the amendments to the discussion are appreciated.