7,226 Matching Annotations
  1. Last 7 days
    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors elegantly combined latent variable models (i.e., HMM, GPFA and dynamical system models) with a calcium imaging observation model (i.e., latent Poisson spiking and autoregressive calcium dynamics (AR)).

      Strengths:

      Integrating a calcium observation model into existing latent variable models improves significantly the inference of latent neural states compared to existing approaches such as spike deconvolution or Gaussian assumptions.

      The authors also provide an open-source access to their method for direct application to calcium imaging data analysis.

      Weaknesses:

      As acknowledged by the authors, their method is dependent on the quality of calcium trace extraction from fluorescence videos. It should be noted that this limitation applies to alternative strategies.

      While the contribution of this study should prove useful for researchers using calcium imaging, the novelty is limited, as it consists of an integration of the calcium imaging model from Ganmor et al. 2016 with existing LVM frameworks.

      Reviewer #2 (Public review):

      Summary:

      This compelling study proposes a framework to implement latent variable models using population level calcium imaging data. The study incorporates autoregressive dynamics and latent Poisson spiking to improve inference of latent states across different model classes including HMMs, Gaussian Process Factor Analysis and nonlinear dynamical systems models. This approach allows for a more seamless integration of existing methods typically used with spiking data to apply on calcium imaging data. The authors test the model on piriform cortex recordings as well as a biophysical simulator to validate their methods. This approach promises to have wide usability for neuroscientists using large population level calcium imaging.

      Strengths:

      The strengths of this study are the flexibility in the choice of models and relatively easy adaptation to user-specific use cases.

      Weaknesses:

      The weakness of the study lies in its limited validation of biological calcium imaging data. Calcium dynamics in a task-specific context in a sensory brain region might be very different from slower dynamics in a region of integration. The biophysical properties of the data would also be dependent on the SNR of the imaging platform and the generation of calcium indicator being used.

      Reviewers 1 and 2 correctly point out that our method depends on the quality of the upstream calcium trace extraction. As they noted, these traces can vary based on the specific indicator used, the signal-to-noise ratio (SNR) of the recording, and the region of the brain being imaged (such as sensory vs. integrative areas). As Reviewer 1 rightly mentions, this is a universal challenge that applies to alternative strategies as well, rather than a limitation unique to our framework. To address these points, we have added a paragraph to our Discussion section.

      Reviewer #3 (Public review):

      Summary:

      S. Keeley & collaborators propose a computational approach to infer time-varying latent variables directly from calcium traces (for instance, obtained with 2p imaging) without the need for deconvolving the traces into spike trains in a preliminary, independent step. Their approach rests on 1 of 3 families of latent models: GPFA, HMM and dynamical systems - which they augment with an observation model that maps latent variables to fluorescence traces. They validate their approach on simulated and real data, showing that the approach improves latent variable inference and model fitting, compared to more traditional approaches (although not directly compared with the 2-step one; see below). They provide a GitHub repository with code to fit their models (which I have not tested).

      Strengths:

      The approach is sound and well-motivated. The authors are specialists in latent variable models. The manuscript is succinct, well-written, and the figures are clear. I particularly liked the diversity of latent models considered, in particular latent models with continuous (GPFA) vs.

      discrete (HMM) dynamics, which are useful for characterizing different types of neural computations. The validation on both simulated and real data is convincing.

      Weaknesses:

      One advantage … is that one can inspect the quality of the deconvolution step independently from the latent variable inference step. For instance, if the inferred latent variables are not interpretable, how can one determine whether this is due to a poor choice of latent model (e.g., HMM with too few states), or a poor fit of the observation model (e.g., wrong parameters for the calcium dynamics)?

      We agree with the reviewer that integrating the calcium likelihood introduces additional parameters that require careful diagnostics. However, for the vast majority of imaging datasets, there is no simultaneous electrophysiology to verify the deconvolution step. If the final latent states are not interpretable, it remains impossible to determine whether the error originated in the initial spike inference from deconvolution or the subsequent model fitting.

      Our framework addresses this by maintaining the raw fluorescence as the fixed observation. We suggest for those using this model to employ cross-validation using this data to select model parameters. We outline how to do this below, but because the data itself does not change with each model fit, you can compare P(data | λ) across any model configuration. In contrast, different deconvolution methods change the data itself (the spiketimes) making comparison across models impossible.

      Could the authors comment on whether their approach allows for instance to compare different forms of latent models (e.g., HMM vs. GPFA) in terms of model evidence, cross-validated log-likelihood or other model comparison metrics?

      We thank the reviewer for highlighting this. In short: yes. Because our framework integrates the calcium observation likelihood with various latent variable models, we can assess held-out prediction P(data | λ) irrespective of the specific latent structure.

      However, because fitting the LVM requires inferring the latent state z to determine the firing rate λ, proper cross-validation across models involves holding out both neurons and timepoints. A principled approach—which our framework supports—is as follows:

      (1) Train both the latent states z and the model parameters (e.g., the mapping from latent space to observations) on a training portion of the recording.

      (2) On a held-out test segment, withhold a subset of "test" neurons and infer the latent states using only the "held-in" neurons.

      (3) Calculate the likelihood of the observed fluorescence for the test neurons given the inferred rates.

      We clarify this procedure in the revised manuscript. While a comprehensive benchmarking across all possible LVM architectures is beyond the scope of this study, we provide the statistical infrastructure for users to perform such comparisons. Furthermore, we would like to emphasize that while predictive likelihood is a rigorous metric for model selection, the primary utility of these LVMs often lies in the interpretability of the latent states themselves, which can remain biologically informative even if cross-validated performance is not the sole optimization target.

      While it certainly makes sense that models accounting for the full transformation of latent => spikes => fluorescence data should outperform the two-step (1) deconvolution => (2) latent variance inference approach, the amount of improvement is not clear. A direct comparison … would be useful

      We thank the reviewer for this point. Figure 4 was designed to address this comparison directly. By using a biophysical simulator, we generated a pseudo-realistic spiking network with ground-truth latent trajectories governed by a Gaussian Process. This allowed us to explicitly compare our unified approach against the traditional deconvolution-then-Poisson-GPFA pipeline. While a first-order (AR1) calcium likelihood did not show improvement over the two-step deconvolution method in recovering the ground-truth latents, the second-order (AR2) process demonstrated an improvement. Because there are no ground-truth parameters in the model, we use the reconstruction error of the latent values as our primary metric for recovery. These results suggest that when the observation model sufficiently captures the underlying calcium kinetics, the unified approach offers a more accurate estimation of the neural state.

      It would be useful to discuss the possible extension of the approach to other types of data that … have different observation models.

      We thank the reviewer for this helpful comment. We agree that the general framing of the likelihood has potential use in a wider range of data modalities.

      Specifically, all sensors (aside from some voltage sensors) have a rise and decay time in line with our model. Thus the autoregressive (AR) nature of the calcium likelihood we utilize makes the current implementation particularly well-suited for a broad range of fluorescence-based sensors with similar temporal profiles. The specific use and extension would depend heavily on the biological target of the sensor. For example Glutamate, dopamine, and similar indicators can be thought of as having a similar underlying Poisson firing model, as the release of these products is tied to neural firing. Other sensors that might relate to other biological processes, such as hemodynamics (via imaging or ultrasound) or broader neuromodulation (Norepinephrine imaging with nLight) might be more continually varying and therefore would require changing the Poisson with an appropriate alternative, for example a Gaussian Process or similar.

      Voltage imaging is the one exception that may require more complex observation models. However, the challenge in voltage imaging is not the ability to identify individual spikes, but more that the speed and scope of imaging is inherently limited by the speed of the voltage process and signal-to-noise ratios induced by the low quantum efficiency and membrane-bound nature of these indicators. If imaged well, single spikes would be clearly visible and the two-stage likelihood would not be necessary—one could simply use the spike times in a Poisson model just as with electrophysiology. We have added a paragraph in the discussion highlighting these points.

    1. Author response:

      Reviewer #1 (Public review):

      In this work, Frey and colleagues have carried out a very large study of α-synuclein polymorphism as a function of aggregation conditions and sample preparation. They provide valuable insight into the many critical factors affecting α-synuclein polymorphism, thereby illuminating the need for detailed reporting in the literature as well as both rigorous and detail-oriented protocols when working with this protein. Indeed, their observations are in line with the difficulties of reproducing structural outcomes across different laboratories and experiments. The authors must be complemented on their openness about the difficulties experienced and the thoroughness of their work. Efforts like these are going to be crucial to achieve an understanding of the unparalleled structural plasticity of α-synuclein amyloid fibrils. It is particularly notable that the authors have managed to optimize protocols to form single-polymorph aggregation reactions with high reproducibility.

      In this work, the authors focus on the influence of α-synuclein purity in aggregation reactions. They find that using reverse-phase HPLC to purify the protein significantly alters the aggregation behaviour. It is interesting, and rather uncommon in the field, to use reverse-phase HPLC as a final purification step, rather than SEC, which is commonly used in many laboratories. It would be useful to compare this new protocol even more directly and extensively with the commonly used SEC protocol. When mentioning their previously published work, it should be mentioned explicitly how the protein was purified in these previous studies.

      On the point of protein purification, the authors lyophilize their protein prior to storage. In their work, they also find that pre-aggregation oligomer formation alters the aggregation pathway of α-synuclein. While they demonstrate that this can be solved by appropriate filtration, it should be discussed why the lyphilizaiton step was not reconsidered/omitted given that this process is known to facilitate oligomer formation. Do the authors have experience with aggregation studies using α-synuclein that has not undergone lyophilisation and are able to comment on the influence of this step in the protocol?

      We will include details in our revised version on how samples were prepared in our previously published work. However we do not plan to do a comparison between SEC and HPLC as final purification steps because these are truly orthogonal separation methods. SEC is the gold standard for oligomer removal but not particularly useful for removing degradation products of similar size (which we believe to affect aggregation outcomes). In our hands, the highest purity and reproducibility come from HPLC-purified material followed by a stringent oligomer removal. Because oligomers in the solubilized sample are a concern, SEC as a final post-solubilization/pre-aggregation step might be ideal. However, as we mentioned in the manuscript, SEC dilutes the sample to the point that much of the sample is too dilute for our purposes (aggregation without seeds) and adding an additional concentration step would risk promoting the formation of new oligomers in the concentration device. That is why we resorted to using a 100 kD MWCO filter to remove oligomeric species.

      We considered skipping the lyophilization step and dialyzing the HPLC-purified sample into the buffer of choice but stuck with lyophilization because it provides an easy control over the protein concentration in the solubilized sample.

      To clarify the logic in this choice of sample preparation steps, we will add a section to the revised manuscript listing/explaining our suggested protocols for preparing alpha-synuclein samples for reproducible aggregation experiments.

      The authors make note of several degradation products affecting their aggregation reactions, which is why they employ a much more thorough purification protocol. However, they also point out that some of the degradation products found at the end of their reaction could form during the reaction itself. Unfortunately, they never investigate this further. In particular, it would be very useful to know if different aggregation conditions (pH, salt, agitation) lead to different and characteristic degradation patterns. HPLS/mass spec of the supernatant at the end of each aggregation reaction would have been a very insightful thing to do.

      In retrospect, we agree that this could have been important from the standpoint of understanding how in situ degradation during the aggregation at 37º C could also play a role in polymorph selection. We have begun to save frozen aliquots of our aggregation samples for subsequent MS analyses of the interesting samples in order to be able to address this in the future. However, it was outside the scope of our original search for the PD polymorph.

      With respect to degradation products, in this work a NΔ4-variant is produced to mimic a disease-relevant degradation product and indeed it is found to alter the structural outcome even at low relative concentrations (5%). Have the authors investigated the minimal fraction of the NΔ4 variant necessary to still influence the structural outcome of the predominant WT protein? This type of analysis could have significant relevance to disease-related analysis, where several variants (truncations and PTM variants) are present in trace amounts.

      We did not try lower than 5% because it was our goal to test if impurities at this level (which usually go undetected) could influence the aggregation outcomes. We think that directly relating the precise impurity level in these in vitro experiments to in vivo aggregation would be difficult due to the many factors we do not yet understand that appear to guide in vivo polymorph selection.

      While on this topic, the authors note that in several of their type 5 fibrils, they find unresolved peptide fragments in their cryo-EM structures. Can the authors speculate if these fragments are indeed peptide degradation products or residues wrapping around the fibril originating from the fibril-incorporated protein?

      We will add this speculation to the revised manuscript. The unassigned peptide density most likely originates from the C-terminal residues of the intact chains rather than from a degradation product. The levels of degradation products observed by MS in our other samples were far too low to account for the amount of peptide that would be required for >50% of the fibrils in sample 23 to contain this extra density. Furthermore, in the 5A polymorph the extra peptide density is sometimes present (e.g. sample 52) and sometimes absent (e.g. sample 3), despite the fact that in both samples the coexisting type 5 polymorphs (5m and 5B, respectively) have the peptide bound. Therefore, it appears that subtle differences between 5A polymorphs determine the presence or absence of this density, whereas for 5m and 5B it is consistently present.

      In this work, the authors have performed an extensive study of α-synuclein polymorphism. However, despite generating what is likely the largest single data set of fibril structures, they perform very little quantitative analysis of their data. It would be interesting to analyse the relative abundance of fibril polymorphs produced in each reaction. Perhaps from the particle-picking data it could be estimated the relative abundance of each polymorph as well as non-resolved fibrils to generate a more nuanced view of polymorphism beyond overall classifications of the resolved structures. Indeed, from such data it could also be studied if certain protofilaments are more prone to pair in asymmetric fibril structures over others or if fibril asymmetry can be attributed to random pairing of protofilaments in accord with their abundance. This latter point is particularly interesting for the type 1 fibrils. Perhaps, the propensity of α-synuclein to form specific symmetries could also be estimated.

      We agree that Cryo-EM datasets contain considerable information that could potentially be used to better understand polymorph populations. We have previously used particle counts as an approximate measure of polymorph abundance; however, the biases introduced during particle picking and subsequent curation are substantial, and we therefore do not consider these counts sufficiently reliable for quantitative comparison of polymorph populations.

      Regarding the symmetry of paired filaments, it is clear that the overwhelming preference of all filaments is to pair as symmetric dimers. Among the in vitro polymorphs, type 1 is the most commonly observed to form an asymmetric dimer, either with itself or with the new type 7. However, the number of observations is too small to establish that this represents a statistically meaningful preference. Types 2 and 3 have also been observed to pair in asymmetric fibrils. Overall, we do not think that our dataset is large enough to add statistical weight to previous observations.

      The authors note that pH is a strong factor in determining polymorph selection. This does indeed appear to be the case, but other parameters do not appear to show any clear trend. Have the authors investigated the influence of aggregation parameters (agitation, duration) on structural outcomes systematically or quantitatively, such as with principal component analysis? Indeed, they also find that some fibril types that otherwise are not compatible at the same pH appear to co-exist when the shaking parameter is modified. Are the authors then confident in the claim that pH is a deterministic parameter?

      We did not systematically vary the agitation but in two cases where it was either intentionally or accidentally varied, we found surprising polymorph outcomes. We felt that these observations were worth reporting, but on their own they do not constitute a thorough study. We would rather conclude that pH is a strong selector and can be deterministic under certain conditions: pure sample, no seeds, continuous or intermittent moderate agitation. We will revise the manuscript to make this clearer.

      It is evident from this and other work that amyloid aggregation is highly sensitive to kinetic effects. It is therefore curious that the effect of protein concentration and reaction time has not been systematically investigated. The authors have some data studying dilution series (Figure 6A) and different reaction times (reaction 56&57). Could the authors comment on the effects of these two parameters, and might there be more information touching upon this that could be highlighted in this work?

      We did not collect sufficient data to draw conclusions about the effects of protein concentration or aggregation time, and therefore do not think that a quantitative analysis of these parameters is justified by the present dataset.

      The authors state that this work likely underreports fibril polymorphisms in samples due to population size or data quality challenges. Could the authors, based on their extensive experience, try to quantify this statement?

      Precise quantification would be difficult. The literature has many mentions of amyloids that could not be solved by Cryo-EM due to a lack of twist. In our hands it is very common to have a small subset of non-twisted filaments in a sample and some samples appear to be exclusively non-twisted. Low-abundance or low-quality fibrils are also rather common in our data but also difficult to quantify. Nevertheless, we agree that it would be useful to place a lower bound on this estimate, and we will re-examine our datasets to determine whether this can be quantified in the revised manuscript.

      Additionally, the authors point out that the current framework for classifying α-synuclein fibril polymorphism is not sufficient to describe the real complexity of this protein system. However, they do not seem to address some of the recent literature aiming to solve such issues (see Scheres 2026, Connor et al. 2025, Milchberg et al. 2025 & Price et al. 2025).

      We agree that we should have discussed this literature in greater detail and will do so in the revised manuscript.

      Was any biophysical/biochemical analysis performed of the many structures produced here, such as CD spectroscopy, Proteinase K digestion, dye binding or FTIR, which could act as low-resolution structure identifiers and might help to retrospectively explain some findings in the older literature? Such data would be very useful for the vast majority of researchers, who do not have access to cryo-EM.

      We did not perform these analyses precisely because they are low resolution. Retrospectively, such analyses might have been useful for interpreting past data, but most of these methods (particularly Proteinase K resistance) are difficult to compare between laboratories and work best with side-by-side controls. This is why developing a facile method for polymorph identification is one of our main future research goals.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes insights gained during efforts to reproduce disease-relevant alpha-Synuclein (aSyn) fibrils using recombinant protein in vitro. It follows up on a similar article from this team published in 2024. Although the authors have not been able to produce fibrils with the structure of ex vivo fibrils isolated from patients, they share insights gained into which factors influence the formation of specific fibril polymorphs.

      Strengths:

      This is quite an unusual manuscript because it goes into minute detail about sample preparation that are usual just mentioned in the Materials and Methods sections of other manuscripts (if at all). This makes it very valuable for the scientific community working on exactly the problem of reproducing disease-relevant aSyn fibrils in vitro (which will be a major breakthrough in the field). The authors present an impressive array of cryo-EM fibril structures, some of which have not been described before.

      Weaknesses:

      A major concern with the manuscript is that its story and messaging are a bit murky. The authors describe a few new polymorphs, show that some polymorphs (type 1) have small variations, show that sample purity and fragmentation will influence polymorph formation, and present a helical-symmetry mystery. This all reads like a loose collection of findings without any major takeaway. Looking at Table 1, it still seems that the authors do not have a good control over any of these polymorphs. Are they able to make any of these polymorphs reliably? I think the impact of this work could be strengthened if it ended with a reliable protocol for the production of any of the polymorphs described.

      It is true that the initial results were a collection of findings compiled while searching for the PD polymorph. However, the trends that we saw in these data inspired us to pursue a more systematic approach which was used to show how impurities play a crucial role in polymorph selection even at low levels. We agree that the impact of the work will be enhanced by summarizing the protocols that can be used to obtain the types 1, 2, 3 and 5 polymorphs and will add this to the revised manuscript.

      A second major concern is the quality of the aggregation kinetics and their interpretation. Are these kinetics just done once per concentration? The figure caption talks about 'three independent samples' but it is unclear if this refers to the three different concentrations or NΔ4 percentages used, or true repetitions. Looking at the curves themselves, it seems that only one of the conditions was actually done in triplicate, which should be the minimum to draw conclusions. Further, it would have been helpful to characterize the kinetic data quantitatively. Finally, because there are only kinetic data for a fraction of the conditions tested, it is not clear what they add to the overall manuscript. My recommendation is to either remove the kinetics from the manuscript or substantially expand this section.

      In Figure 6A, the three curves represent three independent samples; in panels B–D, each curve similarly represents one independent sample. We will try to clear up the ambiguity in the revised version. We agree that the kinetic data have limited utility for determining kinetic parameters of the aggregation. The data were collected primarily to determine when the aggregation reactions were complete so that samples could be prepared for cryo-EM. Nevertheless, we think that the kinetic traces provide two qualitative observations that are relevant to the structural results: 1) In panels A and C of figure 6, the type 5 polymorphs are associated with shorter lag times, suggesting that they either nucleate faster, or as we propose, arise from a small amount of oligomeric protein in the original sample. 2) The longer lag phase associated with the NΔ4 construct was reproducible (panels B, C and D), suggesting that its intramolecular self-chaperoning effect is enhanced due to the increased positive charge in its N-terminal region. Future studies will determine whether this same electrostatic change also gives rise to enhanced secondary nucleation.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      From a big picture viewpoint, this work aims to provide a method to fit parameters of reduced models for neural dynamics so that the resulting tuned model has a bifurcation diagram that matches that of a more complex, computationally expensive model. The matching of bifurcation diagrams ensures that the model dynamics agree on a region of parameter space, rather than just at specially tuned values, and that the models share properties such as qualitative features of their phase response curves, as the authors demonstrate. A notable point is the inclusion of extracellular potassium concentration dynamics into the reduced model - here, the quadratic integrate-and-fire model; this is straightforward but nonetheless useful for studying certain phenomena.

      Strengths:

      The paper demonstrates the method specifically on the fitting of the quadratic integrateand-fire model, with potassium concentration dynamics included, to the Wang-Buzsaki model extended to include the potassium component. The method works very well overall in this instance. The resulting model is thoroughly compared with the original, in terms of bifurcation diagrams, production of various activity patterns, phase response curves, and associated phase-locking and synchronization properties.

      Weaknesses:

      It is important to note that the proposed method requires that a target bifurcation diagram be known. In practical terms, this means that the method may be well suited to fitting a reduced model to another, more complicated model, but is not likely to be useful for fitting the model to data. Certainly, the authors did not illustrate any such application. Secondly, the authors do not provide any sort of general algorithm but rather give a demonstration of a single example of fitting one specific reduced model to one specific conductance-based model.

      We thank the reviewer for this critical assessment. It is true that we demonstrate our approach using as target the bifurcation diagram of a more realistic model. In principle, the method would be applicable to experimental systems if parameter space is sampled under appropriate experimental conditions, e.g., by recording at different extracellular potassium concentrations. However, this is challenging: generating reliable” experimental bifurcation diagrams” would require tightly controlled experimental repetitions under various conditions, particularly for reconstructing two-dimensional bifurcation diagrams. We hope that this work encourages the development of such methods. We have added a discussion of this point; see the paragraph starting at line 549.

      We have also included a general algorithm (Table I) and provide a second example illustrating the procedure (Supplementary Fig. S2).

      Finally, the main idea of the paper seems to me to be a natural descendant of the chain of reasoning, starting from Rinzel - continuing through Bertram; Golubitsky/Kaper/Josic; Izhikevich; and others - that a fundamental way to think about neuronal models, especially those involving bursting dynamics, is in terms of their bifurcation structure. According to this line of reasoning, two models are “the same” if they have the same bifurcation structure. Thus, it becomes natural to fit a reduced model to a more complicated model based on the bifurcation structure. The authors deserve credit for recognizing and implementing this step, and their work may be a useful example to the community. But the manuscript should have described and cited this chain of works to put the current study in the correct context.

      We have added a paragraph in the Discussion section (starting at line 517) to better situate the manuscript within the relevant literature and to explicitly acknowledge the chain of work.

      Reviewer #1 (Recommendations for the authors):

      Please see my public review. In line with my comments, I recommend that the authors either (a) provide a general algorithm for fitting at least a class of reduced models (i.e., those that satisfy some general assumptions) to a class of bifurcation diagrams, or (b) provide at least one more example of implementing their method. Step (b) would not need to be done to the same degree of thoroughness as the example they provided (e.g., the PRCs and synchrony need not be considered), but to me, this step would be very important if (a) is impractical. Otherwise, the paper should probably be rewritten to de-emphasize the message that this is a general method; instead, this should be a paper about specifically fitting the QIF (with potassium dynamics) to the Wang-Buzsaki model (with potassium dynamics).

      We provide a general algorithm for deriving a quadratic integrate-and-fire model with dependence on a biophysical parameter by fitting the bifurcation structure of a given class I conductance based neuron model near an SNL bifurcation induced by this parameter; see Table I.

      In addition, we provide a second example of the reduction procedure: motivated by Hesse et al. (Nature Communications, 10.1038/s41467-022-31195-6, 2022), we derive a QIF model that captures dependence on temperature instead of potassium concentration; see Supplementary Fig. S2.

      Not surprisingly, I also think it’s essential that the authors describe and cite the chain of works on thinking of neuronal models in equivalence classes based on bifurcation diagrams, and make clear that this paper builds on the ideas set forth in that chain.

      We thank the reviewer for this comment. As mentioned above, we have added a paragraph in the Discussion section, starting at line 517, to acknowledge this chain of work.

      Also, the authors should make clear that their method is not one for fitting a model directly to data, which will require rewriting at least the first paragraph of their Discussion section.

      We thank the reviewer for helping us make our manuscript clearer. To avoid confusion, we have clarified this point already in the Introduction (see lines 52-57) and have included a new paragraph in the Discussion (starting at line 549).

      Other specific corrections are:

      (1) Typos should be fixed, as the paper has several. The first line of the abstract has one (“concentrations” → “concentration”), for starters. “Original model” on pg. 3 is missing “be” in “can defined”. “ceases” → “cease” on pg. 13. “nerons” → “neurons” on pg. 19. “standart” → “standard” pg. 24.

      Done. Additional typos were also corrected.

      (2) The abstract mentions “consequences in networks” in its second sentence. This is misleading because studying network dynamics is not at all the main emphasis of the paper, but rather a corollary application of the main ideas, so some restructuring of the abstract is needed. Similarly, the final abstract sentence overstates somewhat what was done with studying synchronization and should be rewritten more precisely.

      We have restructured the abstract accordingly.

      (3) For readers who are interested in the ideas here but not familiar with the QIF model, it will be very difficult to follow the first paragraph of Results. Elementary aspects of QIF dynamics should be explained here (e.g., what is the saddle-node bifurcation), and a basic figure panel about this should be included in Figure 1.

      We added a supplementary figure (Figure S1) adapted from Izhikevich for readers who might not be familiar with the QIF model.

      (4) Bottom lines of page 3 should be reworded to make clear that the slow variables are averaged over each member of a family of fast subsystem limit cycles. Also, “one limit action potential cycle” is an awkward phrase.

      We rephrased this sentence (see paragraph starting at line 140).

      (5) Text under system (1) – why isn’t c mentioned? Also, references to Figure 3 should be to Figure 2 here. And authors should state what they mean by “target model” and be clear about whether it includes potassium dynamics and/or pump current.

      c scales the parabola corresponding to the branch of fixed points, given by c(I<sub>app</sub> − I<sub>SN,0</sub> − I<sub>pump</sub>) = −a(ν − ν<sub>SN</sub>)<sup>2</sup>. Consequently, it also affects the position of the homoclinic bifurcation: In the previous version of the manuscript, c was inadvertently omitted from the expression for the branch of fixed points; this has now been corrected. See paragraph starting at line 154.

      Figure references have been corrected.

      By target model, we mean the conductance-based model that includes potassium dynamics and a pump current, in our case System 4. We clarified this point at the beginning of the Results section (see paragraph starting at line 116). Throughout the manuscript, we now explicitly indicate when we refer only to its fast subsystem and whether the pump current is included. In particular, note that the parameter derivation shown in Figure 4 is performed on the fast subsystem of the target model, in the absence of the pump current. I<sub>pump</sub> can be considered as a potassium-dependent contribution to the applied current, and can be added a posteriori to the QIF model. See paragraph starting at line 162.

      (6) Next par: is the “saddle-node bifurcation” that with I<sub>app</sub> as bifurcation parameter? Please clarify.

      Yes, it is the saddle-node bifurcation with I<sub>app</sub> as bifurcation parameter. We have clarified this in the manuscript; see the paragraph starting at line 170.

      (7) Bottom pg. 5: does “beyond” mean above? below?

      We meant above (larger values of ). In the text, we have replaced “beyond” with “larger than”. See paragraph starting at line 178.

      (8) Formula for v<sub>r</sub> at top of page 6: Please specify what formula for I<sub>pump</sub> is being used here.

      The formula for I<sub>pump</sub> is given in Eq. 5d. We are using the same formula throughout the paper.

      Note that to clarify the reduction procedure, we derive QIF parameters to match the bifurcation diagram with respect to the applied current of the fast subsystem of the target model when I<sub>pump</sub> = 0. Reintroducing I<sub>pump</sub> produces the same horizontal shift in this bifurcation diagram for both the QIF and Wang–Buzsáki versions.

      We have restructured the paragraph starting at line 178 to clarify these aspects.

      (9) Three lines below this: I don’t understand what “matching...is appreciable” and “in the continuity of...” mean. Please revise and also explain why a closer matching of v<sub>r</sub> to the min voltages in Figure 4d was not used, and exactly how the v<sub>r</sub> that is shown was chosen.

      With “matching...is appreciable”, we meant that values assigned to v<sub>r</sub> should be close to the minimum voltage values reached during spiking. With ”in the continuity of...”, we meant that when .(SNIC case), we choose v<sub>r</sub> by extrapolating the linear fit performed on the values of v<sub>r</sub> assigned when , (homoclinic case).

      When , v<sub>r</sub> was chosen so that the homoclinic bifurcation occurs at the same value of applied current as in the fast subsystem of the target model. This is explained in the paragraph starting at line 178 (see Eq. 2). This criterion also allows the minimum voltage values reached during spiking to be captured reasonably well (compare the green dotted line and the purple solid line in panel e of Figure 4).

      We have rewritten the paragraph starting at line 186 to clarify these aspects.

      (10) Bottom page 6 - reference to Figure 3e should be 4e. Also, the text mentions the shrinkage of spike amplitude, but the figure shows that vth increases over most of the K+ range before decreasing, so a correction is needed.

      We corrected the figure reference.

      The maximal voltage of the limit cycles of the target model’s fast subsystem (upper purple curve in Figure 4e) increases slightly between and , by less than 1mV. It then decreases by about 27mV before the fold of limit cycles. The sigmoidal function vth () allows us to capture this substantial decrease in the QIF model. The small preceding increase is not captured. We reformulated the text to avoid confusion (see paragraph starting at line 209).

      (11) Figure 4d: Why is E<sub>K</sub> plotted here? It should be mentioned in the caption and text. More generally, the caption for Figure 4e should be expanded to mention what the purple curves are, what is the black curve for K < K<sub>SNL</sub>, and what the other structures shown are. Finally, the text describes that theSNIC/SNL/Hom is determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>, so it’s not clear what is I<sub>app,SNL</sub> in the caption - please clarify.

      In conductance-based models, higher weakens the potassium concentration gradient, thereby raising E<sub>K</sub>. The sodium and potassium reversal potentials typically bound voltage oscillations during spiking (see for example Chander and Chakravarthy, PLOS ONE, 10.1371/journal.pone.0048802, 2012), so the minimum voltage of spikes is expected to be higher when is larger. We plotted E<sub>K</sub> in Figure 4d to show that the increase of the reset voltage v<sub>r</sub> at larger reflects this effect in the QIF version of the model. We clarified this in the caption of Figure 4 and in the text (see paragraph starting at line 199).

      Purple curves show families of limit cycles, while black curves, including the one for , show families of fixed points. The green curve shows the linear fit of v<sub>r</sub> from panel d. All these have now been included in the legend.

      In the QIF model, the onset bifurcation (SNIC, SNL, or homoclinic) is indeed determined by the choice of v<sub>r</sub> relative to v<sub>SN</sub>. Figure 4e shows the bifurcation diagram of the fast subsystem of the conductance-based model (Wang-Buzsáki). This is a bifurcation diagram with respect to , for a fixed value of applied current. We chose to fix I<sub>app</sub> at its value at the SNL bifurcation, denoted I<sub>app,SNL</sub>. I<sub>app,SNL</sub> is near 0.22 (see Figure 3a).

      (12) Eqn. (2a): Shouldn’t v<sub>SN</sub> depend on potassium like I<sub>SN</sub> and I<sub>pump</sub> do? What is the formula for Ipump there? Why isn’t the RHS of (2b) dependent on v as in the original model? Please clarify.

      For simplicity, we did not include a potassium dependence for v<sub>SN</sub> in the QIF model. Instead, we set it to its value at the SNL bifurcation (Figure 4b). This is explained in the paragraph starting at line 199: “We notice that v<sub>SN</sub> and the normal form coefficient a are relatively conserved in this interval. We fix them to their value at [K<sup>+</sup>]<sub>o,SNL</sub>.”

      The formula for I<sub>pump</sub> is given in Eq. 5d. The potassium dynamics depends on the voltage via the reset rule in Eq. 3d. This allows us to capture the small increments in at each action potential in the original model (see for example Figure 6c,h). We have clarified these two points in the manuscript (see paragraph starting at line 223).

      (13) Figure 3: Please indicate the criticality of the Hopf bifurcations shown.

      The legend of Fig. 3 now indicates that the Hopf bifurcations are subcritical. The same clarification has been added to the following figures as well.

      (14) Bottom pg. 9: Why isn’t there an I<sub>pump</sub> term as in eqn. (6a)? Please clarify. 

      You are correct, the I<sub>pump</sub> term should be included in the equation for the averaged slow subsystem of the QIF model (see paragraph starting at line 243); it was accidentally omitted in the manuscript. Thank you for pointing this out.

      (15) Top pg. 11: It’s important to reference the slow averaged dynamics here, which allows K+ to increase. Also, this first paragraph should already explain that this averaged dynamics is only relevant along the family of FS periodic orbits, not during the recovery when the FS has a branch of stable equilibria.

      The averaged slow subsystem is indeed only relevant along families of limit cycles of the fast subsystem. Along families of equilibria, averaging is not necessary and the standard slow subsystem can be used. We now explicitly define this standard slow subsystem for both Wang-Buzsáki and the QIF model (see paragraphs starting at lines 241 and 628). In panels e and j of Figure 6, both systems are now represented.

      In the paragraph starting at line 273, we now refer to the averaged slow subsystem to explain the overall increase of during bursts (purple curves in Fig. 6e,j), and to the standard slow subsystem to explain the decrease of during quiescent phases (black curves).

      (16) Pg. 11, par 3: This is unnecessarily confusing. Please try to reword and clarify this paragraph.

      We have simplified this paragraph (starting at line 293). The key point is that the reduction to the averaged slow subsystem is not valid too close to the homoclinic bifurcation.

      (17) Pg. 13, end of Scenario 2: Is there any evidence this is a canard effect and not a noise effect? If so, please mention the evidence; otherwise, perhaps take this out.

      What happens there appears to be a noise-induced canard effect: in the beginning of the burst, the system follows a family of stable limit cycles of the fast subsystem, i.e. a stable object. However, at some point, noise induces a transition to a portion of trajectory where the system evolves near the saddle branch, i.e. a repelling object, for a substantial amount of time. Such phenomena have been thoroughly investigated in the literature, and can also be obtained in a deterministic way; see for example Marin et al. (Physical Review E 90, 042718, 2014). Bursting traces similar to the one in Figure 7c are observed experimentally (see, for example, Figure 4c of Marin et al.), which is why we considered it worth mentioning. We have revised the paragraph starting at line 332 to clarify this point.

      (18) I only see 4 curves in Figure 9a,c, but the legend has 5. Are two on top of each other? Please clarify.

      Yes, the curve for = 7.21mM lies beneath the curve for = 5.21mM. This has been clarified in the figure caption.

      (19) Text should note that the QIF iPRC does not develop a negative region at high K+ and high phase, as WB iPRC does.

      We have updated the paragraph starting at line 409 to mention this.

      (20) Pg. 16, line 4: “at the network scale” is cryptic - a more precise phrase would be preferable.

      We have reformulated the sentence to clarify its meaning (see paragraph starting at line 380).

      (21) Pg. 16, line 9: Reordering of words could make this clearer.

      Done (see paragraph starting at line 385).

      (22) Pg. 16: I am confused by line 14 because the big changes in the iPRC in Figure 9c do not align with the spike phase in Figure 9d. Please clarify what is meant here.

      In the QIF model, at a given phase, the iPRC is the inverse of the slope of the voltage trace as a function of phase. Flatter slopes in Figure 9d therefore correspond to larger iPRC values in Figure 9c. This is illustrated in Figure S5. We have revised the paragraph starting at line 393 to make this point clearer.

      (23) Pg. 16: Please clarify what is meant by a “delta synapse”.

      By “delta synapse,” we meant a configuration in which each spike induces an instantaneous voltage jump in the postsynaptic neuron, modeled using the Dirac delta distribution. We have replaced the term “delta synapse” with “pulse-coupled,” which is more commonly used in the literature, and have added a clarification at its first occurrence in the manuscript.

      (24) Discussion, line 2: Delete comma.

      Done.

      (25) Importantly, as noted above, the first par. needs to be rewritten since the presented method won’t work directly from data or from a target model for which most of the parameters, and hence the bifurcation diagram, are not known.

      As mentioned above, we have included a new paragraph in the Discussion, starting at line 549, to clarify this point.

      (26) Pg. 18: ”Originally” → ”Typically”, perhaps?

      Done.

      (27) Note the work of Marder et al. on temperature-related neural variability.

      We thank the reviewer for this comment. We have added two relevant references from the work of Marder and colleagues addressing temperature-dependent neural variability and ionic concentrations in our manuscript (see the sentence starting on line 537).

      (28) Pg. 20: Cut the ”Potassium dynamics and network models” subsection since it does not add anything substantive as written (or else expand it and include it in the subsection below).

      We have expanded this paragraph and incorporated it into the subsequent subsection, as suggested by the reviewer.

      (29) Finally, it’s a bit confusing that the authors refer to the potassium concentration as a slow variable yet have an instantaneous jump in this quantity at reset in their QIF model (i.e., instantaneous is VERY fast). Some explanation about this should be provided. Do they make the general assumption that ∆<sub>K</sub> is small, for example, such that this reset reflects the slow nature of K+ evolution (i.e., during the reset period, K+ would only change slowly, and hence by a small amount)?

      Yes, ∆<sub>K</sub> is chosen to be small, to capture the behavior of the original model (compare for example panels c and h of Figure 6). As a result, in the QIF model, in the same way as in the original model, despite the fact that the dynamics of includes a fast component, on average evolves slowly. By using the averaging method, we can determine whether overall increases or decreases.

      We clarified this in the manuscript, in the paragraph starting at line 243.

      Reviewer #2 (Public review):

      Summary:

      The authors derive an integrate-and-fire model to describe the dynamics of a more complex Wang-Buzsaki model and compare the two models. A detailed discussion of bifurcation schemes in both models is convincing and allows us to evaluate the simpler model.

      Strengths:

      The idea is interesting, and the mathematical approach appears to be convincing. In addition, differences between the simple and original models are also discussed.

      Weaknesses:

      A comparison to experimental data is necessary to support the theoretical work.

      As mentioned above in our answer to Reviewer 1, we demonstrate our method using as target the bifurcation diagram of a more realistic neuron model. Ideally, one would want to derive phenomenological models that capture bifurcation structures obtained from data. However, this is challenging and beyond the scope of the present study. We hope that this work encourages the development of such methods. We have revised the Introduction (see lines 52-57) and added a paragraph in the Discussion (see the paragraph starting at line 549) addressing this point.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is well-structured; however, it appears that it has been edited with less care. Please see comments below:

      (1) Page 2: “A third bifurcation, the saddle-homoclinic orbit (HOM) bifurcation,”: provide a reference for the bifurcation.

      We have added a reference to the book by Izhikevich (see paragraph starting at line 63).

      We have added additional references in the Introduction that we considered helpful.

      (2) Page 3: “while a larger concentrations it is mediated by...”: remove “it”.

      There was indeed a typo in this sentence. The intended phrasing is: “while at larger concentrations it is mediated by...”. We have corrected it accordingly (paragraph starting at line 134).

      (3) Figure 2, caption: “dashed lines for unstable branches”: this is a dotted line.

      Corrected to “dotted lines”. Thank you.

      (4) Page 4: “is smaller than vSN (Fig. 3c),”: this figure panel does not exist, as well as the Fig.3d referred to afterwards. Please correct.

      We intended to refer to Fig. 2. Figure references have been corrected. See paragraph starting at line 154.

      (5) Page 6: “Fig. 3a-d shows ISN,0, vSN and a for [K]+o between 4 and 16 mM”: Fig.3c+d do not exist, please correct. Similar comment to “the absence of pump (Fig. 3e).” on the same page.

      We intended to refer to Fig. 4. Figure references have been corrected (paragraphs starting at lines 199 and 209).

      (6) Page 8: “(panel A)” → ”panel (a)”.

      Done.

      (7) Figure 4e: What is the meaning of the green dotted curve?

      This curve represents the linear fit of the reset voltage v<sub>r</sub> from panel d of Fig. 4, to show that the minimal voltage values of the limit cycles are also well captured. We have added this curve to the legend and included a brief explanation in the figure caption.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants. The authors provide solid evidence that the two symbionts can generate contrasting phenotypes: Rickettsiella increases plant damage while reducing wing formation and dispersal, whereas Regiella reduces aphid population growth and feeding damage. The successful establishment and stable transmission of these novel symbiont-host associations, combined with experiments spanning individual, whole-plant, population and mesocosm scales, are notable strengths of the work; however, the mechanisms underlying these effects remain unresolved, evidence for horizontal transmission is indirect, and some population-level conclusions rely on relatively small sample sizes or effects that are not consistently detected across time points, and therefore the potential application of these findings to pest management remains promising but speculative. The study will be of broad interest to researchers working on insect symbiosis, plant-insect interactions, and biologically based approaches to pest management.

      We have made revisions to the manuscript to cover issues raised around sample numbers and mechanisms. We appreciate that we have not been able to finalize the exact mechanism underlying plant damage effects. We note that while comparisons of population performance were limited by the number of populations we could feasibly maintain; sample sizes were substantial in some of the other experiments. We also do substantiate effects through a combination of experimental approaches that start with controlled conditions and then encompass more complex environments.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

      We thank the reviewer for recognizing the strengths of the work and for providing extensive comments.

      Weaknesses:

      There are some major aspects of this paper that I thought could be strengthened. My main concern is that the manuscript feels broader than it is conceptually focused. A wide range of outcomes is measured, which gives the study breadth, but it also makes the central question harder to identify. As written, the paper reads more strongly as a proof-of-principle demonstration of ecologically relevant phenotypes than as a tightly framed test of a specific biological idea.

      The broad range of tests in our work was intentional. Rather than relying on a single experimental approach or scale, we designed the study using multiple complementary approaches to independently evaluate the effects of endosymbionts. Thus, our conclusions are supported across multiple experimental contexts rather than by a single experiment or scale. For example, the effects on plant damage were consistent across different host plants, including wheat and barley (Figures 1 & S1), and across different experimental scales, ranging from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level, using single aphids maintained on individual plants under favorable conditions (Figures S5 & S6), and at the population level under crowding conditions (Figures 3 & 4).

      Despite the breadth of measurements, the study is focused on establishing robust evidence for the contrasting effects of the two endosymbionts on aphid dispersal and plant feeding damage. To help address this concern, we have moved the summary table from the Supplementary Materials to the main text (now Table 1), which provides an overview of the experimental approaches and main findings and should help readers more clearly see how the different experiments are connected.

      A second issue is that the biological basis of the reported phenotypes remains less developed than the phenotypic description itself. The authors make a genuine effort to address mechanism through JA, JA-Ile, SA, and metabolomic profiling, but these analyses only partially explain the main results. The negative result for the canonical defense markers is informative, yet it still leaves a substantial gap between the observed variation in plant damage and the processes responsible for it.

      Our analyses of JA, JA-Ile, SA, and the metabolomic profiles provide some initial insights, but we appreciate that they do not fully explain the differences in plant damage observed between treatments. The primary aim of this study was to evaluate the phenotypic effects of the endosymbionts and their potential for agricultural application, rather than to provide a comprehensive mechanistic explanation. We have accordingly avoided overinterpreting the mechanistic results and now explicitly state that elucidating the underlying biological mechanisms will be an important direction for future research in Discussion section.

      I also think some caution is needed in how the two symbionts are compared. The authors explain why some follow-up experiments were designed differently for Rickettsiella and Regiella, and that rationale is understandable. Still, because the downstream assays were not fully matched, the paper is strongest when each symbiont is interpreted on its own terms rather than as a strict comparison.

      Our initial plant-damage experiment was designed as a first comparison to test whether different endosymbionts can have diverse and contrasting effects on their aphid host population and plant damage. We then investigated the individual phenotypes of each endosymbiont in greater detail, particularly in relation to their potential agricultural applications. Specifically, our results suggest that Regiella may reduce plant damage, whereas Rickettsiella may reduce dispersal. Based on these early findings, some further experiments were conducted with slightly different experimental setups. Nevertheless, many of the experiments conducted for the two endosymbionts were broadly comparable. We designed the additional experiment carried out only with Rickettsiella to test whether reduced alate production observed in our earlier experiments also translated into reduced dispersal at the population level. An equivalent experiment with Regiella was not undertaken because we failed to detect an effect of this endosymbiont on alate frequency in our preceding experiments. We have clarified this rationale in the Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”).

      Overall, I would suggest softening the Significance Statement so that it more clearly reflects what is directly shown here, namely that introduced symbionts can alter plant damage and dispersal-related phenotypes under controlled conditions, rather than implying that the study directly tests management utility in agricultural settings.

      We have done this in the Significance Statement. We appreciate that the current experiments have been carried out under controlled conditions, rather than in agricultural settings. Pending permit approval, we are currently planning to extend this work to contained field settings to test whether effects on plant damage and dispersal ability are also observed under less controlled conditions.

      Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      One thing that I struggled a little with was the rapid spread of Regiella in the shared plant experiments. Possibly this could be attributed to an increased reproductive output (due to faster developmental time, and/or an increase in fecundity) or efficient horizontal transmission. However, the other experiments performed indicate a slight negative impact (Figure 4a) or no influence (Figure 4C, 4D, Figure S6) of Regiella infection on host fitness. Given these other results, it seems that Regiella must spread fairly efficiently between hosts, which comes as a surprise, and there are very few examples of horizontal transmission of Regiella like this in the literature. The manuscript would benefit from a clear and direct demonstration of horizontal transmission, rather than it being inferred indirectly. The similar spread observed in the Rickettsiella mixed cages is less surprising, because there are several examples where this has been demonstrated.

      We agree that the rapid spread of Regiella in the shared-plant experiments cannot be readily explained by host fitness alone, though we have noted fitness benefits of Regiella in a different transinfection in oat aphids (Yu et al., 2025). In a previous study with transinfected green peach aphids and despite a substantial fitness cost, we found that Rickettsiella can spread relatively rapidly in a population and show high stability (Gu et al., 2023) and perhaps transmission of Regiella follows a similar pathway. However, whereas Rickettsiella may spread through plant tissues, Regiella showed relatively low levels of horizontal transmission through this pathway.

      We certainly agree that more work is required to establish the mechanism and dynamics of horizontal transmission in this system. Rather than focusing on mechanism, our objective here was to examine endosymbiont spread where intact plants were available and where there was a mixed aphid population. Note that we also conducted an additional experiment in which Regiella-infected aphids were present at a frequency of only 10% of the initial population, and in this situation Regiella nevertheless still increased in frequency including to a low Cp value in most (8/10) replicates, highlighting its persistence and potential to increase in populations.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Yu et al., A persistent bacterial Regiella transinfection in the bird cherry-oat aphid Rhopalosiphum padi increasing host fitness and decreasing plant virus transmission. Pest Manag Sci 81, 2791-2799 (2025).

      It is also a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

      We conducted the final mesocosm-dispersal experiment specifically with Rickettsiella because our earlier individual-plant experiments had already shown a clear reduction in alate production in Rickettsiella-infected aphids, together with effects on plant damage (Figure S1I) and population growth (Figure 3C). In contrast, we did not detect a significant difference in alate frequency between Regiella-infected and wild type aphid strains in the similar set up experiments (Figure S1I and Figure 4C). We therefore designed the additional experiment to test whether the reduced alate production observed with Rickettsiella also translated into reduced dispersal at the population level. We did not conduct the same experiment with Regiella because there was no difference in alate frequency in our earlier experiments. We have also added this explanation in Materials and methods section (“Aphid dispersal ability and plant feeding damage in mesocosms”). We do appreciate however that future experiments on dispersal of Regiella will be worthwhile resources permitting.

      Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      We acknowledge that more replication is always desirable, but we would also argue that significant effects were not marginal and replication was substantial in many cases (e. g. 9-10 replicate plants per damage treatment evaluation). We were also focused on using multiple experimental approaches and scales to independently and repeatedly evaluate the effects of endosymbionts on aphid fitness, wing development, plant damage, and aphid dispersal. Thus, conclusions are not based on a single experiment or experimental scale. For plant damage, for example, we observed consistent effects across different host plants, including wheat and barley (Figure 1 and Figure S1), as well as across different experimental scales, from individual plants maintained in small cages (Figure 1) to a dispersal experiment involving 24 plants (Figure 5). Similarly, aphid fitness was evaluated both at the individual level using single aphids maintained on individual plants under favorable conditions (60 replicates per treatment) (Figures S5 & S6) and at the population level under crowding conditions (Figures 3 & 4). These complementary experimental designs allowed us to examine whether the observed phenotypes were consistent across different environmental and population contexts. We did face challenges in high levels of replication of independent aphid strains in population cage experiments but attempted to replicate as much as possible given the resources that were available.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their dlimitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      (1) Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.

      (2) Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.

      (3) Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.

      (4) James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      We agree that qPCR-based measurements of endosymbiont gene copy number have limitations even if they are the standard approach used in most studies. In the current set of experiments, qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples and the only one feasible given the number of monitoring events and samples required. Nevertheless, we acknowledge the limitations of this approach (and have pointed this out ourselves in a recent COIS paper – Hoffmann et al 2026). We now mention this under further work and provide a couple of references (see Discussion).

      Reference:

      Hoffmann, A. A., Yang, Q. and P. A. Ross. Aphid endosymbionts revisited: molecular detection, diversity, and population dynamics. Curr Op Insect Sci (in press). (2026)

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Overall, I found the study interesting and worthwhile, particularly because it shows that novel symbiont associations can generate contrasting phenotypes in an important pest species. My main suggestion would be to sharpen the framing of the paper and to bring the mechanistic and applied discussion into slightly closer alignment with the current evidence base. With those points addressed, I think the manuscript would read as a clearer and more balanced contribution.

      We have rephrased the Discussion around the mechanistic component of the work to sharpen this and link more directly to evidence. For instance:

      “Despite this, our metabolomic analyses suggest that endosymbionts in D. noxia may influence some other aspects of wheat metabolism in a spatially structured way. Within aphid-feeding areas, wheat exposed to Rickettsiella aphids showed increased valine and decreased 3-phenyllactic acid. Changes in valine have been reported in plants responding to herbivory and other forms of stress (43-45), while 3‑phenyllactic acid has been associated with antimicrobial and defense-related activity (46). In non-feeding areas, wheat exposed to Regiella aphids showed reduced urea and increased pantothenic acid. These changes may reflect differences in nitrogen metabolism and allocation (47), and in metabolic processes involving pantothenic acid (48) respectively. However, the present metabolomic data do not establish the functional consequences or causal mechanisms underlying these changes. They indicate that aphids carrying different endosymbionts are associated with some distinct metabolic responses in wheat, including responses that differ between aphid-feeding and non-feeding areas. Our findings are consistent with previous work showing that phloem‑feeding insects can induce changes in plant metabolites in response to herbivory (49, 50) and provide a basis for future studies to determine how endosymbionts influence aphid-induced plant responses and contribute to contrasting plant phenotypes.”

      We have not undertaken a complete reframing of the paper but further emphasized the focus on phenotypic contrasts in a few places including incorporating some changes to the comments below.

      (2) Lines 130-135: Please clarify more explicitly whether the main aim of the paper is to test a specific biological hypothesis about endosymbiont-mediated aphid-plant interactions or to provide a broader proof-of-principle survey of symbiont-associated phenotypes. As it stands, the framing moves between multitrophic biology and pest-management relevance, which makes the central conceptual contribution harder to identify.

      We have rephrased this sentence as “By integrating these factors, we investigate how endosymbiont infection influences aphid fitness and aphid–plant interactions, providing insights into the ecological consequences of novel microbial associations and their potential relevance to sustainable pest management.”

      (3) Line 205: For the metabolomic analysis, please consider adding a formal multivariate test for strain effects, especially for the comparisons shown in Figure 2C and 2D, or otherwise interpret the PCA more cautiously as descriptive rather than inferential.

      We have emphasized the descriptive component and rephrased this sentence as “However, within either area, there was no clear separation among aphid strains (Figures 2C & 2D), suggesting broadly similar metabolomic profiles among strains of the same aphid clone carrying different symbionts.”

      (4) Please clarify how the top and bottom feeding leaves were handled analytically in the analyses, and explain the rationale for collapsing them into a single "feeding area" category. If possible, it would be helpful to show whether leaf position itself influenced the plant-response patterns.

      We combined the upper and lower leaves together to provide a representative measure of the plant-level responses, rather than focusing on responses at a particular leaf position. This approach was consistent with the main objective of our study, which was to investigate whole plant responses to aphids hosting different endosymbionts, rather than differences in responses among different plant parts. In addition, combining the two portions provided sufficient plant material for the GC-MS analysis and helped ensure reliable metabolite measurements from the same material. Because the two leaf positions were combined prior to GC-MS analysis, we were unable to separately test the effect of leaf position on the metabolomic response in this dataset. We have clarified this point in the revised manuscript in Materials and methods section (“Plant defense responses”).

      (5) Line 228: The use of 19{degree sign}C and 25{degree sign}C is not unusual in aphid work, but it would still help the reader if the manuscript stated more explicitly why these two temperatures were chosen in this study.

      We selected 19°C and 25°C because they represent contrasting temperature conditions within the range suitable for Russian wheat aphid development, allowing us to assess whether temperature influences Rickettsiella transmission and population dynamics. In particular, our previous observations indicated differences in the rate of Rickettsiella spread between these temperature conditions (Gu et al., 2023).

      Reference:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      (6) Lines 390-392: It would help to discuss more explicitly how the relatively modest effects in the individual life-history assays relate to the clearer signals seen at the whole-plant and population level.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with endosymbiont infection. In contrast, we also examined aphid performance at the population level (Figures 3 & 4), where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress. Under these conditions, the effects of endosymbiont infection differed from those observed in the individual assays, with Rickettsiella-infected aphids showing greater population growth and Regiella-infected aphids showing reduced population growth.

      These results suggest that the effects of endosymbionts on aphid fitness are context-dependent and may become more pronounced as population density increases and density-dependent stress develops. Thus, relatively modest effects observed at the individual level where experiments are often undertaken may not translate to differences at the population level, and plant level effects may subsequently influence feeding pressure and plant damage. We have added this perspective in the Results and Discussion sections.

      (7) Line 678 & 688: In both whole-plant experiments, please explain how 3 and 4 replicate plants were selected.

      The replicate plants were randomly selected from the available plants for each treatment to minimize potential selection bias. We have clarified this procedure in the revised manuscript in Materials and methods section.

      (8) Lines 686-690 / Figure 4: In the Regiella whole-plant experiment, the Methods state that 16 plants were established per treatment and that 4 plants per treatment were removed at each time point (days 4, 8, 12, and 16). However, in Figure 4B-D, day 12 appears to include 5 data points. Please clarify this apparent mismatch between the described sampling scheme and the data shown in the figure.

      We thank the reviewer for pointing out this and we have corrected this mistake. We initially established 16 plants for the wild type and 17 plants for the Regiella treatment. Four plants per treatment were originally planned to be sampled at each time point (Days 4, 8, 12, and 16). However, because Day 12 was a key time point at which an obvious difference in plant damage was observed between the treatments, we selected one additional plant each treatment for measurement at Day 12, resulting in five data points for this treatment at that time point. The remaining plant was therefore measured at Day 16. We have clarified the sampling procedure in the revised Materials and methods section and figure legend.

      (9) Figure 2A and Figure 4A: These schematics are helpful overall, but the brown supporting sticks stand out quite strongly and may make the panels a little harder to interpret at first glance. I wonder whether they could be simplified, made less prominent, or replaced with photographs of the actual setup if those are available.

      We have revised Figures 2A and 4A to simplify the supporting structures and reduce their visual prominence.

      Reviewer #2 (Recommendations for the authors):

      Some of the statistics were not entirely clear to me, particularly the tests reported in the results which differ from what is described in the figure legends:

      (1) Lines 306-308: "Total nymph numbers were higher on Rickettsiella aphids from Day 21 to Day 25" and indicates that this is based on GLM testing, but the figure does not indicate statistical significance, and the figure legend states independent sample t-tests were used. Similar comment for the following paragraph and corresponding figure.

      The GLMs were used to test the overall patterns in nymph numbers across the relevant time periods, including Days 21–25, rather than testing each time point independently. We also conducted independent-sample t-tests to assess differences between treatments at individual time points. We have clarified this distinction in the Statistical section.

      (2) Line 43: Does not seem like the appropriate reference (reference is on plant virus transmission, not salivary toxins).

      We thank the reviewer for pointing this out. We have removed this reference and replaced it with reference 34 (Luna et al., 2018) that directly supports the statement regarding aphid salivary toxins.

      Reference:

      Luna et al., Bacteria associated with Russian wheat aphid (Diuraphis noxia) enhance aphid virulence to wheat. Phytobiomes J 2, 151-164 (2018).

      Reviewer #3 (Recommendations for the authors):

      Minor editorial comments:

      (1) Figure S10 - the figure legend needs improvement as the current version does not help the reader understand the figure. Please also include a key.

      We have changed the figure legend with reference to our aim, and also explained use of the Cp values. “Figure S10. Rickettsiella Cp values in (A) routine screening of laboratory Rickettsiella colonies and (B) the mixed cage experiment assessing endosymbiont frequency changes over time at 19 °C and 25 °C. The dark red area represents overlap between the 19°C and 25°C experiments. Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA. This figure shows the typical range of Cp values observed in our laboratory Rickettsiella -infected aphid colonies. We used this range as a reference for identifying aphids that acquired Rickettsiella through horizontal transmission, as horizontally infected aphids generally showed much higher Cp values than vertically infected aphids.”

      (2) Define Cp.

      We have explained it as “Cp values represent the quantification cycle values obtained from qPCR, with lower Cp values indicating a higher amount of Rickettsiella target DNA.”

      (3) Supplemental Information: Line 113 - T is missing from Table.

      This has been added.

      (4) Move Table S2 to the main paper - this table provides a useful summary of the work.

      This has been moved.

      Main Manuscript:

      (1) Line 145: Serratia was not detected at G0 - was it detected later? It seems possible that titer could be very low to begin and increase in later generations; please clarify.

      We have previously monitored the aphid populations for the presence of Serratia transinfected from the same donor resource across subsequent generations, and Serratia was not detected at any later generation. We also did not detect Serratia in the other aphid species we examined, including green peach aphids (Gu et al., 2023 & 2025) and oat aphids (Yang et al., 2026). Therefore, we believe that the absence of Serratia at G0 was not due to a very low initial titer followed by an increase in later generations but instead that Serratia had been lost from the aphid population.

      References:

      Gu et al., A rapidly spreading deleterious aphid endosymbiont that uses horizontal as well as vertical transmission. Proc Natl Acad Sci USA 120, e2217278120 (2023).

      Gu et al., Transinfections of the endosymbiont Rickettsiella viridis in different Myzus persicae (Hemiptera: Aphididae) clones show consistent deleterious effects and stable transmission. J Econ Entomol 118, 1544-1552 (2025).

      Yang et al., A Rickettsiella transinfection in Rhopalosiphum padi reduces fitness and alate production but not plant virus transmission. Pest Man Sci 82, 3894-3906 (2026).

      (2) Line 206: "different aphid strain" - I learned from the manuscript that a single clone of D. noxia is found in Australia. Further, from my reading of the manuscript, I understand that one clonal isolate was propagated and then infected with symbionts. I think it is important to reword this sentence so that it is clear that the aphid genetic background is held constant, and that the only differences here are the presence or absence of the different symbionts. My reaction to this sentence was that you are working with the same aphid strain hosting different symbionts.

      We have clarified it by adding “the same aphid clone carrying different symbionts” after the different aphid strains.

      (3) Measuring symbiont "density" is a tricky thing to do; I explain this above in the public review. I suggest considering some revisions to the manuscript to be sure that you accurately reflect what has been measured and what can reasonably be inferred from using qPCR to measure gene copy number.

      We agree that qPCR-based measurements of symbiont gene copy number have limitations. In this experiment, we had a relatively large number of samples, and qPCR provided a practical and efficient approach for assessing variation in endosymbiont abundance and infection status among samples. While we acknowledge the limitations of this approach, the relative differences in gene copy number can still provide an indication of variation in endosymbiont abundance and infection status among treatments. Unfortunately, other approaches remain challenging given resource and expertise limitations.

      (4) I think that you may be undervaluing the results of the mixed infection experiments; I find them to be compelling. To me, the data suggest that the symbionts increase aphid fitness.

      In our experiments, we assessed aphid fitness under different environmental conditions. In the individual fitness assays conducted on cups with a single wheat plant (Figures S5 & S6), aphids were maintained under relatively favourable conditions, with limited environmental stress and without substantial crowding. Under these conditions, we observed an increase in fitness associated with symbiont infection. However, we also examined aphid performance under population-level conditions, where populations were established from a small number of aphids and subsequently experienced increasing crowding and density-dependent stress (Figures 3 & 4). Under these conditions, the effects on fitness were different from those observed in the individual assays. We therefore agree that our results suggest that symbionts can increase aphid fitness under some conditions, but that this effect may be context-dependent and can differ under population-level conditions where density-dependent stress occurs. We have also added this information to our Results section to make this clear to readers.

      (5) Lines 288-291: This sentence doesn’t make sense to me. What "minor fitness costs" are being referred to? If the infected lines are increasing in representation relative to the uninfected lines, that suggests that there are not fitness costs, but fitness advantages under the experimental conditions.

      We have rephrased it to “minor fitness effects” which we refer to the fitness test under favourable conditions.

      (6) The section that begins on line 293 - I find this part to not be particularly robust and suggest dropping it from the paper.

      We appreciate the reviewer’s concern regarding the robustness of this section. We included this experiment to monitor changes in aphid population over time and to help explain the differences in plant feeding damage observed between aphids carrying different endosymbionts. Importantly, we conducted this experiment using whole wheat plants to evaluate whether the patterns observed in our other experiments were also evident. We believe that these results provide important complementary evidence for interpreting the differences in plant damage among the aphids hosting different endosymbionts and therefore are relevant to the overall conclusions of the study. For this reason, we prefer to retain this section in the manuscript.

      (7) Line 298: three plants per time point - I do not think this sample size is reflected in the methods of the paper.

      We have mentioned in the method part with “At day 14, three replicate plants were randomly selected from the available plants for both treatments and we counted the total number of nymphs, alate adults, and apterous adults. This was repeated again at Days 21 and 25”.

      (8) Line 304: "the frequency of alates decreased as plant damage increased" - this seems to be counterintuitive!

      As plant damage increased, the total aphid population also increased, resulting in an increase in the absolute number of alates. However, the frequency (proportion) of alates decreased because the increase in the total aphid population was greater than the increase in the number of alates.

      (9) Line 404: "horizontal transmission through plant tissue and/or transfer via aphid contact or honeydew" - what evidence is there that this happens? I have not kept on top of the literature with respect to transmission of secondary symbionts in aphids, but back when I was very familiar with that literature, the data did not support transmission by any of those routes. If there is now evidence supporting transmission by these routes, please cite it here.

      Previous studies have provided experimental evidence that horizontal transmission of aphid-associated endosymbionts can occur through plants. For example, plant-mediated transmission has been demonstrated for direct detection of secondary endosymbiont in the plant tissue including Hamiltonella defensa (Li et al., 2018), Rickettsia (Shi et al., 2024) and Serratia symbiotica (Pons, et al., 2019a). Our previous research also demonstrates that Rickettsiella endosymbionts were detected in uninfected aphids after feeding by infected aphids regardless of physical contact (Gu et al., 2023). Endosymbionts have also been detected in aphid honeydew (Darby and Douglas, 2003) and some primary transmission route appears to be horizontal, through honeydew (faeces) and host plant phloem (Pons, et al., 2019a & 2019b; Perreau et al., 2021). Appropriate references have been added to the revised manuscript in the Discussion section.

      References:

      Li et al., Plant-mediated horizontal transmission of Hamiltonella defensa in the wheat aphid Sitobion miscanthi. J Agric Food Chem 66, 13367-13377 (2018).

      Shi et al., Rickettsia transmission from whitefly to plants benefits herbivore insects but is detrimental to fungal and viral pathogens. mBio 15, e02448-23 (2024).

      Pons et al., Circulation of the cultivable symbiont Serratia symbiotica in aphids is mediated by plants. Front Microbiol 10, 764 (2019a).

      Darby and Douglas, Elucidation of the transmission patterns of an insect-borne bacterium. Appl Environ Microbiol 69, 4403-4407 (2003).

      Pons et al., New insights into the nature of symbiotic associations in aphids: infection process, biological effects, and transmission mode of cultivable Serratia symbiotica bacteria. Appl Environ Microbiol 85, e02445-18 (2019b).

      Perreau et al., Vertical transmission at the pathogen-symbiont interface: Serratia symbiotica and aphids. mBio 12, e00359-21 (2021).

      (10) Line 511: revise to "to measure their relative densities relative to a host gene".

      We have revised it.

      (11) Throughout the manuscript, please replace "five aged-matched" with an accurate description of the aphids used in the experiment. Please pay particular attention to the figure legends. Simply state e.g. "five 10-day-old apterous females" etc.

      This has been replicated in the main manuscript and figure legends.

      (9) Line 727 - lowercase t for Tests.

      This has been added.

      (10) Line 738 - what happens when you don’t exclude the early time points? Do your significant results go away? Also, please define what is meant by "early time points".

      When the early time points (Day 14 for Rickettsiella and Day 4 for Regiella) were included in the analysis, the GLM still showed a significant effect of strain on alate production for Rickettsiella (F<sub>1,12</sub> = 26.434, P < 0.001). We excluded these early time points from the analysis presented in the manuscript because, at these stages, aphid population sizes were still similar between treatments. We also conducted independent-sample t-tests at individual time points and observed differences at the later time points, when aphid population sizes began to diverge. We therefore considered the later time points to be more informative for assessing fitness effects and population sizes under increasing crowding conditions. We have now clarified this as “The earliest time points in the experiments for Rickettsiella (Day 14) and Regiella (Day 4) were excluded” in the manuscript.

      (11) In the legends of all figures, please be explicit about sample sizes.

      We have added the relevant information about replicate number or sample sizes in the main and supplementary figures.

      (12) Line 1010 - replace "each leave" with "each leaf" - there was at least one other place, I think in the supplemental information, that leave was used instead of "leaf".

      We replaced this.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Weaknesses:

      Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

      (1) Bias and representations of data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

      (2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TErelated mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

      (3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated. Comparisons with other noncoding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

      (4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in drosophila, worms, mouse, and human. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied, for their functions and corss-species conservations. The authors should explicitly show what is new here in their analyses.

      (5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

      (6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

      (7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

      Comments on revisions:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We thank Reviewer #1 for the constructive comments and recognition of the value of our ncORF atlas. We have addressed the key concerns by strengthening the analyses and clarifying the framing and limitations of our conclusions. We appreciate the reviewer’s support for publication and believe these revisions have further improved the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

      (2) Some analytical methods and standards were not clearly presented in the manuscript.

      We thank Reviewer #2 for the positive assessment of our comprehensive dataset and standardized analytical framework. We have clarified the analytical methods and criteria throughout the manuscript and better defined the scope and limitations of our bioinformatics-based analyses. We appreciate the reviewer’s constructive suggestions, which have helped improve the clarity and rigor of the manuscript.

      Recommendations for the authors:

      Reviewing Editor:

      We have evaluated the revision together with the reviewers' second-round assessments and your responses. The reviewers agree that the manuscript has improved in clarity and that the standardized integration of large-scale Ribo-seq datasets provides a valuable resource for the field. However, several important concerns remain insufficiently resolved. In multiple cases, the revision relies primarily on acknowledgment or reframing of limitations rather than additional analyses or clearer methodological justification, leaving the evidential support for several conclusions incomplete.

      Because the study is entirely computational, all analytical procedures, criteria, and thresholds should be explicitly defined and adequately justified to meet the expected standard of technical rigor. In particular, key components of the analytical framework require clearer description, including the definitions and criteria used for co-translation and ncORF classification. The limitations of the dataset should also be discussed more explicitly, especially regarding the heterogeneity of public Ribo-seq datasets, technical factors influencing detection sensitivity, and the interpretation of variable detection frequencies across samples.

      In addition, several conclusions remain largely descriptive, and the distinction between novel findings and confirmation of previous observations should be clarified more carefully. Conclusions should be framed strictly within the limits of the presented data and positioned appropriately relative to prior work, with suitable citation to avoid overstating novelty.

      We therefore ask the authors to refine the technical descriptions, ensure that all methods and analytical criteria are presented unambiguously, and expand the Discussion to clearly articulate the limitations of the dataset and analysis. The conclusions should also be revised to reflect an appropriately cautious interpretation of the findings.

      The primary strength of this study lies in the scale and standardization of the dataset as a community resource. Given this substantial resource value, we believe the manuscript could become suitable for publication provided that the issues outlined above are addressed clearly and transparently. With these revisions, the work will provide a useful foundation for future studies in this area.

      We therefore invite you to submit a final revised version addressing the points described above.

      We thank the Editor for the careful assessment and constructive guidance. In the final revision, we have clarified all key methodological definitions and analytical criteria, expanded the Discussion of dataset and detection limitations, and revised the conclusions to avoid overstatement. We believe these changes improve the technical rigor, transparency, and overall value of the manuscript as a community resource.

      Reviewer #1 (Recommendations for the authors):

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

      We appreciate the reviewer’s support for publication.

      Reviewer #2 (Recommendations for the authors):

      While the authors have made commendable efforts to revise the manuscript and address previous concerns, the revised version still falls short of fully resolving several critical issues regarding data interpretation and methodological transparency. I recommend the following modifications before the manuscript can be considered for publication:

      (1) Although the authors have annotated the detection frequencies of individual sORFs in the revised

      Supplementary Table 3 and Fig. S1B, the biological and technical implications of these data require further clarification:

      (a) As the data demonstrates, even the most abundant sORFs were detected in no more than half of the samples. If the authors attribute this low detection rate to technical limitations (e.g., batch effects, sequencing depth, or threshold stringency), this must be explicitly discussed and annotated in the main text to prevent readers from misinterpreting this as low biological penetrance.

      We thank the reviewer for this constructive comment. We have added a paragraph to the Discussion explicitly addressing the potential technical factors underlying the variable detection frequencies to avoid overinterpreting these frequencies as biological penetrance.

      (b) The authors did not fully address my previous query regarding tissue specificity. Given the diverse and complex origins of the analyzed ribo-seq datasets, it is crucial to know whether any of these sORFs are tissue-specifically translated. The authors should analyze and state whether certain sORFs are exclusively detected in specific sample categories (tissues/organs), and whether this translation pattern aligns with the tissue-specific expression of their corresponding host transcripts.

      We thank the reviewer for raising this important point. Strict tissue-exclusive translation is difficult to establish from heterogeneous public Ribo-seq datasets, as gene expression is inherently stochastic and most genes have some probability of being expressed across tissues, although expression levels may vary substantially between tissues. Moreover, failure to detect an ncORF in a given tissue may reflect low expression or insufficient sequencing depth rather than true biological absence. We therefore avoid making definitive claims about tissue-specific translation and instead quantify variation in ncORF expression across tissues using a tissue specificity index.

      (2) Echoing the concerns raised by Reviewer #1, I remain concerned that the observed uORF-CDS cotranslation might be a computational artifact or false positive. The current Methods section lacks sufficient detail on how "co-translation" was strictly defined and quantified. I strongly recommend that the authors include a schematic diagram (e.g., in Figure 6 or supplementary figures) that explicitly details their analytical strategy, statistical thresholds, and evaluation criteria for defining co-translation. As was pointed out, it is well-established that uORFs typically exert an inhibitory effect on the translation of the main CDS. To validate the accuracy and robustness of their analytical pipeline, the authors should use their collected dataset to demonstrate the prevalence and nature of this canonical inhibitory phenomenon. Successfully capturing this expected repression would serve as a crucial positive control for their methodology. While the authors provided a theoretically acceptable mechanistic model in the text to reconcile co-translation with uORF-mediated repression, this hypothesis currently lacks literature support. The authors must cite relevant prior studies that support this specific regulatory dynamic to strengthen their argument.

      We thank the reviewer for this important comment. We have further clarified the definition, detection criteria, and statistical framework for co-translation in the Methods, and have made the complete analysis code publicly available to facilitate reproducibility and independent evaluation. We have also added relevant literature supporting the proposed regulatory interpretation clarifying the relationship between our observations and the established inhibitory effects of uORFs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents valuable findings regarding cardiac and autonomic effects of seizures and epilepsy, with relevance to sudden unexpected death in epilepsy (SUDEP). They present solid evidence that genetic deletion of the potassium-chloride cotransporter in hypothalamic corticotropin-releasing hormone (CRH) neurons exacerbates bradycardia and enhances autonomic disturbances in a mouse model of temporal lobe epilepsy. However, the evidence that this deletion produces chronic hyperexcitability of the hypothalamic-pituitary-adrenal axis was incomplete, leaving a mechanistic gap. This work will be of interest to neuroscientists working on epilepsy, the HPA axis, and autonomic control.

      We thank the editors and reviewers for their feedback. Although the loss of Kcc2 from CRH neurons in Kcc2/Crh mice has been confirmed (Melon et al., 2018) and leads to HPA axis hyperexcitability in response to stress or seizures, it does not chronically drive HPA axis hyperactivity in unstressed conditions (Basu et al., 2024).

      The following details have been added regarding the Kcc2/Crh model. We now describe the Kcc2/Crh as “hyperreactive” rather than “hyperexcitable/hyperactive” throughout the manuscript.

      -In the Introduction: “This loss of Kcc2 in CRH neurons has been previously confirmed and shown to cause an exaggerated HPA axis response to stress that is absent in baseline conditions (Basu et al., 2024; Melon et al., 2018)”

      -In the Discussion: “This aligns well with lack of elevated plasma corticosterone at baseline in Kcc2/Crh mice, compared to WT, because elevated PVN<sup>CRH</sup> neuron activity should otherwise increase this signal (Basu et al., 2024).”

      -In the Discussion: “Most notable, our model utilizes a developmental strategy to knock out Kcc2 from CRH neurons, which has been confirmed previously (Melon et al., 2018).”

      Public Reviews:

      Reviewer #1 (Public review):

      This study could be improved with a more thorough assessment of heart rate, blood pressure and breathing during and following the seizures, and in particular the fatal event.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2. Overall, pronounced bradycardia that occurred near seizure termination was followed by recovery of HR to pre-ictal baseline in the early post-ictal period (30 sec). In Results: “Independent of genotype, HR recovered to baseline levels during the immediate post-ictal period (0-10 sec, Fig. 2G; 10-30 sec, Fig. 2K).

      In pilot work, we determined that HR during spontaneous seizures were fundamentally different than HR during status epilepticus (see Author response image 1). Therefore, it was critical for us to examine HR during spontaneous seizures. We previously published that Kcc2/Crh+KA mice have a rate of 1-2 seizures per day. To limit additional stressors and seizure provocation, we employed radio telemetry (over tethered systems) and opted out of carotid instrumentation for BP as well as restricted environments of plethysmography chambers for these assessments of HR. After determining that heart rate was different between our mouse lines, we tested whether this change was mediated by central circuits regulating HR (ie: baroreflex, Bezold-Jarisch reflex) (Fig. 3, 4, 5).

      Author response image 1.

      In Discussion: “Although the present study includes only non-fatal seizures, our report of HR during spontaneous seizure events supports work suggesting physiological events during non-fatal seizures predict SUDEP risk (Lamrani et al., 2023; Ryvlin, Nashef, & Tomson, 2013; Schuele et al., 2011). It remains to be determined if ictal events during fatal and non-fatal spontaneous seizures are different and future studies could help clarify any distinctions. Our examination focused on HR (and not respiration or BP). Whether exaggerated BJR-mediated HR response co-occurs with greater magnitude BJR-mediated apnea and hypotension remains to be determined.”

      It is unclear if the bradycardias were spontaneous or a result of preceding central or obstructive apneas, oxygen desaturations, hypercapnia, arrhythmias, or other possible triggers.

      Our work demonstrates the occurrence of ictal bradycardia whereas identifying precipitating factor(s) and interaction(s) of this phenomenon will require alternate approaches. Normal activation of hypoxic and hypercapnic ventilatory responses would be expected to increase HR. However, whether these chemoreflex circuits undergo remodeling in Kcc2/Crh mice remains unknown. Obstructive apnea via laryngospasm can cause reflex bradycardia, but we did not record airflow or respiratory EMG in these studies to determine the existence of obstructive apnea. More testing is merited.

      In Discussion, “…Additional seizure-related disturbances such as central or obstructive apneas may contribute to BJR activation. Hypoxia, which could result from ictal apnea, is known to increase excitatory neurotransmission to cardiac vagal motor neurons that cause vagal bradycardia (Griffioen et al., 2007) and induces platelet activation (Tyagi et al., 2014) which is considered the main source of circulating serotonin for the BJR. As such, hypoxia resulting from apnea may increase likelihood of exaggerated BJR during seizures. Consistent with this…”

      Considerable prior work in the literature suggests SUDEP could be mediated, in some patients, by a burst of parasympathetic activity to the heart. Were the heart rate changes in these animals during seizures inhibited or blocked by atropine or atenolol?

      With this study targeting spontaneous seizures we were unable to test acute pre-treatment with atropine or atenolol. We did observe reduced mortality in Kcc2/Crh+KA mice that underwent chronic parasympathetic blockade via osmotic minipump of methylscopolamine (Figure 5).

      In Discussion:

      “Chronic inhibition of vagal parasympathetic motor output (the driver of BJR reflex bradycardia) improved mortality by 10% in Kcc2/Crh mice. Although this improvement provides some hope for patients at high risk for SUDEP with no treatment options, additional avenues of investigation are needed to more directly link seizure-related bradycardias to vagal parasympathetic motor output.”

      The injection of the 5HT agonist phenylbiguanide into the right jugular is not a selective approach for activating the Bezold Jarisch Reflex (BJR), which is caused by increased activity of intracardiac sensory neurons (generally activated with is chemia or a combination of low preload with high contractility). The results should be interpreted more cautiously, as a response to systemic administration of phenylbiguanide only.

      BJR can be experimentally triggered with intravenous infusion of various compounds including veratrum alkaloids (Cramer, 1915), 5HT (Fozard 1983), or 5HT3R agonists (Verberne & Guyenet, 1992). We now specifically refer to BJR in our study as that induced by PBG (a 5HT3R agonist), as others have done (Yamano et al., 1995; PMID: 8786638) and acknowledge endogenous BJR activation in the discussion.

      Added to the results: “As dysfunction of serotonergic signaling is implicated in the pathophysiology of SUDEP (Richerson & Buchanan, 2011), we investigated the Bezold Jarisch Reflex (BJR) (Fig. 5), a cardioinhibitory reflex that is reliably triggered experimentally by activation of cardiopulmonary vagal afferents containing serotonin type 3 receptors (5HT3R) (Fozard 1983; Yamano et al., 1995).”

      Added to Discussion: “Although our report is the first to link BJR to SUDEP, serum serotonin levels are elevated following generalized seizures (Murugesan et al., 2018), likely via release from activated platelets (Cloutier et al., 2018). This surge in serum serotonin could lead to endogenous activation of BJR, as bolus intravenous infusion of serotonin reliably triggers BJR experimentally (Fozard 1983; Whalen et al., 2000).”

      Reviewer #2 (Public review):

      Some of the conclusions may be a bit overstated as is and would benefit from more discussion and perhaps additional data.

      Post-ictal HR data are now included in the Results, Table 2.1 and Figure 2.

      The Discussion now includes more details regarding respiration, BP, obstructive apnea, properties of the BJR, and distinctions in seizure type based on whether evoked or lethal.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Weaknesses:

      (1) In addressing the question of residual body participation in sorting of organelles, a clear definition of this structure is required including when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. The authors' definition is as follows: 'The RB originates from the collapse of the maternal parasite during daughter cell budding and occupies the space previously occupied by the mother cell.' As such, a clear marker of the mother cell 'collapse' is required, but such a marker is not identified or used in the study to separate what might be considered an active part of the mother cell during early daughter formation, and the residual body. This might seem like moot a point, but it would help to give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? The authors elegantly show that MyoF is necessary for segregation of micronemes and rhoptries into daughters, and that MyoF depletion leads to accumulation of these organelles within the residual body. Moreover, restored expression of MyoF can then recover these organelles. This clearly demonstrates the activity of the residual body as part of the syncytium space that participates in the maintenance of the vacuole. But does it imply that this space necessarily handles all inherited micronemes and rhoptries as a 'trafficking hub'? My concern with the lack of a clear definition could provide some misinterpretation or overinterpretation of the contribution residual body.

      We thank the reviewer for raising this conceptual point. We agree that, in the absence of a molecular marker that uniquely defines the nascent RB, the precise transition between posterior maternal cytoplasm and a morphologically distinct RB cannot be determined during early daughter formation. We have therefore clarified our terminology in the revised manuscript and define the RB operationally as the posterior compartment/connection between daughter parasites. We also avoid assigning early posterior trafficking events unambiguously to a fully formed RB. We further agree that the current data do not establish that all inherited micronemes and rhoptries must transit through the RB. We have therefore revised the Results and Discussion and softened terminology such as “central trafficking hub” and now conclude that the RB represents an important dynamic compartment in organelle recycling and redistribution, without implying that it is an obligatory intermediate for every inherited secretory organelle.

      (2) A further, remarkable conclusion is that maternal micronemes are evenly segregated into daughters through an active process for 'balanced microneme inheritance'. The proportion of maternal micronemes is quantified up to the 8-cell stage and shown to be not significantly different between cells. But would this result be expected with random assortment at this stage? The authors model the probability of a 32-cell stage vacuole occurring with each daughter having within 0-3 maternal micronemes and this is considered unlikely. However, the authors neither present the modelling for the 8cell stage or show quantification of 32-cell vacuoles. They do show some images of large vacuoles, but it is not possible to determine the distribution of maternal micronemes in these images. A regulated process of segregation would require a complex mechanism where some form of microneme counting would be required to create the proposed balance. It is, therefore, important to have strong data supporting such a hypothesis, but this is not currently presented.

      We agree that the previous wording implied a mechanistic conclusion beyond what can be established from the present dataset. We have therefore revised the manuscript so that the relatively even distribution of maternal micronemes at the 8-cell stage is presented as an observation that is consistent with a non-random or regulated partitioning process, rather than evidence for an established microneme-counting mechanism. The 32-cell model is now presented as supportive rather than definitive evidence, and we explicitly acknowledge that the quantitative experimental dataset was obtained at the 8-cell stage. 

      Reviewer #2 (Public review):

      (1) The second half of the paper focuses on a more detailed characterization of microneme and rhoptry recycling. The authors strongly argue that the RB is a central hub for recycling micronemes and rhoptries; however, this conclusion is not fully supported by the data. For example, the authors state…Thus, the model that all microneme and rhoptry trafficking is RB-dependent is based primarily on the MyoF depletion phenotype (which results in RB accumulation) together with the observation that a relatively small amount of maternal microneme and rhoptry material is detectable in the RB of wild-type parasites. Although the authors' interpretation-that recycling is RB-dependent-is one possible explanation, alternative models are not discussed. For example, an alternative possibility is that the majority of micronemes and rhoptries are trafficked directly from the apical end of the mother parasite to the daughter cells without passing through the RB. In this scenario, only a subset of the organelles would enter the residual body, perhaps reflecting imperfect trafficking efficiency rather than an obligatory recycling step. Loss of MyoF would impair this trafficking pathway, resulting in the accumulation of secretory organelles within the RB. In other words, RB accumulation could be a consequence of MyoF depletion rather than evidence that all trafficking in wild-type parasites normally proceeds through the RB. This alternative interpretation seems particularly relevant for the rhoptries, given that the authors themselves state that "M-RON2 was integrated into daughter rhoptries prior to mother cell collapse and formation of the RB."

      We agree with the reviewer that the current data do not establish obligatory transit of all maternal micronemes and rhoptries through the RB. We have revised the manuscript throughout to make this distinction explicit. In particular, we now emphasize that maternal MIC2 can be directly observed entering the RB, whereas most maternal RON2 is incorporated into daughter rhoptries before mother-cell collapse and formation of a morphologically distinct RB. This observation leaves open the possibility that a substantial fraction of maternal rhoptries is transferred directly from the mother to developing daughters. We have also revised our interpretation of the MyoF-depletion phenotype. The accumulation of maternal MIC2 and RON2 in the RB following MyoF depletion demonstrates that MyoF is required for efficient redistribution of both organelle populations, but does not by itself demonstrate that both normally follow an identical spatial route through the RB. We now explicitly state that their precise trafficking routes and timing may differ.

      We nevertheless retain the conclusion that the RB is a dynamic compartment involved in organelle recycling because maternal MIC2 can be directly observed entering and leaving this compartment, and material accumulated there following MyoF depletion can subsequently be redistributed after restoration of MyoF.

      (2) Figure S10C. To determine whether microneme degradation occurs in the RB, the authors quantified the fluorescence intensity of individual micronemes in control parasites and following auxin washout, showing that after redistribution the fluorescence intensity of individual vesicles is unchanged. However, this is not the appropriate analysis to address the question being asked. To conclude that micronemes are not degraded, the authors would need to quantify the total fluorescence intensity within the entire vacuole. For example, if half of the micronemes were degraded, the remaining micronemes would be expected to retain the same fluorescence intensity as those in the control parasites. Thus, unchanged fluorescence intensity of individual vesicles does not exclude the possibility that degradation has occurred.

      We agree with this criticism and have revised the interpretation of the experiment accordingly. We no longer conclude that the analysis excludes microneme degradation. We now explicitly acknowledge that analysis of individual recovered micronemes cannot exclude degradation of a fraction of the total microneme population during RB retention. Thus, the experiment supports preservation of MIC2 signal in the recovered organelles but is no longer presented as evidence that no microneme degradation occurs.

      Reviewer #3 (Public review):

      Weakness:

      The inability to achieve higher temporal resolution due to phototoxicity precluded tracking of single micronemes, thus it remains possible that some micronemes follow a path similar to rhoptries and enter daughter cells before development of the residual body while others are recycled via the residual body. 

      We agree and have incorporated this limitation into the revised interpretation. Our live imaging demonstrates that maternal MIC2 can enter the RB and subsequently redistribute to daughter parasites, but it does not establish that every individual maternal microneme follows this route. We thank the reviewer for highlighting this distinction, which has helped us clarify the model presented in the Results and Discussion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Carruthers and Sibley, 1997, is still the cited ref for Tic20 in the apicoplast despite the authors saying that would correct this to a van Dooren publication.

      Corrected

      (2) In Figure 2B three biological replicates were used however there are no error bars shown. The only reason for performing replicates is the observe the variance in the data, so if this is not shown the replicates are effectively meaningless. I strongly advice that error bars are given to indicate this seeing that the data in this figure form the basis of the major conclusions of the study. It might be necessary to show this in supplemental forms with fewer proteins if the error bars are too difficult to see in the combined figure.

      We agree and have corrected Figure 2B to display the variability between the three independent biological replicates. Error bars now represent the standard deviation.

      (3) Line 136: Can you conclude that these inheritance patterns are 'organelle-specific' when each organelle is only sampled with one or two proteins. Isn't it better to conclude that these are protein-specific, with the hypothesis that they might represent the orgnalle as a whole. I imagine that some proteins in organelles such as the apicoplast have shorter half-lives than others, and therefore some apicoplast proteins might behave like 'Group 3' proteins.

      We thank the reviewer for raising this point. We agree that individual proteins within the same organelle may differ in their turnover kinetics and that analysis of one or two markers cannot establish that every molecular component of an organelle behaves identically. However, we do not think that describing the observations exclusively as protein-specific inheritance would fully reflect the biological process investigated here. The proteins analysed are established markers of defined organelles, and our conclusions are based not only on changes in fluorescence intensity, but also on the localization, morphology, partitioning, and spatial relationship between maternally inherited and newly synthesized organelle populations.

      This is particularly evident for micronemes and rhoptries, where maternal and de novo material remain spatially separated and individual organelles can be followed during inheritance.

      In addition, the microneme phenotype observed with MIC2 was confirmed using AMA1, MIC4, and MIC8.

      We therefore retain the terminology of organelle inheritance, while acknowledging that individual proteins within a given organelle may exhibit different turnover kinetics and that the markers analysed may not represent the behaviour of every molecular component of the organelle.

      (4) Line 222: It is an odd phrase to suggest that the Golgi, ER etc 'bypass' the residual body, which suggests an active avoidance mechanism. Would the authors also conclude that the nucleus 'bypasses' the RB? Moreover, the ER and mitochondria are actually known to be present in the RB forming continuous organelles between daughters in a vacuole. So again, this might be an overstatement that mispresents how the RB participates in vacuole functions.

      We agree and have removed the term “bypass.” The revised text now states only that we did not observe comparable accumulation of the analysed Golgi, ER, or apicoplast markers in the RB during inheritance. This avoids implying an active avoidance mechanism and is compatible with the known continuity of ER and mitochondria through the RB.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      Figure 8F and Video S6: The authors should specify the time point after IAA washout at which live imaging was initiated. Does time 0 in the video correspond to the point at which IAA was removed?

      We have clarified this in the Results and Methods. Auxin was removed after 24 h of replication, and live imaging was subsequently initiated. Time 0 in Figure 8F and Video S6 corresponds to the first acquired frame after auxin washout.

      Figure S10 should read auxin, not auxine.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      I have no further suggestions. Congratulations to the authors for a lovely study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      The major revisions include:

      (1) Conceptual framing and scope: defined canalization in the Introduction and clarified that our conclusions are restricted to limited NuRE plasticity during one growing season under a single moderate salinity treatment.

      (2) Methods and classification: clarified the substrate composition and elemental measurements, specified that the ecotype analysis included only Chinese populations, and explained the partial association between ecotype and phylogeographic group.

      (3) Interpretation: expanded the discussion of K resorption and inverted nutrient limitation and tempered the interpretation of latitude and the substantial unexplained variation.

      (4) Robustness and presentation: added Supplementary Figure S7 showing that carbon standardization did not alter the main conclusions, added significance symbols to Table 1, corrected the unit in Figure 2b, and revised repetitive wording in the Discussion.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      (R1-P1) First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long-term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study’s main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      We agree that our evidence is limited to the absence of NuRE plasticity during one growing season under the imposed salinity treatment and does not resolve chronic or multigenerational responses or their physiological basis. We therefore revised Discussion 4.1 to delimit the canalization inference, identify the proposed mechanisms as untested, and specify the longer-term and mechanistic studies needed to distinguish among them.

      Discussion 4.1, fourth paragraph, inserted immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied ”.

      “Our inference of canalization is therefore limited to the absence of a plastic NuRE response during one growing season under the imposed salinity treatment. Chronic, more severe, or multigenerational salinity exposure may produce acclimatory, epigenetic, or transgenerational responses that cannot be evaluated here. Moreover, although altered phloem loading, disruption of senescence-associated remobilization, and reallocation towards osmotic adjustment are plausible explanations for the observed response, we did not directly measure these mechanisms. Long-term experiments combined with targeted molecular and transport measurements are needed to distinguish among these possibilities.”

      (R1-P2) Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermines the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

      We agree and have clarified all three evidential limits. First, the resorbed N: P and N: K analyses are now described as indirect evidence consistent with nutrient-limitation control, not as a causal test; direct nutrient-addition experiments would be required for causal inference. Second, metabolomics is identified as validation of physiological stress rather than a genotype-specific mechanistic analysis. Third, because ecotype and phylogeographic group are partly associated, they were fitted in separate models. The genotype random effect accounts for paired measurements but does not remove confounding between the classification schemes; ecotype differences are therefore interpreted as complementary rather than independent evidence.

      Discussion 4.2, third paragraph.

      “Within this context, the consistent ‘inverted’ nutrient limitation pattern (i.e., a slope significantly >1 for the relationship between log Resorbed N:P and log Green N:P) provides indirect evidence consistent with nutrient-limitation control at the intraspecific level in P. australis, but it cannot establish causal nutrient limitation. Direct factorial nutrient-addition experiments would be required to determine whether the observed resorption patterns are driven by the relative limitation of N, P, or K.”

      Discussion 4.1, first paragraph, inserted immediately after the sentence ending “providing a robust foundation to evaluate NuRE responses”.

      “In this study, metabolomic profiling was used primarily to confirm that the salinity treatment induced broad physiological stress, rather than to resolve genotype-specific metabolic mechanisms underlying NuRE variation. Integrating metabolite profiles with genotype-level NuRE responses would be a valuable direction for future mechanistic research.”

      Methods 2.4.

      “Phylogeographic group and ecotype were analysed in separate linear mixed-effects models because ecotype classifications were available only for Chinese populations and were partly associated with phylogeographic structure. Each model included salinity treatment and either phylogeographic group or ecotype as fixed effects, with genotype fitted as a random effect to account for the paired experimental design in which each genotype was exposed to both control and salt conditions.”

      Discussion 4.4, first paragraph, inserted immediately after the sentence ending “governed by geographic origin (phylogeographic group and ecotype)”.

      “The separate-model approach avoids including the two correlated classification schemes as simultaneous independent predictors, but it does not fully disentangle deep phylogeographic history from recent habitat-associated differentiation. We therefore interpret the ecotype analysis as complementary evidence of habitat-associated differentiation rather than as an effect independent of phylogeographic history.”

      Reviewer #2 (Public review):

      (R2-P1) The experiment covers only one growing season, with salinity applied in June and measurements in December. While the stress is clearly effective, longer-term or multi-year stress might reveal acclimation or epigenetic effects that are not captured. Given the author team’s expertise in parental and transgenerational effects in clonal plants, this limitation is particularly relevant and warrants more thorough discussion in the manuscript.

      We agree. This concern overlaps with Reviewer #1’s temporal-scope comment. We have revised Discussion 4.1 to state explicitly that our inference is restricted to the absence of a plastic NuRE response during one growing season. We also acknowledge that chronic or multigenerational exposure could induce acclimatory, epigenetic, or transgenerational responses that were not captured by the present design.

      See the full revised text under R1-P1 above.

      (R2-P2) The salinity treatment uses a single moderate level of 10 ppt, which does not allow assessment of whether more extreme stress might trigger a plastic response. A dose-response design across a gradient would have provided stronger inference about the threshold at which NuRE canalization might be overcome. Additionally, the ecotype analysis in Figure 4 applies only to Chinese populations, as classification was not available for non-Chinese populations, which should be stated more explicitly in the Results.

      We agree with both points. We have revised Discussion 4.1 to acknowledge that the single 10 ppt treatment does not exclude the possibility of a plastic NuRE response at higher salinity or along a broader dose-response gradient. We therefore frame the identification of a possible response threshold as a priority for future experiments.

      We have also revised Results 3.2 to state explicitly that the ecotype analysis in Figure 4 included only Chinese populations because ecotype classifications were unavailable for non-Chinese populations.

      Discussion 4.1, fourth paragraph, at the same revision point as R1-P1: immediately after the sentence beginning “This discrepancy may stem from differences in the type and duration of stress applied”.

      “Because only one moderate salinity level (10 ppt) was tested, our results do not exclude plastic NuRE responses at higher salinity or along a dose-response gradient. Future experiments should determine whether a threshold exists beyond which the apparent stability of NuRE is overcome.”

      Results 3.2, second paragraph, inserted immediately after the sentence reporting the ecotype effects and ending “(Figure 4; Figure S5)”.

      “It should be noted that the ecotype analysis here was based only on Chinese populations, as ecotype classification was not available for non-Chinese populations.”

      (R2-P3) The variation partitioning shows latitude as a significant predictor, but the R<sup>2</sup> values are relatively low, indicating that much variance remains unexplained. The manuscript should avoid overinterpreting latitude’s explanatory power and more openly acknowledge the role of unmeasured factors. The interpretation of slopes greater than 1 for the resorbed N:P versus green N:P relationship, labeled as “inverted limitation”, also needs further explanation regarding its functional significance.

      We agree that the original wording overemphasized latitude. Results 3.4 and Discussion 4.3 now describe latitude as the largest contributor among the measured predictors while emphasizing its modest individual R<sup>2</sup> and the substantial unexplained variation. We also expanded Discussion 4.2 to explain that slopes greater than 1 indicate disproportionate recovery of P or K relative to N and to present nutrient balance, P conservation, K mobility, and constraints on N remobilization as non-exclusive hypotheses rather than established mechanisms.

      Results 3.4, first paragraph, replacing the sentence beginning “Furthermore, variation partitioning analysis indicated that latitude”.

      “Variation partitioning indicated that latitude had the largest individual contribution among the measured predictors, but its contribution was modest for N, P, and K resorption (individual R<sup>2</sup> = 0.092, 0.057, and 0.094, respectively; Table 1). The very small contributions of green-leaf P concentration (R<sup>2</sup> = 0.006) and stoichiometry (R<sup>2</sup> < 0.001) further indicate that most variation in P resorption was associated with factors not represented in the present models.”

      Discussion 4.3, first paragraph.

      “Although latitude explained more variation than the other measured predictors, its individual contribution remained modest. The substantial unexplained variance indicates that additional climatic, edaphic, demographic, or genetic factors also contribute to NuRE variation. We therefore interpret latitude as a significant but limited correlate of NuRE rather than as a dominant determinant.”

      Discussion 4.2, first paragraph.

      “Functionally, slopes greater than 1 indicate that changes in green-leaf N:P or N: K are accompanied by disproportionate changes in the corresponding resorbed ratio, consistent with relatively greater recovery of P or K than of N across the observed nutrient gradient. This pattern may contribute to maintaining internal N:P:K balance during regrowth. It may also reflect stronger conservation of P, the high mobility of K, or constraints on the remobilization of N retained in structural or metabolic compounds. Because the experiment did not include nutrient additions or direct measurements of remobilization costs, these explanations remain hypotheses.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) Clarify the concept of “canalization” early. Consider adding one sentence in the Introduction or Discussion explicitly defining canalization in this context (i.e., NuRE varies among populations but this variation is genetically determined and shows little plasticity to salinity). This will help readers less familiar with evolutionary biology terminology.

      We thank the reviewer for this helpful suggestion. We agree that canalization should be defined when it is first introduced. We have therefore added a sentence to the Introduction, immediately after introducing the evolutionary canalization hypothesis, to clarify how this concept is used in our study.

      Introduction, fourth paragraph.

      “Here, canalization denotes genetically based differences in NuRE among populations, coupled with limited phenotypic plasticity of this trait under the short-term salinity treatment.”

      (R1-R2) Discuss potassium more deeply. In the Discussion (section 4.2 or 4.4), speculate on why K resorption lacks concentration control. Does Na<sup>+</sup> accumulation functionally substitute for K in osmotic adjustment, thereby decoupling resorption from green leaf K concentration?

      We appreciate this suggestion and have expanded Discussion 4.2. We now propose partial functional substitution of K<sup>+</sup> by accumulated Na<sup>+</sup> during osmotic adjustment as one possible explanation for the weak coupling between green-leaf K concentration and K resorption. We explicitly frame this explanation as tentative because it was not directly tested in the present experiment.

      Discussion 4.2, second paragraph.

      “One possible explanation is the partial functional substitution of K<sup>+</sup> by Na<sup>+</sup> during osmotic adjustment. Na<sup>+</sup> can replace part of the nonspecific vacuolar osmotic function of K<sup>+</sup> in plants and has been reported to become a major osmoticum in P. australis from higher-salinity habitats (Wakeel et al., 2011; Zhao et al., 1999). Increased Na<sup>+</sup> accumulation may therefore reduce reliance on K<sup>+</sup> for osmotic adjustment, potentially contributing to the weak relationship between green-leaf K concentration and K resorption efficiency observed here.”

      Wakeel, A., Farooq, M., Qadir, M., & Schubert, S. (2011). Potassium substitution by sodium in plants. Critical Reviews in Plant Sciences, 30(4), 401–413. https://doi.org/10.1080/07352689.2011.587728

      Zhao, K. F., Feng, L. T., & Zhang, S. Q. (1999). Study on the salinity-adaptation physiology in different ecotypes of Phragmites australis in the Yellow River Delta of China: Osmotica and their contribution to the osmotic adjustment. Estuarine, Coastal and Shelf Science, 49(Supplement 1), 37–42. https://doi.org/10.1016/S0272-7714(99)80006-7

      (R1-R3) Address the potential confounding of ecotype and phylogeography. The dual classification (Figure 3 vs 4) is elegant. However, Chinese ecotypes (freshwater, coastal, inland saltmarsh) may be partially confounded with phylogeographic lineages. A brief sentence explaining how the analytical approach (separate models, random effects) helps separate deep evolutionary history from recent local adaptation would strengthen the interpretation.

      We agree. This concern is addressed in detail under R1-P2. We clarified why phylogeographic group and ecotype were fitted in separate models, what the genotype random effect accounts for, and why the ecotype results cannot be interpreted independently of phylogeographic history.

      See R1-P2, Locations C1 and C2.

      (R1-R4) In section 2.1, specify whether the soil mixture ratio (2 soil: 1 peat moss: 1 river sand) is by volume or by mass.

      We thank the reviewer for identifying this ambiguity. The ratio was based on volume, and we have revised Methods 2.1 accordingly.

      “The plants were planted in barrels (total volume 25 L; top diameter 32.5 cm, bottom diameter 28.4 cm, height 38.5 cm) containing 20 L of a substrate composed of soil, peat moss, and river sand in a 2:1:1 volume ratio (Figure S1).”

      (R1-R5) In the Materials and Methods section, the element potassium (K) is described twice. Specifically, in line 176, the sentence “Besides C, N and P, other eight elements (K, Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” should be revised to “Besides C, N, P and K, other seven elements (Cu, Zn, Fe, Mn, Mg, Si, Na) were quantified in leaf tissues” to avoid redundancy.

      We thank the reviewer for identifying this redundancy. We have corrected the sentence in Methods 2.2.

      “Besides C, N, P and K, seven additional elements (Cu, Zn, Fe, Mn, Mg, Si, and Na) were quantified in leaf tissues.”

      (R1-R6) Citations should be formatted and ordered alphabetically or by year of publication.

      We thank the reviewer for pointing out this issue. We standardized multi-reference citations, reordered the reference list alphabetically by first-author surname, and verified correspondence between in-text citations and reference entries.

      Full-manuscript citation and reference-list audit.

      Reviewer #2 (Recommendations for the authors):

      (R2-R1) In the discussion of the “inverted” nutrient limitation, elaborate on why P and K are recovered more relative to N. Consider whether this reflects a strategy to maintain an optimal N: P: K ratio or arises from higher costs or lower availability of N.

      We agree. This issue is addressed in detail under R2-P3, where we explain the meaning of slopes greater than 1 and present the possible functional mechanisms as hypotheses rather than established explanations.

      See R2-P3, Location C.

      (R2-R2) Expand Table 1 to include confidence intervals or p-values for individual effects. The extremely low R<sup>2</sup> for concentration and stoichiometry on P resorption should be more explicitly noted as evidence for the dominance of latitude.

      We agree and have expanded Table 1 by adding significance symbols (*) to the individual R<sup>2</sup> values. Significance was assessed for the corresponding fixed effects in the full linear mixed-effects models using Type III tests with Satterthwaite’s approximation for degrees of freedom. For P resorption, latitude had the largest individual contribution and was significant (R<sup>2</sup> = 0.057, p = 0.005), whereas the contributions of green-leaf P concentration and stoichiometry were very small and nonsignificant (R<sup>2</sup> = 0.006 and < 0.001, respectively). We revised the Results to emphasize this contrast while acknowledging that latitude explained only a modest proportion of the total variation.

      See R2-P3, Locations A and B, for the corresponding Results and Discussion revisions.

      (R2-R3) Standardize the abbreviation to “NuRE” throughout, correcting the use of “NRE” in the introduction.

      We thank the reviewer for identifying this inconsistency. We standardized the abbreviation to NuRE throughout the manuscript.

      (R2-R4) Acknowledge in the discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response, justifying future dose-response experiments.

      We agree. This limitation is addressed under R2-P2, where we state that a single 10-ppt treatment cannot exclude plastic responses at higher salinity or along a broader dose-response gradient.

      See R2-P2, Location A.

      (R2-R5) In the Results section, explicitly state that the ecotype analysis in Figure 4 is based only on Chinese populations, not all 110 genotypes, because ecotype classification was not available for non-Chinese populations.

      We agree. This clarification is provided under R2-P2, where Results 3.2 is revised to state that the Figure 4 ecotype analysis includes only Chinese populations.

      See R2-P2, Location B.

      (R2-R6) Consider adding a supplementary figure comparing raw and carbon-standardized NuRE values to show whether the correction altered main conclusions.

      We agree and have added Supplementary Figure S7 comparing raw and carbon-standardized NuRE. The two estimates were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999), and analyses using raw NuRE retained the same conclusions for phylogeographic group, ecotype, salinity, their interactions, and latitude. Thus, carbon standardization slightly shifted the absolute values without altering the main conclusions.

      Methods 2.3.

      “Raw NuRE was calculated without carbon standardization and compared with carbon-standardized NuRE using Pearson correlations; the main linear mixed-effects analyses were also repeated using raw NuRE.”

      Results 3.4, inserted after the existing paragraph reporting the latitude effects and referring to Figure 6 and Table 1.

      “Raw and carbon-standardized NuRE were strongly correlated for N, P, and K (Pearson’s r = 0.997–0.999; Figure S7), and analyses using raw NuRE did not alter the conclusions for phylogeographic group, ecotype, salinity, their interactions, or latitude.”

      Supplementary Figure S7 legend

      “Figure S7 Comparison of raw and carbon-standardized nutrient resorption efficiency (NuRE) for N, P, and K. Points represent genotype-by-treatment observations (n = 206), coloured by treatment. Grey dashed lines indicate the 1:1 relationship, and black lines show ordinary least-squares fits. Pearson’s r and the mean standardized-minus-raw difference (Δmean, percentage points) are shown.”

      (R2-R7) In Figure 2b, check the y-axis label; the text reports mg/kg, but the axis shows g/kg, which needs correction.

      We thank the reviewer for identifying this unit discrepancy. We corrected the y-axis label in Figure 2b from g/kg to mg/kg so that it matches the units reported in the text.

      (R2-R8) In the Discussion, rephrase the sentence “This indicates that NuRE is a conservative trait…” to avoid repetition with the Results summary, for example, “This finding highlights the conservative nature of NuRE.”

      We thank the reviewer for this helpful wording suggestion. We have rephrased the sentence to avoid repetition.

    1. Author response:

      On the eLife Assessment. We appreciate the assessment’s recognition that this study addresses a long-standing question concerning hippocampal and prefrontal contributions to working memory. We think the broader significance lies in constraining how causal manipulations in working-memory tasks are interpreted. Working memory encompasses many processes distributed across a trial and showing that a region is required for a delayed-response task does not establish when its contribution is necessary. The observation that hippocampus and mPFC are required while navigating to a goal, but not during stationary delay, challenges a common assumption and, in our opinion, has implications across the broader field of working-memory research.

      We agree that the original manuscript did not adequately report effect sizes or convey the uncertainty associated with some smaller samples. However, the dataset does include duration-matched stationary and running conditions, including a hippocampal long-delay control matched to early-running stimulation in both duration and elapsed time after cue onset. The new effect-size and within-animal analyses support a robust hippocampal epoch difference, while the corresponding mPFC comparison is less precisely estimated and should be interpreted more cautiously.

      We therefore think the central finding remains well supported: hippocampal function, and potentially mPFC function, is required during the active navigation period of this memory-guided task but not detectably during the stationary delay. This does not establish the specific computation disrupted during running. Neural recordings would certainly provide further insight into the underlying mechanism, but their absence does not detract from the value of the behavioral result itself.

      Reviewer #1 (Public review):

      The simplicity of the behavior makes it difficult to resolve how exactly the hippocampus and mPFC contribute to working memory.

      The task was designed to combine the temporal precision of cue-based delayed-response paradigms with the behavioral richness of freely moving navigation. Few tasks combine a fixed cue-presentation period, an explicit delay, and a subsequent navigation phase involving extended running. This structure creates well-defined behavioral epochs that can be targeted with temporally precise perturbations, allowing us to ask when hippocampal and mPFC contributions are required within an ongoing memory-guided behavior. How these regions contribute is the harder question, and one we are pursuing next. Identifying when perturbations disrupt behavior is an important step toward understanding how these regions support memory-guided navigation

      Also, the language does not always reflect the trends in the data: the authors claim that optogenetic perturbations cause mice to repeat previous choices, but the data show that perturbations increase the likelihood of choosing a preferred side (which is left for most mice). A side bias is not the same as choice repetition. This has implications for interpreting the nature of the behavioral effects.

      We agree that our results do not clearly distinguish a directional bias from a tendency to repeat the previous choice. To examine whether mice consistently favored a particular direction, we compared their side preferences during silencing across sessions. Mice generally favored the same side across silencing conditions, although some switched direction in individual sessions (Author response image 1a). Within sessions, the preferred side was maintained from no-stimulation to stimulation trials in 13 of 19 cases and reversed in six (Author response image 1b). These observations are consistent with a directional preference that can sometimes reverse during stimulation, and cannot be disambiguated from perseveration. In the revision, we will describe the effect as increased directional bias and revise the language concerning choice repetition and perseveration throughout the manuscript.

      Author response image 1.

      Silencing increases directional bias. (a) Fraction of choices made to the right in each session, for every mouse (rows) and each silencing condition (symbols). Open symbols, no-stimulation trials; filled symbols, stimulation trials from the same session; blue and red denote a left or right preference during silencing. Filled symbols falling predominantly on the same side of 0.5 within a row indicate that a mouse generally favored the same direction across silencing conditions, although some mice switched direction. One session per mouse and condition, hippocampal silencing only; T2 and T6 did not perform the 0.8 s condition. (b) The same sessions expressed as signed bias, from no stimulation to silencing. Black lines mark the six sessions in which the preferred side reversed; grey lines the thirteen in which it was maintained.

      Reviewer #2 (Public review):

      A major weakness of the study is the lack of a balanced design with equal time periods of silencing during the stem running period and temporal delay period in many of the animals, which precludes any conclusion about distinct functional roles of the regions during these two phases of the task. The main conclusion of distinction between temporal delay and stem running delay periods is therefore not adequately tested for the prefrontal cortex, and the statistics in terms of number of animals for this important control for hippocampal inactivation are also not comparable to the main experiment.

      We agree that matching stimulation duration is an essential control and recognize that the relevant comparisons and statistics were not sufficiently clear in the original manuscript. Four hippocampal conditions used closely matched stimulation durations of 2 s: Cue+Delay, long delay, early running, and running with a 0.8 s onset (Author response image 2a). Crucially, the long-delay control matched both stimulation duration and elapsed time after cue onset to the early-run condition, while the mouse remained stationary.

      Author response image 2b–c shows the estimated impairment and its 95% confidence interval for each condition. In the hippocampal experiments, the duration-matched stationary conditions showed effects close to zero, whereas the running conditions showed large impairments. Despite the smaller sample, the upper confidence limit for the long-delay impairment was approximately 10 percentage points, substantially below the observed early-running impairment. These estimates establish that despite the smaller sample, the data support a lack of effect compared to early running.

      Author response image 2.

      Duration-matched stimulation produces different behavioral effects across task epochs. (a) Stimulation timing relative to cue onset. Numbers within bars indicate calculated median stimulation duration in seconds; black ticks indicate door opening. Bold labels identify conditions with 2 s stimulation. (b-c) Mean impairment in choice accuracy for hippocampal and mPFC manipulations. Impairment is accuracy during baseline minus accuracy with stimulation, in percentage points. Error bars show 95% confidence intervals across animals; numbers indicate mice. Open circles denote single-animal observations, and arrows indicate confidence intervals extending beyond the plotted range.

      For mPFC, the duration-matched Cue+Delay condition likewise showed an effect close to zero, whereas early-running stimulation produced substantial impairment. However, the long-delay condition included only one mouse. The later-running effects in both regions were also less precisely estimated. We will distinguish these limitations from the more informative stationary-condition results.

      To directly test whether the duration-matched effects differed across epochs, we next compared impairment within the same mice, including only animals tested in both conditions (Author response image 3). For the hippocampus, every mouse showed greater impairment during early running than during either duration-matched stationary condition, and the confidence intervals for both paired differences excluded zero. These within-animal comparisons support an epoch-dependent effect that cannot be explained by stimulation duration alone. The mPFC comparison showed the same direction of effect, although the confidence interval for the Cue+Delay versus running difference narrowly included zero. We will therefore distinguish the stronger evidence for the hippocampal epoch difference from the more limited evidence for mPFC.

      Author response image 3.

      Within-animal comparisons of duration-matched stimulation effects. (a–b) Impairment during Cue+Delay or long-delay stimulation compared with early-running stimulation for hippocampal (a) and mPFC (b) manipulations. Points represent individual mice, lines connect observations from the same mouse, and black bars indicate mean. Annotations report the mean paired difference in impairment (running minus stationary), and its 95% confidence interval calculated using the t distribution. Positive differences indicate greater impairment during running. The single-mouse mPFC long-delay comparison is descriptive. Blue indicates stationary epochs and orange indicates running.

      The central question of this study is not new, with many previous studies investigating distinct and overlapping roles of hippocampus and prefrontal cortex in spatial working memory tasks and memory-guided navigation, using inactivation of one or both regions, crossed inactivation approaches, as well as targeting direct and indirect connections between the regions (PMIDs: 20074655, 9030646, 17045348, 30179661, 27511010, 10491611, 26017312, etc.), in addition to several physiology studies.

      We agree that hippocampal and prefrontal contributions to spatial working memory have been extensively studied. That’s precisely why we find these results impactful when placed in the rich context of the field. The requirement for these regions in delayed working memory tasks has often been interpreted in terms of their contributions during the delay period. However, a requirement during a delay-based task does not itself demonstrate a requirement during the delay. Previous manipulations have not isolated delay periods of waiting from the subsequent navigation within a trial.

      Our experiments extend this work by separately targeting cue presentation, the delay, and different portions of navigation within the same task. This allows us to test whether the behavioral consequences of perturbation depend on the particular epoch in which it occurs. This distinction is important: knowing that a region is required for a memory-guided task does not establish when its contribution is needed. Identifying those periods constrains how we interpret the deficits produced by longer-lasting inactivation. We believe that this is an important result that should be considered when interpreting these broader findings. We will ensure that the appropriate literature and discussion are included in the revision.

      It is not clarified why such short delay periods were used compared to long periods of ~10s in T-maze spatial alternation tasks with delay, and whether the temporal delay period of 1 sec is strictly distinct from the spatiotemporal delay period during stem running in terms of short-term memory function.

      The 1 s delay was chosen to maintain reliable task performance, as some mice could not perform the task with longer delays. Indeed, one reason the longer-delay condition includes fewer animals is that two mice could not perform reliably with the 3 s delay (one additional mouse was not tested in this condition). Delays on this timescale have also been used in rodent cued delayed-response tasks, including a 0.5 s delay in Kopec et al. (2015), a 1.3 s delay in Guo et al. (2014), and a 1.2 s delay in Inagaki et al. (2019). Like these tasks, our paradigm requires mice to remember an externally presented cue specifying the upcoming response, rather than their own previous arm choice as in spatial alternation. It therefore combines spatial navigation with a cued delayed-response requirement, and the delay durations tolerated in alternation tasks are not necessarily directly comparable. We will clarify this rationale in the manuscript.

      We agree that the stationary delay and the subsequent run both require retention of information after cue offset. Our experiments distinguish these behavioral epochs, but do not establish that they involve separate short-term memory processes. The different effects of perturbation suggest that the contribution of these regions changes as the animal moves from waiting to navigating. Determining what accounts for this change is an important direction for future work.

      The choice of time windows for inactivation needs to be better justified, which currently appears to be rather random (2s initial running period, run after 0.8s, run after 1.6s; for an average reported running period of ~2.4-2.5s).

      We sought to target different portions of the central-arm run. Given the typical traversal time of approximately 2.4–2.5 s, stimulation onsets at 0, 0.8, and 1.6 s sampled the beginning, middle, and later portions of the run. Our setup allowed precise control of stimulation timing, and, as shown in Figure 4b, these onset times correspond approximately to the start, middle, and end of the central arm. Each condition used a nominal 2 s stimulation window, truncated if the mouse reached the choice point sooner. The windows therefore overlap, with the later-onset condition generally producing shorter stimulation. We will clarify this rationale and the distinction between stimulation onset and duration in the revised manuscript.

      The mixture of PV-Cre animals (5 animals), Dlx targeting (2 animals), and one WT animal is also suggestive of a fragmented approach, and inactivation efficacy cannot be assumed to be similar for different animals. Importantly, there is no physiological evidence for confirmation of suppression in the optogenetic experiments, even in exemplar animals.

      The Dlx animals were included to improve regional specificity through local viral expression and to test whether the behavioral effects were consistent across targeting approaches. Both approaches produced comparable impairments during running, supporting their inclusion in the same analysis. We will clarify this rationale in the manuscript.

      Optogenetic activation of inhibitory interneurons is an established approach for suppressing local principal-cell activity, with physiological validation in previous studies (Guo et al., 2014; Li et al., 2019, Zutshi et al., 2022). In our experiments, running-period stimulation produced robust behavioral impairments that were consistent across animals and targeting approaches. Furthermore, comparisons across epochs were performed within animals, using the same preparation and stimulation parameters. Differences in efficacy between animals therefore cannot readily account for the observed epoch dependence. Although direct recordings would establish the magnitude and spatial extent of suppression in our preparation, the central behavioral finding is supported by these within-animal comparisons.

      Reviewer #3 (Public review):

      The findings are potentially important because they challenge the common assumption that hippocampal and prefrontal contributions to delayed-response tasks are centered on delay-period maintenance.

      We thank the reviewer for describing our findings as “potentially important” and for highlighting their implications for how hippocampal and prefrontal contributions to delayed-response tasks are understood. We appreciate the constructive suggestions and address the public comments below.

      Most notably, the study lacks a non-memory control task, such as a visually guided version of the maze, making it difficult to determine whether the observed deficits specifically reflect disruption of memory-guided behavior or more general impairments in action selection, behavioral flexibility, or movement planning. The observed increase in perseverative responding and delayed commitment to a turn are consistent with either interpretation.

      We agree that leaving the cue on throughout the trial would provide an important control for distinguishing memory-specific effects from broader effects on action selection or movement planning. The senior author is currently setting up a new laboratory, so implementing this control may take some time. We hope to include it in the revised manuscript.

      The running-period deficit indicates a disruption of processes that enable the animal to act on a remembered cue. Whether this reflects disruption of memory itself, movement planning, or another component of translating the cue into a choice requires further clarification but does not detract from the observed dependence on task epoch. Uncertainty about the mechanism of the running-period deficit also does not change the observation that the same manipulation produced no detectable impairment during the stationary delay, when the cue was absent and still had to be remembered. This will be clearly discussed in the revision.

      In addition, several experimental conditions rely on relatively small numbers of animals, limiting confidence in some negative findings, particularly for the longer-delay and later-run manipulations.

      We agree that small sample sizes limit the interpretation of some negative findings. We now report animal-level effect estimates and 95% confidence intervals for each condition (Author response image 2b–c; Author response table 1).

      Of the nine conditions with no detectable impairment and more than one mouse, seven had confidence intervals that excluded effects as large as the observed mean early-running impairment in the same region. This included the hippocampal long-delay condition, despite its smaller sample. These results argue against similarly large impairments in these conditions, although smaller effects remain possible. The two later-run conditions remained too uncertain to exclude such impairments. These and the two single-animal conditions are marked in Author response table 1 and will be interpreted cautiously.

      Author response table 1.

      Animal-level impairment estimates and uncertainty

      Finally, while the Discussion proposes that hippocampal-prefrontal circuits become engaged during the transformation of stored information into action, this mechanistic interpretation remains speculative because no neural recordings accompany the causal manipulations.

      We will clarify in the Discussion that the proposed transformation of stored information into action remains speculative. Nevertheless, we believe this is an exciting possibility raised by our findings that warrants further investigation. Neural recordings would help test this interpretation but are beyond the scope of the current paper. We plan to explore this question in future work.

      References

      Guo ZV, Li N, Huber D, Ophir E, Gutnisky D, Ting JT, Feng G, Svoboda K (2014). Flow of cortical activity underlying a tactile decision in mice. Neuron 81:179–94. doi:10.1016/j.neuron.2013.10.020. PMID 24361077.

      Kopec CD, Erlich JC, Brunton BW, Deisseroth K, Brody CD (2015). Cortical and subcortical contributions to short-term memory for orienting movements. Neuron 88:367–77. doi:10.1016/j.neuron.2015.08.033. PMID 26439529.

      Inagaki HK, Fontolan L, Romani S, Svoboda K (2019). Discrete attractor dynamics underlies persistent activity in the frontal cortex. Nature 566:212–217. doi:10.1038/s41586-019-0919-7. PMID 30728503.

      Li N, Chen S, Guo ZV, Chen H, Huo Y, Inagaki HK, Chen G, Davis C, Hansel D, Guo C, Svoboda K (2019). Spatiotemporal constraints on optogenetic inactivation in cortical circuits. eLife 8:e48622. doi:10.7554/eLife.48622. PMID 31736463.

      Zutshi I, Valero M, Fernández-Ruiz A, Buzsáki G (2022). Extrinsic control and intrinsic computation in the hippocampal CA1 circuit. Neuron 110:658–673.e5. doi:10.1016/j.neuron.2021.11.015. PMID 34890566.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      Using sci-L3-Strand-seq, this study shows genome-wide, single-cell evidence that DSBs triggered by CRISPR/Cas9 are frequently resolved through sister chromatid exchange (SCE), which is undetectable with conventional whole-genome sequencing approaches. This important work provides helpful metrics quantifying the occurrence of SCE events at targeted and repetitive genomic loci that are of particular interest to those working in gene editing, DNA repair, genome instability, or repetitive genomic loci. However, the evidence supporting the proposed involvement of under-replicated region/replication termination-zone resolution and TRAIP/URR-like pathways is currently incomplete and could be strengthened with an increased number of reciprocal daughter-cell pairs and by genetic or molecular perturbation, or alternatively, this can be addressed by changing the discussion.

      We appreciate the editor’s and the reviewers’ recognition of Cas9-induced SCE as an important previously invisible repair outcome and of RDCP analysis as a notable feature of the study. As proposed, we incorporated the two recent Science studies and now tone down the Discussion of our RDCP observations as consistent with, rather than definitive evidence for, the TRAIP-dependent pathway. We also detail how our single-cell genomic observations complement and extend these two studies in terms of the biological significance of the CDK1-TTF2-TRAIP axis.

      Specifically, these changes are in:

      Discussion. We changed

      “Two recent studies revealed how the CDK1-TTF2-TRAIP axis is cell-cycle regulated to trigger mitotic CMG helicase disassembly and fork cleavage: one study showed a two-fold SCE reduction in mouse ES cells (Fujisawa and Labib, 2026), while the other showed that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions (Can et al., 2026). Our observation of the "WWC-or-WCC/deletion pair" signature in wild-type cells provides, to our knowledge, the first genetic evidence of linking a deletion with SCE and revealing both W and C unreplicated template strands present in the reciprocal daughter cell, consistent with this mechanism at single-cell genomic resolution (illustrated in Fig.3), although this is limited by the observation of only one RDCP.”

      While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells.

      Fig.3. Title and legends. We changed

      “Haplotype-aware analysis of observed RDCP (Pair 4, chr1) shows SV patterns at the SCE junctions consistent with the predicted RDCP signature of SCE mediated by URRs or replication termination zones (green shaded area), although the lagging strands, rather than the leading strands, must be resolved to generate these mitotic breaks.”

      Reviewer #1 (Public review):

      Summary:

      This manuscript uses sci-L3-Strand-seq to map sister chromatid exchange events following CRISPR/Cas9-induced DNA damage. Because exchanges between identical sister chromatids are largely invisible to conventional sequencing, the study addresses an important blind spot in the assessment of genome editing outcomes. The authors compare single-locus Cas9 cleavage, simultaneous targeting of 237 repetitive genomic sites, and Cas9 nickase variants. They further use reciprocal daughter-cell pair analysis to ask whether Cas9-associated SCEs are copy-neutral or linked to larger structural alterations. Overall, this is a valuable study that introduces an important additional layer to the analysis of CRISPR/Cas9 repair outcomes. The central finding that Cas9-induced DSBs can trigger frequent local SCE is well supported and likely to be of broad interest. The evidence for structural complexity associated with some induced SCEs is intriguing, but the mechanistic interpretation should either be tested directly or presented more cautiously.

      Strengths:

      The major strength of the manuscript is the application of a strand-resolved, single-cell method to a question that is difficult to address with standard genome sequencing. The evidence that a single Cas9induced DSB can trigger strong local SCE is compelling in concept and supported by multiple guide RNAs targeting distinct loci. The reported on-target SCE frequencies, reaching up to 41%, suggest that inter-sister exchange is a substantial and underappreciated outcome of Cas9 cleavage.

      Of particular interest is the comparison between single-site and multi-site targeting. The finding that 237 programmed Cas9 targets produce only mild bulk enrichment of on-target SCE but stronger enrichment in a subpopulation of cells with elevated SCE burden is interesting and may have wider biological implications, particularly if the findings extend beyond Cas9-induced SCE to spontaneous SCEs. Given that potential, the current manuscript would benefit greatly from any experiments characterizing this sub-population: are these cells in a particular cell cycle state, experiencing changes in gene expression, or do they have other unique biological properties?

      The reciprocal daughter-cell pair analysis is another notable feature of the study. The observation that some Cas9-associated SCEs are accompanied by structural alterations could challenge the assumption that SCE after a programmed break reflects error-free homologous recombination.

      Thank you very much for this assessment.

      Weaknesses:

      The number of informative RDCPs is limited, and the mechanistic interpretation of the "WWC-orWCC/deletion" signature is more suggestive than definitive. In particular, the manuscript invokes (even though only in the Discussion section) URR or replication-termination-zone resolution and discusses TRAIP-dependent CMG unloading, nuclease cleavage, and polymerase theta-mediated joining, but these pathway components are not directly tested herein. A more conservative conclusion that some Cas9-associated SCEs coincide with structural alterations is more appropriate, particularly in the Discussion and Conclusion. For example, the statement that this work provides "direct genetic evidence" for a URR-type mechanism is overstated unless supported by additional experiments or a more extensive analysis of alternative models. Similarly, while the authors explain the limitations of acute Cas9 disruption of LIG3, LIG4, XRCC1, and XRCC4, the manuscript should clarify what biological questions this experiment can and cannot answer.

      Please see response to the eLife Assessment as this is a common point raised by multiple reviewers.

      Additionally, we clarified what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) by adding the following text in the “Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell” section:

      “One additional complication is that, although Cas9 RNP achieves >90% knockout efficiency in bulk assays and is therefore used as a substitute for siRNA, sorting BrdU-labeled cells in the subsequent G1 enriches for cells that escaped frameshift editing, particularly for essential genes; thus, 100% knockout in 90% of cells is not equivalent to 90% knockdown in every cell, representing a unique challenge for single-cell assays.”

      Reviewer #1 (Recommendations for the authors):

      (1) Temper the mechanistic claims about URR/TRAIP-type resolution.

      The RDCP data support the conclusion that some Cas9-associated SCEs are accompanied by structural alterations and may arise through non-classical mechanisms. However, claims about TRAIPdependent CMG unloading, URR resolution, or polymerase theta-mediated joining should be framed as a model unless directly tested.

      We cited mechanistic dissection of the CDK1-TTF-2-TRAIP axis, which was published since the review of the paper. While each study highlights the biological significance of the CDK1-TTF2-TRAIP axis individually for SCE and deletions, we show the coupling relationship (and further evidence of unreplicated template strands) in reciprocal daughter cells but qualified that this observation is in only one RDCP. We further tempered our claims by changing “direct evidence” to “consistent with” as we did not perturb genes involved in these processes.

      (2) Clarify the impact of the small RDCP sample size.

      The manuscript would be strengthened by explicitly stating how many total RDCPs were analyzed, how many SCE events were informative, and how much confidence can be placed in the estimated fraction of SCEs associated with SVs. A short table summarizing RDCP counts, SCE counts, copy-neutral events, and SV-associated events would be helpful.

      A total of 15 RDCPs were recovered from close to 4,000 single cells analyzed across all conditions. We added Tab.S3 detailing SCEs in all 15 RDCPs in addition to SCE and SV breakdown in Tab.S2 (originally Tab.S1). The new Tab.S3 is cited in the “RDCP analysis reveals large-scale SVs on chromosomes with induced SCE, as well as structural alterations at Cas9-induced SCE junctions” section.

      (3) Provide more detail on the "rescued" SCE calls.

      Because the central conclusions rely on SCE detection, the criteria for breakpoint R-based calls versus rescued calls should be explained clearly in the main text or methods. It would be useful to know how sensitive the main conclusions are to the inclusion or exclusion of rescued calls. Is this laid out in greater detail in an additional manuscript?

      We previously included an “On-target SCE identification” section in the Methods, where we described in detail the rescue of on-target SCEs missed by the initial breakpoint R calls. We also depicted calls and calls+rescues for all the conditions in Fig.S1B.

      We agree with the reviewer and now expanded the description of rescued SCE calls in the main text (under the “A single Cas9 DSB induces potent local SCE” section). In brief, 50-92% of SCEs (typically >70%) across the four single-targeting sites were directly called rather than rescued, with the exception of LIG4, where only 30% were direct calls. This is because LIG4 is located only 6 Mb from the telomere and is therefore particularly prone to missed breakpointR calls in low-coverage cells. On-target rescue at individual sites is self-contained in this manuscript because our lab primarily focuses on spontaneous SCE, for which there are no expected SCE sites. However, the rescue methodology was previously implemented in the original development of sci-L3-Strand-seq to identify SCEs at centromeres.

      (4) Clarify the biological interpretation of the repetitive-target enrichment.

      The high-SCE subset analysis is interesting, but the manuscript should explain whether these cells have evidence of higher RNP uptake, altered cell-cycle state, greater DNA damage, or lower sequencing quality. If these possibilities cannot be distinguished, the text should state this clearly.

      We agree with the reviewer. The high SCE subset does not have lower sequencing quality by coverage or background (0.3% coverage for high-SCE vs. 0.28% coverage overall, and 3% background for both high-SCE and overall). However, our current data do not allow us to distinguish among biological explanations for the elevated SCE. The original manuscript acknowledged this limitation (“Whether this reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment, remains to be determined.”). To make this limitation more explicit and to address the possibility of sequencing quality raised by the reviewer, we have revised the text as follows: “The high-SCE subset did not show evidence of lower sequencing quality, based on either sequencing coverage (p=0.17) or background SCE levels (p=0.13). However, we cannot distinguish whether the elevated on-target SCE reflects a cell-cycle state more permissive to both cutting and recombination, or stochastic variation in RNP uptake coupled with a recombination-prone chromatin environment.”

      This question may be better explored by future co-assays with sci-L3-Strand-seq; currently we cannot enrich for cells with high SCEs to characterize the molecular features of this subset of the cells using other omics approaches.

      (5) Reconsider the framing of the DNA repair gene targeting experiment.

      The current data do not strongly test whether LIG3, LIG4, XRCC1, or XRCC4 regulate Cas9-induced SCE, because functional protein loss is delayed and essential-gene targeting introduces selection. This section may be better framed as a negative/control observation rather than as a pathway analysis.

      We agree and please refer to Public Reviews for a single-cell assay-specific explanation.

      (6) Consider including some additional control experiments, for example, Cas9 without sgRNA, nontargeting sgRNA, or mock-transfected cells to make sure that some phenotypes (for example, cell-cycle arrest) directly result from DNA cleavage rather than from the transfection procedure.

      We thank the reviewer for this suggestion. We have carefully considered these additional controls but have chosen not to add further experiments. Our existing Cas9 nickase experiments provide a control that directly addresses whether the observed arrest is attributable to DSB formation rather than RNP delivery/transfection. Both the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9, but neither produced the cell-cycle arrest observed following DSB induction by wild-type Cas9. Thus, these experiments control for Cas9 RNP delivery while altering the nature of the DNA lesion and support the interpretation that the observed arrest is associated specifically with Cas9-induced DSBs rather than the transfection procedure itself.

      (7) Figure 1: the fonts should be increased. The majority of the labels are impossible to read in a printed copy of this manuscript.

      We thank the review for pointing this out. We enlarged Fig.1 fonts.

      (8) Figure 1C. The pileup plots should be described and interpreted in a clear way. In its present form, it is unclear how the interpretations and conclusions are made.

      We added explanation of the pileup analysis immediately following mentioning the Fig.1C pileup: “We next examined … SCEs using genome-wide pileup analysis (Fig.1C, Fig.S1B), in which we plot the total number of SCEs detected across all single cells within each 1 Mb window.”

      Reviewer #2 (Public review):

      Summary:

      In this short paper, a clever single-cell Strand-seq method was used to study the number and location of sister chromatid exchange events (SCEs) in cells after CRISPR/Cas9-induced DNA double-strand breaks (DSBs). Unique as well as multiple genomic loci were targeted. Cas9-induced cuts at unique genomic locations led to statistical enrichment of SCEs at the target site, whereas Cas9 targeted at repetitive targets revealed only mild enrichment of on-target SCEs unless analysis was restricted to a subset of cells with >8 SCEs per cell. Interestingly, reciprocal daughter-cell pair analysis revealed largescale structural alterations on some chromosomes. Whereas disruption of DNA repair genes, including LIG3, LIG4, XRCC1, and XRCC4, did not measurably alter SCE frequency per cell within 24 hrs, consistent with delayed functional loss following editing and selection against essential genes. Together, these findings demonstrate that Cas9-induced DSBs are potent local triggers of SCE at unique loci and can be associated with structural alterations, highlighting the influence of lesion type and genomic context on recombination outcomes during genome editing.

      Strengths:

      The data in this paper represent a very rich resource of how parental DNA template strands are distributed in paired daughter cells after various treatments. Abnormalities observed in only one of such paired daughter cells provide a novel and exciting approach to study mechanisms of DNA instability and DNA repair at a genome-wide level in general and following Cas9-induced DSB in particular.

      Thank you very much for this assessment.

      Weaknesses:

      The effect of Cas9-induced DSBs in the cells that are used will depend on the cell cycle stage of the cells that are targeted, as well as the number of times cuts are made. The latter could happen before, during, and after DNA repair reactions on one or both alleles in a diploid cell. As a result, it is very difficult to extrapolate the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements. Novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome some of these limitations. The language and logic in the paper can be improved, and some of the claims seem incorrect. For example, the abstract reads "A single Cas9 cut at a unique genomic locus led to strong local enrichment of SCE at the break site, reaching up to 41% in the same cell cycle and 17% in the subsequent division, indicating that DSB repair frequently engages non-local inter-sister repair." The evidence that only a single Cas9 cut was made is lacking (see my earlier comment); it is not clear how local enrichment or non-local inter-sister repair are defined.

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We also agree that novel approaches are needed to limit the number and timing of Cas9-induced breaks to overcome these limitations, perhaps by using vfCas9 but more importantly, if new approaches to turn off Cas9 are developed. We have added a brief discussion on this limitation (in the Limitation section) and the resulting constraints on extrapolating mechanisms of DNA instability and repair from the observed genomic rearrangements. We thank the reviewer for pointing out the distinction between “a single Cas9 cut” vs. "Cas9 targeting of a single genomic locus." We went through the manuscript and revised where cutting only once was implied. We also explicitly acknowledge the possibility of multiple rounds of cutting at the same sites.

      We thank the reviewer for pointing out that “non-local repair” is a non-standard term. We use it operationally to distinguish repair confined to the broken chromatid (e.g., fill-in synthesis or end joining in cis) from repair involving exchange between sister chromatids. We have added a schematic (Fig. S1A) illustrating this distinction and cited this figure immediately before where we operationally defined SCE as a “reciprocal strand switch between sister chromatids, without implying a single mechanistic pathway.” This distinction is important because, particularly for two-ended Cas9 DSBs, an SCE-like outcome could potentially arise through either HR-mediated crossover or NHEJ of DNA ends across sister chromatids; the latter may involve different genetic requirements from classical NHEJ at least in end-tethering. We have revised the manuscript to define “non-local repair” explicitly at its first use.

      Reviewer #2 (Recommendations for the authors):

      References to relevant earlier studies using Strand-seq to study SCEs are missing (PMID: 27185886 and PMID: 29348659).

      We thank the reviewer for pointing this out. We added these references in the 3rd paragraph of the Introduction where we briefly review Strand-seq methods.

      Reviewer #3 (Public review):

      Summary:

      Chovanec and Yin used their newly developed sci-L3-Strand-seq powerful method to characterize SCE after Cas9 cleavage in a human cell line, using either a single target site or an element repeated 237 times in the genome. SCE are often neglected in DNA repair analyses since they are « genetically silent ». Interestingly, the authors found enrichment of SCE at unique Cas9 sites, but only a modest enrichment of SCE when Cas9 targets 237 sites in the genome. The genetic control of SCE formation at Cas9 sites is not deliberately addressed in this paper. However, the authors found that targeted SCE seem to be enriched in a subpopulation of cells, particularly « permissive » for SCE, but the determinants of such a population are unknown. Finally, the power of the sci-L3-Strand-seq allowed the authors to characterize a specific type of SCE based on the analysis of reciprocal daughter-cell pairs' genomes that is associated with a specific type of chromosomal rearrangement compatible with the ones observed in HR defective BRCA1/2 deficient cells.

      Strengths:

      This is an interesting paper that molecularly explores sister chromatid exchanges, which represent an important challenge in molecular biology since they are genetically silent.

      Thank you very much for this assessment.

      Weaknesses:

      A complexity of the current paper is that it heavily relies on a recently published paper (Chovanec et al 2026, NAR) describing the powerful but complex technique sci-L3-Strand-seq. Knowledge of this paper is a prerequisite to understanding the current manuscript because no reminder is provided. In addition, the current manuscript presents the use of the sci-L3-Strand-seq technique in the study of SCE after Cas9-induced DSBs, while a companion study is referred to several times for containing results about SCE in XRCC1 KO. At some point, one questions the relevance of splitting the use of sci-L3-Strandseq in different papers instead of making a single integrated one.

      We appreciate this concern. The original sci-L3-Strand-seq study is an extensive methodology paper that establishes and validates various computational framework, whereas the companion study focuses on the genetic regulation of spontaneous SCE. The experimental designs and biological questions of the companion study and the present work are therefore distinct, although we draw on selected results from the companion study where they provide useful comparisons and contrasts between spontaneous and Cas9-induced SCE. The present study addresses a distinct biological question, the response to Cas9-induced DSBs, and we therefore believe that combining these studies would make the resulting manuscript unnecessarily broad and obscure their different biological questions.

      We nevertheless agree that the present manuscript should be understandable without requiring detailed knowledge of either paper. We have therefore added a brief description of the sci-L3-Strandseq approach (3rd paragraph of Introduction, Fig.S1A legends, and Fig.1B legends) and clarified the relevant methodological concepts where they are first introduced. We hope to improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract

      "Identical sisters ": redundant

      "non-local" inter-sister repair: the meaning is not clear. Do the authors refer only to "inter-sister" and therefore "non-local" is redundant, or do they imply something specific by "non-local", in which case it needs to be clarified?

      "237 repetitive targets": at least a slight description of this target is needed. Is it a "random" repeat, a satellite sequence, a sequence related to a transposable element ?...

      We agree with the reviewer that “identical sisters” is technically redundant. However, we have retained “identical” here to emphasize the distinction between sister chromatids vs. homolog, as SCE is sometimes misconstrued as exchange between homologs and as potentially causing loss of heterozygosity. We prefer the slight redundancy here for conceptual clarity.

      We thank the reviewer for pointing out that the meaning of “non-local” was unclear. As discussed in our response to Reviewer #2, we use “non-local repair” operationally to distinguish repair confined to the broken chromatid in cis from repair involving exchange between sister chromatids. We have added a schematic (Fig.S1A) illustrating this distinction and explicitly define the term at its first use in the revised manuscript. Please see our response to Reviewer #2 above for the detailed rationale.

      We thank the reviewer for asking us to clarify the nature of the 237 repetitive targets. The sgRNA targets an Alu sequence and was selected from a larger screen of >20,000 sgRNAs targeting repetitive sequences occurring at >200 genomic sites. In that screen, cellular toxicity did not simply scale with the number of predicted target sites; we therefore selected this sgRNA because its intermediate phenotype allowed us to introduce a large number of programmed DSBs without either minimal perturbation or excessive loss of cells. Thus, the 237-site guide was not an arbitrarily selected Alu-targeting sgRNA. The full repetitive-element screen is beyond the scope of the present study, but we have clarified in the Abstract that these 237 sites are Alu targets and added a brief description of the guide selection in the Methods.

      (2) Introduction:

      "non-local outcome / non-local repair processes": The use of "non-local" is not standard and is obscure for the reader. Specify if it has any meaning or remove it.

      Please see our response above to both Reviewers #2 and #3 regarding our definition and use of “nonlocal repair”.

      The authors mention that replication through a DSB generates four broken ends. However, in case the DSB is reached by one replication fork before the converging one, there are only three broken ends for at least the time required for the converging fork to reach the DSB from the other side. This may influence the repair outcome.

      We agree with the reviewer. If one replication fork encounters the DSB before the converging fork, a transient three-ended intermediate can exist before the second fork reaches the break. This temporal asymmetry could influence repair pathway choice, including engagement of HR, end joining, or BIR-like repair. Our assay captures the resulting SCE outcome but cannot distinguish the order in which replication forks encounter the DSB or the repair pathway engaged at these intermediate stages. We have revised the text (2nd paragraph of the Introduction) to clarify that four broken ends represent the eventual configuration after replication through the DSB, rather than necessarily a simultaneous intermediate.

      (3) Results

      Cell cycle arrest experiment: it seems that a control condition with no Cas9 is missing to conclude better about what looks like a G2-M arrest, but that is not clearly mentioned.

      Please see our response to Reviewer #1, Recommendation 6, regarding additional controls for the cell-cycle arrest experiment. Briefly, the D10A and H840A Cas9 nickases were delivered under the same experimental conditions as wild-type Cas9 but did not produce the cell-cycle arrest observed following DSB induction, providing an internal control for RNP delivery/transfection and supporting the association of the arrest with Cas9-induced DSBs. We would also like to clarify that the observed cell-cycle arrest is primarily a G1/S, rather than G2/M, basing on the FACS signal (see revised Fig.1B legend). This is consistent with the strong G1/S checkpoint in mammalian cells and the predominantly G1 cell-cycle distribution of BJ-5ta cells.

      Note that the font size in Figure 1 is too small for readability.

      We have enlarged the font sizes throughout Figure 1 to improve readability.

      Figure 1B, D10A and H840A conditions:

      The authors mention that nicks can be converted into DSBs through the passage of the replication fork, but do not see any cell cycle defect in the conditions tested. Is it possible that the absence of effect results from the fact that the analysis is done prior to nicks being converted into DSBs? This remark notably applies to the 237 target sites experiment. It seems that controlling for cell cycle delays for longer times is needed to conclude clearly about this aspect. In case a clear absence of cell cycle delay is observed in the 237 target sites in the Cas9 nicking condition, this would suggest that replication born DSBs behave differently from "classical" two-ended DSBs and do not trigger cell cycle arrest.

      We agree that the timing of nick conversion during replication could contribute to the absence of a detectable cell-cycle delay under the conditions examined. However, extending the duration of Cas9 nickase treatment or labeling would not necessarily resolve this question, because persistent Cas9 activity permits repeated rounds of nicking across successive cell cycles, making it difficult to relate a later cell-cycle phenotype to a defined replication-born lesion. More generally, we believe that the relationship between replication-associated nicks, SCE formation, and cell-cycle progression is better addressed in the context of spontaneous SCE, which is the focus of our companion study. The present study is focused on SCE following programmed Cas9-induced DSBs, and analysis of replication-born nick lesions would require precise temporal control (ideally vfCas9 nickases that can be turned off) of individual nicking events relative to replication, for which an appropriate experimental system is not currently available to us.

      Figure 1C should mention somewhere the genomic location of the four targets to clearly show that they correspond to the four major SCE peaks. In addition, there is no legend for the vertical pink stripes. Finally, it might be wise to keep the same y-axis scale for better comparisons.

      Figure 1 overall: it might be wise to clearly show a no Cas9 condition to clearly set the SCE baseline and show that it is independent of Cas9 induction. As of now, it is not clear whether the non-targeted SCE comes from a specific cleavage of Cas9 or not. Such an aspect could benefit from putting Figure S1C in the main Figure 1. Alternatively, results from Chovanec et al 2026 (NAR) should be better restated because the reader does not necessarily have them in mind.

      We thank the reviewer for these suggestions. The expected Cas9 target positions were already indicated by vertical bars in Fig. 1C; however, we agree that this was not sufficiently clear. We have therefore revised the figure legend to explain that the vertical bars indicate the expected Cas9 target positions. The Chovanec et al. (2026, NAR) study focused entirely on spontaneous SCE, which we simultaneously map here as the background signal, rather than the on-target SCEs induced by Cas9. We hope that explicitly identifying the target locations in the revised legend makes this distinction clear and ensures that prior knowledge of the NAR study is not necessary to interpret Fig. 1C.

      We have retained the individual y-axis scales because the magnitude of SCE enrichment differs substantially among conditions. Using a common y-axis scale would make several of the on-target SCE peaks difficult to visualize.

      The section « Disruption of DNA repair genes at the cut site did not measurably alter SCE frequency per cell » is questionable in the results section for the following reasons:

      (i) The DNA repair genes are used here as target sites for Cas9 cleavage, but are not the object of the study, but may be the object of a companion paper. This aspect is slightly misleading.

      (ii) As first mentioned in this section, there is evidence strongly suggesting that inactivating DNA repair genes will not affect SCE, and this is what the authors observed.

      (iii) As an alternative, one could put the emphasis on the fact that the effect of Cas9-mediated inactivation of DNA repair genes (ie LIG3) starts to be detectable only in the washout condition ie after at least one cell cycle. But in this case, this is addressing the role of DNA repair genes in Cas9-induced SCE, which is not the point of the current paper.

      We thank the reviewer for this comment and agree that the original framing of this section could give the impression that these experiments were intended to test the functions of the targeted DNA repair genes in SCE. This was not our intent. Rather, these genes provided defined genomic target sites for Cas9 cleavage, and the primary purpose of the experiment was to characterize SCE associated with Cas9-induced DSBs at these loci.

      As discussed in our response to the Editor Assessment, there are important limitations to using these experiments to infer the consequences of loss of the targeted proteins, including the delay between Cas9 cleavage and depletion of pre-existing protein and, and particularly in the single-cell assay, selection for cells that escape disruptive editing at essential genes. We have added text to the Results explicitly describing these limitations.

      We therefore agree with the reviewer that the delayed effects observed under the washout condition should not be interpreted here as establishing a role for individual DNA repair genes in Cas9-induced SCE. We have revised the section title to “Cas9 targeting of DNA repair gene loci did not immediately alter overall SCE frequency per cell” to clarify the scope of this experiment and to avoid implying that testing the functions of the targeted DNA repair genes is a major objective of the present study.

      The conditions in Table 1 need to be homogenized and better explained:

      - 237 sites and 237 cuts are used: homogenize?

      We thank the reviewer for spotting this. We revised both to be “237 sites”.

      - May explain better the rationale for putting BrdU simultaneously with Cas9 or after 24 h and a wash.

      For the single-targeting sites, we observed more SCE when BrdU was added simultaneously with the Cas9 for the same 24 hours, compared to adding BrdU in the subsequent division after a wash. Therefore, for the 237 sites, we analyzed both conditions.

      - Typo in the text: 237cuts_24ws_40BrdU instead of 237cuts_24ws_BrdU

      We apologize for the lack of clarity in these labels and have substantially revised the Table 1 legend. In brief, the “40” is not a typo. In the 237 sites experiments, wild-type Cas9 considerably prolonged the cell cycle. Therefore, rather than labeling with BrdU for 24 hours as in the other conditions, we extended BrdU labelling to 40 hours in the last two conditions to allow more cells to progress into the subsequent G1 for successful Strand-seq analysis. We clarified that “237 sites 24 + 16hrs BrdU” refers to the condition in which Cas9 RNP and BrdU were added simultaneously. After 24 hours of Cas9 RNP treatment, BrdU labeling was continued for an additional 16 hours (a total of 40 hours of BrdU). The “237sites 24ws40BrdU” condition is the corresponding washout condition, in which Cas9 RNP was removed after 24 hours and cells were then labeled with BrdU for 40 hours post-washout.

      - 237 cuts: Are some sites more enriched in SCE than others?

      Yes, some sites showed greater SCE enrichment than others. We tested whether this variation correlated with chromatin accessibility but found no significant association. This was not unexpected, as the sgRNAs predominantly target Alu elements.

      - Table 1: There is a difference between 237 sites 24ws24BrdU and 237cuts24ws40BrdU, with a significant enrichment of on-target SCE for the latter condition only. Could the increase in SCE rise even more with longer BrdU exposure? In other words, does the low enrichment in SCE at target sites in the 237 sites experiment result from a non-optimal timing for the analysis?

      Yes, this is possible. We did not systematically test additional treatment or labeling durations. A 24-hour Cas9 RNP treatment is typically used for Cas9 RNP-mediated knockout experiments, and we therefore initially used this duration to assess gene-editing outcomes. For Strand-seq, BrdU labeling is ideally limited to approximately one cell division. Because BJ-5ta cells have an approximately 24-hour cell cycle, extending BrdU labeling substantially beyond 40 hours could allow some cells to undergo a second round of replication and become double-labelled. We therefore did not extend BrdU labeling beyond 40 hours. Thus, the lower enrichment in the 24ws24BrdU condition may in part reflect the timing of the assay.

      - Figure 2 / RDCP analysis:

      Interpretation of this figure relies exclusively on the 2026 NAR paper from the authors. This, at least, should be mentioned to help the reader understand it. Once the legend restates, this figure misses clear identification of the SCE and other genomic rearrangements. For readability, maybe the full genome should be kept for the supplementary data, and only the rearranged chromosomes should be kept in the main figure so that the rearrangements are clearly visible and annotated.

      We thank the reviewer for this suggestion. To make the Strand-seq plots interpretable without relying on our 2026 NAR paper, we have added an explanation of Strand-seq orientation in the third paragraph of the Introduction and in Fig.S1A. We have revised Fig.2 legends to improve readability. We have retained the whole-genome view because Strand-seq data are conventionally presented in this format and it provides important genome-wide context for interpreting the observed events. The rearranged chromosomes and events were annotated in Fig.S3.

      (4) Discussion

      - Most DSB never formed or did not produce SCE: how to understand this better? What would be the argument in favor of one or the other possibility?

      We agree that these are two possible explanations that cannot be distinguished by the current experiment. The absence of an SCE at a targeted site could reflect either inefficient DSB formation or repair of a DSB through a pathway that does not generate an SCE. Distinguishing these possibilities would require direct measurement of cutting efficiency at individual target sites, which was beyond the scope of this study.

      - The conclusion about the effect of the Cas9 nickases needs to be toned down as long as the proper timing for SCE analysis has not been performed (see comment above).

      We agree and have toned down this conclusion by specifying that no significant on-target SCE enrichment was detected under the conditions tested and acknowledging that we cannot exclude SCE formation at other time points (Discuss, first paragraph).

      - As much as possible, avoid the use of non-conventional acronyms like URR.

      We agree and have reduced the use of non-conventional acronyms where possible. We have retained URR (under-replicated region), as the term appears seven times throughout the manuscript, but have ensured that it is clearly defined at first use.

      - The discussion about the RDCP analysis in the second paragraph of the discussion should refer to Figure 3.

      Thank you for pointing this out. We added this reference to Fig.3

    1. Author response:

      We thank the Editors and Reviewers for their encouraging evaluation and constructive feedback. We are glad that they recognized the value of this systematic synthesis in reconciling disparate findings across many studies on an important question, and in providing new insights that help address long-standing debates around the interpretation of CP values. They also acknowledged our rigorous data curation involving many original study authors, and our transparency in reporting limitations and null results.

      To address their constructive recommendations, we will provide additional hierarchical regression analyses where possible, state some limitations more explicitly, and condense the Discussion section to improve focus and readability.

      Regarding the hierarchical nature of the data, we distinguish three potential hierarchical levels: studies, monkeys, and neuronal samples. Because individual monkeys contribute roughly one observation per study and cannot be tracked across publications, animal-level random effects are statistically unidentifiable. We will address the rare cases where identical neuronal pools were evaluated across multiple task conditions by providing a sensitivity analysis restricted to one data point per unique neuronal sample. At the study level (median 2, range 1–7 observations per study), we will present linear mixed-effects models with random study intercepts.

      On the bistability findings, we agree with Reviewer 1's concern about the limited number of studies. We noted in the Results and Discussion that this effect currently relies on rotating-cylinder paradigms and emphasized the need for CP to be quantified with other forms of bistable stimuli. We will make this limitation explicit in the Abstract and Figure 8 caption. We will revise the Discussion to emphasize the three primary drivers (neuronal sensitivity, brain area, and stimulus duration) and treat the bistable stimulus effect separately as a distinct finding.

      Regarding Reviewer 1's concern about reaction-time experiments: because reaction-time (RT) paradigms are heavily confounded with detection tasks in the existing literature (89% of detection tasks are RT tasks, and 67% of RT tasks are detection ones), including RT as a separate factor introduces near-complete collinearity. We will make this constraint and the inability to statistically disentangle them explicit in the main text.

      Addressing Reviewer 2's concern regarding the CP–duration predictions, we agree that they depend on specific modeling assumptions. While our predictions—for the feedforward framework in particular—reflect standard models from the literature, we will explicitly acknowledge that alternative feedforward assumptions—such as duration-dependent response covariance—could alter the expected relationship.

      In response to Reviewer 3’s comments on the rationale and future utility of measuring CP, we will revise the Discussion to emphasize that while the interpretation of CP has evolved from feedforward readout to include feedback mechanisms, our results confirm that CP remains a robust neural correlate of subjective perception. Although the exact mechanisms linking CP to perception remain unresolved, this ambiguity does not justify abandoning the metric; rather, CP remains an indispensable tool, provided it is supplemented with additional analyses as outlined in our recommendations.

    1. Author response:

      eLife Assessment

      This study presents an open-source reinforcement learning framework for the real-time, closedloop optimization of spatiotemporal electrical stimulation in engineered neuronal networks. Using single-spike-resolution activity as continuous feedback, the authors provide solid evidence that their platform can identify stimulation patterns that drive specific activity motifs within a structurally constrained four-node circuit. While this reproducible system offers a valuable and accessible tool for interacting with biological neural networks in an adaptive manner, further validation is needed to determine how well these stimulation strategies generalize to larger, unstructured, or more conventional network architectures.

      We thank the editors and reviewers for their kind and insightful words, as well as their constructive feedback and assessment. We agree that in its current stage, the manuscript does not convincingly argue, that specific stimulation strategies applied to one network generalise well to other network architectures. We intend to address this problem by adjusting the focus of the manuscript to be more on the platform than on any specific neuroscience claim, as we believe that this is the more useful angle for the community at large. We also intend to discuss the transferability of our results between the presented culture system and other neuroscience systems.

      Based on the reviewer comments, we will further give a more approachable introduction to reinforcement learning to make the concept more accessible to a wider audience. We will also motivate the algorithm selection in more depth.

      In the following, we will discuss the comments put forward by the reviewers and how we intend to adapt our manuscript to address them. In this provisional response, we will only focus on major concerns. Minor comments, where we are following the suggestions made by the reviewers directly, will not yet be discussed.

      Reviewer #1 (Public Review):

      The manuscript appears undecided about whether it wants to be about RL control, about a technical implementation of long-term stimulation in vitro, or about the properties of neuronal networks and interaction with them. The introduction is well written, comprehensive and insightful, focusing on biological aspects. Methods are then extensively about RL algorithms, without explaining why several were used or why these in particular, but with specialist language hard to understand for neuroscientists. The results then quantitatively compare the performance of the RL but do not really explain what this teaches us about neuroscience or what we learn about the networks beyond that they can be stimulated for longer propagation patterns. Extensive supplementary material almost advertises the hardware built by the team. Some figures suggest, though I’m not 100% certain about this, that different algorithms find different optimal stimulation patterns in the same network - which I find puzzling. What then does this tell us about the stimulation patterns and the variability of the responses?

      We thank the reviewer for pointing out these issues with our manuscript. The goal of our work is the hardware framework. We plan to highlight this more in the revised version. The RL part will be revised with less specialist language and will get less emphasis in the next version of our manuscript. We will also discuss why different algorithms seem to find different solutions, which can be traced back to having different networks and agents finding different local maxima.

      By clearly setting the scope of our manuscript on the hardware aspects and revising the RL part to be seen more as one possible closed-loop control paradigm, we further intend to address the concerns put forward by the reviewer regarding our focus on RL in other parts of the manuscript.

      Reviewer #2 (Public Review):

      An important limitation is that the framework was evaluated using a relatively small and highly structured network. Four stimulation electrodes were positioned around a single network, and stimulation was delivered to microchannels in which axons were concentrated. This configuration is well suited to the initial demonstration, but it remains uncertain whether the same approach will perform similarly in conventional monolayer dissociated cultures, larger networks, or systems with different numbers and spatial arrangements of electrodes. Additional validation across a broader range of network structures and experimental configurations would therefore be needed to establish the general applicability of the framework.

      We thank the reviewer for their feedback. They are right to point out that the manuscript here focuses purely on small and highly structured networks. With the presented experiments, our manuscript does not and cannot make any biological claims about how neurons communicate. However, with this manuscript, we also do not intend to do so, as such an endeavour would be out of scope. In the next version of our manuscript, we will better highlight the focus of our manuscript, which lies on the hardware framework itself. Furthermore, we will discuss in the outlook the generalisability and scalability concerns raised by the reviewer in more detail.

      Reviewer #3 (Public Review):

      Several aspects of the analysis and interpretation require further clarification. In particular, the definition and computation of directional propagation and reward are not always clear. The generalizability of the optimized stimulation patterns across cultures also remains unclear.

      We thank the reviewer for their insightful feedback. We agree with them and will discuss (1) the currently implemented reward based on directional propagation of activity and (2) the limitations and aspects influencing reward algorithm selection in more detail in our revised version of the manuscript. We further believe that we cannot make any claims about generalisability between cultures or when changing the experimental paradigm.

    1. Author response:

      Reviewer #1:

      Some results (or lack of) cast doubt about the ability of the used technique (3T fMRI) to detect the desired effects (bottom-up vs top-down activity). 

      As the reviewer notes, 3 T MRI alone cannot partition the relative contributions of bottom-up and top-down signals, since both are present during movement in an intact system. This reflects the premise of our study design but is also an important caveat when interpreting the group-level results, which characterise the net task-related response. Our central inference therefore focuses on the persistence of activation in PT01, in whom hand movement was absent. PT01 therefore lacks bottom-up signals, and any observed activity must be driven by top-down processes. In the revised manuscript, we provide measures of signal quality and activation magnitude to better characterise the sensitivity of these measurements and will draw on existing evidence that our approach resolves task-specific responses within these nuclei.

      Reviewer #2:

      (1) Attribution of the observed activity to corticocuneate projections: My main reservation concerns the inferential step from "not peripheral" to "corticocuneate". The data establish the former convincingly; the latter may not. The central claim rests on an argument by elimination: because bottom-up drive is excluded in PT01, the residual activity must be top-down and, by extension, corticocuneate. Two distinct gaps should be addressed. First, "top-down" is not equivalent to "direct corticocuneate". Descending influence could reach the cuneate nucleus through multiple indirect pathways. Second, the activity observed at the three levels (cuneate, VPL, S1) need not be serially propagated, since layer 6 corticothalamic projections, for example, could drive VPL independently of any cuneate contribution. I would ask the authors to either provide evidence bearing on the routing, or to consistently use a route-neutral term (e.g. "descending" or "top-down") and reserve "corticocuneate" for the discussion of candidate mechanisms. 

      We agree with the reviewer that our previous attribution of the observed brainstem effects to corticocuneate processing was speculative and should have been presented as such. Our findings support a non-peripheral, top-down contribution but do not allow us to attribute this descending influence to a specific anatomical route. We have therefore revised the manuscript throughout to use “top-down” processing as a more route-neutral term, reserving the corticocuneate pathway for discussion of possible candidate mechanisms. We have also clarified in the revised discussion that activity observed across the cuneate nucleus, VPL, and S1 does not necessarily imply serial propagation through these structures.

      (2) Afferent input arising above the lesion level: The EMG control in PT01 was restricted to hand and forearm muscles. Musculature innervated above C4 (cervical paraspinals, trapezius, and to a variable extent the shoulder girdle) remained available to this participant, and attempted hand movement is frequently accompanied by increased proximal co-contraction, postural stabilization, and altered respiratory effort. Afferent to the upper cervical cord is known to project to the ipsilateral cuneate nucleus, and its activity would produce lateralized, ipsilaterally dominant cuneate input - that is, precisely the pattern reported. This alternative is not excluded by the present control and should be addressed directly, ideally with proximal EMG in PT01 (and, if possible, in the other participants), or at minimum with an explicit discussion. Relatedly, the authors recorded respiratory and cardiac signals: please report whether respiratory volume or heart rate differed between movement and rest blocks, and between groups, since the dorsal medulla lies adjacent to cardiorespiratory nuclei.

      The reviewer raises an important point. To test whether proximal muscle activity could account for the cuneate response in PT01, we collected additional EMG data during attempted hand movements, focusing on muscles innervated above the lesion level. These included the anterior (AD) and middle deltoid (MD), upper (UT) and middle trapezius (MT), and cervical paraspinals (CP), along with two of the original distal recordings (thenar eminence, TE; extensor digitorum, ED). Nonetheless, we found no significant difference in EMG activity in any recorded muscle during attempted left- or right-hand movement versus rest (Supplementary Figure 3A and 3B). To confirm that the chosen electrode montage could detect proximal muscle activity, we further instructed the participant to perform left or right shoulder shrugs during the same session. This produced a clear increase in activity during movement across the trapezius, cervical paraspinal and middle deltoid recordings (Supplementary Figure 3C and 3D).

      In addition, we analysed the cardiac and respiratory recordings acquired during fMRI to determine whether movement-related physiological changes could explain the observed brainstem activity. Importantly, any physiological change in heart rate during movement is global and therefore cannot explain the hand-dependent lateralisation of the cuneate response. Furthermore, cardiac and respiratory nuisance regressors were included in all first-level models. Heart rate showed a small but significant increase during movement compared with rest (controls: +0.31 bpm; SCI: +0.69 bpm; main effect of condition F(1,33) = 10.22, p = 0.003, η<sup>2</sup></sub>p</sub> = 0.24, BF<sub>10</sub> = 8.44), whereas respiratory volume per time (RVT) did not differ between conditions (F(1,33) = 0.35, p = 0.56, η<sup>2</sup></sub>p</sub> = 0.01, BF<sub>10</sub> = 0.24). Given the autonomic consequences of cervical injury, we also tested whether these changes differed between groups. Neither measure showed a Group × Condition interaction (heart rate: F(1,33) = 1.41, p = 0.24, BF<sub>10</sub> = 0.56; RVT: F(1,33) = 1.60, p = 0.22, BF<sub>10</sub> = 0.61), suggesting that they cannot account for the group differences we report.

      We have added the additional EMG control and the cardiorespiratory analyses to the Supplementary Material of the revised manuscript. Together, these controls suggest that neither proximal muscular nor cardiorespiratory factors explain our findings.

      (3) Functional significance of the preserved top-down signal: The discussion establishes that top-down input persists but says relatively little about why it should. If the principal role of descending input to the cuneate nucleus is the gating of incoming afferent traffic, then in the absence of afferents there is nothing left to gate, and one might have expected the signal to be lost. Its persistence is the most interesting aspect of the finding and deserves fuller discussion. Candidate accounts the authors may wish to consider include: an efference copy or predictive signal delivered to a comparator that no longer receives its input, in the framework the authors already invoke (references 27, 28); engagement of the non-lemniscal outputs of the dorsal column nuclei (e.g., cuneocerebellar, cuneo-olivary projections); attempted movement engages motor imagery and attention, in which case the relevant question becomes what distinguishes these from movement-related gating. A related interpretational point: in behaving primates, movement-related modulation of cuneate transmission is bidirectional and includes prominent suppression (refs. 7/12). Note also that BOLD increases are compatible with increased inhibition, so they do not indicate facilitated throughput. 

      We thank the reviewer for their comment and agree that the persistence of this descending signal despite profound loss of peripheral input is a very interesting aspect of the findings, and that our manuscript will benefit from a more extended discussion of this result. We will expand on this and the candidate accounts raised in the revised discussion. We will also clarify that movement-related modulation of cuneate processing may include both facilitation and suppression, and that our finding of increased BOLD activity does not necessarily imply facilitated sensory throughput.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work provides a valuable toolkit for endogenous isolation of projection neuron subtypes. With further validation, it could present a solid method for low-input ribosome affinity purification using a ribosomal RNA (rRNA) antibody. The experimental evidence for the distinct ribosomal complexes is limited to this method and indirect support from complementary analyses of preexisting data.

      However, with additional experimental data to support the specificity of ribosomal complex pulldown and confirmation of the putative ribosomal complex proteins of interest, the study would provide compelling evidence for translation regulation of neuronal development through compositional ribosome heterogeneity. 

      This work would be of interest to neuroscientists, developmental biologists, and those studying translational networks underlying gene regulation.

      Strengths

      (1) This in vivo labeling of specific projection neurons and ribosomal rRNA affinity purification method accommodates a low input of <100K somata per replicate, which is useful for the study of neuronal subtypes with limited input. In principle, this set of techniques could work across different cell types with limited input, depending on the molecule used for cell type labeling.

      (2) The authors are also able to isolate endogenous neurons with minimal perturbation up to the point of collection, preserving the native state for the neuron in vivo as long as possible prior to processing. 

      (3) This study identified over a dozen potential non-ribosomal proteins associated with SCPN ribosomal complexes, as well as a ribosomal protein enriched in CPN.

      We appreciate the reviewer's thoughtful and detailed review. We especially appreciate the positive evaluation of its strengths, including the use of rRNA affinity purification to access ribosomal complexes in low-input neuronal subtypes in vivo with minimal perturbation, and the resulting identification of distinct ribosomal complexes in SCPN and CPN with associated non-ribosomal proteins. We are also pleased by the recognition of its significance in advancing our understanding of neuronal subtype-specific post-transcriptional gene regulation. We have carefully addressed the limitations below.

      Limitations

      (1) In this study, the authors address the advantages of their ribosomal complex isolation method in SCPN and CPN against RPL22-HA affinity purification. While this does show more pull-down of the ribosomal RNA by the Y10B rRNA antibody, the authors claim this method identifies cell-type-specific ribosomal complex proteins without demonstrating a positive control for the method's specificity. 

      There are very limited experiments to truly delineate how "specific" this method is working and whether there could be contamination from other complexes bound by the antibody. I see this as the major limitation that should be addressed. To boost their claims of capturing cell-typespecific ribosomal complexes, the authors could consider applying their rRNA affinity purification pipeline to compare cell types with well-characterized ribosome-associated proteins, like mouse embryonic stem cells and HELA cells.

      The reviewer can completely appreciate the elegance in the neural characterization here, but it seems there needs to be a solid foothold on the specificity of the method, perhaps facilitated by cell types that can be more readily scaled up and tested.

      We thank the reviewer for the opportunity to further clarify how our experimental design addresses the question of specificity of ribosomal complex pulldown. The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or non-specific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.

      (2) The authors followed up on their differentially enriched ribosomal complex proteins by analyzing the ribosome association of these proteins in external datasets. While this analysis supports the ribosome-association of these proteins, there is limited experimental validation of physical association with the ribosome, much less any functional characterization.

      The reciprocal pulldown of PRKCE is promising; however, I would recommend orthogonal validation of several putative ribosomal complex proteins to increase confidence. 

      Specifically, the authors could use sucrose gradient fractionation of SCPN and CPN, followed by a western blot to identify the putative interaction with the 80S monosome or polysomes. This would also provide evidence towards the pulldown capturing association with mature ribosome species, which is currently unclear. This experiment would provide substantial evidence for the direct association of these non-ribosomal proteins with subtype-specific ribosomal complexes.

      We thank the reviewer for the feedback and suggested future directions for candidate validation. We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.

      We appreciate the reviewer's suggestion of sucrose gradient fractionation for polysome profiling. We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph.

      (3) The authors state interest in learning more about the differences underlying translational regulation of projection neuron development. This method only captures neuronal somata, which will only capture ribosomes in the main cell body. There are also ribosomes regulating local translation in the axons, which may also play a critical role in axonal circuit establishment and activity. These ribosomal complex interactions may also be rather transient and difficult to capture at only one developmental stage. Therefore, this method is currently limited to a single developmental snapshot of ribosomal complexes at P3 within the main cell body. It would be exciting to see the extended utility of this method to sample neurites and additional developmental stages to gain further resolution on the developmental translation regulation of these projection neurons.

      We thank the reviewer for the opportunity to further highlight the foundational significance of this work in translational regulation of projection neuron development. Here, we identified subtype-specific differences in ribosomal complex composition within somata at a critical developmental time window for two PN subtypes. This work provides foundation for future investigation of local translation and its regulation in growth cones and axons. It will also enable direct comparison with potential future results from axons and growth cones, developmental subcellular specializations at axon tips that implement pathfinding and circuit formation (such work is not yet feasible due to exceptionally low available input). Our lab has significant ongoing work regarding subtype-specific axon and growth cone biology, including recent investigations of growth cone-localised RNA and protein molecular machinery that regulate circuit formation, maintenance, and function of distinct cardinal PN subtypes (Poulopoulos*, Murphy* et al. Nature 2019; Engmann*, Hatch* et al. Nature Prot 2022; Itoh et al. Cell Rep 2023; Veeraraghavan*, Engmann* et al. Nature Neurosci 2026; Durak*, Kim* et al. bioRxiv 2023; Veeraraghavan*, Tillman* et al. bioRxiv 2025; Tillman et al. bioRxiv 2026). Combining subtype-specific growth cone purification with ribosomal complex investigation represents a logical and exciting future direction. We appreciate the reviewer's encouragement of this future line of investigation and now highlight it in the revised Discussion.

      Likely impact of the work on the field, and the utility of the methods and data to the community:

      The authors introduce a unique pipeline of techniques to identify cell-type-specific ribosomal complex compositions. With more validation, there is certainly potential for those studying neuronal translation to leverage this method in limited primary cells as an alternative to existing methods that do not rely on ribosomal protein tagging, such as ARC-MS (Bartsch et al., 2023), RAPIDASH (Susanto and Hung et al., 2024), and RAPPL (Nature Communications, 2025).

      Reviewer #2 (Public review):

      Summary:

      This study presents a sophisticated molecular dissection of ribosome-associated complexes (RCs) in two well-defined cortical projection neuron subtypes (ScPN and CPN) during early postnatal development. 

      The authors develop and optimize an rRNA immunoprecipitation-mass spectrometry (rRNA IPMS) workflow to recover RCs from FACS-purified, retrogradely labeled neurons, achieving remarkable subtype specificity and biochemical resolution. Through proteomic profiling, they reveal both shared and distinct ribosome-associated proteins between ScPN and CPN, with a focus on non-core RC components and their potential functional relevance. The work advances our understanding of cell-type-specific translation regulation, moving beyond the transcriptome to explore the proteome-level complexity in neuronal subtypes.

      Strengths:

      This work stands out for its technical sophistication and innovation. The authors combine retrograde labeling, FACS purification, and an optimized rRNA IP-MS approach (low input) to isolate ribosome-associated complexes from highly specific neuronal subtypes in vivo, a challenging issue that they execute with impressive rigor. The methodological pipeline is both elegant and well-controlled, yielding high-quality, reproducible data. The depth of proteomic coverage is remarkable, with nearly all known cytoplasmic ribosomal proteins identified, along with hundreds of ribosome- associated proteins (RAPs), including translation factors, chaperones, and RNA-binding proteins. 

      The analysis not only reveals shared components between ScPN and CPN RCs but also uncovers subtype-specific differences in associated proteins. Particularly notable is the integration of this new proteomic dataset with previously published transcriptomic and ribosome footprinting data, which helps to validate the specificity and relevance of the findings. Overall, the clarity of the writing, the robustness of the data, and the transparency of the methods make this a strong and compelling contribution.

      Weaknesses:

      Despite the depth and high quality of the dataset, the study remains descriptive. While the identification of subtype-specific RC components is intriguing, the current version of the manuscript does not explore their functional roles or the biological consequences of their alterations. There is no perturbation, causal testing, in vitro or in vivo manipulation to demonstrate whether these proteins are necessary for ScPN or CPN identity, specific axonal targeting, metabolism, or synaptic function. One important point highlighted by the authors in the discussion - and critical for establishing the subtype specificity of the identified proteins - is that some ribosomal complexes may be specialized for specific developmental stages, rather than exclusively for the subtype-specific needs of projection neuron development. The work presented here provides a valuable starting point for further investigation into such RC specialization. 

      However, it will be essential to determine to what extent these RCs exhibit true subtype specificity, independently of their temporal maturation context. As a result, key mechanistic insights remain a bit speculative. Although several of the identified proteins have known roles in processes like synaptogenesis or metabolism, their relevance to the specific neuronal subtypes under study is not experimentally addressed. 

      That said, given its rich content and the comprehensive early postnatal dataset, the manuscript represents an extremely valuable resource for the community. While primarily exploratory, it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.

      We thank the reviewer for their excellent summary and for their very positive assessment of our work. We are pleased that the methodological rigor, proteomic depth, and integrative analyses were well-received. We thank the reviewer for highlighting that our “work presented here provides a valuable starting point for further investigation into such RC specialization” and that “it lays a strong foundation for future functional studies aimed at uncovering the biological impact of the identified ribosomal complexes.” This is exactly how we view this contribution – as a foundation to share with broader field so such functional investigation and investigation of developmental dynamics can be pursued by multiple groups in the broader related field.

      We again thank the reviewer for the very positive and insightful comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) As listed in the limitations, I would recommend that the authors consider applying their rRNA affinity purification to additional cell lines to confirm the specificity of the method as a positive control, where just demonstrating the technology may be easier to carry out than with more limited samples.

      As we noted in our response above to Limitation 1: “The rRNA affinity purification pipeline was applied identically to both SCPN and CPN, with the analysis focused on comparative, differential analysis between the two subtypes. We employed this approach to subtract out potential background signal or nonspecific binding to the Y10b antibody present in both subtypes. We have now clarified this experimental design in the text, at the end of the first result section.”

      (2) Also, as listed in the limitations, I recommend that the authors provide orthogonal experimental evidence for the putative SCPN and CPN ribosomal complex proteins of interest (e.g., sucrose gradient > western blot).

      As we noted in our response to Limitation 2 above: “We appreciate the recognition of our analysis of external datasets from independent approaches that provides support for physical ribosome association of these candidates. We respectfully submit that the scope of this work is an unbiased comparison of ribosomal complexes between SCPN and CPN to identify and nominate candidates for future investigation. We agree that future work to advance this direction of inquiry would optimally include further characterization of their subtype-specific physical interactions with ribosomes to further elucidate functional implications.”

      Regarding sucrose gradient fractionation (for polysome profiling), we also noted in our response to Limitation 2: “We carefully considered this approach midway through this work, and we pursued pilot experiments to test feasibility. These pilot experiments reinforced findings in the existing literature that it typically requires on the order of 10⁷ cells, well beyond the feasible scope of these low-input purified neuronal subtypes. We respectfully submit that imaging-based approaches, such as proximity ligation assays and super-resolution microscopy, are likely more applicable to such very limited material, though would require extensive candidate-specific optimization beyond the scope of this project. We have now expanded future experimental consideration in the Discussion’s penultimate paragraph”.

      (3) The authors are interested in preserving the native state of the projection neurons to isolate ribosomal complexes; therefore, they may consider using biotin-conjugated CTB (Thermo Fisher) sorting through anti-biotin MACS columns (Miltenyi Biotech) as opposed to FACS in the future. While it does not allow for the same visualization as using a CTB-FP, this slight pipeline adjustment could help save time and physical processing of the projection neurons, helping preserve their endogenous state.

      While we appreciate the reviewer’s constructive suggestion for this theoretically alternative approach, we respectfully submit that this approach is unlikely to effectively isolate projection neurons from the living brain as effectively as FACS approach employed here. We considered this approach. We respectfully offer that CTB standardly enters neurons by binding GM1 gangliosides at the axon terminal, after which it is internalized and undergoes retrograde transport to the soma, several millimeters or more away depending on the projection. By the time CTB reaches the soma, we further respectfully offer that it is standardly fully internalized and is no longer surface-exposed for the theoretically suggested antibiotin capture for MACS. We used magnetic bead-based pull-down for the ribosomes themselves and find those molecular approaches very beneficial. Despite our use and openness to magnetic-conjugate-based separation approaches, we judge that fluorophore-conjugated CTB and FACS remain the most efficient and feasible approach for isolation of these exceptionally polarized projection neurons while maintaining cell viability. 

      (4) While the authors provided a comparison of overall protein detection levels between SCPN and CPN, I recommend an additional analysis and potential normalization for the average core ribosomal protein abundance across samples. I will note that this is more accessible when samples are prepared using TMT-labeling methods, which might be considered for future experiments.

      We thank the reviewer for these suggestions. We appreciate the opportunity to address them together, to further clarify our deeply considered choice of MS-based proteomic analysis and corresponding primary normalization approach. To complement our initial normalization approach, we have now also implemented the reviewer’s suggested approach of normalization to the average intensity of core ribosomal proteins. Notably, these new results agree with those from our initial normalization approach. This insightful suggestion has further strengthened the paper’s results and interpretation. 

      Here, we employed label-free quantification (LFQ), wherein samples are assayed sequentially rather than simultaneously, to most rigorously establish which proteins are truly present in some neuronal subtypes but absent in others. As the reviewer is aware, this capability to determine absence vs. presence of MS-detectable peptides distinguishes LFQ from approaches that assay samples simultaneously, such as isobaric tandem mass tag (TMT) labeling, which standardly offers advantages in relative quantification studies. Advances in sample preparation, instrumentation, and data analytical algorithms (from recent developments in single-cell proteomics and related approaches) now enable application of LFQ in quantitative differential analysis, even in the ultra-low-input regime. We now include both detailed discussion of these points and relevant citation in the text. 

      As the reviewer is also aware, in LFQ, normalization is crucial to ensure the protein quantification is accurate and comparable across sequential runs. For quantitative differential analysis of proteins detected in both subtypes, we have implemented a primary normalization approach employing “median-of-ratios” normalization across all detected proteins for robustness against outliers and technical variability. This primary approach results in similar overall distributions of protein abundances across CPN and SCPN samples (Figure S2A), providing confidence in quantitative comparison between SCPN and CPN in the ultra-lowinput regime. 

      Following the reviewer’s suggestion, to further ensure rigor of identification of differential proteins, we have also implemented a second normalization approach, rescaling each sample to the average intensity of its core ribosomal proteins alone (new Figure panels S2B, C). Subsequent differential analysis reveals equivalent CPN > SCPN enrichment of RPS30/eS30, GUCY1A1, and CELF3 (proteins identified as CPN-enriched with the primary normalization approach). These three proteins rank among the five proteins with lowest p-values, though false-discovery-rate-corrected significance is reduced. This confirmation by a second normalization approach further strengthens the findings.

      We have included clarifications in the main text, added the second normalization approach and subsequent analysis in both the main text and Figure S2. In addition, we have now noted in the discussion that TMT labeling with correspondingly appropriate normalization approaches might better define relative quantitative differences between functional candidates present in multiple subtypes.

      Recommendations for improving the writing and presentation:

      (1) I recommend this as a Tools or Resource article, seeing as the biological conclusions are limited.

      We respectfully submit that this work investigated biological questions and identified biological answers beyond pure development of Tools or offering a dataset as a Resource. We further respectfully submit that the question of differential neuronal subtype-specific translation of shared transcripts has become an emerging area of interest in regulation of precise neuronal and circuit development, maintenance, and function, as well as the neurobiological basis of disease. This has been quite hard to study, and this paper brings a first level of answers to that biological question. Of course, the biological results of this paper are not the complete answer, but as with all biological discovery papers, it provides a foundation for many further studies by multiple labs. 

      (2) In the rationale for studying ribosomal complex machinery, it may be helpful to say that ribosome composition and associated proteins that are present in the cytoplasm provide a way in which ribosomes can tune translation rapidly. This is especially important, seeing as this affords post-mitotic neurons the opportunity to remodel and repair by using readily available proteins while also avoiding the energetic demands of producing new ribosomal complex proteins.

      We thank the reviewer for this insightful comment and fully agree. We have now added this rationale to the Introduction. 

      (3) Is there a need for a CTB injection control? Does GM1 binding affect translation pathways? Please list citations, if possible.

      We respectfully submit that a CTB injection control is not necessary. As noted in our response to Limitation 1 in the Public Review portion, both PN subtypes underwent retrograde labeling with CTB in this work. While it remains unknown whether CTB-GM1 binding affects translation, potential effects of CTB are expected to apply equivalently to both subtypes and would therefore not confound these between-subtype comparisons. Please also see our response to recommendation 4 immediately below, in which we further clarify that retrograde tracing with CTB is a long-standing, well-accepted method shown to cause minimal damage and not interfere with continued neuronal development.

      (4) Do the traced/labeled PNs keep developing normally? In other words, does retrograde tracing inhibit proper PN development? Please list citations, if available.

      We thank the reviewer for encouraging us to further clarify that these are longstanding and well-accepted methods in the field, found to cause minimal damage and not to interfere with continued neuronal development. These and related retrograde labeling methods have led to the identification of the field's cardinal regulatory genes and molecules of axonal connectivity. This includes substantial work from our own lab (PMID in parentheses): Arlotta*, Molyneaux* et al. Neuron, 2005 (15664173); Lai*, Jabaudon* et al. Neuron, 2008 (18215621); Molyneaux*, Arlotta* et al. J. Neurosci., 2009 (19793993); Galazo et al. Neuron, 2016 (27321927); and more recently Sahni et al. Cell Rep., 2021a (34686320); and Sahni et al. Cell Rep., 2021b (34686337). These methods have also employed by other groups, such as Bin Chen (e.g. McKenna et al. PNAS, 2015 (26324926)) and Marta Nieto (e.g. De León Reyes et al. Nat. Commun., 2019 (31591398)). We have now clarified this in the text and included references for the benefit of the readers.

      (5) It may be helpful to mention that RPS30 associates with immature ribosomes during biogenesis (PMID: 25706898).

      We thank the reviewer for the opportunity to further clarify background knowledge of RPS30. As the reviewer is aware, RPS30/eS30 is definitively part of the mature 80S ribosome, as established, e.g., by cryo-EM of human 80S ribosomes (PMID: 25901680). Like multiple other ribosomal proteins, RPS30/eS30 also associates with immature ribosomes during biogenesis (PMID: 25706898). Intriguingly, RPS30/eS30 is produced by cleavage of a fusion protein comprising ubiquitin-like FUBI and RPS30/eS30, with cleavage recently identified as a late step in cytoplasmic 40S maturation (PMID: 34318747). As noted in the text, we confirmed that the peptides used for RPS30/eS30 identification appropriately map only to the amino acid sequence of RPS30/eS30 and not FUBI. We now mention that it is a core component of the mature 80S ribosome and its immature ribosomal association. 

      (6) Please clarify how the rRNA-IP is pulling down mature ribosomes. If not, this should be incorporated into the discussion.

      We thank the reviewer for raising this interesting point. As the reviewer notes, it is well established in the ribogenesis field that ribosomes are continuously produced and therefore exist at various stages of maturation. Our protocol removes a major source of immature ribosomes by subjecting FACS-purified cells to two centrifugal spins that remove the nucleus, the site of ribogenesis and early steps of maturation. In pilot experiments, nuclear removal was confirmed by the absence of a contaminating genomic DNA peak on Bioanalyzer electropherograms of total RNA extracted from input samples (without genomic DNA removal) immediately prior to rRNA-IP. However, ribosomes also undergo cytoplasmic maturation steps, and various functional states of ribosomes have been found to be present in the cytoplasmic fraction. For these reasons, we have referred to what we pulled down as "ribosomal complexes" throughout the manuscript. We now explain this nuclear/immature ribosome depletion and cytoplasmic ribosomal enrichment in the text.

      (7) Would have been interested to see some discussion of the most enriched CPN RAPs or why these might not exist in most replicates (inter-subtype heterogeneity?)

      The reviewer asks an interesting question. As the reviewer is aware, when considering a single sample in isolation, absence of mass spectrometry-based detection is not definitive proof of absence, especially not within this work’s ultralow input regime. One of us (B. Budnik) has substantial experience with ultra-low-input samples across multiple cell types outside of the nervous system, and identifying a protein in three of four identical samples is not uncommon. We therefore used detection in three or four samples as an indicator of presence, while absence across all samples was taken to indicate true absence.

      That said, the reviewer is correct that further diversity and heterogeneity within both CPN and SCPN subtypes additionally might be involved. This is an interesting question for future research, and we have added relevant text to the Discussion. 

      (8) It may be helpful to mention that RPL22, while not stoichiometric, is known to have extraribosomal functions and can pull down independently from assembled ribosomes (PMID: 17381311, PMID: 28575669. Moreover, RPL22/eL22-3xFLAG has been previously used as a control for ribosome affinity-based pulldowns and could be added to citations for Figure S1 (PMID: 28625553, PMID: 28625553).

      We thank the reviewer for this helpful recommendation. We have added relevant text and citations. 

      Minor corrections to the text and figures:

      (1) Want to confirm that in Figure 2, P adj is <0.1 is correct? 

      Yes

      (2) It is not necessary to show MS spectra in Figure 2.

      While we understand the spectra are not strictly necessary, we respectfully submit that they enhance the figure and aid readers in assessing data quality. 

      Reviewer #2 (Recommendations for the authors):

      To strengthen the impact and interpretation of the authors' findings, we encourage consideration of the addition of functional validation experiments for at least one (or more) of the ribosomeassociated proteins that are differentially enriched in ScPN. This could include genetic manipulation (e.g., knockdown or overexpression) to test whether these proteins influence subtype-specific features or neuronal function. Even a limited set of perturbation experiments, such as targeting PRKCE, which is particularly interesting due to its known role in synaptogenesis, would help move the study from descriptive to a more mechanistic nature of the work.

      There appear to be no issues related to data availability, ethics, or compliance, assuming all raw proteomic data and associated code for differential analysis are made publicly available.

      We again thank the reviewer for this encouragement and highlighting that our work provides the foundation for further functional and developmental dynamic investigations. We view this work as providing that foundation for further investigation by multiple labs in the broader fields, beyond the scope of this paper.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of an experiment that demonstrates a disruption in statistical learning of room acoustics when transcranial magnetic stimulation (TMS) is applied to the dorsolateral prefrontal cortex in human listeners. The work uses a testing paradigm designed by the Zahorik group that has shown improvement in speech understanding as a function of listening exposure time in a room, presumably through a mechanism of statistical learning. The manuscript is comprehensive and clear, with detailed figures that show key results. Overall, this work provides an explanation for the mechanisms that support such statistical learning of room acoustics and, therefore, represents a major advancement for the field.

      Strengths:

      The primary strength of the work is its simple and clear result, that the dorsolateral prefrontal cortex is involved in human room acoustic learning.

      Weaknesses:

      A potential weakness of this work is that the manuscript is quite lengthy and complex.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how listeners adapt to and utilize statistical properties of different acoustic spaces to improve speech perception. The researchers used repetitive TMS to perturb neural activity in DLPFC, inhibiting statistical learning compared to sham conditions. The authors also identified the most effective room types for the effective use of reverberations in speech in noise perception, with regular human-built environments bringing greater benefits than modified rooms with lower or higher reverberation times.

      Strengths:

      The introduction and discussion sections of the paper are very interesting and highlight the importance of the current study, particularly with regard to the use of ecologically valid stimuli in investigating statistical learning. However, they could be condensed into parts. TMS parameters and task conditions were well-considered and clearly explained.

      Weaknesses

      (1) The Results section is difficult to follow and includes a lot of detail, which could be removed. As such, it presents as confusing and speculative at times.

      (2) The hypotheses for the study are not clearly stated.

      (3) Multiple statistical models are implemented without correcting the alpha value. This leaves the analyses vulnerable to Type I errors.

      (4) It is confusing to understand how many discrete experiments are included in the study as a whole, and how many participants are involved in each experiment.

      (5) The TMS study is significantly underpowered and not robust. Sample size calculations need further explanation (effect sizes appear to be based on behavioural studies?). I would caution an exploratory presentation of these data, and calculate a posteriori the full sample size based on effect sizes observed in the TMS data.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a well-designed and insightful behavioural study examining human adaptation to room acoustics, building on prior work by Brandewie & Zahorik. The psychophysical results are convincing and add incremental but meaningful knowledge to our understanding of reverberation learning. However, I find the transcranial magnetic stimulation (TMS) component to be over-interpreted. The TMS protocol, while interesting, lacks sufficient anatomical specificity and mechanistic explanation to support the strong claims made regarding a unique role of the dorsolateral prefrontal cortex (dlPFC) in this learning process. More cautious interpretation is warranted, especially given the modest statistical effects, the fact that the main TMS result of interest is a null result, the imprecise targeting of dlPFC (which is not validated), and the lack of knowledge about the timescale of TMS effects in relation to the behavioural task. I recommend revising the manuscript to shift emphasis toward the stronger behavioural findings and to present a more measured and transparent discussion of the TMS results and their limitations.

      Strengths:

      (1) Well-designed acoustical stimuli and psychophysical task.

      (2) Comparisons across room combinations are well conducted.

      (3) The virtual acoustic environment is impressive and applied well here.

      (4) A timely study with interesting behavioural results.

      Weaknesses:

      (1) Lack of hypotheses, particularly for TMS.

      (2) Lack of evidence for targeting TMS in [brain] space and time.

      (3) The most interesting effect of TMS is a null result compared to a weak statistical effect for "meta adaptation"

      Reviewer #4 (Public review):

      Summary:

      Several behavioral experiments and one TMS experiment were performed to examine adaptation to room reverberation for speech intelligibility in noise. This is an important topic that has been extensively studied by several groups over the years. And the study is unique in that it examines one candidate brain area, dlPFC, potentially involved in this learning, and finds that disrupting this area by TMS results in a reduction in the learning. The behavioral conditions are in many ways similar to previous studies. However, they find results that do not match previous results (e.g., performance in anechoic condition is worse than in reverberation), making it difficult to assess the validity of the methods used. One unique aspect of the behavioral experiments is that Ambisonics was used to simulate the spaces, while headphone simulation was mostly used previously. The main behavioral experiment was performed by interleaving 3 different rooms and measuring speech intelligibility as a function of the number of words preceding the target in a given room on a given trial. The findings are that performance improves on the time scale of seconds (as the number of words preceding the target increases), but also on a much larger time scale of tens to hundreds of seconds (corresponding to multiple trials), while for some listeners it is degraded for the first couple of trials. The study also finds that the performance is best in the room that matches the T60 most commonly observed in everyday environments. These are potentially interesting results. However, there are issues with the design of the study and analysis methods that make it difficult to verify the conclusions based on the data.

      Strengths:

      (1) Analysis of the adaptation to reverberation on multiple time scales, for multiple reverberant and anechoic environments, and also considering contextual effects of one environment interleaved with the other two environments.

      (2) TMS experiment showing reduction of some of the learning effects by temporarily disabling the dlPFC.

      Weaknesses:

      While the study examines the adaptation for different carrier lengths, it keeps multiple characteristics (mainly talker voice and location) fixed in addition to reverberation. Therefore, it is possible that the subjects adapt to other aspects of the stimuli, not just to reverberation. A condition in which only reverberation would switch for the target would allow the authors to separate these confounding alternatives. Now, the authors try to address the concerns by indirect evidence/analyses. However, the evidence provided does not appear sufficient.

      The authors use terms that are either not defined or that seem to be defined incorrectly. The main issue then is the results, which are based on analysis of what the authors call d', Hit Rate, and Final Hit rate. First of all, they randomly switch between these measures. Second, it's not clear how they define them, given that their responses are either 4-alternative or 8-alternative forced choice. d', Hit Rate, and False Alarm Rate are defined in Signal detection theory for the detection of the presence of a target. It can be easily extended to a 2-alternative forced choice. But how does one define a Hit, and, in particular, a False Alarm, in a 4/8-alternative? The authors do not state how they did it, and without that, the computation of d' based on HR and FAR is dubious. Also, what the authors call Hit Rate, is presumably the percent correct performance (PCC), but even that is not clear. Then they use FHR and act as if this was the asymptotic value of their HR, even though in many conditions their learning has not ended, and randomly define a variable of +-10 from FHR, which must produce different results depending on whether the asymptote was reached or not. Other examples of usage of strange usage of terms: they talk about "global likelihood learning" (L426) without a definition or a reference, or about "cumulative hit rate" (L1738), where it is not clear to me what "cumulative" means there.

      There are not enough acoustic details about the stimuli. The authors find that reverberant performance is overall better than anechoic in 2 rooms. This goes contrary to previous results. And the authors do not provide enough acoustic details to establish that this is not an artefact of how the stimuli were normalized (e.g., what were the total signal and noise levels at the two ears in the anechoic and reverberant conditions?).

      There are some concerns about the use of statistics. For example, the authors perform two-way ANOVA (L724-728) in which one factor is room, but that factor does not have the same 3 levels across the two levels of the other factor. Also, in some comparisons, they randomly select 11 out of 22 subjects even though appropriate test correct for such imbalances without adding additional randomness of whether the 11 selected subjects happened to be the good or the bad ones.

      Details of the experiments are not sufficiently described in the methods (L194-205) to be able to follow what was done. It should be stated that 1 main experiment was performed using 3 rooms, and that 3 follow-ups were done on a new set of subjects, each with the room swapped.

      We sincerely thank the Editor and the Reviewers for their careful evaluation of our manuscript and for their constructive and insightful comments. We greatly appreciate the time and expertise invested in reviewing our work. The feedback has been invaluable in improving the clarity, rigor, and overall presentation of the manuscript.

      In response to the reviewers’ comments, we have carefully revised the manuscript throughout. The revisions include clarification of the study hypotheses, re-analysis of the TMS data using a mixed ANOVA framework, additional methodological details regarding the TMS procedures and behavioural analyses, expanded justification of the statistical modelling approach, clarification of the room-acoustics paradigm, revision of figures and figure legends, additional discussion of study limitations, and a more balanced interpretation of the TMS findings. We have also substantially revised the Discussion section and improved the overall structure and readability of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is understood that this topical area is necessarily detail-heavy, but if there are ways to streamline the manuscript to more quickly arrive at key results (Figure 3?), the work might have an even greater overall impact.

      We appreciate the reviewer’s feedback and have carefully revised the manuscript to address all comments from all reviewers. However, we have retained the existing order of figures and results to maintain consistency and avoid extensive structural changes that could compromise the clarity and flow of the manuscript.

      (2) Minor point: I believe Equation 1 should be d' = z(H) – z(F).

      Thanks for noticing this, we have fixed the equation.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 115: Towards the end of the introduction, the hypotheses for the current study remain unclear. Please explicitly outline each hypothesis before the Methods section.

      We have now specified the hypotheses tested at the end of the Introduction.

      (2) Line 229: Please state the minimum MEP amplitude criterion used during TMS thresholding (usually 50 µV). Please also reference the EMG hardware and software used to measure MEPs.

      We are grateful to the reviewer for noticing that a few details regarding our TMS procedure were missing in the “Continuous theta-burst stimulation” section of the Methods. We have extensively revised this section to include the required details. We did not, however, use MEP amplitude as a criterion for estimating motor thresholds. For single-pulse TMS-induced motor threshold determination, we used visual observation of the first dorsal interosseous (FDI) muscle twitch (Varnava et al., 2011).

      (3) Line 248: The Coordinate Response Measure corpus contains combinations of callsigns, colours, and numbers. What was the rationale in asking participants to only identify the colour and the number spoken, and not the callsign?

      The rationale for reporting only Color and Number was to ensure an equal number of keyword identifications across phrase lengths—for example, CP0 and CP1 do not contain a callsign. This has now been clarified in the Procedure section of the Methods.

      (4) Lines 327–334: The analyses outlined here are unclear – it appears that there are multiple different statistical tests being conducted on the same outcome variable in this study. Given that the design includes a between-subjects factor (TMS condition) and two within-subjects factors (Rooms, Carrier Length Phrase), could the analyses be simplified by employing a mixed ANOVA as opposed to separate repeated-measures and univariate ANOVAs? If this is, in fact, the analysis that was conducted, please improve the wording. Clarification/correction is also needed surrounding the use of the term “univariate” ANOVA, as this can refer to different statistical tests.

      We appreciate the reviewer pointing this out. We have re-analysed the TMS data using a mixed ANOVA design and reported it as such in the Results section. Although the numerical values have changed, the significant findings and conclusions remain the same.

      (5) Line 340: Whilst the authors identify that an alpha value of 0.05 and Bonferroni corrections were used for statistical inference in two-tailed t-tests, there is no indication of the alpha value used in the interpretation of the ANOVA results. Please include this before the Results section. Furthermore, given that multiple tests are being conducted in this project, the alpha value used to infer statistical significance should be corrected in accordance with the number of hypotheses, to reduce Type I error rates (e.g., 4 hypotheses would result in an alpha inference criterion of α = 0.0125).

      We appreciate the reviewer’s observation. The corrected alpha value was not explicitly reported because all statistical analyses were conducted using IBM SPSS Statistics for Windows, Version 29.0.2.0 (IBM Corp., Armonk, NY; RRID: SCR_002865). In SPSS, Bonferroni corrections are applied by adjusting the p-values rather than the alpha threshold itself. For transparency, we have now clarified in the manuscript the factors included in each statistical comparison to make the tested hypotheses fully explicit.

      (6) Line 346: The sample size calculation used in the current study could be improved. It is unclear why a sample size estimation of ≥18 is used when the actual sample recruited is significantly greater than this (62). Is this because 62 participants were divided across multiple experiments in this paper? Further clarification is needed. The alpha value used in this calculation should also be corrected to account for multiple statistical models.

      We thank the reviewer for noting this. We have clarified that the initial sample size estimation (n = 18) referred to individual ANOVA analyses. In the revised manuscript, we specify in each experimental section (Identity of Sound Environments and Continuous Theta Stimulation) the exact number of participants recruited per experiment, which together sum to a total of 74 participants across all experiments (also clarified in the Participants section of the Materials and Methods).

      (8) Line 438: Many of the statistical tests presented in the Results section have not been outlined in the Methods section or had the rationale explained in the Introduction – this makes the analyses feel confusing and speculative. Explicitly identifying the core hypotheses earlier in the manuscript and clearly stating which hypotheses are confirmatory or exploratory would be an important improvement.

      We appreciate the reviewer noticing this. We have clarified the hypotheses tested in the Introduction (4th and 5th paragraphs) and provided a more detailed account of the statistical analyses in the Materials and Methods (Statistical Analysis section).

      (9) Line 565: It is unclear how the 62 participants recruited in the study were divided across each of the experimental conditions. I would recommend expanding the Participants section in Methods to outline the number of participants involved in each stage of the study.

      Thanks for noticing this discrepancy—this calculation was indeed confusing and incorrect. We ultimately tested a total of 74 participants. We have specified the sample size per experiment in both the “Identity of selected sound environments” and “Continuous theta-burst stimulation” sections of the Methods, and reiterated this in the Results section to prevent confusion.

      (9) Line 792: The Results section as a whole is incredibly complex and lacks structure. This could be condensed significantly – details regarding previous research should be removed from Results, as this should already be outlined in the Introduction as rationale for the current project. It would be clearer to outline each confirmatory and exploratory hypothesis in the Introduction, then identify at each stage in the Results section where it is being tested.

      We appreciate the reviewer’s point. We have rewritten the 4th and 5th paragraphs of the Introduction to outline the confirmatory and exploratory hypotheses. While we have streamlined parts of the Results section to enhance clarity, we have retained contextual information for each analysis to help readers follow the logic of the findings, given the complexity and scope of the study.

      Reviewer #3 (Recommendations for the authors):

      (1) Introduction – Overinterpretation and Hypothesis Clarity. The final paragraph of the Introduction (page 4) discusses the experimental findings and their interpretation, which belongs in the Discussion. This section should instead clearly state hypotheses for both the behavioural and TMS experiments. In particular, the TMS experiment lacks a clear rationale: what mechanism is being tested, and what behavioural outcome is predicted? Please revise this section to focus on the theoretical motivation, clearly defined hypotheses, and expected results, and less on summarizing the results.

      We appreciate this point and have accordingly deleted the final paragraph of the Introduction, as it is already incorporated into the Discussion section. We have stated the hypotheses/rationale more clearly in the Introduction.

      (2) TMS – Mechanism, Timeframe, and Clarity. The manuscript does not adequately explain how TMS produces long-lasting effects relevant to the task, which occur minutes (or possibly longer) after stimulation.

      (a) What is the specific timeframe between stimulation and behavioural testing?

      The timeframe between stimulation and behavioural testing has been clarified in the Methods section, e.g.: “Participants exposed to ‘real’ or ‘sham’ TMS completed the familiarization and behavioural task right after cTBS procedures.”

      (b) What is the evidence that TMS to the prefrontal cortex affects function on this timescale?

      We have added the following paragraph to the Discussion to address the timescale of continuous theta-burst stimulation (cTBS): cTBS, as used in our study, typically induces aftereffects lasting 20–50 minutes (Huang et al., 2005; Wischnewski & Schutter, 2015). While these effects are well established in the motor cortex—with motor-evoked potential changes persisting for up to one hour—recent evidence suggests that similar durations of cortical modulation can also occur in the prefrontal cortex (Taylor et al., 2025). Specifically, studies applying inhibitory rTMS to the dorsolateral prefrontal cortex (dlPFC) during cognitive tasks have demonstrated functional effects lasting up to one hour in healthy participants (Wagner et al., 2006). Furthermore, Tupak et al. (2013) showed that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels, reflecting decreased cortical activity, for at least 45 minutes—the same duration as the experimental task in our study. Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), the fNIRS-measured alterations in local cerebral blood oxygenation provide an indirect but reliable indicator of TMS-induced neural modulation within this timescale.

      (c) Can post-stimulation effects be objectively measured or confirmed?

      Although no objective post-stimulation measures were collected in the present study, we acknowledge this as a limitation. However, previous research has shown that inhibitory rTMS to the dlPFC leads to reduced oxygenation levels—reflecting decreased cortical activity—for at least 45 minutes (Tupak et al., 2013). Given that changes in cerebral haemoglobin concentration closely correspond to neuronal activation (Liao et al., 2013), these findings indicate that fNIRS can serve as an indirect but reliable method for confirming TMS-induced neural modulation. We plan to incorporate such objective measures in future studies.

      We thank the reviewer for raising these important points and have addressed them by updating the Methods and Results sections and adding a paragraph to the Discussion. We would like to clarify, however, that the TMS effects observed in our study are not weak: the statistically significant differences between sham and TMS conditions were accompanied by large effect sizes, indicating that bilateral inhibitory stimulation of the dlPFC produced a robust and consistent effect across participants—specifically, a reduction in performance, reflecting decreased improvement in speech understanding with increasing exposure to the reverberant environment.

      (3) Figure 3A – Anatomical Specificity and Interpretation. Figure 3A implies precise stimulation of dlPFC and its projections to auditory cortex (A1), but the authors cannot actually target dlPFC or its connections specifically with this approach. Rather, the TMS protocol disrupts an undetermined region of PFC, with diffuse downstream effects. This should be clearly acknowledged in the figure legend and main text.

      We appreciate this comment and have acknowledged this point in the figure legend and Discussion, e.g., Figure 3 legend: “Although the TMS protocol was intended to target the dlPFC, it likely affected adjacent prefrontal regions, leading to diffuse downstream effects that may have included modulation of A1.”

      (4) Additionally, the Discussion overstates the evidence for a specific dlPFC → AC role in reverberation learning. The weak and poorly localized TMS effect does not support strong claims about this pathway. Please scale back this interpretation and focus more on the robust psychophysical results, which are the manuscript's stronger contribution.

      We thank the reviewer for this comment. We have revised our interpretation to clarify that the proposed dlPFC–auditory cortex link is speculative, and have added caveats regarding the limited spatial precision of TMS targeting and individual variability in its effects, in the Discussion section “A role for dlPFC in statistical learning of room acoustics.”

      (4a) Line 152: Extra comma after “of”? Also, why are there square brackets around “callsigns” etc.?

      Fixed.

      (4b) Line 156: “RRID:SCR_001622” is unexplained and likely unnecessary—consider removing.

      Removed.

      (4c) Line 159: Why was no ramping applied at the end of the noise? Please clarify.

      Similar to Brandewie & Zahorik (2013), no ramping was applied. This has been clarified in the Methods (“Acoustic Stimuli”).

      (4d) Line 207: Methods do not describe the sham TMS protocol—please add this information.

      Thanks for noticing this. Information on the sham TMS protocol has been added.

      (4e) Line 352: Fix bracket formatting.

      Fixed.

      (4f) Line 382: Sentence is grammatically incorrect—please revise.

      Fixed.

      (4g) Line 394: Unclear use of square brackets—clarify or standardize.

      Fixed.

      (4h) Line 401: It is unclear how interleaving the talker and length ensures a different room each trial. Aren't these variables independent?

      The reviewer is correct: the only variable that was pseudorandomized was room order, to prevent carry-over effects, similar to Brandewie & Zahorik (2013). This has been clarified in the Methods section.

      (4i) Figure 1: Clarify that AI-generated images refer only to the room images, not other components.

      Fixed.

      (4j) Figure 2A: Confirm that “overall” includes all speakers and durations—clarify in legend.

      Fixed.

      (4k) The interesting duration effects in Figure 1D are not discussed in the text and appear before overall room effects in Figure 2A—please reorder and comment on these results.

      Fixed.

      (4l) Supplementary Figure 2: Caption contains a typo (“Lecture Room/Open-Plan Office”).

      Typo has been fixed.

      Also, I recommend adding this result to the main figure set—e.g., include overall d′ for all six talkers in Figure 2 alongside rooms (2A) and lengths (2D).

      We thank the reviewer for the suggestion. Including overall performance for all six talkers in Figure 2 would require substantial restructuring and risk making the figure crowded. We have therefore retained these results in the Supplementary Materials (Supplementary 2 and 3), as originally presented.

      (4m) Line 555: Phrase “to better understand” could be clearer—consider rewording.

      This section has been reworded.

      (4n) Lines 587–592: The lack of main effect of TMS is helpful, but more important is whether interactions between TMS and room/length variables occur. Please report these interactions, as they are central to interpreting the TMS effects.

      We appreciate the reviewer highlighting this. We have reviewed this analysis and reported the interaction Condition x CP length as follows: “A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], and no significant interaction Condition x CP length was observed: [F (3,93) =0.48, p=0.69, ŋp2 = 0.01]; confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population”. 

      (4o) Line 601: Reiterate the timeframe of the TMS-behaviour gap. Is there supporting evidence that TMS can affect behaviour over this duration? Could null effects reflect fading TMS efficacy?

      We appreciate the reviewer pointing this out. We have clarified the timeframe of TMS stimulation in both the Methods and Results, e.g.: “The procedure began with TMS manipulation, and although the behavioural task lasted 45 minutes, the inhibitory effects of TMS extended for at least 60 minutes post-stimulation (Huang et al., 2005; Gamboa et al., 2010; Hoogendam, Ramakers, & Di Lazzaro, 2010; Romero et al., 2022).” However, we cannot dismiss individual differences in the duration of TMS effects, nor differences in efficacy duration between anatomical areas (motor cortex vs. dlPFC). We have noted this in the Discussion section “A role for dlPFC in statistical learning of room acoustics” (Pallant, 2011).

      (4p) Figure 3I: The “meta-adaptation” effect is marginal in both Exp 1 (p = 0.03) and Exp 2 sham (p = 0.04). These should be interpreted cautiously, given their statistical fragility.

      We appreciate this comment. We have now calculated effect sizes for all Wilcoxon signed-rank tests (Pearson’s r) and report them. For the two comparisons noted by the reviewer, the effect sizes are medium (Exp 1) and large (Exp 2). We are therefore confident that, even where the p-values are not extremely low, the statistical differences are reliable.

      (4q) Line 696: Reverberation is described as “common,” but it is nearly universal. Consider rephrasing to reflect this.

      We appreciate this suggestion and have rephrased this line.

      (4r) Line 816: The authors state that TMS reduced overall performance, but the earlier ANOVA (lines 587–592) shows no such effect. Please correct this discrepancy.

      We appreciate the reviewer noticing this. This section has been clarified: the lack of statistical significance at lines 587–592 relates to the comparison between a subset of ‘no-TMS-exposed’ listeners and ‘sham’-TMS-exposed listeners, made only to demonstrate the absence of placebo effects in the sham sample. Following the reviewers’ suggestions, we also re-analysed the data using a one-way ANOVA; this slightly changed the numerical values of the reported main effect but did not change the statistical outcome.

      Reviewer #4 (Recommendations for the authors):

      (1) Lines 201–202: It's not clear what is meant by combination and by carrier length here.

      This section has been rewritten for clarity.

      (2) Line 330: What is meant by “Univariate” here? I think this was a mixed ANOVA, with a betweensubject factor of TMS exposure and the remaining factors within-subject.

      We appreciate this suggestion; we have re-analysed this section to use a mixed-ANOVA design. The numerical results differ, but the statistical outcome remains the same.

      (3) Lines 335–343: This is impossible to follow if one does not understand that there were 3 followup experiments.

      Thank you for highlighting that this section was confusing. We have rewritten it to clarify the following: “Three follow-up experiments were performed (univariate ANOVA) to assess whether speech understanding was affected by room context (i.e., the third room in which Open-Plan Office and Underground Car Park were learnt), with one between-subjects factor: room context (levels: Anechoic Room, Living Room, Lecture Room, and Highly Reflectant Room).” We have also added the following earlier in the Methods: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).”

      (4) Lines 414–416: The review of Tsironis et al. (2024) (doi:10.1177/23312165241273399) does not provide strong evidence that there are multiple scales for adaptation to room reverberation (most adaptation effects stabilize within 1 sec).

      We apologize for this mistake, which arose from an issue with our reference manager. It has been corrected to: Robinson, Harper, & McAlpine (2016), Nature Communications, and Simpson, Harper, Reiss, & McAlpine (2014), Journal of Neuroscience.

      (5) Line 427: The supplementary figure shows that many subjects did not achieve asymptotic performance.

      We appreciate the reviewer pointing this out. The fittings have been extensively reviewed; please see the Methods and Results for the new fitting analysis. Indeed, some participants, although very slowly, keep improving over time without reaching clearly asymptotic behaviour. This section, however, referred specifically to the point at which performance stabilised within ±10% of final performance.

      (6) Line 428: The ±10% statistic is random (as discussed below). And why switch to HR now? And what is its meaning when the HR did not converge by the end of the run?

      We appreciate the reviewer raising the inconsistent use of d′ versus HR. d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis; time-course analysis does not allow us to calculate d′ at each trial or time point, owing to the lack of HR and FA values for single trials. We have clarified this in the Methods: “Given how d′ was calculated for |Color| and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data for each |Color| and |Number| affecting the temporal resolution of any generated curve. We therefore analysed the development of individual and average performance in each acoustic environment using cumulative hit rates, applying a 5-point moving average (~7 seconds) to each trace and plotting performance as a function of mean cumulative exposure time.”

      We have also extensively reviewed our fits, following these steps: (i) we compared single- vs. double-exponential fits across 22 participants (Supplementary Figure 1), which showed that double exponentials provided a better R<sup>2</sup> for the majority of participants; (ii) we re-ran all analyses forcing the fits to the HR endpoint; (iii) characterizing taus was not informative for our sample, given flat-like performance for some participants (Supplementary Figure 1)—in these cases, taus do not aid understanding of how performance stabilizes over time, particularly given the use of two taus; and (iv) cutting initial points differs by participant.

      (7) Line 434: Or that they learned/adapted to other characteristics that were fixed.

      Thank you—we have added “adapted to” in the sentence.

      (8) Supplementary tables often show differences, but the actual values are not shown. Also, the tables and figures randomly switch between d' and HR.

      Supplementary tables are intended only to show additional detail not reported in the main text or figures, to avoid redundancy; means (referred to in the Supplementary tables) are always shown in the main figures. We appreciate the reviewer raising the inconsistent use of d′ versus HR, addressed above, and have clarified this throughout the Methods and Results.

      (9) Lines 436–442: There seem to be a lot of issues with the fitting shown in Supplemental Figure 1 and Figure 2B:

      (a) It does not seem to converge, especially for the green line. So, presumably, the asymptotic value obtained for tau_slow was the upper bound set to 2000 s for many subjects' conditions. But those values are never shown—they should be in Supplemental Figure 1.

      (b) Then the FHR value, derived from that, is completely dependent on what the bound was set to, and is therefore arbitrary. And its value of 10% is also arbitrary. Why do this when tau itself of an exponential model represents the time it takes to reach 67% of the asymptotic value, from which one can derive whatever time it should take to reach the final 10%?

      (c) Even the use of the model specified by Equation 2 seems arbitrary. Average data in Figure 2B do not provide strong evidence for two time scales. If the authors are worried about the instability of the data at the beginning, a simple exponential with a weighted fit that prioritizes the later portions seems sufficient.

      (d) Lines 430–434: This conclusion seems wrong, based only on the arbitrary measure chosen for “global likelihood learning.” Looking at Figure 2B, there is no evidence that the green graph reached any asymptote, while for the yellow and blue it appears to have. The authors should try fitting a simple exponential function to it to show that tau is larger.

      We appreciate the reviewer’s comments and have significantly revised these sections of the Methods and Results. In summary, we fitted the data with single- and double-exponential functions. Double exponentials were fitted to the full time course. Single exponentials were fitted to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) to mitigate initial variability (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments. Compared against the truncated single-exponential fit, the double-exponential model retained a lower mean AIC (−420.7 vs. −399.1) and remained the preferred model for approximately half the subjects. The double-exponential model was preferred not only for its automated nature (requiring no manual truncation) but also for the magnitude of improvement: when the single model was superior, the advantage was marginal (ΔAIC = 9.7 ± 1.4), whereas the double model’s advantage was substantial (ΔAIC = 52.8 ± 8.9). We therefore used the double-exponential fit for further analysis.

      (10) Lines 436–446: How can this analysis be performed if asymptotic performance was not achieved in any of the conditions (nothing has plateaued in Figure 2B)? Also, the 10% FHR measure is dependent on the FHR estimate; correlating two measures based on the same measure is, by definition, expected to be correlated. This result seems to reflect that if one's learning is faster within a fixed number of trials (150), one has more opportunity to reach a higher final PCC even if asymptotic performance is identical.

      To clarify, the variables being correlated are not the FHR and ±10% of the FHR values themselves, but rather the time points at which each participant reached ±10% of their individual FHR during the task. This analysis therefore does not involve two measures derived directly from the same estimate. The timing of reaching ±10% of the FHR reflects the learning-settling trajectory rather than the FHR magnitude, so there is no a priori reason for the two measures to be intrinsically correlated. While asymptotic performance was not reached within 150 trials for some participants, the estimated FHR still provides a consistent individual marker of learning rate, allowing comparison of relative learning dynamics across participants and conditions.

      (11) Lines 448–461: Brandewie & Zahorik (2013) show that a large portion of that improvement is due to tuning to the voice and location of the speaker. Also, in the current study, there are some issues with the anechoic condition (see below).

      We thank the reviewer for this comment. This section refers specifically to results related to carrier phrase length, not to speaker identity (addressed separately below) or location, both of which were fixed in our study and therefore unlikely to account for the observed effects. We address the reviewer’s concerns about the anechoic condition in our responses below.

      (12) Line 491: What were the average trial numbers for the steady and initial trials? Also for the anechoic condition?

      We appreciate the reviewer raising this. We analysed the average trial number at which initial and steady trials occurred across a total of 360 trials: for all 22 participants, initial trials mean = 6 ± 4 and steady trials mean = 37 ± 9; sham TMS: initial trials mean = 5 ± 4, steady trials mean = 34 ± 6; real TMS: initial trials mean = 10 ± 9, steady trials mean = 38 ± 12; anechoic condition: initial trials mean = 6 ± 3, steady trials mean = 36 ± 8; and the 11 randomly selected subjects: initial trials mean = 6 ± 5, steady trials mean = 38 ± 7. This information has been added to the relevant Results sections.

      (13) Line 498: It's still not clear when the anechoic condition was performed. Lines 200–205 talk about combinations in which the Lecture Room was swapped, but it's impossible to follow when and how often that occurred. Given that the anechoic room was not included in the same way as the main three rooms, the conclusion at lines 500–504 is questionable.

      Thank you for noticing this. We have rewritten the relevant section of the Methods (“Identity of sound environments”) as follows: “Additionally, we performed three follow-up experiments in different groups of listeners, assessing performance across combinations of three rooms: (1) Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners); (2) Living Room/Open-Plan Office/Underground Car Park (11 naïve listeners); and (3) Highly Reflectant Room/Open-Plan Office/Underground Car Park (10 naïve listeners). These conditions were used to determine whether a specific room combination was required to observe improvements in speech performance with increasing exposure to room acoustics—i.e., with increasing carrier phrase length—and were assessed in the same way as Brandewie & Zahorik (2013).” The anechoic room was therefore explored in the same manner as the main three rooms.

      (14) Lines 508–521: Tuning to the talker's voice/location would not predict that the effect would be different for a different voice.

      We thank the reviewer for this point. If the improvement in speech understanding were due to tuning to a specific talker’s voice or location, we would expect the effect to differ across talkers. However, our analysis across six talkers (three female, three male) showed no significant interaction between talker, carrier phrase length, and room (F(30,630) = 0.73, p = 0.85, ηp<sup>2</sup> = 0.034). Although overall performance differed across talkers (main effect of talker: F(5,105) = 27.19, p < 0.001, ηp<sup>2</sup> = 0.56), these differences did not modulate the carrier phrase effect. We therefore conclude that the improvement in speech understanding with increasing carrier phrase length is consistent across talkers.

      (15) Lines 523–536: Neither of these tests addresses the question directly. That would require switching the talker randomly between the carrier and target (or throughout the sentence).

      We thank the reviewer for this comment. We respectfully disagree that our analyses fail to address the question. While our experiment was not specifically designed to test the effect of switching talkers between the carrier and target segments, we examined whether adaptation to a talker could explain the improvement in performance with increasing carrier phrase length through three complementary analyses: (1) a repeated-measures ANOVA testing for interactions between talker and carrier phrase length (see response above); (2) an analysis of potential carry-over effects across consecutive same-talker trials; and (3) an assessment of talker-learning effects in the absence of reverberation (anechoic condition). As detailed in the Results section “Improvements in performance are explained by exposure to the environment, not talker idiosyncrasies,” none of these analyses revealed evidence that talker identity influenced the observed improvement in speech understanding. We therefore conclude that the performance improvements with increasing carrier phrase length are better explained by adaptation to the acoustic environment than to specific talkers.

      (16) Line 544: Why is FHR used in this measure when d' is used for the standard analysis in Figure 2D?

      As noted above, d′ is referred to only in statistical analyses that do not bear a specific relation to the time-course analysis. Given how d′ was calculated for | Colour | and |Number|, it was not considered a useful metric for describing performance across time in different environments, owing to the paucity of data affecting temporal resolution. We therefore used cumulative hit rates, with a 5-point moving average (~7 seconds), plotted against mean cumulative exposure time. This has been clarified in the Methods.

      (17) Also, why is FHR, as opposed to HR (which I assume is really PCC), computed across the whole experiment?

      FHR refers to the Final Cumulative Performance. This naming was used to distinguish it from trial-by-trial Hit Rate used in the time-course analysis. FHR is the final data point of the cumulative hit rate—i.e., after all responses have been accumulated in that listening environment. This has been clarified throughout the manuscript.

      (18) Still worse, it's also not clear when these anechoic trials were measured.

      We appreciate the reviewer noting a lack of clarity here. We performed three follow-up experiments in different, naïve groups of listeners, assessing performance across combinations of three rooms, including Anechoic Room/Open-Plan Office/Underground Car Park (10 naïve listeners), assessed in the same way as Brandewie & Zahorik (2013). The anechoic room was therefore explored in the same manner as the main three rooms; this has been clarified in the Methods and reiterated in the Results.

      (19) And specifically, from Supplemental Table 4, it looks like the improvement was considerable (up to 15%), supporting that the effect is occurring. Also, note that there seems to be something numerically wrong in Supplemental Table 4: the improvement CP0–CP1 is −6, CP1–CP2 is −9.667, and CP2–CP3 is −2.5. Based on this, CP0–CP2 is expected to be −15.667 (which it is), but CP0–CP3 is expected to be −18.167, yet it's stated as −13.167.

      We appreciate the reviewer pointing this out. Our statistical analysis does not match the calculations the reviewer derived from Supplemental Table 4. For transparency, we report below the means for each carrier phrase in the anechoic room, exported directly from SPSS, which are the values reported in the manuscript. We have reviewed this section to improve clarity and have included a link to the raw supplemental data.

      CP0: Mean 40.500, SE 4.548, 95% CI [30.212, 50.788]

      CP1: Mean 46.500, SE 5.296, 95% CI [34.519, 58.481]

      CP2: Mean 56.167, SE 3.777, 95% CI [47.624, 64.710]

      CP3: Mean 53.667, SE 3.966, 95% CI [44.695, 62.638]

      (20) Lines 555–557: This sentence seems grammatically incorrect.

      It has been corrected.

      (21) Lines 559–560: The sentence “a brain region implicated in listening performance in noise (Houtgast & Steeneken, 1973; Knudsen, 1929; Lochner & Burger, 1961)” seems to imply that the cited studies support dlPFC being the brain region implicated in hearing in noise. None of these studies does that.

      Thank you for noticing this—this was an error introduced by our reference manager and has been corrected.

      (22) Lines 586–592: What was the “overall performance” measure—d′, HR, or PCC? Also, what is “univariate” analysis here? A mixed ANOVA with a between-group factor of condition and withingroup factors of room and CP length would be appropriate, and the whole group of 22 subjects should be used for the “no-exposure” group, rather than a random selection of an 11-subject subgroup.

      We appreciate this comment and we have revised this analysis to include a mix ANOVA as suggested by the reviewer. It reads as follows in the Manuscript: “Given the potential placebo effects of a ‘sham’ TMS stimulation, we first tested whether our sample of 11 ‘sham’ TMS participants exhibited similar behavioural performance to the larger sample of 22 participants who had not been exposed to any TMS manipulation. A mixed ANOVA with a between-group factor of Condition and within-group factors of room and CP length confirmed that performance in these two populations was comparable, with no significant main effect of TMS conditions observed (‘sham’ vs. no exposure to TMS): [F (1,31) =0.01, p=0.91, ŋp2 = 0.00], confirming that participants experiencing ‘sham’ TMS did not perform significantly differently from the ‘no exposure to TMS’ population.”

      (23) Lines 600–601: By “univariate ANOVA” is meant one-way ANOVA? And why wasn't it a two-way ANOVA with factors of room and sham/real TMS? More importantly, the 10% of FHR measure is arbitrary and should be replaced by standard fitting, as discussed earlier. Looking at Figure 3B, the black line appears near an asymptote while the red one is still growing toward the end, and that should be reflected in tau.

      The ANOVAs in this and other sections have been revised based on the reviewers’ suggestions; mixed ANOVAs have instead been performed and reported, yielding similar results. The fittings and 10% FHR calculations have also been extensively revised. We now show that double exponentials are better suited to our dataset, and that two-tau parameters are not informative about when performance reaches a stable point during the task.

      (24) Also, why is the exposure time on the x-axis different in Figure 3B from Figure 2B (150 vs 500)? And it would be good to see where the across-room average no-TMS data would lie here (or show the equivalent average in Figure 2B).

      We appreciate the reviewer noticing this mismatch. Figure 2B shows the time course for each environment (150 s of exposure to each), whereas Figure 3B shows all environments collapsed (150 s × 3). This is because, for the 22 listeners without TMS exposure, a Rooms main effect was observed, justifying separate time courses per room; however, for listeners exposed to sham and real TMS, no Rooms × TMS interaction was observed, so separating time courses per room was not statistically justified. As the only significant effect was TMS condition, we grouped the time spent across all environments by TMS condition.

      (25) Lines 606–617: This analysis and Figure 3C have the same issues as described for Figure 2C— asymptotic performance was not achieved for many conditions, so the 10% measure is arbitrary, as is the resulting correlation.

      We appreciate the reviewer raising these fitting issues. This part of the manuscript has been extensively revised, including new analyses and figures, although our results have not changed. Additional detail has been added to the Methods (“Speech Performance Analysis and Timecourse Fittings of Mean Cumulative Hit Rates”), and the following summary has been added to the Results (“Statistical learning of reverberant environments occurs over long and short time courses”): we fitted data with single- and double-exponential functions; double exponentials were fitted to the full time course, and single exponentials to both the full time course (Supplementary Figure 1) and a truncated version excluding the first 10 points (~14 s) (Supplementary Figure 2). AIC comparisons heavily favoured the double-exponential model for the full time course (mean AIC: double −442.8 vs. single −376.4), providing a better fit for ≥20/22 subjects across all environments, and remained preferred when compared against the truncated single-exponential fit (−420.7 vs. −399.1, preferred for roughly half the subjects). The double-exponential model was preferred for both its automated nature and the magnitude of improvement (marginal ΔAIC = 9.7 ± 1.4 when the single model won, versus substantial ΔAIC = 52.8 ± 8.9 when the double model won). We therefore used the double-exponential fit for further analysis.

      (26) Lines 619–620: Was d′ really calculated using FHR (the final value) and a non-final False Alarm Rate? This would be arbitrary. It is still unclear how HR and FAR are defined here. There is no apparent benefit to switching between HR (Figure 3B/C, presumably overall percent correct, PCC), d′ (D, E, F), and back to HR (H, I).

      We appreciate the reviewer raising this. We have clarified in the Methods (“Speech performance analysis and time-course fittings of mean cumulative hit rates”) how Hit Rate, False Alarm Rate, and d′ were calculated.

      (27) Line 655: Figure 3E should be Figure 3F.

      Corrected.

      (28) Lines 661–670: Why was Number only analyzed for initial trials, while Color was analyzed for both initial and steady trials? Also, the choice of trials 9–10 for “steady” is arbitrary and should be shown somewhere in Figure 3B.

      We analysed performance for |Number| on initial trials (1–2) only, for CP0, because performance for this speech token could only improve if positively influenced by short-term, within-trial accumulation of information (acknowledging that | Colour | precedes |Number|). To determine how much knowledge accumulated over repeated exposures — i.e., metaadaptation—we instead needed to compare performance on a speech token whose improvement could only stem from knowledge gained across previous trials, not within a single carrier phrase. We compared | Colour | performance on CP0 between initial trials (1–2) and later, steady trials (9–10). If this hypothesis is supported, it suggests that | Colour | performance for CP0 benefits from meta-adaptive information conveyed across trials as knowledge of the environment’s global structure accumulates—our proxy for meta-adaptation (Figure 3G). The choice of trials 9–10 as “steady” follows work on animal models of meta-adaptation (Robinson, Harper, & McAlpine, 2016), which described a faster adaptation rate after the eighth presentation of an environment. This has been clarified in the manuscript.

      (29) Lines 724–728: This description is confusing. The main effect of “Lecture Room” vs. “Highly Reflectant” context is that performance is very good in the Lecture Room (green line) and poor in the Highly Reflectant Room (purple). Averaging that with OPO and CP and reporting “mean difference = 20.09” (in what units?) as “overall performance” distracts from the main point. Moreover, how can that be entered into an ANOVA when the room contexts differ (LR+OPO+CP vs. HR+OPO+CP)? That ANOVA seems incorrect; it should only be performed on OPO+CP across the two contexts.

      This section has been revised and re-analysed as suggested. Redundant and unnecessary statistical comparisons were removed, retaining only those that show the effect of context on OPO and CP when comparing the different contexts in which these common environments were learned.

      (30) Lines 740–750: Again, it is not surprising that when Living Room replaces Lecture Room—and performance in Living Room is worse than in Lecture Room—the average of LiR+OPO+CP is lower than LER+OPO+CP, if OPO+CP performance is unchanged. The interesting question is whether anything changed in OPO+CP performance, as suggested for the previous point.

      This section has been revised as suggested by the reviewer.

      (31) Lines 752–771: Again, the same issue—the main effect is that performance in Anechoic trials is worse than in Lecture Room or Living Room trials, which alone explains the group difference. I am also sceptical of the finding that Anechoic performance is worse than reverberant performance, contrary to typical spatial-release-from-masking results, where reverberation degrades performance by adding noise energy at the better ear and reducing binaural benefit through decorrelation. This may be an artefact of how target and noise levels were normalized after convolution with HRTFs/BRIRs (or the use of Ambisonics); no acoustic analysis of the stimuli is provided. At minimum, the total received level at the two ears for target and masker in every environment should be reported. Brandewie & Zahorik (2013), using equivalent anechoic and reverberant conditions, never observed reverberant performance to exceed anechoic, contrary to what is stated here (lines 754–755).

      We appreciate the reviewer raising these points. The statistical analysis in this section has been revised: only the common rooms across the three-room conditions (Open-Plan Office and Car Park) were directly compared. Performance in Living Room and Lecture Room was not statistically different (Results, paragraph 4, “Statistical learning of room acoustics is tuned to universally experienced reverberation times”). However, Anechoic and Lecture Room performance remained significantly different (mean difference = 11.08, t(9) = 2.66, p = 0.013, Cohen’s d = 0.84), as did performance in the common rooms when learned in the context of Lecture Room versus other contexts (mean difference = 9.7, F(1,41) = 13.24, p < 0.001, ηp<sup>2</sup> = 0.26).

      While this setup resembles many masking studies, Brandewie & Zahorik (2013) tested four rooms simultaneously, whereas we tested three-room conditions explicitly designed to test environment-mix adaptation. We observed a synergistic relationship between performance in ‘good’ reverberant rooms (Lecture Room, Living Room) and the common but less favourable rooms (Open-Plan Office, Car Park, with longer RT60): in the absence of a ‘good-reverb anchor,’ performance in the common rooms improved less over time, possibly because participants had less to leverage in anechoic environments.

      We agree that verifying at-ear acoustic levels is critical to ruling out a normalization artefact. Our stimuli were normalized in the 41-channel sound field, not at the listener’s ears: source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses, the 41-channel noise energy was scaled to a target of 70 dB, and the 41-channel speech field was scaled to the target SNR; the 41-channel signals were then rendered to two channels using a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related acoustic effects such as head shadow.

      To verify that this did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (see supplied table, Summary Reverb Data). In the anechoic condition, the noise (positioned to the left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, raising noise level at the right ear to 55–56 dB depending on room, substantially lowering the ear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming this is not a normalization artefact but rather a genuine perceptual spatial release from masking, likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added the at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2, and updated the Discussion to clarify this mechanism.

      (32) Lines 755–759: Describing anechoic spaces as “rare” and as rooms whose “walls are treated” states the facts backwards. Open spaces (e.g., a grass lawn) are largely anechoic fields, and people spend considerable time in such environments. An anechoic room may be artificial, but an anechoic (or near-anechoic) space is very common, and the room is simply an attempt to simulate that within an enclosure.

      We appreciate this point and have rewritten this section to reflect it.

      (33) Lines 808–812: This sentence appears incorrect. It refers to “the ability to correctly report keywords spoken in environments with the more extreme—lower or higher—RIRs,” presumably meaning OPO and CP, but these are not the environments with extreme RIRs; or, if referring to An and HR, those were not “encountered in experimental blocks also containing the moderately reverberant Lecture Room or Living Room.”

      Thank you for noting this. We have rephrased this section to refer only to the extreme high-RIR environments encountered.

      (34) Line 843: Should Fig 3Ai be Fig 3A? Also, in that figure there are arrows between dlPFC and A1, and between A1 and (the cerebellum?)—it's unclear what these represent.

      The arrows were intended to represent feedforward and feedback information flow to lower auditory brain centres. We acknowledge they were confusing and have removed them from the figure.

      (35) Lines 929–949: The authors did not account for listeners tuning to voice and location (as now cited via Best et al.), and their own and Brandewie & Zahorik's data show improvement due to carrier phrase even in the anechoic case (with the inconsistency in Supplemental Table 4 noted earlier). A direct test—switching the environment between carrier phrase and target phrase, as in Brandewie & Zahorik and Vlahou et al.—would be needed to fully attribute the effect to reverberation rather than other factors.

      We thank the reviewer for this detailed comment and agree that directly manipulating the environment between carrier and target phrase would provide the most direct test of environment-specific adaptation. While our study did not implement this manipulation, our data provide converging evidence: (1) listeners showed improvement with longer carrier phrases even in the anechoic condition, consistent with previous reports, but this improvement did not interact with talker identity, carrier phrase length, or room, indicating it is not driven by tuning to specific voices or locations; and (2) regarding Supplemental Table 4, the calculations suggested by the reviewer do not match our statistical analysis—we have reported the SPSS-exported means directly (shown above) and reviewed this section for clarity, including a link to the raw supplemental data. Taken together, while we cannot fully quantify the proportion of adaptation attributable to reverberation versus other factors without the direct environment-switch manipulation, our results indicate that the observed improvements primarily relate to exposure to the environment rather than talker-specific effects.

      (36) Lines 951–953: It is unclear what about “understanding speech in background noise” distinguishes this study from previous studies of adaptation to reverberation, many of which also examined speech in noise (as reviewed in Tsironis et al., 2024). Rather than reviewing pertinent studies on adaptation to reverberation for speech tokens, the authors cite abstract noise-texture studies that are only partially relevant, given the prevalence of speech in everyday listening (lines 955–958).

      We appreciate the reviewer raising this point. Our intention was to refer specifically to statistical learning of implicit environmental acoustic features such as reverberation, rather than to speech-in-noise perception per se. We have revised the text accordingly: “A key feature of our study, which distinguishes it from previous investigations of statistical learning of acoustic features in human listeners, is the use of an ethologically valid listening task—understanding speech in background noise while listeners implicitly learn repeated acoustic features.”

      (37) Lines 978–980: In what way? For environments with large T60, a simpler explanation than “ecological validity” is that there is more late reverberant energy in the target acting as a masker, predictable from DRR.

      This section has been rewritten to clarify that it is the decline in performance at longer RT60 that is reminiscent of the decline observed under rTMS.

      (38) Line 981: What does “the better to understand speech in reverberant background noise” mean?

      This sentence has been revised.

      (39) Lines 987–988: When did “performance decline over the course of an experimental session”? Figures 2B, 3B, and 4A all show performance improving over the session.

      We have rephrased this sentence to refer to a decline in overall performance.

      (40) Lines 990–992: Many previous studies report better adaptation to reverberation for some rooms than others (e.g., Brandewie & Zahorik, 2010; Vlahou et al., 2021), but none have reported decreased performance for an anechoic space relative to a reverberant one. This anomaly should be explained and reconciled with the existing literature before invoking ecological explanations such as “ethologically relevant environments.”

      We thank the reviewer for raising this important point. We agree that verifying the at-ear acoustic levels is critical to ruling out a normalization artefact, particularly given our finding that reverberant performance exceeded anechoic performance. To address this directly: our stimuli were normalized in the 41-channel sound field, not at the listener’s ears. The source speech and noise were convolved with the 41-channel anechoic or reverberant impulse responses; the 41-channel noise field was scaled to a target of 70 dB and the 41-channel speech field scaled to the target SNR; the 41-channel signals were then rendered to two channels via a Higher Order Ambisonics-to-binaural decoder (hoa2bin), preserving natural head-related effects such as head shadow because normalization preceded binaural rendering.

      To verify that this sound-field normalization did not create an at-ear artefact, we extracted simulated at-ear RMS energy for speech and noise after hoa2bin rendering and mapped these to approximate dB SPL using the 70 dB sound-field anchor (Summary Reverb Data table). In the anechoic condition, the noise (positioned left) is strongly attenuated at the right ear by head shadow (dropping from ~57 dB to ~51 dB), giving the frontal target speech a highly favourable SNR at the better ear. In the reverberant condition, room reflections fill in the head shadow, increasing right-ear noise level to 55–56 dB depending on room, substantially worsening the atear SNR relative to the anechoic condition. The improved behavioural performance in reverberation therefore occurred despite a poorer acoustic SNR at the better ear, confirming the finding is not a normalization artefact but instead reflects a genuine perceptual spatial release from masking—likely driven by early reflections aiding target integration and late reverberation decorrelating the noise binaurally. We have added these at-ear acoustic details to the Methods (“Stimulus Normalization and Binaural Rendering”) and Table 2 to clarify this mechanism.

      (41) Discussion: Given the questions about the results, the discussion might need to be rewritten to only discuss claims that are actually supported.

      The Discussion section has indeed been extensively revised.

      (42) The hippocampus and other areas have been proposed for statistical learning, and studies also show that disruption of DLPFC can boost statistical learning (https://doi.org/10.1016/j.jml.2020.104144).

      We appreciate the reviewer raising this. We have cited Ambrus et al. (doi:10.1016/j.jml.2020.104144) in the Discussion (line 888) as evidence of opposing effects of dlPFC stimulation on statistical learning, and have further revised lines 891–903 of the Discussion to more clearly describe the known projections and functional interactions between dlPFC, hippocampus, striatum, and basal ganglia that support implicit and statistical learning.

      (43) Line 1739: What is “cumulative” here?

      “Cumulative” has been deleted.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study by Luden et al. seeks to elucidate the molecular functions of AHL15, a member of the AT-HOOK MOTIF NUCLEAR LOCALIZED (AHL) protein family, whose overexpression has been shown to extend plant longevity in Arabidopsis. To address this question, the authors conducted genome-wide ChIP-sequencing analyses to identify AHL15 binding sites. They further integrated these data with RNA-sequencing and ATAC-sequencing analyses to compare directly bound AHL15 targets with genes exhibiting altered expression and chromatin accessibility upon ectopic AHL15 overexpression.

      The analyses indicate that AHL15 preferentially associates with regions near transcription start sites (TSS) and transcription end sites (TES). Notably, no clear consensus DNA-binding motif was identified, suggesting that AHL15 binding may be mediated through interactions with other regulatory factors rather than through direct sequence recognition. The authors further show that AHL15 predominantly represses its direct target genes; however, this repression appears to be largely independent of detectable changes in chromatin accessibility.

      In addition to the AHL protein family, the globular H1 domain-containing high-mobility group A (GH1-HMGA) protein family also harbors AT-hook DNA-binding domains. Recent studies have shown that GH1-HMGA proteins repress FLC, a key regulator of flowering time, by interfering with gene-loop formation. The observed enrichment of AHL15 at both TSS and TES regions, therefore, raises the intriguing possibility that AHL15 may also participate in regulating gene-loop architecture. Consistent with this idea, the authors report that several direct AHL15 target genes are known to form gene loops.

      Overall, the conclusions of this study are well supported by the presented data and provide new mechanistic insights into how AHL family proteins may regulate gene expression.

      However, it is important to note that the genome-wide analyses in this study rely predominantly on ectopic overexpression of AHL15 at developmental stages when the gene is not usually expressed. Moreover, loss-of-function phenotypes for AHL15 have not been reported, leaving unresolved whether AHL15 plays a physiological role in regulating plant longevity under native conditions. It therefore remains possible that longevity control is mediated by other AHL family members rather than by AHL15 itself. In this regard, the manuscript's title would benefit from more accurately reflecting this broader implication.

      The ahl15 loss-of-function phenotype has previously been described in Karami et al., 2020 (Nat. Plants), Rahimi et al., 2022a (New Phyt.), and Rahimi et al., 2022b (Curr. Biol.), showing that ahl15 loss-of-function among others results in accelerated vegetative phase change and flowering, a reduced number of leaves produced by axillary meristems in short day grown plants and reduced secondary growth in the inflorescence stem. The dominant-negative ahl15 delta-G allele, expressing a mutant protein lacking the conserved G motif in the PPC domain, shows these phenotypes more clearly in the heterozygous ahl15 +/- background, and is embryo lethal in the homozygous ahl15 background (Karami et al., 2021, Nature Comm.). In addition, we recently show that leaf senescence is significantly accelerated in the ahl15 loss-of-function mutant (Luden et al., 2025, BioRxiv). These results show that AHL15 is involved in several aspects of ageing in Arabidopsis, and we have adjusted the introduction to discuss these previous findings more explicitlyWe agree with reviewer 1 on the possibility that multiple AHLs could have an effect on longevity, which is partially supported by the delayed flowering time observed in the AHL20, AHL27, or AHL29 overexpression lines (Karami et al., 2020, Street et al., 2008). However, the induction of the AHL15-GR fusion alone by DEX shows a clear delay of developmental phase transitions and the aging process in general, indicating that AHL15 by itself is able to extend longevity as other AHLs are not affected by DEX treatment (proven by the fact that their expression is not significantly changed in our RNA-seq analysis of DEX-treated 35S:AHL15-GR seedlings).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Luden et al. investigates the molecular function and DNA-binding modes of AHL15, a transcription factor with pleiotropic effects on plant development. The results contribute to our understanding of AHL15 function in development, specifically, and transcriptional regulation in plants, more broadly.

      Strengths:

      The authors developed a set of genetic tools for high-resolution profiling of AHL15 DNA binding and provided exploratory analyses of chromatin accessibility changes upon AHL15 overexpression. The generated data (CHiP-Seq, ATAC-Seq and RNA-Seq is a valuable resource for further studies. The data suggest that AHL15 does not operate as a pioneer TF, but is likely involved in gene looping.

      Weaknesses:

      While the overall message is conveyed clearly and convincingly, I see one major issue concerning motif discovery and interpretation. The authors state that because HOMER detected highly enriched motifs at frequencies below 1%, they conclude that "a true DNA binding motif would be present in a large portion of the AHL15 peaks (targets) and would be rare in other regions of the genome (background)."

      I agree that the frequency below 1% is unexpectedly low; however, this more likely reflects problems in data preprocessing or motif discovery rather than intrinsic biological properties of the transcriptional factor that possesses a DNA-binding domain and is known to bind AT_rich motifs. As it is, Figure 2 cannot serve as a main figure in the manuscript: it rather suggests that the generated CHiP-Seq peakset is dominated by noise (or motif discovery was done improperly) than that AHL15 binds nonspecifically.

      Since key methodological details on the HOMER workflow are missing in the M&M section, it is not possible to determine what went wrong. Looking at other results, i.e. the reasonably structured peak distribution around TSS/TTS and consistent overlap of the peaks between the replicas, I assume that the motif discovery step was done improperly.

      Therefore, I recommend redoing the motif analysis, for example, by restricting the search to the top-ranked peaks (e.g. TOP1000) and by using an appropriate background set (HOMER can generate good backgrounds, but it was not documented in the manuscript how the authors did it). If HOMER remains unsuccessful, the authors should consider complementary methods such as STREME or MEME, similar to the approach used for GH1-HMGA (https://pmc.ncbi.nlm.nih.gov/). If the peakset is of good quality, I would expect the analysis to identify an AT-rich motif with a frequency substantially higher than 1%-more likely in the range of at least 30%. If such a motif is detected, it should be reported clearly, ideally with positional enrichment information relative to TSS or TTS. It would also be informative to compare the recovered motif with known GH1-HMGA motifs.

      If de novo motif discovery remains inconclusive, the authors should, at a minimum, assess enrichment of known AHL binding motifs using available PWMs (e.g. from JASPAR). As it stands, the claim that "our ChIP-seq data show that AHL15 binds to AT-rich DNA throughout the Arabidopsis genome with limited sequence specificity (Figure 2A, Figure S2-S4)" is not convincingly supported.

      Another point concerns the authors' hypothesis regarding the role of AHL15 in gene looping. While I like this hypothesis and it is good to discuss it in the discussion section, the data presented are not sufficient to support the claim, stated in the abstract, that AHL15 "regulates 3D genome organization," as such a conclusion would require additional, dedicated experiments.

      The motifs discovered by HOMER are ranked by their enrichment over background, of which the highest-scoring motifs are very rare in the AHL15-bound targets, but even rarer in the background, which is why they score highly on the percent enrichment score. As expected by reviewer 2, we identified AT-rich motifs that were present in a larger percentage of AHL15 targets (found in 3-18% of targets, depending on the motif, see for example motif #5 in figure S4A), which can be seen at the right tail of the histograms shown in figures 2B-C and figures S2-S4 B-C. However, these motifs were also common in the background and were therefore not considered as significantly enriched in the AHL15-bound regions, with a target:background ratio of <2. As most of these motifs were flagged by HOMER as possible false-positives, and to limit the size of the (supplemental) figures, we did not show each of the motifs identified by HOMER in table form, but the full tables of de novo motifs identified by HOMER, including possible false-positive results are included in Additional file 3.

      Although the identification of AT-rich motifs shows that AHL15 (and very likely most other AHL proteins as well) binds AT-rich regions, it does not sufficiently explain the binding of AHL15 to its target genes, as these motifs are found at almost equal frequencies in non-AHL15-bound regions. In addition, a sequence found at this frequency in the genomic background is, in our view, too unspecific to be considered as a transcription factor binding site. Based on this, we concluded that AHL15 lacks a specific binding motif that can define the genes it binds.

      We have updated the methods section to include more details on the HOMER analysis and have also run the analysis in the top1000 shared peaks as suggested by reviewer 2 for both AHL15 and AHL29 ChIP-seq peaks, which showed that unlike in AHL29, a clear AT-rich motif cannot be found for AHL15 (Additional file 1: Figure S5).

      Reviewer #3 (Public review):

      Summary:

      This study investigated the role of AHL15 in the regulation of gene expression using AHL15 overexpression lines. Their results do show that more genes are downregulated when AHL15 is upregulated, and its binding does not affect the chromatin accessibility. Further, they investigated AHL15 binds in regions depleted in histone modifications and other epigenetic signatures. Subsequently, they investigated the presence of AHL15 in the gene chromatin loops. They found overlaps with both upregulated and downregulated genes. The methods are appropriately described, but could be improved to include the analysis of self-looping gene boundaries.

      Strengths:

      Their study clearly showed a lack of any specific sequence enrichment in the AHL15 binding sites, other than these being AT-rich, suggesting that AHL proteins do not recognize a specific DNA sequence but are recruited to their AT-rich target sites in another way. The study does suggest significant enrichment of AHL15 binding sites at TSS and TES, and AHL15 sites are depleted of any histone marks. They also identified that AHL15 binding sites overlap with self-looping gene boundaries.

      Weaknesses:

      The claim that AHL15 acts as a repressor and genes regulated by it are downregulated needs to be investigated based on AHL15 binding sites, to show enrichment/ depletion of AHL15 binding sites in overexpressing genes and repressed genes. The authors should provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title. Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression may be helpful to understand the significance of the study. Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful. It is not clear how the overlap of AHL15 peaks with self-looping genes has been carried out.

      A metagenome plot of AHL15 binding around genes that are differentially expressed upon DEX treatment can be found in Figure 3F. This analysis shows that AHL15 binding near differentially expressed genes is more pronounced compared to all AHL15-bound genes, and that AHL15 binding near the TSS is especially enriched for upregulated genes.

      As also suggested by reviewer 2, we ran a motif enrichment analysis on the differentially expressed genes that are bound by AHL15 to see if any motifs are enriched compared to the background and overrepresented in the AHL15-bound genes. Again, this did not reveal an AT-rich motif nor a motif that was conserved between up- and downregulated AHL15-bound genes (Additional file 1: Figure S6).

      Plant longevity in 35S:AHL15-GR Arabidopsis plants treated with DEX has been reported previously.. DEX treatment extended vegetative development after flowering resulting in polycarpy (Karami et al., 2020, Nature Plants), enhanced secondary growth resulting in woody stems (Rahimi et al., 2022, Current Biol.) and recently we showed that it delays leaf senescence in Arabidopsis (Luden et al., 2025, bioRxiv). All these observations have now been incorporated in the results section where the p35S::AHL15-GR plants are first presented. In addition, we show that 35S:AHL15-GR plants treated a single time with DEX at 10 days after germination show a significantly delayed flowering time in figure 4C-D of this manuscript.

      The enrichment of AHL15 ChIP-seq peaks in self-looping genes will be analyzed as suggested and compared to a random set of genes as a control, and the methods section will be updated to clarify how the analyses on self-looping genes were carried out.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors need to correct Line 28 on page 10:

      By comparing the AHL15 ChIP-seq data with 1766 previously identified self-looping genes -> By comparing the AHL15 ChIP-seq data with 1792 previously identified self-looping genes (Liu et al., 2016)

      Note: 1792 genes with self-loops were first reported by Liu et al. (2016)

      Liu, C., Wang, C.M., Wang, G., Becker, C., Zaidem, M., and Weigel, D. (2016). Genome-wide analysis of chromatin packing in at single-gene resolution. Genome Res 26, 1057-1068.

      These numbers will be corrected in the text.

      Reviewer #2 (Recommendations for the authors):

      (1) The newly generated datasets have been deposited only as raw sequencing reads. For reusability and reproducibility, the authors should also provide processed data accompanied by detailed metadata.

      We will upload the processed data and corresponding metadata to GEO.

      (2) Figures 4B and 5E are of low quality; can they be improved?

      We have submitted the original high-quality images to the publisher, which should resolve the issue.

      (3) Supplementary Table 5 lacks description (columns do not have names).

      This has been fixed.

      Reviewer #3 (Recommendations for the authors):

      Suggestions:

      (1) Motif enrichment in upregulated and downregulated genes separately to identify binding sequence preferences may be useful.

      This analysis will be performed and included in the revised manuscript as Additional file 1: Figure S6.

      (2) Investigate AHL15 binding sites to show enrichment/ depletion of AHL15 binding sites in overexpressing genes than repressed genes.

      This analysis has been done, please see figure 3F.

      (3) Provide data to support plant longevity with AHL15 overexpression using the DEX-induced system to support the claims in the title.

      For the effect of AHL15-GR induction by DEX on vegetative phase change and flowering time, please see figure 4C-D and Rahimi et al., (2022, New Phytologist). For other phenotypic changes induced by DEX treatment of 35S:AHL15-GR plants, please see Karami et al. (2020; Nature Plants), Rahimi et al., (2022, Current Biology) and Luden et al., (2025; BioRxiv). Text has been added to the results section where the 35S:AHL15-GR line is first introduced to refer to these previous publications.

      (4) Calculation of the enrichment score of AHL15 peaks in the self-looping genes that are upregulated or downregulated, and discussion about the different effects of AHL15 binding on self-looping regions to regulate gene expression, may be helpful to understand the significance of the study.

      This analysis has been done and included in the revised manuscript as Additional file 1: Table S1.

      (5) Describe how the overlap of AHL15 peaks with self-looping genes has been carried out.

      The methods section has been updated with detailed information on this analysis.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      The authors initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LTG products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another. The authors have responded well to most points of the previous review.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted in the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. The conclusions concerning the PG degradation during sporulation needs to be clarified, as described below. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation". This needs to be reflected also in title of the paper.

      We have addressed the reviewer’s concerns and updated the text. 

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 100-125: I am still concerned about clarity in the description of the muropeptides from spores. The claim is that 90% of the recovered muropeptides are anhydro LTG products. If LTGs had truly cleaved so many of the glycosidic bonds such that 90% of the muramic acid was now in the anhydro form, then all of the PG in the spores would be in small fragments (dimers and trimers containing 4-6 sugars and 8-12 amino acids), and these would be soluble and lost during purification. Whatever the form of the PG, it must not be easily soluble, so it must be either larger than that or bound to something else. The fact that the muropeptides are solubilized by muramidase indicates that there are some NAM-NAG bonds remaining, and cleavage of these should release some non-anhydro muropeptides. The spore muropeptide chromatograms have a large, late "mound" of UV-absorbing material that was released by muramidase digestion, and two of the identified anhydro products are present in this mound. It is not clear how these two anhydro products were quantified within this mound and what other (presumably) muropeptide species might be present in this mound. Some explanation of how these two muropeptides were quantified is needed.

      We thank the reviewer for this careful point. As the reviewer notes, our purification procedure recovers only sedimentable PG, and any PG fragments solubilized by LTG activity would be lost during the washing steps and therefore not represented in our analysis. We have now explicitly acknowledged this limitation and discussed its implications in the revised text:

      “Consequently, any PG fragments solubilized by LTG activity during sporulation would be lost at this stage, and the muropeptides we detect derive from the more highly crosslinked material that survives the procedure.

      “Despite the overall decline in identified muropeptides, anhydro-muropeptides a minor component of vegetative PG were enriched in both spore types (Figure 1).”

      I feel that the authors need to address some of this uncertainty in the results and discussion. They might be able to say that anhydro-muropeptides represent 90% of the "identified muropeptides" but need to acknowledge that there is a great deal of unidentified material released by the muramidase digestion. The data might also indicate that the LTG activity solubilizes much of the PG, which is lost, and results in recovery of only highly cross-linked muropeptides that might survive the PG purification process.

      We agree that two of the identified anhydro species elute within a broad, unresolved region of the chromatogram that accounts for a substantial fraction of the muramidase-released material. Because reliable assignment and quantification of individual components within this region is not possible, we concluded that expressing anhydro-muropeptides as a fraction of the total identified muropeptides could be misleading. We have therefore removed the quantitative statement and now describe the enrichment of anhydromuropeptides in spores relative to vegetative cells qualitatively:

      “The identified muropeptides, however, represent only a part of the material released by muramidase from spores: a substantial, as a late-eluting portion of the chromatogram could not be assigned, and its composition remains unknown.”

      This would not eliminate the conclusion that "their abundance in spores indicates that certain LTGs must play essential roles (perhaps change to "might play important roles") in M. xanthus sporulation", which leads to the remaining studies in the paper.

      Following the reviewer's suggestion, we have softened the conclusion of this section to state that LTGs "may play important roles" in sporulation.

      (2) The authors have changed the statement about a "new mechanism of sporulation" at the beginning of the discussion, but this language is still in the paper title. Something more like "Programmed peptidoglycan degradation plays an important role in Myxococccus sporulation"

      Following the reviewer's suggestion, we changed the title to “A novel mechanism for morphological change during bacterial sporulation based on programmed peptidoglycan degradation”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The experiments rely on three groups: CS-only WT, CTA WT, and CTA KO. Can the authors provide a rationale for not having a CS-only KO group?

      We did not include the CS-only (KO) group because longitudinal in vivo recordings in behaving animals are technically demanding, and our primary goal was to follow and compare the dynamics across learning in WT and Shank3 KO mice. However, we recorded an additional habituation water session one day before CST1, which allows us to address (1) the decoding performance in naïve KO animals, and (2) whether higher correlated noise is already present in KO animals before learning. To this end, we first trained and tested the classifier on cross-registered data from the habituation (HAB) water session and the CST1 saccharin session across the CSonly (WT), CTA (WT), and CTA (KO) groups. Because this comparison is made across sessions, it is not exactly comparable to discrimination within a session, but it does show that GC responses in naïve KO animals can discriminate water from saccharin at baseline. These data are now included in Figure 6 – figure supplement 2 and are described in lines 329-338.

      Interestingly, we also observed a significant reduction in the amplitude of suppressed responses during the habituation water session in KO compared to WT animals, as well as a trend toward higher stimulus-evoked neuronal coactivity (data included in Figure 2 – figure supplement 2). This suggests that reduced suppression is already present in GC of KO animals before learning, potentially contributing to slower CTA acquisition (now described in lines 209-215).

      (2) The authors design an effective behavioral paradigm comparing consumption of water and saccharin and tracking extinction (Figure 3). This paradigm shows differences in licking across distinct behavioral conditions. For instance, during T1, licking to water strongly differs from licking to saccharin for both WT and KO. During T2, licking to water strongly differs from licking to saccharin only for WT (much less for KO), and licking to saccharin in WT differs from that in KO. These differences in taste sampling across conditions could contribute to some of the effects on neural activity and discriminability reported in Figures 5 and 6. That is, sucrose and water trials may be highly discriminable because in one case the mouse licks and in the other it does not (or licks much less). The author may want to address this issue.

      This is an important point. As noted by the reviewer, active licking can modulate neuronal activity in the gustatory cortex independent of taste identity (Neese et al., 2022). Because our paradigm required animals to voluntarily sample tastants of different valences, motivated differences in licking are inherently tied to taste value, making it difficult to fully disentangle taste-evoked responses from lick-related activity.

      However, taste exposure is known to induce prolonged neural responses that persist beyond the sampling phase (Juen et al., 2024). In our recordings, we included a 10second post-sampling epoch. We thus trained and tested classifiers using calcium traces during this post-sampling period; in particular, we divided the 10-second duration into five 2-second bins, matching the length of the tastant delivery phase, and analyzed the classifier built within each bin. We found that in both WT and KO animals, decoding performance was consistently above chance throughout the postdelivery period (now added to Figure 6 – figure supplement 1, lines 321-329), suggesting that decoding accuracies during sampling likely reflect taste rather than licking.

      (3) Are there any omission trials following CTA? If so, they should be quantified and reported. How are the omission trials treated with regard to the analyses?

      On the day following each CST session, animals underwent a water-only session (i.e., saccharin was omitted) to minimize context–malaise association. During these sessions, animals resumed licking both in the total lick counts and in the number of trials they engaged in, to levels comparable to pre-conditioning behavior. We did not observe significant differences between the WT and KO groups during these omission sessions. This point has been mentioned in the Methods section of the revised manuscript (lines 544-547).

      (4) The authors describe the extinction paradigm as "alternative choice". In decision-making, alternative choice paradigms typically require 2 lateral spouts to report decisions following the sampling from a central spout. To avoid confusion, the authors may want to define their paradigm as alternative sampling.

      We have revised this terminology to “alternative sampling” to avoid confusion with the classical alternative-choice paradigms.

      (5) Figure 4 reports that CTA increases the proportion of neurons that consistently respond to saccharin and water across days. While the saccharin result could be an effect of aversive learning, it is less clear why the phenomenon would generalize to water as well. Can the authors provide an explanation?

      Water and saccharin activated an overlapping population of neurons in GC. When we further quantified their tuning properties in the lifetime plots (Figure 4), we found that neurons responsive to both stimuli showed the most stable responsiveness across days, compared to neurons that responded only to saccharin or only to water (Author response image 1). Because the water-responsive and saccharin-responsive groups in Figure 4 both include this subset of dual-responsive neurons, this likely explains why both plots show increased reliability. This effect on reliability of single-cell responses is thus distinct from changes in the ability to discriminate between tastants at the population level (Fig. 6).

      Author response image 1.

      GC neurons responding to both water and saccharin are more stable during CTA extinction. Lifetime plot showing significant responses of the same neurons responding to only water (blue), only saccharin (magenta), and to both saccharin and water (gold) across test sessions (T1-5) in the CTA (WT) group.

      (6) The recordings are performed in the part of the anterior insular cortex that is typically defined as "gustatory cortex" (GC). Given the functional heterogeneity of the anterior insular cortex (AIC) and given that the authors do not sample all of the anteroposterior extent of AIC, I would suggest being more explicit about their positioning in GC. Also, some citations (e.g., Gogolla et al, 2014) refer to the posterior insular cortex, which is considered more inherently multimodal than GC. GC multimodality is typically associative in nature, as only a few neurons respond to sound and light in naïve animals.

      Our stereotaxic coordinates targeted the conventional gustatory region within AIC (see revised manuscript Methods section, lines 489-490). We have revised the terminology throughout the manuscript to more explicitly reflect this anatomical positioning.

      (7) It would be useful to add summary figures showing the extent of viral spread as well as GRIN lens placement.

      Revised Manuscript Figure 1B shows a representative example of confirmed GRIN lens placement and the viral spread of GCaMP. In most cases, GCaMP expression is confined to GC, with minimal spread to the piriform cortex and along the injection track. Depth and GCaMP expression in GC were further validated during two-photon imaging.

      (8) I encourage the authors to add Ns every time percentages are reported. How many neurons have been recorded in each condition? Can the authors provide the average number of neurons recorded per session and per animal?

      We now included these numbers in the revised manuscript (lines 157-158, 162-163, 253, 268, 271-272).

      (9) It looks like some animals learned more than others (Figure 1E or Figure 3C). Is it possible to compare neural activity across animals that showed different degrees of learning?

      We thank the reviewer for this suggestion – we now show a significant correlation between the magnitude of CTA and the coactivity metric in Figure 1 Figure supplement 3; we elaborate in our Response to Reviewer #3 Public Review 1.

      Reviewer #2 (Public review):

      (1) Causality: The paper infers that increased correlated variability causes learning deficits, but no causal tests (e.g., optogenetic modulation of inhibition or interneuron rescue) are presented to confirm this.

      Although we now provide data showing that correlated variability prior to learning is significantly correlated with the magnitude of CTA (see Response to Reviewer #1 Public Review 1above), we agree that we cannot infer causality without additional manipulations. While it might be possible to manipulate correlated variability by targeting inhibition within GC, optogenetic and chemogenetic manipulations of inhibition are likely to impact behavior through multiple mechanisms; for example, enhancing PV-interneuron activity in visual cortex profoundly impairs vision-dependent learning (Bissen et al. 2026, Leman et al. 2025). Thus, testing this would require finding a paradigm that specifically restores synchronization to WT levels without over-inhibiting the network, which is beyond the scope of the current study. We have rewritten the manuscript throughout to remove the inference of causality, and instead describe these two findings as being “associated” (see e.g. lines 95, 219-220, 379-382).

      (2) Behavioural scope: The study focuses exclusively on taste aversion; generalisation to other flexible learning paradigms (e.g., reversal or probabilistic tasks) is not addressed.

      Our study is focused on conditioned taste aversion (CTA) acquisition and extinction, which provides a well-established model for examining the formation and updating of aversive associative memories. We agree that cognitive flexibility encompasses a broad range of behavioral paradigms, and in the revised manuscript have sought to confined our conclusions to CTA. Whether the mechanisms identified here extend to other forms of flexible learning, such as reversal or probabilistic learning, or even to other sensory-stimulus-guided behavior, will require future investigation.

      (3) Mechanistic insights: While providing interesting findings of altered sensory perception and extinction of learning-related signals in AIC, it offered nearly no mechanistic insights. This makes the interpretation, especially on how generalisable these findings are, difficult. Also, different reported findings are "potentially" connected, but the exact relation between increased correlated variability and faster loss of taste selectivity cannot be assessed.

      In a new analysis we find that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3). This new piece of data (now added to revised manuscript lines 215218) provides a link between these two findings, and suggests that baseline coactivity levels in GC can influence the speed of CTA learning. We agree that this association does not imply causation, and have taken pains to avoid stating this.

      Reviewer #3 (Public review):

      (1) The authors don't make a causal link between the behaviour and AIC neurophysiology, both the percentage of suppressed cells and the coactivity measurements. For the % of suppressed cells, it seems that both WT and KO cells are suppressed in the transition between CST1 and CST2 (Figure 1L), yet only the WT mice exhibit CTA (at least by CST2). For the taste-elicited coactivity measure, it seems that there is an increase in coactivity from CST1 to CST2 in WT (Figure 2C - blue, although not statistically tested?), but persistently higher coactivity in KO. Is this change of coactivity in WT important for the expression of CTA? Plotting behavioral performance (from Figure 1G) against coactivity (from Figure 2C) for each animal would be informative.

      This is a good suggestion (also made by the other reviewers), and we now show that the coactivity level during CST1 showed a significant, positive correlation with the lick ratio (CST2/CST1); that is, the higher the initial coactivity during learning, the slower the CTA acquisition is (Figure 2 – figure supplement 3).

      (2) Shank3 KO cells already show an increase in baseline coactivity (Figure 2- figure supplement 1), and the authors never examine CS-only responses in the KO group, therefore making it difficult to determine whether elevated coactivity and noise correlations reflect a generalized AIC abnormality in Shank3 KOs (perhaps through impaired PV-mediated inhibition in insular cortex - Gogolla et al, 2014) that is not directly responsible/related to CTA?

      We agree that the increased coactivity in KO animals likely reflects a general network defect in AIC before CTA learning. This baseline change is unrelated to the taste stimuli, because (1) the coactivity is already elevated before the animals receive the taste stimuli (that is, a baseline abnormality), (2) the neuronal responses to taste delivery during CST1 is indicative of CS-only responses, as it occurs before the injection of LiCl, when the associative learning process is initiated, and (3) a trend toward higher coactivity is already present in the habituation water session before CST1 (Figure 2 and Figure 2 – figure supplement 2).

      We have clarified our description of these findings to avoid claiming that the increased coactivity “causes” poor learning performance (lines 379-382).

      (3) How do the authors interpret the large range of lick ratios (Figure 1G) for WT (almost bi-modal distribution)? Is there a within-subject correlation with any of the neurophysiological measurements to suggest a relationship between AIC neurophysiology and behavioural expression of CTA?

      See response to Point 1 above.

      (4) Indeed, CTA appears to be successfully achieved for Shank3 KO mice delayed by 1 day, as the level of saccharin aversion during the first retrieval session (T1) is comparable between Shank3 KO and WTs. In this context, not extending the first part of the paradigm to include CST3 seems to be a missed opportunity. Doing so would have allowed for within-cell and within-subject comparison of taste-elicited pairwise correlation across the learning and to investigate the neural mechanism of delayed extinction in KOs more effectively.

      We did not include a third CST session because when we analyzed the lick counts, KO animals already formed robust CTA after CST2 that was indistinguishable from WT animals. This suggests that the faster loss of CTA memory during extinction is due to a faster extinction process, rather than a weaker CTA memory from the outset. Adding a third CST could potentially lead to a memory that is harder to extinguish. Whether Shank3 KO mice would exhibit faster loss of memory in this scenario is an open question that would be interesting to explore in a future study.

      (5) How to interpret Figure 5F: Absolute discriminability is lower for T5 for CTA WT and CTA KO compared to CS-only? Why would AIC neurons have less information on taste identity by the end of extinction than during the unconditioned (CS-only) condition? And if that is the case, how is decoding accuracy in Figure 6C higher in T5 for CTA WT vs CS-only?

      We appreciate the reviewer's confusion about the discrepancy between our single-cell and population-level discriminability results. We speculate that in the CS-only state, individual AIC neuronal responses mostly reflect taste identity. However, after learning (in the CTA group), these neurons develop “mixed selectivity” (Tye et al., 2024), encoding not only identity but also the learned valence and extinction history. The lower single-cell discriminability after extinction (T5) in Figure 5F suggests that, although taste identity may remain constant, the learned history (e.g., "this taste used to be dangerous, but now it's safe") has shifted. This mixing of information makes each cell a weaker discriminator on its own.

      However, the higher population decoding accuracy in Figure 6C demonstrates that the entire population of neurons can work together more effectively. The learning process could reorganize the neural ensemble in such ways that our support vector classifier (SVC) is able to identify and combine the relevant signals within the population, even when the valence of taste stimuli has changed, to better decode stimuli and outperform the non-learned state. This suggests that the brain shifted to a more robust, population-based coding strategy for complex, learned information, which is resistant to changes in selectivity at the single-cell level. The finding that population coding is robust to single-neuron variability has also been reported in other cortical regions (Montijin et al., 2016).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Mechanistic experiments: Consider inhibitory neuron-specific imaging or manipulation (e.g., optogenetic enhancement of interneuron activity) to test whether restoring inhibition rescues learning flexibility.

      We have addressed the limitation and potential issues for manipulating cortical inhibition in Response to Reviewer #2 Public review 1.

      (2) Clarify limitations: Explicitly acknowledge the correlational nature of neuralbehavioural relationships in the Discussion.

      We have removed language that implies a causal relationship throughout, and have emphasized the correlational nature of our findings in the Discussion section of our revised manuscript (lines 379-382)

      (3) Enhance clarity: Simplify some dense methodological sections and expand figure legends to guide interdisciplinary readers.

      We have adjusted the Methods section and figure legends as needed for better readability.

      Individual Comments for Authors:

      (1) L83-90: Confusingly written, not easy to understand for someone not knowing the paradigm in detail.

      - What are the different stages? Memory encoding? Leaning? Extinction

      - More reliable in taste responsiveness - what does that mean?

      We have emphasized the behavior stages where each finding was observed in the revised manuscript (lines 82-91)

      (2) L112: Not sure if these references support the "crucial", since they do not seem to be causal.

      We have reworded this for accuracy (line 111).

      (3) Figure 1: panels h and i in the heat maps, it looks like that in the KO animal, activity is more suppressed from CST1 to 2?

      Panels j, l, m, and Figure 2: Neuronal suppression is already higher in CST1; therefore, there is no CTA effect but a general "perceptual" issue in the Shank3 model. The only effect seems to be a potential reduction in activation in CST2 in KO animals.

      This point has been discussed in Reviewer #3 Public Review 2.

      (4) Clarify in text. Especially with the sentence in the next paragraph, it might be confusing: "We wondered what other features of AIC activity during CTA acquisition might differ between WT and Shank3 KO mice."

      We have rewritten this in the revised manuscript (lines 169-170).

      (5) Clarify which are CTA-dependent and which are general (e.g., if writing suppression during CTA acquisition, it implies that it is related. But these changes were present before CTA.

      We have clarified this in the revised manuscript (lines 209-215).

      (6) Figure 4: Mainly shows a CTA-related increase in reliability in their taste responsiveness. This is not addressed anywhere else in the document and is not taken up in the discussion. How could it be related to the other findings, and what is its relevance? Please elaborate (e.g., in the discussion) or potentially remove?

      We measured response reliability, as stabilization of stimulus-evoked responses has been reported in other sensory cortices across different learning tasks. Yet, it remained unclear whether CTA learning would induce similar changes in AIC. We took advantage of our longitudinal recording to address this question and believe that this piece of evidence will contribute to the research community that studies taste and learning in general. In addition, what is striking to us is that while the taste selectivity is degraded faster in KO animals, their response reliability is largely preserved. This suggests that these two sensory stimulus-related neuronal properties may involve distinct cellular and/or circuit mechanisms.

      (7) Figure 5: Problematic to compare T5 between both groups, since T5 is lower than T4 in WT (against the trend) and T4 is an outlier in KO. e.g., if compared at T5, completely different results? Or why is there significance between T1 and T2 but not between T1 and T4 in KO? Could the authors address this point?

      In Figure 5B, the slightly lower average for WT animals at T5 was driven by a single outlier, and there was no statistically significant difference between T4 and T5 (corrected post hoc t-test, WT, T4 vs. T5, p = 0.4097). Therefore, it does not contradict the trend toward an overall increase in nonselective neurons during CTA extinction. For KO animals, the lower average at T4 than T5 (corrected post hoc ttest, KO, T4 vs. T5, p = 0.0082) was intriguing, and one possible explanation is that neurons in the KO group might undergo more dynamic and variable changes in their responsiveness during CTA extinction, fluctuating before finally stabilizing.

      Comparing T5 instead of T4 thus ensures that neuronal responsiveness is stabilized and reflects an “extinct” CTA memory more truly.

      General Comments:

      (1) While changes in SNR were observed in Shank3 models, the mechanism underlying decreased correlated variability has not been reported to date. Since decreased variability is usually associated with improved SNR ratio, it might be worth highlighting the distinction between "signal" and "noise" as separated in your analyses to make it more understandable for the reader.

      We have described in the Results section what signal and noise correlations indicated and how they were separated in our analyses in both the Results and Methods section of the revised manuscript (lines 185-193, lines 673-681).

      (2) What is the origin of the increased correlated variability?

      We have discussed that reduced cortical feedback inhibition could be a potential source of increased correlated variability in the Discussion section of our revised manuscript (lines 370-375).

      (3) Is the variability generally increased between trials (bigger fluctuations between trials for each neuron), or is the variability of each neuron similar, but they are just more correlated (more synced)?

      Our pilot analysis did not detect any evident changes in the response variability for each neuron across trials; thus, we think that in KO animals, neuronal responsivity becomes more correlated and synchronized.

      Reviewer #3 (Recommendations for the authors):

      (1) Point in line 422-424: Rephrase the closing statement of the discussion as you have shown that mutant mice are actually able to update their behaviour (in fact faster) when the valence of the sensory input changes.

      The “reduced ability to update behavior when the valence of a sensory input changes” refers to the finding that KO animals learned CTA more slowly; i.e., they were unable to timely adjust their behavior after malaise. We have rephrased this for clarity (line 448-449)

      (2) The Figure 6 legend does not correspond to panels D and E in the figure. Νο I, J in figure.

      We have fixed this mismatch in the revised manuscript.

      Minor concerns:

      (1) Cue/lick/taste-responding neurons greatly overlap and are not exclusively selective (Figure 1- figure supplement 2). Is there a genotype difference for the % of selective neurons (i.e., ones that only respond during cut/lick/taste) or the % of overlap?

      When we quantified the stimulus responsivity in KO animals, we also identified neurons that were activated by cues, licks, or tastes. Their respective percentages and overlap did not differ significantly from those in the WT group, indicating that the modality of KO neurons across different sensorimotor cues is not compromised in the KO condition (Author response image 2).

      Author response image 2.

      Neurons in WT and Shank3 KO animals show comparable responsiveness to sensorimotor stimuli during conditioning. (A) Percentage of neurons activated by the cue (left), lick movement (middle), and the tastant (right) in the CTA (KO) group (B) during the first conditioning session (CST1). (B) Venn diagram showing the overlaps among cue-, lick-, and tastant responsive neurons in (Figure 1 - figure supplement 2 C) and (A).

      (2) For Figure 1: The authors could also express consumption as a % of consumed (trial-averaged licks) over the number of trials. It is mentioned that mice undergo daily training sessions consisting of 'approximately 30 trials' (line 114). This can give an indication of how strong the learning is between cta1 and cta2 and how strong the genotype difference is.

      We are not sure if dividing trial-averaged licks over the number of trials would provide additional information, as the trial-average lick is already normalized to the number of trials.

      (3) Figure 4: Why is there a different number of neurons in C vs G?

      The figures B, C, D showed neurons that were activated by saccharin, and the figures F, G, H showed neurons that were activated by water. In all experimental groups, the numbers of neurons responsive to saccharin and water were different (i.e., B vs. F, C vs. G, D vs. H). The exact numbers were included in the corresponding figure legends in the revised manuscript.

      (4) Figure 5B: The grey background box is moved to the left.

      We kept the current figure format, as it effectively presents the mean, fitted mean, error bars, and individual animal data.

      (5) In line 142: (1-2), (2-3), (3-4), the numbers in parentheses are confusing.

      We have relabeled this as epoch 1-2, epoch 2-3, and epoch 3-4 in both text and figures for clarity (lines 145-146).

      (6) Line 188: Do the authors mean noise correlations?

      Rosenbaum et al. and Khoury et al. indeed measured correlated variability (noise correlation) in their study. On the other hand, Rothschild et al. did not specifically separate the noise from signal activities, which more likely reflect the coactivity measured in our case. We have rewritten this for accuracy (line 196).

      (7) Where mentioning in the CS-only group, please explicitly state the CS-only WT group.

      We have relabeled this throughout our revised manuscript.

      (8) In lines 273-274: if the comparison is the reduction in discriminability being faster for the KO animals that had CTA, the correct comparison should be CSonly KO vs CTA KO.

      We think that the better comparison to test how fast taste discriminability is reduced would be to perform post-hoc tests comparing T1 vs T2 within genotypes. We did not see significant changes between T1 and T2 in either genotype, which was reported in the figure legends of the reviewed preprint (lines 1140-1141).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study that used 7T diffusion MRI in subjects from a Human Connectome Project dataset to characterize the zona incerta, an area of gray matter whose involvement has been demonstrated in a broad range of behavioral and physiologic functions. The authors employ tractography to model white matter tracts that involve connections with the ZI and use clustering techniques to segment the ZI into distinct subregions based on similar patterns of connectivity. The authors report a rostral-caudal organization of the ZI's streamlines where rostrally-projecting tracts are rostrally-positioned in the ZI and caudally-projecting tracts are caudally-positioned in the ZI.

      Strengths:

      The paper presents robust findings that demonstrate subregions of the human ZI that appear to be structurally distinct using a combination of spectral clustering and diffusion map embedding methods. The results of this work can contribute to our understanding of the anatomy and structural connectivity of the ZI, allowing us to further explore its role as a neuromodulatory target for various neurological disorders.

      Weaknesses:

      There should be further discussion of the clustering methods employed and why they are appropriate for the pertinent data. Additionally, the limitations of analyzing solely the cortical connections of the zona incerta should be addressed, as anatomical studies of the ZI have shown significant involvement of the ZI in tracts projecting to deep brain regions.

      We are grateful to the reviewer for recognizing the strengths of our study, as well as for providing constructive suggestions to further strengthen the manuscript.

      In response to the reviewer’s feedback, we have expanded our discussion of the clustering methods employed, including the rationale for using spectral clustering in combination with diffusion map embedding, and clarified why this approach is well-suited to connectivity-based parcellation of the ZI.

      Additionally, we have expanded the Discussion to address the limitations of focusing exclusively on cortical connections. As the reviewer correctly notes, anatomical studies have demonstrated that the ZI has extensive connections with deep brain regions, and our approach therefore represents only a partial view of its connectivity. We have previously demonstrated the feasibility of reconstructing subcortical pathways using in vivo diffusion MRI (Kai et al., NeuroImage, 2022), providing a foundation for extending the present framework beyond cortical connectivity. However, as iterated below in our specific response to reviewer 1, we believe this deserves a separate thorough investigation. Nonetheless, we now explicitly discuss this limitation in the revised Discussion and outline directions for future work incorporating subcortical connectivity analyses.

      Reviewer #2 (Public review):

      Summary:

      Haast et al. investigated the organization of the zona incerta (ZI) in the human brain based on its structural connectivity to the neocortex. They found that the ZI is organized according to a primary rostro-caudal gradient, where the rostral ZI is more strongly connected to the prefrontal cortex and the caudal ZI to the sensorimotor cortex. They also found that the central region of the ZI is differently connected to the neocortex compared with the rostral and caudal regions, and could be important as a deep brain stimulation target for the treatment of essential tremors.

      Strengths:

      I think the overall quality of this work is great, and the results are presented in a very clear and organized manner. I particularly appreciate the effort that the authors put into validating the results using 7T and 3T data, as well as test-retest data.

      Weaknesses:

      That being said, I was left with a couple of concerns after reading the paper.

      - Although the authors discussed animal evidence for a dorsal-ventral organization of the ZI, I thought that the evidence they presented for it in this paper was not so convincing. In Figure S5, the second gradient (G2) shows a clear dorsoventral pattern, but this pattern seems to primarily separate the ZI and H fields rather than show an internal topology of the ZI. This is more likely the case given that there are two bands (superior and inferior) of high G2 values surrounding a single band (middle) of low G2 values. The evidence for the rostrocaudal gradient, on the other hand, is quite convincing.

      - HCP data is still too advanced for clinical translation. Although 3T is becoming more and more prevalent for presurgical planning, the HCP 3T dataset is acquired with a voxel size of 1.25mm, which is a far higher resolution than the typical clinical scan. It would be very useful for clinical readers to see what individual subject replicability looks like if the data were acquired at the more typical voxel size of 2mm. This could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for their positive evaluation of our work and for highlighting the clarity of the results, and our validation efforts across 7T, 3T, and test-retest datasets.

      Regarding the reviewer’s concern about the evidence for a dorsal-ventral organization, we agree that the rostro-caudal gradient is more prominent and convincing in our data, while the dorsal-ventral pattern is less robust. As the reviewer points out, the second gradient (G2) in Figure S5 may primarily reflect differences between the ZI and surrounding H fields, rather than a clear internal subdivision within the ZI itself. We have revised the Discussion to clarify this interpretation, emphasizing that our evidence for a dorsal-ventral organization is more tentative and requires further validation, particularly in light of prior animal literature.

      We also appreciate the reviewer’s important point regarding clinical translation. Indeed, the HCP datasets, both at 7T and 3T, use acquisition parameters (e.g., 1.25 mm voxel size at 3T) that exceed those of typical clinical scans. We therefore assessed the replicability of our findings in data acquired at more clinically representative resolutions (i.e., 2 mm voxel size at 3T). Details concerning this analysis are outlined in our response to reviewer 2 below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) When using "spectral clustering" to segment the ZI per its structural connectivity, it is unclear how k=6 clusters were chosen. Moreover based on previous reports of rodent ZI cytoarchitecture into rostral, dorsal, ventral, and caudal regions there is disagreement with the topographic organization of six clusters presented in this manuscript. Moreover, "diffusion map embedding" was not described or cited. This technique of dimensionality reduction and how it was applied to the data should be specifically described.

      We thank the reviewer for this thoughtful comment. We agree that selecting the optimal number of clusters in data-driven approaches such as spectral clustering is inherently challenging in the absence of a definitive ground truth. To address this, we computed alternative cluster solutions across a range of k values (k=2-8), which are presented in Supplementary Figure 2B. Our decision to focus on k=6 was guided by prior cytoarchitectonic descriptions of the rodent ZI by Romanowski et al., 1985, who delineated six distinct sectors (pars rostropolaris, pars dorsalis, pars ventralis, pars magnocellularis, pars retropolaris, and pars caudalis). Thus, our approach aligns the data-driven clustering with established anatomical subdivisions. Additionally, we found that k=6 provided a meaningful level of granularity for probing location-dependent neuromodulatory effects within the ZI, as discussed in the revised Discussion section (‘Discrete subregions of the zona incerta using spectral clustering’, second paragraph).

      Finally, we acknowledge the lack of clarity concerning “diffusion map embedding” in the original submission. We have now more explicitly mentioned diffusion map embedding in the “Connectivity gradients” paragraph in the Methods section, which includes the relevant references as well as description on how it was applied to the data.

      (2) "Strong correlations with cognitive terms and cortical hierarchies were particularly evident for clusters 3-5 (Figure 7a-b), which have the highest number of connecting streamlines (Figure 3c). Cluster 5, located near the central sulcus, is significantly linked to movement related cognitive processing (Pearson's r = 0.623, p < 0.005) and CogPC1 (Pearson's r = 0.593, p < 0.005). Cluster 4, situated more anteriorly, overlaps with regions involved in working memory (Pearson's r = 0.494, p < 0.001) with a high level of expression of serotonin (5-HT1B) receptors (Pearson's r = 0.413, p < 0.005) (Beliveau et al., 2017). Cluster 3, located further towards the frontal pole, is associated with mood (Pearson's r = 0.473, p < 0.005) and impulsivity (Pearson's r = 0.473, p < 0.001), and the sensorimotor association axis (Pearson's r = 0.583, p < 0.005) (Sydnor et al., 2021). Clusters 1, 2, and 6, characterized by the least number of connecting streamlines (Figure 3c), were relatively weakly associated (i.e., low Spearman's coefficient and/or within the spatial autocorrelation range) with cognitive terms or cortical hierarchies."

      The validity of identifying correlations with the spectral clusters and functional connectivity in tasks related to "keywords", and reporting similarities between these data and the regions most strongly connected via tractography, is a bit questionable. These results are based on studies that may not report all relevant findings. There is a bias towards regions that are more commonly studied with fMRI. Moreover, claims can be made about assigning functions specific to a brain region for almost any structure/function relationship (with some exceptions).

      We agree that correlations between connectivity-defined ZI clusters and functional annotations derived from external datasets should be interpreted with caution, given potential biases in the available literature (e.g., overrepresentation of well-studied cortical regions in fMRI meta-analyses) and the inherent risk of over-assigning functions to structural subdivisions. Our intention was not to make definitive claims about the functional specialization of individual ZI subregions, but rather to provide an exploratory framework for situating the ZI within broader cortical hierarchies and functional domains. These analyses are intended to generate hypotheses and to offer preliminary insight into how connectivity-based subdivisions of the ZI may relate to cognition and behavior.

      We agree with the reviewer that future studies should be specifically designed to address these questions more directly, for example, by combining connectivity-informed parcellations of the ZI with task-based or resting-state fMRI in the same subjects. Such targeted approaches will be necessary to rigorously establish the integration of the ZI within the brain’s functional organization.

      (3) The ZI's connections to many subcortical structures have also been reported in rodents and non-human primates. Moreover, the authors describe the efficacy of DBS of the caudal ZI in alleviating symptoms in patients with essential tremor, which indicates modulation of the dentato-rubro-thalamic tract fibers that project to subcortical structures such as the VIM thalamus, red nucleus, and cerebellum. The atlas in the study was characterized per the ZI's cortical connections only. These concerns should be addressed in the discussion.

      We agree with the reviewer that incorporating subcortical connections is essential for a comprehensive understanding of ZI connectivity. Building on our prior work demonstrating the feasibility of subcortical tractography (Kai et al., NeuroImage, 2022), we propose a systematic investigation of in vivo subcortico-incertal tractography as a critical next step. We believe that a dedicated investigation is required to systematically evaluate subcorticoincertal tractography, optimize reconstruction of key pathways (e.g., the dentato-rubrothalamic tract), and determine how these subcortical connections contribute to the topographic organization of the human ZI. We now explicitly discuss these considerations and identify them as an important direction for future work. Such (currently ongoing) work will not only clarify how subcortical inputs shape the topography of the ZI, but will also enable targeted optimization of tractography parameters to maximize the reliable reconstruction of specific but key pathways, including the dentato-rubro-thalamic tract. We believe this line of investigation will be important in advancing both the anatomical characterization and translational relevance of the ZI.

      Reviewer #2 (Recommendations for the authors):

      (1) Re: data quality compared to the clinic, this could be achieved by replicating the analysis on a downsampled version of the HCP data that more closely resembles clinical data. This is understandably a large undertaking, so it could be left to future validation work.

      We thank the reviewer for this valuable suggestion. While this was suggested as potential future work, we felt adding this analysis would strengthen the manuscript. In response, we repeated our analyses using diffusion MRI data that more closely approximates a clinical acquisition with a lower spatial resolution (i.e., 2 mm vs. 1.25 mm isotropic) and number of diffusion-encoding directions (i.e., 130 vs. 270, Kasa et al., NeuroImage Clin., 2022). We have included these analyses in the revised manuscript.

      Reassuringly, the principal rostro-caudal gradient of cortico-incertal connectivity was preserved, demonstrating that the dominant organizational feature of the ZI is robust even under clinically representative acquisition conditions. However, finer-grained parcellations were less consistent with the original HCP analyses. In particular, cluster solutions with larger numbers of clusters (k > 3) became increasingly variable, indicating that differentiation of subtle connectivity-defined subregions benefits from the higher spatial and angular resolution afforded by research-grade diffusion MRI.

      We believe these findings provide a more nuanced assessment of the translational potential of our approach. They suggest that the large-scale topographic organization of the ZI can be recovered using clinically realistic diffusion MRI, while also highlighting the current limitations of routine clinical acquisitions for resolving finer anatomical subdivisions.

      We have incorporated these results and their implications into the revised manuscript.

      (2) Figure 6 legend labels: (c) and (d) should be (b) and (c).

      Thank you for highlighting this discrepancy. We have corrected as proposed.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful comments on the manuscript. In response to their suggestions, we have:

      Improved hardware calibration flexibility and documentation (Rev 1).

      Clarified the optical specifications of the system, including axial resolution and working distance (Rev 1).

      Updated Figure 2 and Figure S1 (Revs 1 and 2).

      Corrected typographical errors and clarified terminology throughout (Revs 1 and 2). In addition, we have a new Zapit release (v1.0.4, which includes release notes), that contains many improvements and bug-fixes including suggestions from reviewers. 

      Public Reviews:

      Reviewer 1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      We thank the reviewer for their assessment of the manuscript, particularly the reference to finding the balance between modularity and an integrated package.

      (1) Command signals

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      We thank the reviewer for this suggestion. Zapit uses a linear calibration by default, which works well for high-quality diode lasers with built-in power control. For users with EOMs or AOMs we have now implemented a feature that allows creation of the appropritate sigmoid calibration curve. This is in Zapit version 1.0.4 and the process for generating the calibration curve is documented on our GitBook doc site. and the commit containing most of the changes is here. We also include a third order polynomial fit, which we hope will help users of some cheaper lasers where the control function has non-linearities. The appropriate non-linear fit is chosen automatically. We describe this in the legend and main text associated with Fig. 7F.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      The number of grid lines and their spacing are set via GUI options. The initial galvo-toimage mapping uses a field-centred affine guess based on the user’s "scanners.voltsPerPixel" setting, and the setup process is described in the user guide. The software ships with a suitable default gain value, which is unlikely to require modification. Invert flags for X and Y are also provided. Once beam locations are recorded, a similarity transform is fitted for residual offset, rotation, and scale. The latest Zapit release includes bug-fixes associated with the centering of the initial calibration point grid in the field of view.

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      We kept the mapping fixed for simplicity in both the build instructions and the code. Zapit will require a dedicated DAQ so we do not anticipate re-wiring is a hurdle. Nonetheles, the code is open-source and such a change is possible: the settings file would need to be augmented and the functions that write the analog output waveforms modified. 

      (2) Laser and optics

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      This suggestion mirrors our own thought process, but we opted not to add a ray-tracing rendering to Figure 1, as doing so comprehensively would require illustrating additional optical principles that would detract from accessibility. However, building on the reviewers recommendation, we now provide references explaining the underlying scanning principles for interested readers (Schottdorf et al. 2025, referenced in Fig. 2). The relationship between scanner angle and beam position is also shown diagrammatically in Figure 7.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      This is an excellent point. With our specifications (0.8 mm beam diameter, 473 nm wavelength, 200 mm focal length objective), the effective NA is approximately 0.002, yielding a Rayleigh range of approximately 37.6 mm. The beam must therefore travel nearly 4 cm from focus before the point-spread function doubles in width, making the system highly insensitive to skull curvature. We have added this calculation and noted its practical advantage in the revised manuscript. (Section 2.2).

      What is the working distance?

      The Plossl scan/objective lens is housed at the end of the lens tube, giving a working distance of approximately 20 cm for the 200 mm focal length objective. We have added this information to Figure 2.

      Reviewer 2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank the reviewer for their positive assessment of the manuscript, and their vision for Zapit as an important community resource.

      Reviewer 3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

      We thank the reviewer for their detailed assessment of Zapit.

      Concerns & comments:

      (1) While the authors argue that it offers the best utility to affordability trade-off - faster than motorized drivers and require much less power than DMDs (100X) and less expensive/easier to use compared to SLMs, in the current form, the manuscript does not clearly list the limitations of the approach. At such, in my opinion, the authors should include side by side comparisons (perhaps as a table). For example, clear statements should be included with respect to comparisons in lateral (x-y), axial (z) spatial resolution, as well as temporal sequential aspect of Zappit and other photo-stimulation techniques involving DMDs or SLMs.

      Section 3.2 (Comparison to other approaches) compares the scanner-based approach to related techniques. Whilst this is brief, we believe it is adequate because the resolution is ultimately limited by tissue scattering. Indeed, we demonstrate that the radius of neural inhibition ~10 times larger at 2 mW than the lateral PSF (Figure 8D). The size of a DMD pixel on the brain will likely also be smaller than the excitation area, but it does depend on the imaged size of the DMD on the brain. Since that can vary from system to system, a comparison of even theoretical resolution is not straightforward. In terms of spatial patterning, DMDs and SLMs allow for arbitrary shapes to be created on the brain and we point this out in section 3.2. 

      (2) Is power really a limitation in terms of the laser sources? Or is this a disadvantage mainly because using less power has beneficial effects on the tissue health? It may be useful to provide metrics of comparisons along these lines between Zappit and DMD-based approaches.

      The reviewer highlights an important distinction. Too much laser power can cause phototoxicity, and can make neurons more excitable due to heating. It also results in a larger region of stimulation and increases the chance of off-target effects. However, when activating multiple sites with a galvo-based system like Zapit, dwell time goes down ~linearly as the number of “simultaneous” stimulation sites increases. Therefore, even if the same power at the sample is maintained (and the same risk of phototoxicity), peak power must go up to provide the same average power at each site. For example, stimulating 20 points at 40 Hz requires 10 times the peak power compared to stimulating 2 points at 40 Hz.

      The same is true of DMD-based approaches. For example, the Mightex recommend a 1 to 4 W laser to run their Polygon DMD-based photostimulation system over an area the size of the mouse dorsal cortex (personal communication); Kauvar, et al. 2020 used a 5 W laser to cover an area 7 mm across using a Polygon system. The cost of such a laser and the Polygon alone likely exceeds 60,000 USD. 

      Other than prices and logistics, the laser powers needed at the sample and their duty cycles are essentially the same across approaches and so there are no meaningful comparisons we can provide in this domain. However, based on the reviewer’s comments, we now clarify the effects of heating from high laser powers on neural excitability in the discussion (Section 1.1).

      (3) Arbitrary-scanning vs random scanning may be more appropriate to describe to strategy.

      We agree with the confusion surrounding “random-scanning” and have chosen the phrase “laser-scanning” rather than “random access”, which is the term used in the original pre-print. 

      Reviewer 1 (Recommendations for the authors):

      Consider noting that most other galvo scanner models can probably be used.

      We have added: "Other scanners could also be substituted, as can other lasers, lens combinations, and laser wavelengths etc.” (Section 4.3).

      The basic version of Zapit requires MATLAB (although alternatives are possible and guidance/code is provided), which is not unreasonable but may limit adoption.

      We acknowledge this point. MATLAB is widely available in academic neuroscience laboratories, and we provide a Python-based interface, but a complete conversion is beyond the scope of this manuscript. We hope that the open-source code will be adapted to other languages by the community as needed.

      Where laser power is mentioned (e.g., Discussion: "we recommend using 1-2 mW time-averaged light power..."), it is not always clear if this is at the laser or in the specimen plane; clarifications would be helpful.

      Thank you, we have clarified throughout that reported laser powers refer to measurements at the specimen plane.

      Abstract: "causally manipulating" - the "causal" part is redundant and can be dropped, or replaced with "transient" (more relevant).

      Changed to "transiently".

      Section 2.2: "The modular and open-source nature of Zapit means that exciting new configurations" - strike "exciting".

      Corrected.

      Figure 7D - y-axis label text is clipped.

      Corrected.

      Reviewer 2 (Recommendations for the authors):

      (1) In Figure S1 - Our paper (Heindorf et al.) is incorrectly listed as not making code available - the "camera controlled laser stimulation" software we used is part of Iris2p (that is freely available on SourceForge and linked as such in the paper). Granted, it's not user-friendly or easy to find in the large Iris2p software package - but it is technically "available".

      We apologise for this oversight. We have updated the Figure S1 legend to note that “available” code in this context refers to a dedicated and documented standalone package. We have also added links to both Pinto and the Keller-lab software packages.

      (2) In Figure 3: "and THE beam goes"

      Corrected.

      (3) In Figure 4: "the useR will be prompted"

      Corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Yuan and colleagues present a thorough study of gene activation before and during metamorphosis in sponge larvae, combining in-depth analyses of staged transcriptomes and chromatin accessibility profiling (ATACseq). Amongst several very interesting findings, the study reveals that the acquisition of settlement competence, which arises in response to decreasing light at sunset, is characterized by changes in chromatin accessibility that anticipate strong transcriptional shifts occurring as metamorphosis starts. Another notable finding is a set of transcription factors amongst the genes strongly up-regulated at the onset of metamorphosis. In addition, larvae exposed to constant light, a condition that stalls metamorphosis, were found to activate metabolic pathways that are not normally expressed in swimming larvae. Together, the findings provide a rare level of understanding into how environmental conditions can promote deployment of alternative developmental programs in planktonic larvae.

      Strengths:

      This is a very comprehensive, well-documented and rigorous study of a phenomenon of wide interest. It will inspire researchers working on other species to look for similar, environmentally-driven "anticipatory" epigenetic mechanisms. It also provides a wealth of detailed information on genes, notably transcription factors, that are candidates for involvement in regulating specific metamorphosis transitions - and beyond. The data presented here are thus undoubtedly a rich and valuable resource.

      We thank reviewer #1 for finding our study on sponge metamorphosis interesting and compelling, and that it is likely to inform future studies on gene regulation and activity in environmentally regulated developmental processes and metamorphosis.

      Weaknesses:

      I see no significant weaknesses; however, the documentation of the data is very compressed, with all the findings contained in 4 multi-panel figures with succinct legends. It is not always straightforward to connect the conclusion statements in the text to the figures. Although the relevant data is available in supplementary files, I would appreciate more help in navigating the data to assess the support for key conclusions, if possible, illustrating each text conclusion explicitly in the main figures.

      Thank you sincerely for these suggestions on how to better present the results. We agree that the figures and associated legends are succinct. To rectify this in the hope of improving clarity and accessibility, we have (i) created two new figures by splitting our original four figures into six, and (ii) expanded figure legends to provide more explanatory details. We also made minor additions to the main Results text to more fully explain some results (see also reviewer #3’s comments).

      Specifically, we:

      (1) Removed panel K (heat map of TF expression) from original Fig. 1 and created a new figure (new Fig. 2) that focuses solely on TF expression and emphasises the extraordinarily high expression of many TFs. We also moved into the new Fig. 2 a panel from the original SFig. 1 that documents the larval cell types that express these most highly expressed abundant TF transcripts. This new figure should provide the reader with a clearer perspective on high TF expression in the larval competence and the initiation of metamorphosis.

      (2) Removed panel F from the original Fig. 4 (now Fig. 5) to create a new, expanded Figure 6 that presents a stand-alone summary of the main findings of this work; that is, environmental regulation of competence and early metamorphosis. This allowed us to (i) incorporate the constant light experiment into the summary figure, and (ii) provide a more detail explanation in the legend.

      Reviewer #2 (Public review):

      Summary:

      It is demonstrated that sponge larvae prepare for receiving the environmental cue (sunset) by extensively modifying their chromatin accessibility in the vicinity of genes that are going to be regulated during metamorphosis, in the absence of large gene expression changes. This program can be offset by modifying the cue (making light constant), leading to a novel molecular state.

      Strengths:

      This is a top-notch study of a key lifecycle transition in an organism of great phylogenetic importance, involving concurrent gene expression and chromatic accessibility profiling (to the best of my knowledge, this has never been done in non-bilaterians and likely anywhere outside Vertebrata). The result is highly non-trivial. There is also an additional experiment modifying the key environmental cue (constant light), adding additional insight.

      We thank reviewer #2 for their efforts and for appreciating the approaches we employed to understand environmental induction of sponge metamorphosis. In addition to the phylogenetic importance of sponges, their pelagobenthic life cycle is likely shared with disparate bilaterians (but not with vertebrates and other chordates, whose metamorphoses are probably derived).

      Weaknesses:

      I have only a couple of suggestions.

      (1) Not all new pre-emptively opened OCR regions are associated with genes that are going to be regulated during metamorphosis. Is their association with such genes statistically significant? (Fisher's exact test?)

      Thank you for raising this helpful point. In following your suggestion to statistically test this, we determined that a Fisher’s Exact Test was not appropriate because that test is generally used only for small samples or tables with expected counts below 5; in our data, all four expected cell frequencies are well above 5 (minimum = 228.6) and N = 25,149. Thus we instead tested for an association between chromatin accessibility and differential gene expression using the more appropriate Pearson's chi-squared test. We found no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.18), and have added these details into the Results (lines 344-46) as follows: “Consistent with this interpretation, 62% of all genes that are differentially expressed in 1 hps postlarvae (3032) have proximal chromatin regions already accessible in competent larvae (Supplementary Tables 3 and 9), although statistically we find no significant difference in DEG rate between genes associated with newly opened OCRs and those associated with other OCRs (9.84% vs 11.47%; χ<sup>2</sup>(1) = 1.76, p = 0.184).”

      (2) Re: extended discussion on possible reasons for activation of specific transcription factor families. I feel it is not terribly useful since it is hardly more than guesswork. The authors should consider condensing this part to better emphasize the major (and most unexpected) large-scale regulation patterns.

      We agree with this appraisal and have modified the beginning of Discussion to highlight the large-scale and rapid changes of overall gene expression. This emphasises the regulatory processes – TF expression and chromatin state changes – that must be in place to allow such transcriptional changes to occur. We feel this addition enhances the focus on TF activation and regulation at competence and early metamorphosis, especially given the scale and level of change, with most of TFs being expressed at very high levels (i.e. top 5% of all gene expressed). As outlined above, we created a new figure focussed on highly expressed TFs (new Fig. 2) to hopefully further highlight this phenomenon. It would be of great interest to know if this is conserved amongst animals with a pelagobenthic life cycle and rapid metamorphosis.

      (3) Re: enrichment analysis based on significant genes (Figure 1H): Even though it is a common practice, there is nuance: as we all know very well, many genes pass a significance threshold not because they are highly differentially regulated (i.e., show large fold-change), but because they are more abundantly expressed overall and so the statistical power for them is greater. A good example is ribosomes - before we realized what was happening, they would show up as enriched in almost every experiment of ours, which was not very useful since their fold-change was quite trivial. I see the authors have ribosome enrichment too, and I suspect there are a few more functional groups that made it because they tend to express highly on average. Ideally, we want to see what is enriched among highly regulated genes, not among abundantly expressed genes. Because of this we moved to compute enrichment based only on fold-change, using the GO_MWU package (https://github.com/z0on/GO_MWU). I suggest authors give it a shot, to see if the enrichment results become more interpretable. GO_MWU is also very powerful to analyze enrichment in WGCNA modules, in case the authors want to try that.

      Thank you for this interesting insight and advice. We applied the GO_MWU package to our gene expression dataset. Overall, these new results corroborated the original KEGG enrichment analysis, largely identifying GO biological processes, cellular components and molecular functions consistent with the previously identified KEGG molecular and cellular processes operating at larval competence and 1 hps. These include genomic regulatory processes underlying transcriptional changes and morphogenetic processes that occur in the first hour of metamorphosis, and which are also highlighted in a recent BioRxiv paper (https://doi.org/10.64898/2026.04.23.719999) from our group.

      We have added (i) results from the GO MVU analyses to Supp. Fig. 1 and Supp. Table 4, (ii) the following statement to Fig. 1 legend: “GO-MWU analysis of upregulated genes reveal stage-specific enrichments largely consistent with the KEGG analysis (Supplementary Fig. 1 and Supplementary Table 4).”, and (iii) a brief description of this approach into the Methods.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, Huifang Yan and colleagues perform RNA-seq (CEL-seq) and ATAC-seq experiments to profile the transcriptome and chromatin accessibility of sponge larvae across larval competence, settlement and early postlarval development. Amphimedon, the sponge species that they use, is amenable to lab experiments and can therefore be a convenient model for experimenting with this otherwise difficult to assay ecological parameters and cues. They had previously observed that light conditions (diminished light) at sunset are critical for larvae to enter a pre-settlement stage and prime them for settlement and metamorphosis. In this paper, they report that these conditions induce a gain of accessibility in many genes, including transcription factors, and that altering these conditions by providing continuous light at sunset affects this reprogramming event.

      Strengths:

      The above is a very interesting observation, one that the authors speculate could have a broader significance and be a theme in many more larvae. I agree with the authors that this is an important finding, and I think that the paper will be interesting for a broad readership. If this is the case, the authors open up a new theme of chromatin regulation, extensively studied in mammalian contexts, but severely understudied in pretty much every other context.

      We thank reviewer #3 for their positive assessment of our findings, pointing out the novelty of this research and its broad relevance.

      Weaknesses:

      I think, however, that their paper often reports the data in a difficult-to-follow way, and that other sorts of analyses would have made the results more accessible for a broad readership. Here, I present some suggestions that the authors might want to take into account to improve their results.

      Reviewer #1 also commented on how the results were difficult to follow. Based on your and their comments, we have reworked parts of this section, and the figures and legends. Details of these changes are listed above and can be viewed in the new version with track changes on.

      We note that no further specific suggestions were visible to us in your review.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1) The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy.

      This is correct, our claim is that ECS drives spreading depression dependent IEG expression in cortex. We are not claiming this is the only consequence of ECS. The extracortical response to ECS could contribute to the treatment, and we discuss this in the hippocampus specifically (see Discussion). In current clinical practice, cortical surface EEG is used as a biomarker for predicting treatment outcome. Given that the CSD is necessary to drive Fos expression in cortex, we think it is warranted to speculate that monitoring CSD during ECT is worth exploring as a potentially superior biomarker of treatment outcome. Especially given that this is readily feasible (Discussion).

      Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      The idea that Fos is a marker of neuronal activity is outdated. The primary correlate of Fos expression is neuronal plasticity and learning-related circuit modifications, rather than just neuronal activity – see e.g. (Bolhuis et al., 2001; de Hoz et al., 2018; Fleischmann et al., 2003; Kimpo and Doupe, 1997; Mahringer et al., 2022, 2019; Nakadate et al., 2012; Roy et al., 2016; Ryan et al., 2015; Tanaka et al., 2018; Tyssowski et al., 2018; Watanabe et al., 1996; Yassin et al., 2010). We have now referenced some of those articles in the manuscript (Introduction). But more importantly, our claim is that the CSD is necessary to drive Fos. The fact that CSD can occur in a single hemisphere allows, within a brain, to control for direct ECS response and phase III oscillations. In unilateral CSD, contralateral hemispheres displayed Fos levels in the cortex that were similar to sham mice. Investigating the exact plasticity pathway that is triggered by the CSD is an interesting question we are pursuing in follow-up work, but does not influence our conclusions here.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      The reviewer is correct in that postictal suppression is currently one of the best correlates of positive clinical outcome used in the clinic. However, note that EEG-based correlates are generally weak predictors of treatment outcomes: even recently identified correlates are unreliable between patients cohorts and account for less than 60% of the clinical outcome (Francis-Taylor et al., 2020; Scangos et al., 2019). See also our response to comment (2) of reviewer 3 on this topic.

      The primary problem in the interpretation of postictal suppression of EEG activity is that it is unclear what the source of the EEG signal is in this case. During phase III oscillation following the CSD, cortex is silent and all of the EEG is likely driven by thalamic input to cortex. Note, source localization in EEG does not help as the source of the signal is likely the thalamic axons in cortex (Huels et al., 2023) (or their postsynaptic, and in this case subthreshold, effects). Whatever the source of the cortical surface EEG may be, we know that following CSD it cannot be cortical activity, hence, whatever is driving postictal suppression is also not cortical. If we had to speculate, postictal suppression is likely an exhaustion of thalamus in attempting to drive cortex without getting any excitatory feedback. During an epileptic seizure, thalamus oscillates at delta frequencies and cortex responds at each peak of the oscillation, generating ictal spikes on the EEG (Meeren et al., 2002; Polack et al., 2009). The fact that ictal spikes are missing in an “optimal” phase III oscillation is consistent with our findings that the CSD completely silences cortex for minutes. Thus, during an ECS-driven phase III oscillation, cortex does not respond to thalamic input and at some point thalamus runs out of energy to maintain the oscillation - it is likely the combination of thalamus running out of energy and a silent cortex that gives rise to postictal suppression. Note we do find an asymmetry for the mid-oscillation amplitude (Figure 6A), which in this interpretation can be explained by the thalamo-cortical circuit attempting to compensate for the absence of cortical feedback by increasing its drive to the CSD-affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      This is a misunderstanding. We claim (and show) that the CSD is necessary to drive immediate early gene expression following ECS. Based on this we speculate that the CSD “may serve as a more relevant biomarker for predicting and optimizing therapeutic outcomes of ECT” than EEG biomarkers.

      We do not “state [that there is] a beneficial role of CSD” – we actually address the possibility that SD may not be therapeutically beneficial (Discussion). The suggestion of investigating behavioral consequences of CSD in mice is interesting, but likely not the most relevant follow-up work. This is for two reasons:

      (1) Very different from humans, most mouse behavior does not depend on cortex (Kawai et al., 2015; Pandey et al., 2026). Thus, any potential behavioral consequences of ECS in mice are difficult to interpret in the context of clinical relevance. We don’t think there is any known biomarker based on mouse behavior that has any direct translational value to psychiatric treatments. The reviewer may disagree with this assessment but just consider that no mouse behavior assay has ever been instrumental to the development of a novel psychiatric treatment.

      (2) The more direct – and simpler approach – is to directly test whether CSD correlates with treatment outcomes in patients. Our primary aim with this paper is to inspire exactly this. Note, we are currently pursuing this as well, of course. Monitoring CSDs in humans during ECT is likely possible with fNIRS or other hemodynamic-based measurements (Discussion).

      However, to reiterate – the claim of the manuscript is exactly as stated in the title. Thus, demonstrating the clinical relevance of the CSD is outside of the scope of the current manuscript, but will be trivial to prove by measuring treatment outcomes while routinely monitoring patients for CSD.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

      Reviewer #1 (Recommendations for the authors):

      The results of sham tests are shown only for the immunohistochemical data. However, results from control experiments should be provided for the EEG and calcium data.

      Sham data are included in Author response image 1. We are not sure how they support our conclusions and have left them out of the manuscript. A sham stimulation just characterizes ongoing activity under anesthesia. The more direct comparison is the difference between pre- and post-ECS activity. Baseline widefield calcium imaging data is already shown in Figure 1, 2, 6; baseline mouse EEG data is already shown in Figure 1 and 6. Baseline EEG in patients is not available in the present dataset.

      Author response image 1.

      Sham ECS does not result in phase III oscillations or a CSD. (A) Representative spectrograms (top) and raw traces (bottom) of concurrent EEG (left) and widefield calcium imaging (right) during a sham ECS session. (B) Population raster plot (top) and two example activity traces of neurons (bottom) during a sham ECS two-photon recording.

      Moreover, baseline (pre-ECS) calcium and EEG activity should be shown in Figures 1 and 6 to correctly assess the changes induced by stimulation.

      We are not sure we understand as pre-ECS data are already shown in Figures 1 and 6. We assume the reviewer may mean “more baseline data should be shown”? We have extended the time scale of the relevant panels in Figure 1, 2 and 6 to include 30 s instead of 10 s of pre-ECS data.

      Using the term 'oscillations/phase III oscillations' instead of 'seizures' throughout the text is confusing because the word covers a wide range of brain oscillations - from normal to pathological ones.

      This is indeed confusing – but the confusion arises from the often imprecise usage of the term “seizure” in the ECT literature. The EEG response to ECS is not equivalent to that of an epileptic seizure. The description of a “phase III oscillation” is a more precise description of the EEG signature. It was introduced by (Brumback and Staton, 1982) and is not our terminology. We dedicate an entire paragraph in the introduction to this distinction. We are not sure how to make this clearer in the current manuscript. Continuing to describe the EEG response to ECS as “seizure” is inaccurate, and mechanistically misleading, especially given that cortex is silent during the phase III oscillation following the CSD (i.e. cortex cannot be “seizing” during this phase of the EEG response, see our point above on postictal EEG suppression).

      The presentation of clinical data is scarce and unclear. e.g., the authors claim that the oscillation frequency is similar in mice and humans (lines 206-208). However, in Figure 1F, oscillations during the early post-ECS phase have twice the frequency (6-7 Hz) in the patient EEG recording (6-7 oscillations per 1 second) compared to the mouse EEG (3 Hz, i.e., 6 oscillations per 2 seconds during 73-75 s). It seems a bit odd because Figure 1H shows that the frequency does not exceed 5 Hz, even in patients.

      Please excuse, this is our mistake. The x-axis label in the human data of Figure 1F was incorrect. It should have read 1 to 3 s, not 1 to 2 s. Window sizes were of course matched between mice and human data and are all 2 s. The mistake is now corrected. The frequencies in patients and mice are compared in Figure 1H and are not different.

      A part of the discussion (lines 627-635) is based on factual inaccuracy: cortical SD cannot invade the hippocampus in the in vivo brain (only in slices), although SD can occur in the hippocampus in response to generalized seizures.

      It would be helpful if the reviewer would back up this claim with references. Short of this we are left to speculate - we suspect the reviewer may be referring to earlier work in the rat cortex, that claimed that a cortical SD can only invade hippocampus if glia has been impaired (Largo et al., 1997). This is now an outdated model: there are more recent reports of CSD invading the hippocampus (Bahari et al., 2020; Bonaccini Calia et al., 2022). If the reviewer has specific concerns with any of these papers, we would be happy to discuss in more detail, but the reviewers’ claim seems unfounded here.

      It is reasonable to expect that bilateral stimulation produces bilateral calcium/SD waves. In Rosenthal's experiments, this situation was most common. In the present study, bilateral ECS triggers mostly unilateral calcium waves. Do you have any idea what the reasons for the result are? As stimulation parameters (polarity, intensity, and frequency) have been shown to control the occurrence of SD waves, their unilateral pattern suggests non-uniform stimulation conditions. I am curious whether the uni- or bilateral pattern of calcium/SD waves depends on the ECS parameters.

      The main difference between our patient and mouse data is the polarity of stimulation. For patient data, the polarity of the current alternates with every pulse, while in our mouse data the stimulation was always right unipolar. This is discussed when we mention the asymmetry of SD in our data (Results, Methods) - we suspect the reviewer may have missed this.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      (1) The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      This is correct, our work confirms that the calcium event following ECS is a CSD. The main claim of our paper is that ECS drives immediate early gene expression in a CSD-dependent manner.

      However, note that the Rosenthal paper concluded that the calcium event is a CSD while providing only little evidence for that claim – mind you, we agree with their interpretation, but provide more evidence for the conclusion. We are happy to discuss the limitations of the Rosenthal paper as highlighted below more prominently in the manuscript, if the reviewer thinks this would be helpful, but we think that is likely not necessary.

      Briefly, in the data presented in (Rosenthal et al., 2025), the only support for the calcium response being a CSD is the speed of propagation and the DC shift reported in extracellular recordings. We add to this by showing that the spread of the calcium event follows the pattern expected by a CSD through cortical layers (Figure 4), results in heterogeneous returns to baseline calcium levels (Figure 4), causes vasoconstriction (Figure S5), travels at the speed expected of a CSD regardless of stimulation parameters (Figure S3) and causes Fos expression (Figure 5).

      Most importantly, the ECS as used by Rosenthal and colleagues is not a mouse model of ECT, in the sense that is not a scalp electrical stimulation. The stimulation method is fundamentally different between our two articles: (Rosenthal et al., 2025) implanted stimulating electrodes directly in the mice’s dorsal cranium. Direct cortical stimulation is well known to be able to cause CSD (Leao, 1944). However, it is unclear whether direct cortical stimulation is a useful model for ECT. We suspect that Rosenthal and colleagues were led to believe that auricular stimulation does not work because they saw no evidence of a cortical seizure in calcium recordings following stimulation. (See discussion on the confusion of phase III oscillation and seizure in the ECT literature.) We suspect that EEG recordings following direct cortical stimulation would reveal a very different EEG pattern from that observed in patients. This highlights the importance and novelty of the comparison of mouse and human EEG we present in Figure 1. Note, in our auricular stimulation preparation we do not observe any seizure-like activity in cortex that lasts beyond the stimulation (compare Rosenthal’s Figure 1E vs Figure 1F here). Given the EEG similarity we show between mice and patients, we suspect the cortical seizure Rosenthal and colleagues find is a methodological artifact. This is puzzling to us, as Rosenthal and colleagues do briefly mention a single mouse example with an extracellular electrophysiology recording compatible with CSD following auricular stimulation (Supplementary Figure 2).

      Thus, not only is it necessary to add evidence to the interpretation that the ECS-driven calcium event is indeed a CSD, but also that it can be triggered by a stimulation method that successfully replicates the known EEG response of human patients.

      (2) The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSDCa2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

      We have added additional references to the CSD-Fos literature in the discussion. Regarding the role of Fos as a marker of plasticity rather than activity, we discuss this point in our reply to comment (1) of reviewer 1. Concerning the expression of other genes, we already referenced the TRKB/BDNF pathways (Discussion). We have now added references on RNA-seq following CSD, which show that Fos is one of the most differentially expressed genes following CSD (Dell’Orco et al., 2023). Regarding the use of Fos knockout/knockdown lines, please see our reply to the reviewer’s comment (4).

      Reviewer #2 (Recommendations for the authors):

      (3) The Results and Discussion sections should be revised to better reflect the prior discovery of CSD following ECT in rodents and initial evidence in humans, as discussed in the first point in the Weaknesses section above.

      The prior discovery of the fact that ECS can drive CSD is first mentioned in the third sentence of the abstract “However, this view is challenged by the recent finding that electroconvulsive stimulation (ECS) can trigger a cortical spreading depression (CSD).” (The abstract has no references, but the introduction should make it clear what is meant).

      There is an entire paragraph of the introduction discussing the Rosenthal results.

      The first time we discuss our results (end of the Introduction) we say: “Consistent with previous work (Rosenthal et al., 2025), we observed a slow travelling calcium event that appeared to be a CSD.” The Rosenthal paper is cited in 8 times in total throughout the manuscript.

      We are unsure what the reviewer is asking us to do here. As mentioned above, if any revision is warranted regarding that reference, it should be to clarify that the Rosenthal paper used direct intracranial stimulation rather than ECS, and did not fully confirm the calcium event as a CSD – but they should be credited for finding the first preliminary evidence for CSD following ECS in patients.

      (4) To support the authors' statement that "CSD is the primary driver of plasticity is based on its role in driving Fos expression", additional experiments with Fos knockdown or knockout (or alternative interventions) are needed.

      Our main claim is that CSD causes Fos expression in the cortex following ECS. The argument that Fos can be used as a marker of plasticity follows from the literature, not from the experiments done here – this would require a form of functional plasticity measurement, which is outside the scope of this paper (and might not be the most relevant direction, see our reply to comment (3) of reviewer 1). The only observation that a Fos genetic manipulation would give us is the lack or reduction of Fos expression following CSD, which would be orthogonal to the points we make here.

      (5) It would be helpful to use a more specific term than "Ca2+ event" in Figure 2D and throughout the related Results section. It is assumed that this is the large propagating Ca2+ event attributed to CSD, but the terminology is important, as all the other events in the recording (including during Pre-ECT and the Direct ECT period) are also Ca2+ events.

      We have added a clear definition of what we mean on first usage (Results). The reason to call it a calcium event, and not a CSD, is that we did not want to jump to conclusions. We do think that the event is a CSD, but conclusive proof of that is still lacking (in both our work and that of Rosenthal). It is conceivable that the event propagates via a different mechanism than a CSD.

      (6) The authors should comment on differences among rodent models of ECT stimulation, especially with direct and ear clip methods, as discussed in the context of translational value (Theilmann et al., 2014).

      The Theilmann 2014 paper compares ECS delivered via auricular stimulation and intracranial electrodes. They compare the two stimulation methods, but use stimulation currents, total charges, and stimulation duration that were not matched. In the case of stimulation current, those used in auricular stimulation are approximately 8 fold higher than what they use for intracranial stimulation. Moreover, the stimulations were performed in awake rats. This is scientifically - and ethically, even for 2014 - questionable in light of the fact that this is aimed at developing a model for ECT, which is always done under general anesthesia. They conclude that cortical stimulation has less adverse effects and is more effective in reducing immobility in a forced swim test. The confound in the interpretation of these results is that all stimulation was performed in awake animals. The reason this is no longer done in humans is that it is extremely painful. Direct cortical stimulation is likely much less painful (for the same reason TMS is less painful). This would explain their findings of increased adverse effects with auricular electrodes. Given the differences in stimulation parameters used, the difficulty of calibrating equivalent doses of auricular and intracortical stimulations, as well as the small effect sizes reported, we don’t think the results allow for any solid conclusions as to which method is more effective in reducing immobility in a forced swim test. However, even if one would assume the intracortical ECS is more ‘effective’, this is hardly relevant, as we are interested in using mouse ECS as a model for human ECT. One could speculate that intracranial ECS might also exhibit higher clinical benefit than surface ECS in patients, but that is not the scope of our research, and probably not clinically relevant. Finally, while the authors do include EEG recordings in the rat – these were not compared to patient recordings, and from visual inspection do not resemble patient EEG recordings that we have seen. We would argue that the best rodent model of ECT stimulation is the one that triggers neuronal activity most similar to that observed in patients. We have added a brief discussion of these points to the corresponding Methods section.

      Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      This appears incorrect. It was known that direct cortical stimulation (as was done in the Rosenthal paper) can drive a calcium event that resembles a CSD. It was also known that CSD can drive Fos expression. What was not known is that the Fos expression driven by ECS is fully explained by the calcium event (putative CSD). This is particularly relevant as most people still erroneously assume Fos is a marker or neuronal activity, and prior work has come to the conclusion that ECT does not drive, but likely downregulates Fos expression (Calais et al., 2013; Morinobu et al., 1995; Park et al., 2014; Winston et al., 1990). We have added a more prominent discussion section on this point.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      The importance of Figure 1 is to demonstrate that ECS delivered with auricular electrodes in mice causes an EEG signature that is very similar to that seen in patients. We do not claim we are the first to describe characteristics of phase III oscillations in either patients or mice (we have added the Murakami reference to the manuscript). The reply to comments (1) and (6) of reviewer 2 highlights why this comparison is so important – it was not done in the Rosenthal paper the reviewer mentions below, and it is not clear whether the intracortical stimulation used there even drives a comparable EEG response (given the calcium activity shown, we suspect the answer is no). To the best of our knowledge this direct comparison is novel – but again the key novelty of our work we highlight is the one described in the title.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      It has indeed been demonstrated that ECS delivered using intracranial electrical stimulation can trigger CSD-like events (e.g. Rosenthal et al.). However, the fact that localized intracranial electrical stimulation can trigger a CSD has been shown quite a while ago already (see e.g. (Leao, 1944)). This is not the case for surface stimulation the way it is done in ECT and the way we do it. Nevertheless, note we give full credit to the Rosenthal work for making this connection. Our main contribution – as highlighted by the title – is showing that the Fos expression driven by ECS is fully explained by the CSD.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1.

      If the reviewer has references for this claim, we would be happy to add to the manuscript – we are not aware of any such work. As far as we are aware, this is still an area under active investigation – calcium signals correlate (locally) strongly with shank recordings (Wei et al., 2020), but how this translates into an EEG signal is speculative.

      It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      The reviewer may be jumping to conclusions here. It is correct that this has been shown for a CSD. The more important question (and the reason we did this experiment) is whether the calcium event triggered by ECS is indeed a CSD. We try to be careful to describe it as a calcium event in the results (mind you the calcium event in the Rosenthal is very likely a CSD as their intracortical stimulation (‘ECS’) is likely equivalent to the electrical stimulations used in the discovery of the CSD). We then perform a series of comparisons to see whether the calcium event has the known characteristics of a CSD – and we conclude everything we test is consistent with it being a CSD. Note, once again, that we do not claim novelty in any of this – the primary novelty is the link between ECS, Fos and CSD.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      That is correct. The question however is how much of the Fos expression is explained by CSD. Prior work that has looked at Fos expression in response to ECS has found that ECS results in a slight reduction of Fos expression (Calais et al., 2013; Park et al., 2014). We suspect this is the result of not triggering a CSD. We do not claim to have discovered that CSD induces Fos expression. The novel contribution, which is the main claim of the paper, is that in the context of ECS the entirety of the cortical Fos expression can be explained by the CSD. This links the relative contributions of multiple components of the ECS response to a known marker of neuronal plasticity.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      There is probably a misunderstanding here. We argue and show that there is no cortical seizure following ECS – that is why we refer to the EEG responses as phase III oscillations (characteristic of a silent cortex). And there is little prior evidence that the calcium events triggered by ECS are indeed a CSD (we think this is likely the case, but demonstrating this conclusively will require further work). If the reviewer has references for this, we would be happy to discuss. Note, the Rosenthal et al. paper just assumes (probably correctly) that they are a CSD, but does not demonstrate this. We don’t fully demonstrate this either, we just provide additional evidence. But once again, the novelty is in the title of the manuscript, and we do not claim any other novelty to the best of our reading of our manuscript. If the reviewer has a particularly misleading passage in mind, we are happy to rephrase.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      We are not sure what the reviewer means by “cFOS expression is nonspecific to plasticity”. Does the reviewer mean Fos has other roles beside driving neuronal plasticity? That is very likely correct, but it is unclear how that is relevant. Fos is a key driver of a number of neuronal plasticity pathways (Chaudhuri et al., 2000; Cohen and Greenberg, 2008; Cruz et al., 2015; Durchdewald et al., 2009). Which exact pathways are driven by ECS is an interesting question that we are currently pursuing in follow-up work. Given what we know about Fos expression, it is probably well within reasonable bounds to conclude that Fos increases result in neuronal plasticity. Fos expression directly drives network plasticity (Yap et al., 2021): "our findings indicate that Fos expression has an instructive role in orchestrating persistent circuit modifications”. Likewise, Fos induction is strictly required for experience-dependent representational plasticity during learning (de Hoz et al., 2018): “locally blocking c-Fos expression caused […] decreased cortical experience-dependent plasticity, without affecting baseline excitability or basic auditory processing.”. Calcium activity explains about 15% of the variance of Fos expression (Mahringer et al., 2022). This is likely driven by the correlation between activity and plasticity, not by a direct necessity for Fos expression to maintain neuronal activity. This is consistent with the finding that Fos as a transcription factor does not function to maintain spiking activity; it is part of the gene-regulatory machinery that converts patterned synaptic input into lasting plastic change. Fos is induced by NMDA/Ca2+, ERK, and CREB/Elk signaling rather than by firing alone (Fields et al., 1997; Xia et al., 1996), and those same pathways are required for the transcriptional program that stabilizes long-term potentiation and other durable synaptic modifications (Davis et al., 2000). As an AP-1 transcription factor, Fos drives downstream gene expression, so its appearance is better read as entry into a plasticity-related nuclear program than as a measure of ongoing excitability (Minatohara et al., 2015; Morgan and Curran, 1991; Sheng and Greenberg, 1990). That interpretation is consistent with our work showing that early Fos expression preferentially marks neurons that later undergo the strongest learning-related functional changes (Mahringer et al., 2019) and tracks functional reorganization during learning rather than simple recent activation (Mahringer et al., 2019).

      The link between CSD and therapeutic effect is a speculation we make in the manuscript, not a conclusion. We are of course in the process of performing follow-up work to test whether CSD in patients is a better predictor of treatment outcome than EEG based metrics – no experiment we can do in mice will be able to test the hypothesis that CSD is the mediator of the therapeutic benefit of ECT.

      Regarding the power of EEG metrics to predict therapeutic outcomes, we fully agree with the reviewer. The literature on the reliability of EEG metrics computed from phase III oscillation data is controversial and noisy – this is something we establish in the introduction to motivate our research into other biological processes that could explain how ECT works. Note however, this is certainly not a fringe view – see e.g. comment (2) of reviewer 1: “Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy“ We think the reason for this is that a CSD is necessary for therapeutic benefit, but only has minor effects on the phase III oscillation – note this is a hypothesis based on our results that is trivial to test in patients (which are in currently investigating).

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      Whenever possible, we use mice for multiple experiments in the lab. This is done to reduce the total number of mice used for experiments for ethical reasons. For the mice in question, ChrimsonR was injected in the retrosplenial cortex (Methods), which was originally done to stimulate locally the axons projecting in the imaging area (primary visual cortex here). Thus the labelling is very sparse, and perfectly compatible with two-photon GCaMP6f imaging. Fluorescence emission is far too weak to activate ChrimsonR.

      Qualitatively, one can mentally compare the light power we use to activate optogenetic tools, which tends to be blindingly bright (one shouldn’t look into the optogenetic stimulation laser), with the fluorescence emission from two-photon imaging of calcium indicators, which tends to be barely visible by eye.

      Quantitatively, one can estimate this as follows: At 510nm emission, each photon carries an energy of about . Assuming a neuron that strongly expresses GCaMP6f under two-photon excitation would emit a very high 10<sup>7</sup>photons per second (Har-Gil et al., 2018), its total emission power would be P = N<sub>photons</sub> * E ≈ 4pW. Assuming this is spread over the surface of the cell (sphere with 10 µm diameter), this translates to . This is several orders of magnitude below the value of 1 mW/mm<sup>2</sup> irradiance required to activate ChrimsonR modestly at peak absorption, which is 80nm away from GCaMP6f emission (Klapoetke et al., 2014). Note that this would be true even when ChrimsonR is injected at the site of imaging (Vasilevskaya and Keller, 2026).

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      We suspect that the question is driven by a misunderstanding. While it is likely that the implantation triggers a CSD (and likely so does a standard two-photon window implantation), the implantation surgery and the experiments/imaging are separated by at least 3 weeks (Methods). We have never observed spontaneous CSDs in the days and weeks following an implantation surgery.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      We have added additional details as requested by the reviewer (Methods).

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      We have added two tables (Table S3 and Table S4) that displays the stimulation parameters used for each figure, as well as the distribution of parameters for mice and patients. More importantly, the properties of the travelling calcium event do not depend on the stimulation charge (Figure S3), which removes this confounding factor and supports the idea that the calcium event is a CSD.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      That is correct – only the cortical expression of Fos depends on the cortical CSD (see our reply to comment (1) of reviewer 1). Quantification of the whole-brain Fos expression following ECS is outside the scope of the manuscript, but it is something we are currently pursuing for separate publication. We have now reworked the section describing the Fos expression to make it clear that we are only talking about cortical expression of Fos (Results).

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      That was indeed our speculation based on ongoing work on TMS in the lab – but it was poorly phrased and unnecessary. We have rephrased.

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

      It is speculative indeed – but the speculation is not unfounded and has been made previously. The primary evidence is correlative in that schizophrenia is characterized by a reduction in coupling between thalamus and frontal areas of cortex – see e.g. (Giraldo-Chica and Woodward, 2017; Vinogradov et al., 2023). We have argued in previous work that combining this with computational models of psychosis (Sterzer et al., 2018), it is not unreasonable to speculate that a reduction in the coupling between thalamus and frontal areas of cortex could explain psychosis (Keller and Sterzer, 2024). The value of that speculation here is that it forms a testable hypothesis for the mechanism of action of ECT. Nevertheless, we have rephrased the statement slightly to make it clearer that this is still speculation.

      Reviewer #3 (Recommendations for the authors):

      The authors should design/execute an experiment(s) to attempt to prove causality between CSD, calcium influx, cFOS expression, and therapeutic effect. At a minimum, the authors need to revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field.

      Regarding novelty – see discussion above.

      Regarding therapeutic effects – we are in the process of testing this in patients. Using CSD measurements during ECT to test whether CSD is a better predictor of treatment outcome. I suspect that this will, however, require many more labs to come to firm conclusions. Our aim here is to inspire these experiments. That is why we speculate about clinical relevance in the abstract and the discussion.

      References

      Bahari, F., Ssentongo, P., Liu, J., Kimbugwe, J., Curay, C., Schiff, S.J., Gluckman, B.J., 2020. Seizure-associated spreading depression is a major feature of ictal events in two animal models of chronic epilepsy. https://doi.org/10.1101/455519

      Bolhuis, J.J., Hetebrij, E., Den Boer-Visser, A.M., De Groot, J.H., Zijlstra, G.G.O., 2001. Localized immediate early gene expression related to the strength of song learning in socially reared zebra finches. European Journal of Neuroscience 13, 2165–2170. https://doi.org/10.1046/j.0953816x.2001.01588.x

      Bonaccini Calia, A., Masvidal-Codina, E., Smith, T.M., Schäfer, N., Rathore, D., Rodríguez-Lucas, E., Illa, X., De la Cruz, J.M., Del Corro, E., Prats-Alfonso, E., Viana, D., Bousquet, J., Hébert, C., Martínez-Aguilar, J., Sperling, J.R., Drummond, M., Halder, A., Dodd, A., Barr, K., Savage, S., Fornell, J., Sort, J., Guger, C., Villa, R., Kostarelos, K., Wykes, R.C., Guimerà-Brunet, A., Garrido, J.A., 2022. Full-bandwidth electrophysiology of seizures and epileptiform activity enabled by flexible graphene microtransistor depth neural probes. Nat. Nanotechnol. 17, 301–309. https://doi.org/10.1038/s41565-021-01041-9

      Brumback, R.A., Staton, R.D., 1982. The Electroencephalographic Pattern during Electroconvulsive Therapy. Clinical Electroencephalography 13, 148–153. https://doi.org/10.1177/155005948201300306

      Calais, J.B., Valvassori, S.S., Resende, W.R., Feier, G., Athié, M.C.P., Ribeiro, S., Gattaz, W.F., Quevedo, J., Ojopi, E.B., 2013. Long-term decrease in immediate early gene expression after electroconvulsive seizures. J Neural Transm (Vienna) 120, 259–266. https://doi.org/10.1007/s00702-012-0861-4

      Chaudhuri, A., Zangenehpour, S., Rahbar-Dehgan, F., Ye, F., 2000. Molecular maps of neural activity and quiescence. Acta Neurobiol Exp (Wars) 60, 403–410. https://doi.org/10.55782/ane-2000-1359

      Cohen, S., Greenberg, M.E., 2008. Communication between the synapse and the nucleus in neuronal development, plasticity, and disease. Annu Rev Cell Dev Biol 24, 183–209. https://doi.org/10.1146/annurev.cellbio.24.110707.175235

      Cruz, F.C., Javier Rubio, F., Hope, B.T., 2015. Using c-fos to study neuronal ensembles in corticostriatal circuitry of addiction. Brain Res 1628, 157–173. https://doi.org/10.1016/j.brainres.2014.11.005

      Davis, S., Vanhoutte, P., Pages, C., Caboche, J., Laroche, S., 2000. The MAPK/ERK cascade targets both Elk-1 and cAMP response element-binding protein to control long-term potentiation-dependent gene expression in the dentate gyrus in vivo. J Neurosci 20, 4563–4572. https://doi.org/10.1523/JNEUROSCI.20-12-04563.2000

      de Hoz, L., Gierej, D., Lioudyno, V., Jaworski, J., Blazejczyk, M., Cruces-Solís, H., Beroun, A., Lebitko, T., Nikolaev, T., Knapska, E., Nelken, I., Kaczmarek, L., 2018. Blocking c-Fos Expression Reveals the Role of Auditory Cortex Plasticity in Sound Frequency Discrimination Learning. Cereb Cortex 28, 1645–1655. https://doi.org/10.1093/cercor/bhx060

      Dell’Orco, M., Weisend, J.E., Perrone-Bizzozero, N.I., Carlson, A.P., Morton, R.A., Linsenbardt, D.N., Shuttleworth, C.W., 2023. Repetitive spreading depolarization induces gene expression changes related to synaptic plasticity and neuroprotective pathways. Front. Cell. Neurosci. 17. https://doi.org/10.3389/fncel.2023.1292661

      Durchdewald, M., Angel, P., Hess, J., 2009. The transcription factor Fos: a Janus-type regulator in health and disease. Histol Histopathol 24, 1451–1461. https://doi.org/10.14670/HH-24.1451

      Fields, R.D., Eshete, F., Stevens, B., Itoh, K., 1997. Action potential-dependent regulation of gene expression: temporal specificity in ca2+, cAMP-responsive element-binding proteins, and mitogen-activated protein kinase signaling. J Neurosci 17, 7252–7266. https://doi.org/10.1523/JNEUROSCI.17-19-07252.1997

      Fleischmann, A., Hvalby, O., Jensen, V., Strekalova, T., Zacher, C., Layer, L.E., Kvello, A., Reschke, M., Spanagel, R., Sprengel, R., Wagner, E.F., Gass, P., 2003. Impaired Long-Term Memory and NR2AType NMDA Receptor-Dependent Synaptic Plasticity in Mice Lacking c-Fos in the CNS. J. Neurosci. 23, 9116–9122. https://doi.org/10.1523/JNEUROSCI.23-27-09116.2003

      Francis-Taylor, R., Ophel, G., Martin, D., Loo, C., 2020. The ictal EEG in ECT: A systematic review of the relationships between ictal features, ECT technique, seizure threshold and outcomes. Brain Stimulation 13, 1644–1654. https://doi.org/10.1016/j.brs.2020.09.009

      Giraldo-Chica, M., Woodward, N.D., 2017. Review of thalamocortical resting-state fMRI studies in schizophrenia. Schizophr Res 180, 58–63. https://doi.org/10.1016/j.schres.2016.08.005

      Har-Gil, H., Golgher, L., Israel, S., Kain, D., Cheshnovsky, O., Parnas, M., Blinder, P., 2018. PySight: plug and play photon counting for fast continuous volumetric intravital microscopy. Optica 5, 1104. https://doi.org/10.1364/OPTICA.5.001104

      Huels, E.R., Kafashan, M., Hickman, L.B., Ching, S., Lin, N., Lenze, E.J., Farber, N.B., Avidan, M.S., Hogan, R.E., Palanca, B.J.A., 2023. Central-positive complexes in ECT-induced seizures: Possible evidence for thalamocortical mechanisms. Clin Neurophysiol 146, 77–86. https://doi.org/10.1016/j.clinph.2022.11.015

      Kawai, R., Markman, T., Poddar, R., Ko, R., Fantana, A.L., Dhawale, A.K., Kampff, A.R., Ölveczky, B.P., 2015. Motor cortex is required for learning but not for executing a motor skill. Neuron 86, 800– 812. https://doi.org/10.1016/j.neuron.2015.03.024

      Keller, G.B., Sterzer, P., 2024. Predictive Processing: A Circuit Approach to Psychosis. Annu Rev Neurosci 47, 85–101. https://doi.org/10.1146/annurev-neuro-100223-121214

      Kimpo, R.R., Doupe, A.J., 1997. FOS Is Induced by Singing in Distinct Neuronal Populations in a Motor Network. Neuron 18, 315–325. https://doi.org/10.1016/S0896-6273(00)80271-8

      Klapoetke, N.C., Murata, Y., Kim, S.S., Pulver, S.R., Birdsey-Benson, A., Cho, Y.K., Morimoto, T.K., Chuong, A.S., Carpenter, E.J., Tian, Z., Wang, J., Xie, Y., Yan, Z., Zhang, Y., Chow, B.Y., Surek, B., Melkonian, M., Jayaraman, V., Constantine-Paton, M., Wong, G.K.-S., Boyden, E.S., 2014. Independent optical excitation of distinct neural populations. Nat Methods 11, 338–346. https://doi.org/10.1038/nmeth.2836

      Largo, C., Ibarz, J.M., Herreras, O., 1997. Effects of the Gliotoxin Fluorocitrate on Spreading Depression and Glial Membrane Potential in Rat Brain In Situ. Journal of Neurophysiology 78, 295–307. https://doi.org/10.1152/jn.1997.78.1.295

      Leao, A.A.P., 1944. Spreading depression of activity in the cerebral cortex. Journal of Neurophysiology 7, 359–390. https://doi.org/10.1152/jn.1944.7.6.359

      Mahringer, D., Petersen, A.V., Fiser, A., Okuno, H., Bito, H., Perrier, J.-F., Keller, G.B., 2019. Expression of c-Fos and Arc in hippocampal region CA1 marks neurons that exhibit learning-related activity changes. https://doi.org/10.1101/644526

      Mahringer, D., Zmarz, P., Okuno, H., Bito, H., Keller, G.B., 2022. Functional correlates of immediate early gene expression in mouse visual cortex. Peer Community Journal 2. https://doi.org/10.24072/pcjournal.156

      Meeren, H.K.M., Pijn, J.P.M., Luijtelaar, E.L.J.M.V., Coenen, A.M.L., Silva, F.H.L. da, 2002. Cortical Focus Drives Widespread Corticothalamic Networks during Spontaneous Absence Seizures in Rats. J. Neurosci. 22, 1480–1495. https://doi.org/10.1523/JNEUROSCI.22-04-01480.2002

      Minatohara, K., Akiyoshi, M., Okuno, H., 2015. Role of Immediate-Early Genes in Synaptic Plasticity and Neuronal Ensembles Underlying the Memory Trace. Front Mol Neurosci 8, 78. https://doi.org/10.3389/fnmol.2015.00078

      Morgan, J.I., Curran, T., 1991. Stimulus-transcription coupling in the nervous system: involvement of the inducible proto-oncogenes fos and jun. Annu Rev Neurosci 14, 421–451. https://doi.org/10.1146/annurev.ne.14.030191.002225

      Morinobu, S., Nibuya, M., Duman, R.S., 1995. Chronic antidepressant treatment down-regulates the induction of c-fos mRNA in response to acute stress in rat frontal cortex. Neuropsychopharmacology 12, 221–228. https://doi.org/10.1016/0893-133X(94)00067-A

      Nakadate, K., Imamura, K., Watanabe, Y., 2012. Effects of monocular deprivation on the spatial pattern of visually induced expression of c-Fos protein. Neuroscience 202, 17–28. https://doi.org/10.1016/j.neuroscience.2011.12.004

      Pandey, A., Kang, S., Pacchiarini, N., Wyszynska, H., Masseri, Z., O’Neill, J., Honey, R.C., Fox, K., 2026. Secondary Somatosensory Cortex Is Required for Learning but Not Execution of a Tactile Discrimination. Eur J Neurosci 63, e70390. https://doi.org/10.1111/ejn.70390

      Park, H.G., Yu, H.S., Park, S., Ahn, Y.M., Kim, Y.S., Kim, S.H., 2014. Repeated treatment with electroconvulsive seizures induces HDAC2 expression and down-regulation of NMDA receptorrelated genes through histone deacetylation in the rat frontal cortex. Int J Neuropsychopharmacol 17, 1487–1500. https://doi.org/10.1017/S1461145714000248

      Polack, P.-O., Mahon, S., Chavez, M., Charpier, S., 2009. Inactivation of the Somatosensory Cortex Prevents Paroxysmal Oscillations in Cortical and Related Thalamic Neurons in a Genetic Model of Absence Epilepsy. Cereb Cortex 19, 2078–2091. https://doi.org/10.1093/cercor/bhn237

      Rosenthal, Z.P., Majeski, J.B., Somarowthu, A., Quinn, D.K., Lindquist, B.E., Putt, M.E., Karaj, A., Favilla, C.G., Baker, W.B., Hosseini, G., Rodriguez, J.P., Cristancho, M.A., Sheline, Y.I., William Shuttleworth, C., Abbott, C.C., Yodh, A.G., Goldberg, E.M., 2025. Electroconvulsive therapy generates a postictal wave of spreading depolarization in mice and humans. Nat Commun 16, 4619. https://doi.org/10.1038/s41467-025-59900-1

      Roy, D.S., Arons, A., Mitchell, T.I., Pignatelli, M., Ryan, T.J., Tonegawa, S., 2016. Memory retrieval by activating engram cells in mouse models of early Alzheimer’s disease. Nature 531, 508–512. https://doi.org/10.1038/nature17172

      Ryan, T.J., Roy, D.S., Pignatelli, M., Arons, A., Tonegawa, S., 2015. Memory. Engram cells retain memory under retrograde amnesia. Science 348, 1007–1013. https://doi.org/10.1126/science.aaa5542

      Scangos, K.W., Weiner, R.D., Coffey, C.E., Krystal, A.D., 2019. An electrophysiological biomarker that may predict treatment response to ECT. J ECT 35, 95–102. https://doi.org/10.1097/YCT.0000000000000557

      Sheng, M., Greenberg, M.E., 1990. The regulation and function of c-fos and other immediate early genes in the nervous system. Neuron 4, 477–485. https://doi.org/10.1016/0896-6273(90)90106-p

      Sterzer, P., Adams, R.A., Fletcher, P., Frith, C., Lawrie, S.M., Muckli, L., Petrovic, P., Uhlhaas, P., Voss, M., Corlett, P.R., 2018. The Predictive Coding Account of Psychosis. Biol Psychiatry 84, 634–643. https://doi.org/10.1016/j.biopsych.2018.05.015

      Tanaka, K.Z., He, H., Tomar, A., Niisato, K., Huang, A.J.Y., McHugh, T.J., 2018. The hippocampal engram maps experience but not place. Science 361, 392–397. https://doi.org/10.1126/science.aat5397

      Tyssowski, K.M., DeStefino, N.R., Cho, J.-H., Dunn, C.J., Poston, R.G., Carty, C.E., Jones, R.D., Chang, S.M., Romeo, P., Wurzelmann, M.K., Ward, J.M., Andermann, M.L., Saha, R.N., Dudek, S.M., Gray, J.M., 2018. Different Neuronal Activity Patterns Induce Different Gene Expression Programs. Neuron 98, 530-546.e11. https://doi.org/10.1016/j.neuron.2018.04.001

      Vasilevskaya, A., Keller, G.B., 2026. A functional influence based circuit motif that constrains the set of plausible algorithms of cortical function. eLife 15. https://doi.org/10.7554/eLife.110827.1

      Vinogradov, S., Chafee, M.V., Lee, E., Morishita, H., 2023. Psychosis spectrum illnesses as disorders of prefrontal critical period plasticity. Neuropsychopharmacology 48, 168–185. https://doi.org/10.1038/s41386-022-01451-w

      Watanabe, Y., Johnson, R.S., Butler, L.S., Binder, D.K., Spiegelman, B.M., Papaioannou, V.E., McNamara, J.O., 1996. Null Mutation of c-fos Impairs Structural and Functional Plasticities in the Kindling Model of Epilepsy. J Neurosci 16, 3827–3836. https://doi.org/10.1523/JNEUROSCI.16-1203827.1996

      Wei, Z., Lin, B.-J., Chen, T.-W., Daie, K., Svoboda, K., Druckmann, S., 2020. A comparison of neuronal population dynamics measured with calcium imaging and electrophysiology. PLoS Comput Biol 16, e1008198. https://doi.org/10.1371/journal.pcbi.1008198

      Winston, S.M., Hayward, M.D., Nestler, E.J., Duman, R.S., 1990. Chronic electroconvulsive seizures down-regulate expression of the immediate-early genes c-fos and c-jun in rat cerebral cortex. J Neurochem 54, 1920–1925. https://doi.org/10.1111/j.1471-4159.1990.tb04892.x

      Xia, Z., Dudek, H., Miranti, C.K., Greenberg, M.E., 1996. Calcium influx via the NMDA receptor induces immediate early gene transcription by a MAP kinase/ERK-dependent mechanism. J Neurosci 16, 5425–5436. https://doi.org/10.1523/JNEUROSCI.16-17-05425.1996

      Yap, E.-L., Pettit, N.L., Davis, C.P., Nagy, M.A., Harmin, D.A., Golden, E., Dagliyan, O., Lin, C., Rudolph, S., Sharma, N., Griffith, E.C., Harvey, C.D., Greenberg, M.E., 2021. Bidirectional perisomatic inhibitory plasticity of a Fos neuronal network. Nature 590, 115–121. https://doi.org/10.1038/s41586-020-3031-0

      Yassin, L., Benedetti, B.L., Jouhanneau, J.-S., Wen, J.A., Poulet, J.F.A., Barth, A.L., 2010. An embedded subnetwork of highly active neurons in the neocortex. Neuron 68, 1043–1050. https://doi.org/10.1016/j.neuron.2010.11.029

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      The "multiple-demand" (MD) system is a well-known finding of human brain imaging and is thought to play a central role in cognitive control. To directly compare the MD system in humans and monkeys, Mione et al. used functional magnetic resonance imaging to measure whole-brain activation in a multi-step saccadic maze task. In humans, the authors found a distributed pattern of brain activity close match to the canonical MD network and extends to adjacent regions of dorsal attention and other networks. While there was good correspondence between monkey and human data, differences were also notable in the lateral frontal cortex, the dorsal parietal cortex, and the sensorimotor cortex.

      Strengths:

      Though previous data hint at a corresponding network in the macaque, there has been no direct comparison to human data. This study provides a direct cross-species comparison with whole-brain data from fMRI, and the findings suggest an extended and strongly interconnected brain network recruited by increased cognitive challenge.

      Weaknesses:

      In previous human imaging, the MD system is defined by overlapping activation for many kinds of cognitive demands. In the present work, however, the authors used just a single task. Although there is some overlap between the putative monkey MD network and the canonical MD network identified in human imaging, there should be caution in linking current findings to the MD system based on limited task events.

      In the Discussion, we acknowledge the limitation of using a single task. With this one task, however, the canonical MD network is clearly shown in our human data. Accompanying activation, especially of the canonical dorsal attention network, likely reflects the specific spatial demands of the maze task. In the monkey data, much of this dorsal attention activity is not seen (e.g. superior and medial parietal), likely reflecting limited power. Instead, there is distributed overlap with the previous limited monkey studies than have contrasted higher with lower cognitive demand. Though we agree that task-specific activations may contribute to our results, these arguments suggest that, in large part, our method does successfully identify a distributed set of multiple-demand regions. At the same time we acknowledge the desirability of further work to examine a wider range of task demands.

      Reviewer #1 (Recommendations for the authors):

      (1) Though the whole-brain data obtained by fMRI can provide a direct comparison between species, a single cognitive task might be insufficient to link the findings with the MD system. A cognitively challenging task likely activates multiple regions of the MD system; however, it may also recruit some task-specific regions, which do not belong to the canonical MD network. Furthermore, this is probably the reason that the dorsal attention network showed the strongest activation rather than the core MD for the current visuospatial maze task.

      In the Discussion, we acknowledge the limitation of using a single task (p. 15-16), and the likely contribution of task-specific activations to our data, especially involving the dorsal attention network (p. 15).

      (2) Ideally, a meta-analysis recruiting more fMRI studies on humans and monkeys when they perform various similar cognitive tasks may strengthen the evidence that there is a comparable MD system between species.

      Many human meta-analyses, of course, show the common MD system. For monkeys, however, at least to our knowledge, there are insufficient studies for a similar meta-analysis. Instead we discuss overlaps between the current activation findings and two previous studies of respectively antisaccades and task switching (p. 15), suggesting that, in monkey as in human, there is convergence for different kinds of demand.

      (3) Different from human subjects, monkeys usually require substantial training before fMRI scanning. More details about the training of the two monkeys should be given. In addition, training may reshape cognitive task activations. This potential impact should also be discussed.

      We now address this point in the Discussion (p. 18). As we note, similar results for the two species apparently survive even large differences in protocol. A training summary has been added to Methods (p. 27).

      (4) According to the description in the text, the two monkeys have obvious differences in cognitive task activations. It is necessary to show individual-level brain activations as well.

      Individual results and a conjunction map are shown in Supplementary Figure 6 (see accompanying text on p. 13). As expected, the conjunction map had substantial similarity to the findings from the two animals combined. Similarities and differences between animals are addressed in the Discussion (p. 17).

      (5) Based on the current research content, the title seems to be too general.

      For the reasons given above (see point 2), we think our title is reasonable.

      Reviewer #2 (Public review):

      Summary:

      Mione et al. aim to resolve a long-standing question in comparative neuroscience: whether the macaque brain contains a functional analogue to the distributed human multiple demand (MD) network. To address this, the authors employ a direct cross-species fMRI comparison using a multi-step saccadic maze task in humans and a simplified two-step version in macaques. By contrasting goal-directed navigation against a control condition that requires similar motor responses but no strategic planning, the study isolates the neural signatures of cognitive control across species.

      Strengths:

      The most compelling aspect of this work is its methodological alignment. Previous attempts to compare these systems often relied on comparisons of human BOLD signals and macaque single-unit recordings. By running parallel fMRI protocols, the authors establish a shared measurement basis that allows for a more direct comparison. The resulting activation maps clearly demonstrate conserved network topology across dorsomedial frontal, lateral, and medial parietal, and insula cortices. Combining these results with recent research on functional and structural connectivity further supports the idea that these networks evolved across species and provides a helpful starting point for future comparative studies. The findings will be highly useful for researchers investigating the evolutionary origins of domain-general cognitive control, as well as for neuroimaging methodologists developing cross-species alignment pipelines.

      Weaknesses:

      However, there are several differences in how the two groups were studied that make it harder to compare the results precisely. The human task mixed 2-, 4-, and 6-step trials within the same experimental blocks, whereas macaques performed only 2-step trials. This design difference likely places human participants in a state of sustained proactive cognitive control (Braver, 2012), as they must remain prepared for highly demanding trials at any moment. This elevated baseline arousal may artificially inflate MD network activation during the simpler 2-step trials in humans, making direct magnitude comparisons with the macaque data difficult.

      This is a reasonable concern, which we note in the Discussion. Crucially, as we point out, similarities between species appear to survive this and other differences in procedure.

      Additionally, the general linear model combined correct and error trials into a single regressor. Given that macaques exhibited substantially higher error rates, this approach risks diluting task-specific planning signals with activity related to error monitoring and reward prediction errors. The preprocessing pipeline also applied a 4 mm full-width half-maximum smoothing kernel to macaque data acquired at 1.5 mm resolution. Relative to the smaller size of the macaque brain, this kernel is quite large and likely blurs fine-grained topographical distinctions. This may partly explain why the macaque lateral frontal cortex shows a single dorsal activation patch rather than multiple discrete patches seen in humans.

      These are also reasonable concerns. To address them, we have added supplementary analyses using (a) only correct trials, and (b) a smaller smoothing kernel (see Supplementary Figure 3 and 4). In both cases, there is some loss of power, but otherwise similar results.

      Furthermore, there is concerning inter-individual variability in the macaque data. Normally, a functional network like the MD system is identified by consistent activation across all individuals. In this study, however, the two monkeys show substantially different activation maps and behavioral patterns. This lack of consistency renders the group-level results questionable, as it is unclear whether the group-level map represents a unified biological system or merely an average of disparate individual maps.

      Results for individual animals have now been supplemented with a conjunction map (Supplementary Figure 6). As expected, the conjunction map is similar to the map obtained by pooling data from the two animals.

      Finally, the subcortical activations shown in Figure 7 require more precise anatomical localization to confidently distinguish cerebellar nodes from adjacent brainstem structures.

      The slices shown in Figure 7 have been amended to better show the detail of activation outside cerebral cortex, in particular in cerebellum.

      The authors demonstrate a broad functional correspondence between human and macaque cognitive control networks, moving the field beyond speculative homology. The data suggest that an extended, interconnected network is recruited by cognitive challenge in both species; however, the strength of this claim is limited by the inter-individual variability and methodological constraints noted above. Assertions of precise topological equivalence should therefore be tempered. The absence of ventrolateral prefrontal and strong dorsal parietal activations in the macaque group analysis may reflect genuine biological differences, but could also stem from limited statistical power, excessive smoothing, or task design asymmetries. While the overall conclusions are plausible, they would be significantly strengthened by a more explicit discussion of these limitations and additional analytical clarifications regarding individual-level consistency.

      Indeed, as we note in the Discussion, there is a strong possibility that ventrolateral frontal and dorsal parietal activation are missing in our monkey data because of limited statistical power. Further work would be needed to address these possible limitations.

      Reviewer #2 (Recommendations for the authors):

      (1) Please discuss how the mixed-difficulty block design in humans may induce a state of proactive cognitive control that elevates baseline MD activation compared to the fixed 2step macaque condition. A supplementary analysis comparing early versus late block 2-step trials in humans would help clarify whether activation magnitudes reflect sustained task-set maintenance or transient trial demands.

      We now note the potential importance of this in the Discussion (p. 18), and point out that similarities between species appear to survive this difference in procedure.

      (2) Please clarify the rationale for combining correct and error trials in a single regressor.

      Given the higher macaque error rates, please provide a supplementary GLM restricted to correct trials only for the monkey data. This will demonstrate whether the core activation topography remains consistent when error-related signals are excluded.

      A new analysis addresses this point (Supplementary Figure 3). As we note, restricting analysis to correctly-completed problems somewhat reduces power but leaves major features of the results intact.

      (3) Please justify the use of a 4 mm FWHM smoothing kernel for macaque data, given cortical thickness and brain size differences. If feasible, re-run the analysis with a smaller kernel or surface-based smoothing to assess whether finer topographical distinctions emerge in lateral frontal and parietal cortices.

      This new analysis has also been run (Supplementary Figure 4), again with reduced power but major features of the results intact. We also refer to previous work indicating choice of either 3 mm or 4 mm smoothing for macaque fMRI (p. 13), approximately matching the smoothing needed to align electrophysiological and fMRI maps (Issa et al., 2013, J.Neurosci.).

      (4) Please detail how the human '2-step problems only' analysis in Supplementary Figure 1 was specified in the GLM. Explicitly state whether 4- and 6-step trials were modeled as separate regressors of no interest to prevent hemodynamic bleed-over from contaminating the 2-step beta weights.

      Indeed, 4- and 6-step trials were removed using regressors of no interest, as now specified in Methods (p. 24).

      (5) Please provide higher-resolution slices or probabilistic atlas overlays for the cerebellar and subcortical activations in Figure 7. This will help clearly distinguish cerebellar hemispheres from adjacent brainstem structures and ensure anatomical labeling is accurate.

      To address this question, additional slices have been added to a revised Figure 7, in particular adding detail to cerebellar activation.

      (6) Please address the substantial inter-individual variability in the macaque data, particularly regarding Monkey B's performance. Even after excluding five poor-performing sessions, Monkey B only reached about 65% accuracy on the first step of the maze task. This suggests that the animal was largely guessing rather than following a strategic, goal directed plan. The notably higher accuracy on step 2 (>80%) could simply be a selection bias artifact, as the task terminates immediately following step 1 errors, meaning step 2 trials only occur when step 1 was already correct. Including an animal that likely did not fully grasp the overarching task rule in a cohort of N=2 raises serious concerns about signal dilution and the reliability of the group-level maps. Please explicitly justify why Monkey B's data were retained despite these performance concerns, or consider excluding this animal and acquiring data from a third, behaviorally reliable subject. At minimum, report a conjunction map showing regions strictly active in both animals alongside the combined analysis, and present individual subject maps more prominently to improve transparency.

      Though we appreciate this concern, we do not think the data indicate that monkey B failed to use the maze goal to constrain choices. As shown in Figure 3, the great majority of errors were timing errors (mostly not waiting for go signal). Excluding these, for step 1, of choices directed to one of the two available alternatives, 88% were correct (Figure 4, compare “correct” with “wrong open location”).

      As noted above, we have added a conjunction map (Supplementary Figure 6) to our previous presentation of individual data. We note (p. 13) that, as expected, this map strongly resembles results from the combined-animal analysis.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

      Weaknesses:

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects.

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single-cell labelling”, which admittedly limits precision.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function.

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an effect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately.

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question but a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative.

      Indeed, we do not reach single-cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content.

      We also think that combining our transgenic line with dopamine expression was sufficient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism.

      Thank you for this encouraging comment.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      Comments:

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect:

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/ or motor learning in motor neurons https://pubmed.ncbi.nlm.nih.gov/38779314/ or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest.

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal thresholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal-directed navigation, respectively."

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors claim that several dop-positive cell types resemble cell types described in Drosophila - to be able to better compare the two, it would be great to have the supplement and main figure pictures in one figure.

      We have added Fig. 11 from the main text to suppl. Fig. 3 for comparison

      (2) In Figure 11 and the corresponding supplement, it would be great to have some landmarks to better understand the expression patterns.

      In the legend, we have now referred to Figs. 3, 4 and 5 for depictions of these neurons within the neuropil reconstructions

      (3) For a better understanding of neurotransmitter expression, is it possible to figure out if dopamine is rather coexpressed with Glut- or ChaT-positive neurons?

      Very interesting idea. Unfortunately, the first author of the study has graduated and left the lab, such that we are unable to add this piece of information.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

      We thank the referees and the editors for their efforts and their positive feedback. As requested, we clarified throughout our manuscript that our proposal of an Orco-centered post-transcriptionally controlled feedback loop in the plasma membrane, which we term PTFL clock, is a novel hypothesis that needs to be tested in future experiments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study combines a five-year field experiment with a meta-analysis to quantify the effects of grazing on ecosystem CO<sub>2</sub> fluxes as modified by wetness. The results of this study are potentially valuable, but the Methods description is incomplete and compromises reproducibility and interpretability. A major caveat of this study is that the assessment of CO<sub>2</sub> fluxes in time and space is not complete.

      We appreciate the comments. The incomplete assessment of CO<sub>2</sub> fluxes in time and space has been added in the Limitations and implications for future study Section in the revised manuscript. Method description has been revised as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      Lines 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      Lines 468-470, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that was not monitored.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study integrates long-term (7th-11th year) ecosystem CO<sub>2</sub> flux measurements from a continuously grazed typical steppe in Inner Mongolia with a global meta-analysis of grazing experiments (585 observations) to systematically investigate how the wetness index modulates the effects of grazing intensity on net ecosystem productivity (NEP) in grasslands.

      Key findings include:

      (1) Heavy grazing significantly reduced gross primary productivity (GPP) and ecosystem respiration (ER) in the typical steppe but did not significantly affect NEP.

      (2) Globally, grazing significantly reduced GPP, ER, and NEP, and a higher wetness index and aboveground biomass (AGB) enhance the positive response of NEP to grazing.

      (3) Under light and moderate grazing, the response of NEP to grazing was positively correlated with the wetness index, a relationship potentially mediated by a higher plant relative growth rate (RGR) in wetter years.

      This study holds considerable practical significance. In the context of global change, comprehending the regulatory function of water is of great importance for the adaptive management of grasslands.

      Strengths:

      (1) The study cleverly combines long-term in situ observations (revealing temporal dynamics and potential mechanisms) with a global meta-analysis (testing the generality of patterns), forming a complete evidence chain from "process understanding" to "pattern verification".

      (2) Focusing on the 7th to 11th years of grazing treatments avoids the common "initial disturbance effects" observed in short-term grazing experiments and truly captures the steady-state response of the ecosystem after it has reached a new equilibrium.

      (3) By analyzing plant relative growth rate (RGR) and aboveground biomass (AGB), this study provides mechanistic clues regarding how the wetness index modulates grazing effects. The finding that plant compensatory growth is enhanced in wetter years is a crucial pathway explaining the variation in the NEP response.

      (4) The meta-analysis systematically integrates published literature with a substantial sample size (585 observations) and broad geographical coverage (Figure 1B).

      (5) The finding that light grazing promotes carbon sinks in wet years, while heavy grazing reduces productivity even in wet years, has direct implications for adaptive grassland management under climate change.

      Thank you for the comments. We would like to express our gratitude for your positive and insightful evaluation of this manuscript.

      Weaknesses:

      Although the paper does have strengths in principle, there are weaknesses and areas for improvement in the paper. In particular:

      (1) Incomplete mechanistic chain: While RGR and AGB data offer valuable clues for mechanistic interpretation, the causal chain from "increased wetness → higher RGR → maintained NEP" remains incomplete. Key intermediate processes, such as soil moisture dynamics, nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, are absent, making the mechanistic explanation somewhat speculative.

      Thank you for the comments. The data of the soil moisture dynamics, composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain. However, the structural equation modeling could not be set up if C<sub>3</sub> and C<sub>4</sub> plant communities were included (Fig. S8). Leaf photosynthetic physiological parameters were not collected in this study. We have made following changes in the revised manuscript:

      Lines 339-341, page 11: “In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Lines 472-474, page 15: “Furthermore, key intermediate processes, such as soil nutrient availability, changes in community composition, and leaf photosynthetic physiological parameters, should also be investigated in the future.”

      (2) Inadequate exploration of heterogeneity: The global meta-analysis reveals significant effect heterogeneity. Currently, only a few moderators, such as wetness index, precipitation, and temperature, are analyzed. Potential moderators, including grazing history, grassland type (typical steppe/alpine meadow/savanna), livestock type (cattle/sheep/mixed), and soil type, are not adequately explored.

      Thank you for the comments. Grazing duration was presented in Fig. S13. The heterogeneity analysis of grassland types and livestock types was supplemented as Fig. S11 and Fig. S12 in the Supplementary documents. We have also made following changes in the revised manuscript:

      Lines 367-370, page 12: “Grazing decreased GPP, ER and NEP in desert and temperate grassland, but had no significant effect on GPP, ER and NEP in alpine grassland and savanna (Fig. S11). In addition, cattle and sheep grazing decreased GPP and NEP, but livestock mixed grazing did not show significant effect on GPP, ER and NEP (Fig. S12).”

      (3) Integration of long-term experiment and meta-analysis could be tighter.

      Thank you for the comments. Integration of long-term experiment and meta-analysis has been revised in the revised manuscript as follows:

      Lines 449-463, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term grazing field experiments revealed the mechanisms underlying the effects of grazing on ecosystem CO<sub>2</sub> fluxes. Global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analyses consistently revealed that wetness was an important factor in regulating the ecosystem CO<sub>2</sub> fluxes in response to grazing. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP (Morgan et al., 2016; Owensby et al., 2006). Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness modulated the effects of grazing on NEP in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      Reviewer #2 (Public review):

      Overgrazing in grasslands is a widespread, enormous problem, and its consequences need more research attention.

      However, I have some concerns about the methods in this manuscript.

      (1) The plots were grazed from May to September. For the CO<sub>2</sub> measurements, it states growing season, and that data was collected on sunny days from 9 to 11 am. However, annual values are reported, and no details on how these were calculated. With no measurements outside of 9 to 11 am, and only on non-sunny days, and only during the growing season, I question how accurate the annual numbers are.

      Thank you for the comments. The ecosystem CO<sub>2</sub> fluxes were measured during growing season; the description of annual CO<sub>2</sub> have been revised to “growing-season” CO<sub>2</sub>. The ecosystem CO<sub>2</sub> fluxes were measured on sunny days during the growing season (May-September) between 9:00 and 11:00 a.m. It can represent typical daytime carbon flux dynamics during the growing season (Fan et al., 2011; Li et al., 2010). In addition, previous studies showed that the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. represented the average ecosystem CO<sub>2</sub> fluxes of the day (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), so we use the ecosystem CO<sub>2</sub> fluxes measured at 9:00 and 11:00 a.m. to calculate the total ecosystem CO<sub>2</sub> fluxes of the day. We have also revised it in the limitation for future study section. We have made following changes in the revised manuscript:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      Lines 468-472, page 15: “The assessment of ecosystem CO<sub>2</sub> fluxes was incomplete due to limited measurement time, area and sampling intervals. This may bias the representation of seasonal cycles and spatial heterogeneity, especially in the regions that were not monitored. Multiple time-point sampling should be adopted; the sampling frequency and points should also be increased in the future field study.”

      (2) This study uses one chamber, 0.5 by 0.5 m, for a total area of 0.25 m<sup>2</sup>. This is pretty small, and with no replication, I would like to see more detail on how these are placed, especially when one of the dominant species is a bunchgrass. Sampling on or off a bunchgrass will give very different readings, and neither will be representative of the entire plot. The same for the soil temp and water measurements, there is no detail on how many replicates are in each plot, and working randomly does not work when there are spatial patterns caused by the bunchgrasses in the plots.

      Thank you for the comments. Two replicates of the chambers in each plot were set, and their specific sites were labeled in Fig. S2. The study area was located in a typical steppe dominated by Stipa grandis and Leymus chinensis. Therefore, we selected quadrats mainly containing these two species. We have made following changes in the revised manuscript:

      Line 186-187, page 6: “In each plot, the chambers of two replicates were measured.”

      (3) MAP and MAT are important in this study, and no mention is given if this data is from this site or a neighboring site, and if so, what the distance to this location is.

      Thank you for the comments. The data of MAP and MAT were provided in the revised manuscript as follows:

      Lines 193-195, page 7: “The precipitation and air temperature were collected from the Xilinhot Meteorological Observation Station, which was near the experimental site in this field experiment.”

      (4) Vegetation was sampled in five locations in each plot. I note no detail on the belowground biomass, how many reps, what area, or which depth? This essential data is missing. The same goes for the RGR, and here the authors talk about grazing events. Whereas earlier in section 2.2, it implies continuous grazing every day during the growing season. RGR needs more details, such as how many times, its replication, etc.

      Thank you for the comments. The detailed method for how to collect belowground biomass and how to calculate the RGR has been supplemented in the revised manuscript as follows:

      Lines 205-215, page 7: “Four grazing rotations were conducted each year from 2019 to 2023 in the field study. Before each grazing rotation, two cages (1.2 m × 1.2 m × 1.2 m) were installed in each plot to ensure that the vegetation inside would not be foraged by sheep. Aboveground plant biomass was collected after the end of each grazing rotation. Specifically, in each plot, five 1 m × 1 m quadrats were placed. All aboveground biomass within these quadrats was clipped at ground level and then oven-dried at 65 °C for 48 h to determine the community-level dry biomass. The plant species were classified into C<sub>3</sub> and C<sub>4</sub> groups (Table S1). Belowground biomass (BGB) was quantified by collecting root biomass using soil core with diameter of 7 cm each September from 2019 to 2023 (Fig. S2). Specifically, two soil cores were collected from each 1 m × 1 m quadrat corresponding to aboveground biomass measurements at depths of 0-30 cm. The belowground parts were then extracted from the soil by washing with water and oven-dried at 65 °C for 48 h to obtain dry weight.”

      (5) Hypothesis 2 is vague: "act as key factors". A more precise hypothesis would be better, otherwise it just leads to p-hacking.

      Thank you for the comments. The misleading description has been deleted.

      (6) More generally, neither hypothesis is really based on the introduction, as the authors report mixed results in the literature. Hence, the more specific hypothesis 1 reads like it is HARKED, based on the results that the authors found. This is not a good research practice. See the following papers on this topic:

      Murphy & Aquinis. 2019. HARKing: How bad can cherry-picking and question trolling produce bias in published results? Journal of Business and Psychology 34:1-17

      Bishop. 2019 Rein in the four horses of irreproducibility. Nature 568: 435

      Fraser et al. 2018. Questionable research practices in ecology and evolution. Plos One 13:7

      Parker et al. 2016. Transparency in ecology and evolution: real problems, real solutions. Trends in ecology and evolution 31:9

      Thank you for the suggestion. The hypotheses have been deleted in the revised manuscript to avoid HARKED mistakes.

      (7) Figure 2 mentioned n=3, which is correct. However, the variance around the means in the figures is tiny, which raises questions about the replication that is actually used here. I would like to see the entire statistics tables, including DF and sample sizes, included as an appendix, so that the reader can evaluate this much more.

      Thank you for the comments. In this study, the data presented in Fig.2 were expressed as mean± standard error (SE):

      In addition, we have provided the complete statistical table in Supplementary Table S2 according to the suggestion of the reviewer.

      (8) Figure 3, now only low, medium, and high grazing are reported as changes from the control. This can be misleading as the reader can't see how the control varies along the various gradients. I suggest including a figure with the raw data for each as an appendix. In addition, the sample size in 3d, e, f is much higher, and it looks to me like the authors used both the reps and years together. This is not good practice. In addition, full statistics tables should be included in the appendix.

      Thank you for the comments. The relationships between ecosystem CO<sub>2</sub> fluxes and wetness index have been supplemented in Fig. S6. We have also provided the complete statistical table in Supplementary Table S3.

      (9) In Table S1, since there are already a number of recent meta-analyses on this topic, I think there needs to be a stronger justification for this one.

      Thank you for the comments. The results of our five-year field monitoring experiments showed that carbon fluxes and their components were significantly influenced by the wetness index. Therefore, we explored whether such effects also occur in grazing experiments at the global scale. The results demonstrated that similar patterns indeed exist worldwide, indicating that our meta-analysis is meaningful and necessary. In addition, previous studies have addressed related topics, net ecosystem productivity has mostly been treated as an auxiliary variable rather than the primary focus of investigation (Jiang et al., 2020; Shi et al., 2022; Zhang et al., 2022; Zhou et al., 2019). In previous studies, meta-analysis literatures on grazing and NEP were limited and not adequate. Our meta-analysis is more comprehensive and reflects the reliability of the results. We have made following changes in the revised manuscript:

      Lines 491-492, page 16: “In previous studies, meta-analysis literatures of effects of grazing intensities on NEP were limited and not adequate.”

      Lines 512-514, page 16: “In summary, the meta-analysis of this study presents the first comprehensive assessment of how annual wetness index affects the response of ecosystem CO<sub>2</sub> fluxes to grazing across global grasslands.”

      (10) Regarding the grazing-induced CO<sub>2</sub> fluxes, these are only based on the plants and do not incorporate the animal CO<sub>2</sub> flux, nor the animal litter CO<sub>2</sub> flux, as they were penned at night outside the plots. Thus, the grazing impact is inflated in the data reported here and does not really represent GPP, NET, or RE. This needs to be written about in the discussion section.

      Thank you for the comments. The objective of this study was to evaluate the carbon exchange processes from vegetation and soil under grazing disturbance. Thus, the grazing effects reported in this study were the vegetation-soil ecosystem scale rather than the net carbon balance of the entire grazing system. Consequently, our conclusions regarding the effects of grazing on grassland ecosystem carbon exchange remain reliable and ecologically meaningful.

      We have made following changes in the manuscript:

      Lines 483-490, page 16: “There are also limitations in evaluating the effects of grazing on grassland-livestock ecosystem CO<sub>2</sub> fluxes in this study. Specifically, carbon emissions derived from animal respiration were not included in this study. Therefore, the results of the field experiment and meta-analysis in this study should not be interpreted as a complete carbon budget assessment of the grassland-livestock ecosystem. Our primary objective was to evaluate the carbon exchange processes between vegetation and soil under grazing disturbance. Thus, the grazing effects reported in the field experiment and meta-analysis of this study were responses at the vegetation- soil ecosystem scale rather than the net carbon balance of the entire grazing system.”

      Reviewer #3 (Public review):

      Combining a five-year field experiment with a global meta-analysis, Wu et al. investigate how grazing intensity influences ecosystem carbon dioxide (CO<sub>2</sub>) fluxes in grasslands and how these effects are regulated by environmental conditions such as grazing duration, wetness index, and soil temperature and moisture responses.

      The authors show that the response of net ecosystem productivity (NEP) to light grazing shifts from negative to positive along a wetness gradient, whereas heavy grazing consistently suppresses NEP across wetness conditions. Importantly, this pattern is supported by both the field experiment and the meta-analysis, suggesting that moderate levels of grazing can potentially enhance both plant productivity and carbon sequestration under favorable moisture conditions.

      The integration of experimental data with a global synthesis is a particular strength of the study, allowing the authors to evaluate grazing impacts across both temporal variability (precipitation fluctuations in the field experiment) and spatial variability (wetness gradients across global grasslands). Overall, the conclusions are well supported by the data.

      However, several aspects of data acquisition, analysis, and presentation could be clarified to further strengthen the reproducibility and interpretation of the results:

      (1) A PRISMA-style flow diagram would be helpful for the meta-analysis to clearly illustrate the study selection process and facilitate interpretation of the dataset.

      Thank you for the comments. We have added the PRISMA-style flow diagram in the revised supplementary materials. See Fig. S9.

      (2) Grazing intensity requires a clearer definition and, where possible, standardization between the field experiment and the studies included in the meta-analysis. Key parameters such as the number of stock per area, days per rotation or per year, and total years of grazing should be clearly defined. In addition, the criteria used to classify grazing intensity into LG, MG, and HG in the meta-analysis should be explicitly described.

      Thank you for the comments. In our meta-analysis, the classifications of grazing intensity into light, moderate, and heavy grazing were provided in previous studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the USDA criteria (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf).

      The criteria was that: light grazing: approximately equal to a maximum of 40% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; moderate grazing: approximately equal to a maximum of 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season; Heavy grazing: greater than 50% Utilization (grazing and trampling) of forage standing crop (current and previous years’ growth) at the end of the growing season (November 15).

      We have made following changes in the manuscript:

      Lines 271-277, page 9 “(e) The classifications of grazing intensity (light, moderate, and heavy grazing) were primarily based on the definitions provided in the original studies. For studies in which grazing intensity was not explicitly reported, we classified grazing intensity according to the modified Grazing Intensity Classes proposed by the USDA (https://www.fs.usda.gov/Internet/FSE_DOCUMENTS/stelprdb5109714.pdf) (Yin et al., 2023).”

      (3) In several sections of the manuscript, it is difficult to distinguish whether phrases such as "this study" or "our study" refer specifically to the field experiment or to the overall study, including both the experiment and the meta-analysis. Clearer wording distinguishing these components would improve readability.

      Thanks for the suggestion. The specific distinctions have been revised to clarify whether they refer to the field experiment, meta-analysis or the comprehensive conclusions of the results from both field experiment and meta-analysis.

      Lines 50-53, page 2: “Overall, the meta-analysis and field experiment jointly provide global perspectives on the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity and improve our knowledge of the factors influencing the response of ecosystem CO<sub>2</sub> fluxes to grazing intensity.”

      Lines 191-193, page 6: “The annual precipitation and mean annual air temperature (MAT) for the field study area from 2019 to 2023 were obtained from the China Meteorological Data Service Centre (http://data.cma.cn/).”

      (4) The discussion attributes the non-significant annual NEP response to intra-annual precipitation variability, with grazing enhancing NEP under wet conditions but suppressing it under dry conditions. Another potential explanation may be that grazing affects GPP and ecosystem respiration (ER) at similar magnitudes (i.e., RR(GPP) ≈ RR(ER)), resulting in limited net changes in NEP.

      Thank you for the comments. Another potential explanation has been revised in the manuscript as follows:

      Lines 385-389, page 13: “Grazing decreased the responses of ecosystem CO<sub>2</sub> fluxes and plant biomass in global grasslands, but only NEP was not significantly affected by grazing in our field experiment (Fig. 4). The lack of a significant response in NEP may be because the site in this field study was managed for year-round continuous low-intensity grazing (Liang et al., 2021). In addition, grazing affected GPP and ER at similar magnitudes, resulting in limited net changes in NEP.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see from the above reviews and the specific recommendations below, the first issue you should resolve is a full description of methodological detail such that readers can, in principle, repeat your study. They have to know how you placed the chamber to have a representative measure of vegetation (averaging between high and low biomass patches). They also must know how to calculate grazing intensity and assign the values to the three classes. They want to see your raw data and statistical analyses (including formulae, replication, and degrees of freedom). Consider non-linear relationships of grazing and wetness. Try to better integrate the results from the experiment and the meta-analysis. Finally, make sure that you develop hypotheses from the prior knowledge presented in the introduction. For instance, instead of stating that NEP responses would shift from negative to positive with increasing wetness index, simply hypothesize that wetness mitigates the negative effects of grazing on CO<sub>2</sub> fluxes, even if you find that this is not true under heavy grazing.

      Thank you for the comments. The description of methodological details has been supplemented in the revised manuscript. The results from the experiment and the meta-analysis have been integrated.

      Reviewer #1 (Recommendations for the authors):

      General suggestions:

      (1) The Introduction section and the assumptions should be rewritten and improved. The logicality of the introduction should be revised to better prioritize and contextualize the research problem. The reader is lost since the links between assumptions and previous knowledge are not clear.

      Thanks for the suggestion. The Introduction section and the assumptions have been rewritten as follows:

      Lines 135-147, page 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable? In this study, we investigated the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes by combining a long-term (7-11 years) field experiment conducted in a typical steppe and a meta-analysis of global grasslands. The objectives of this study were to: (i) investigate the effects of grazing intensity with annual wetness fluctuations on ecosystem CO<sub>2</sub> fluxes (GPP, ER and NEP) covering the 7th to 11th years of a continuous grazing experiment in a typical steppe as well as the meta-analysis in global grasslands; (ii) explore how environmental factors (particularly wetness index, soil moisture and temperature, and grazing intensity) regulate the effects of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (2) In the "Materials and methods" section, you need to provide the reason why you chose the "wetness index" in this study, rather than other drought indices (such as Standardized Precipitation Evapotranspiration Index (SPEI), Aridity Index (AI)).

      Thank you for the comments. Although the standardized precipitation evapotranspiration index (SPEI) would be a more appropriate indicator for this study, its calculation requires relatively long and continuous climate data series, which were difficult to obtain in our global meta-analysis. This limitation was particularly important because our study also included a meta-analysis, for which complete climatic datasets were often unavailable from the collected literature. In contrast, the Aridity Index (AI) cannot adequately reflect interannual variability. We have supplemented this limitation in the revised as follows:

      Lines 482-483, page 15: “Furthermore, more drought or wetness indices should be investigated in future studies of grazing on ecosystem CO<sub>2</sub> fluxes.”

      (3) For the results and discussions, the results are currently presented in parallel (long-term experiment first, then meta-analysis). It is recommended to add a dedicated integration paragraph in the discussion.

      Thank you for the comments. The dedicated integration paragraph has been supplemented in the revised manuscript as follows:

      Lines 450-464, page 15: “Our long-term experiment and global meta-analysis jointly provided complementary evidence for understanding the effect of grazing on the ecosystem CO<sub>2</sub> fluxes in grassland ecosystems. Long-term experiments revealed the mechanisms underlying the effects of grazing on plant characteristics and ecosystem CO<sub>2</sub> fluxes under control conditions. And global meta-analysis further clarified the universality of these response patterns across different climatic zones and grassland types. It should be noticed that both field experiment and meta-analysis consistently revealed that wetness was an important factor in regulating the effect of ecosystem CO<sub>2</sub> fluxes. Under wetter conditions, light grazing promoted compensatory plant growth, enhanced leaf turnover and photosynthetic recovery capabilities, thereby maintaining or even enhancing the NEP. Conversely, under drier conditions or heavy grazing pressure, a decrease in aboveground biomass and weakened ecosystem resilience jointly limited the ecosystem carbon uptake capacity. These results suggested that wetness determined whether grazing promoted or inhibited the carbon sink function in grasslands. Moderate grazing may be sustainable under better water conditions, while heavy grazing may exceed the ecosystem resilience threshold even under relatively wet conditions, leading to sustained ecological degradation.”

      (4) Make sure that the whole manuscript has undergone professional proofreading, and check it carefully to avoid language mistakes.

      Thank you for the comments. The manuscript has been revised by professional proofreading to avoid language mistakes.

      Specific suggestions:

      (1) Lines 126-131: It is recommended to explicitly state three levels of research questions: (i) How does grazing intensity affect ecosystem CO<sub>2</sub> fluxes? (ii) Does the wetness index modulate this effect? (iii) Are these relationships globally generalizable? This will provide a clear logical thread for the paper.

      Thanks for the suggestion. We have made following changes in the revised manuscript:

      Lines 135-139, pages 5: “Previous studies on the effects of grazing intensities on ecosystem CO<sub>2</sub> fluxes have been constrained by limited spatial and temporal scales, which has led to an incomplete understanding of how different grazing intensities influence ecosystem CO<sub>2</sub> fluxes. Furthermore, does the wetness modulate the effect of grazing on ecosystem CO<sub>2</sub> fluxes in the typical steppe? Are these relationships globally generalizable?”

      (2) Lines 181: Briefly justify the choice of the De Martonne wetness index (Equation 1) in the introduction (why this index over other aridity indices).

      Thank you for the comments. The De Martonne wetness index is easier to obtain compared to other drought index, as it only requires annual average temperature and precipitation data for calculation, facilitating the statistics of global meta-analysis. The justification of the De Martonne wetness index has been revised in the manuscript:

      Lines 120-125, pages 4-5: “Annual precipitation is one of the climatic parameters, while the wetness index (WI) serves as a more integrative climatic indicator that incorporates both precipitation and temperature, thereby reflecting the overall water surplus or deficit (Song et al., 2019). A higher wetness index (WI > 30) indicates sufficient water availability for plant growth, whereas a lower wetness index (WI ≤ 30) suggests the water availability may be limited (De Martonne, 1926).”

      (3) Lines 173-174: Please specify the exact timing of flux measurements (e.g., "measured three times per month between 9:00 and 11:00 AM" is already stated, but add "on sunny and calm days" to ensure consistent conditions).

      Thanks for the suggestion. We have supplemented the exact timing of flux measurements in the revised manuscript as follows:

      Lines 183-186, page 6: “The gas exchange measurements were conducted on sunny and calm days between 9:00 and 11:00, a time when the ecosystem CO<sub>2</sub> fluxes represented the daily average (Niu et al., 2008; Rong et al., 2017; Yu et al., 2025), with a frequency of three times per month.”

      (4) The RGR and AGB data provide important clues for explaining the moderating role of the wetness index, but the mechanistic chain can be further refined. It is recommended to use the structural equation model.

      Thanks for the comments. The structural equation model has been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      Suggested addition: If data on soil moisture, soil nutrients (e.g., ammonium, nitrate), leaf photosynthetic parameters (e.g., maximum photosynthetic rate, stomatal conductance), or community composition are available, please incorporate them into the analysis to test a more complete mechanistic pathway.

      Thanks for the comments. The data of the composition of C<sub>3</sub> and C<sub>4</sub> plant communities and their species richness have been supplemented to complete the mechanistic chain.

      Lines 338-342, page 11: “Heavy grazing decreased C<sub>3</sub> plant biomass but increased C<sub>4</sub> plant richness compared with other treatments (Fig. S7). In addition, the structural equation model showed that grazing altered net ecosystem productivity by increasing soil temperature and relative growth rate, while decreasing aboveground biomass of plants (Fig. S8).”

      (5) The current meta - analysis has established the moderating role of the wetness index, but there is room for further exploration: Test for non-linearity. Could the relationship between the wetness index and the grazing effect size (lnRR of NEP) be non-linear in the global data? Consider fitting models that include a quadratic term for the wetness index or using generalized additive models (GAMs). If a threshold is identified, report the threshold estimate and its confidence interval and discuss its management implications.

      Thank you for the comments. Following the reviewer's suggestion, we further examined whether there is a nonlinear relationship between the wetness index (WI) and the magnitude of grazing effect (NEP_RR). We compared a linear mixed-effects model including a first-order term for WI with a quadratic mixed-effects model, and set Study ID as a random effect in both models. The results indicated that the AIC value of the quadratic model (256.56) was higher than that of the linear model (241.48), suggesting that adding the second-order term did not improve the model fitting precision. Based on this, we retained the simpler linear model in the revised manuscript.

      Author response table 1.

      The comparison of linear mixed-effects model and quadratic mixed-effects model

      Note: if the AIC value was lower, the model fitting accuracy was higher.

      (6) Section 4.1: This section is quite long. Consider splitting it into 2-3 paragraphs, discussing: (i) overall grazing effects on CO2 fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) differential effects of grazing intensities.

      Thanks for the suggestion. Section 4.1 has been split into four paragraphs according to the suggestion of the reviewer. (i) overall grazing effects on ecosystem CO <sub>2</sub> fluxes; (ii) the moderating role of the wetness index and its mechanisms; (iii) integrated discussion of field study and meta-analysis.

      (7) Integrating the long-term experiment and meta-analysis. Does the effect size observed in the long - term experiment (e.g., an 85.83% increase in NEP under LG in the wettest year) align with the average effect size from the global meta - analysis under similar conditions? If not, what are the potential reasons? (e.g., specificity of the typical steppe, methodological differences between chamber and eddy covariance measurements) Is the mechanism identified in the long - term experiment (e.g., increased RGR) likely to be common globally? What are the joint management implications from both parts of the study? Are there contexts where caution is needed in extrapolating the findings? (e.g., alpine meadows might be more sensitive to grazing).

      Thank you for the comments. The long-term experiment and meta-analysis have been integrated in the revised manuscript. The results of field experiment showed that light grazing increased NEP by 85.83%. However, the response of light grazing was not significant in meta-analysis. Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions. The different results in effect size and statistical significance were mainly because the long-term experiment was conducted in a typical grassland ecosystem, which may possess a strong compensatory growth capacity under moderate water conditions. Therefore, light grazing can enhance NEP by increasing plant photosynthetic rate, promoting new leaf growth, and improving community resource utilization efficiency. However, the global meta-analysis integrated different grassland types, climatic conditions and grazing durations. Consequently, the average effect of global meta-analysis may be diluted by the high heterogeneity among ecosystems. The data to calculate RGR were not available in the original studies of the meta-analysis, so we could not include RGR in the meta-analysis. Therefore, it is still unclear if the mechanism of increased RGR could be common globally. The joint management implications have been revised as follows:

      Lines 393-400, page 13: “Both the results of field experiment and global meta-analysis showed that light grazing did not reduce NEP under wetter conditions (Fig. 3B, 6B and 7H). Light grazing usually stimulates leaf regrowth following defoliation, and these new leaves often are more physiologically active than the older leaves that contribute much of leaf area in ungrazed treatment (Polley et al., 2008), which likely imply a stronger leaf photosynthesis and C sink (Reich et al., 2007). Considering factors such as different grassland types (Fig. S11), livestock grazing modes (Fig. S12), climatic conditions, and grazing durations, the future grazing studies on NEP should focus more on ANPP and RGR, and extrapolate the results cautiously.”

      (8) Figure 8: The conceptual diagram is clear and effectively summarizes the main findings. Briefly explain the meaning of the arrows in the caption.

      Thank you for the suggestion. We have revised Fig. 8 accordingly:

      “Schematic summary of the effects of grazing intensity on ecosystem carbon dioxide (CO<sub>2</sub>) fluxes and biomass in global grasslands in the meta-analysis. The blue and black arrows indicate negative effects on grazing and grazing intensities, respectively. Asterisks (<sup>*</sup>) indicate significant effects on variables at P < 0.05. GPP, gross primary productivity; ER, ecosystem respiration; NEP, net ecosystem productivity; LG, light grazing; MG, moderate grazing; HG, heavy grazing.

      (9) Please check all references for consistency with eLife style. Some entries currently have inconsistent formatting (e.g., some include issue numbers, while others do not; page number formatting varies).

      Thank you for the comments. We have checked the references one by one to revise them consistent with eLife style.

      Reviewer #3 (Recommendations for the authors):

      (1) Method citation:

      (a) Please cite the original method references rather than studies that applied the methods. For example, the original publication introducing the wetness index (WI) is: De Martonne, E. Une nouvelle fonction climatologique: l'indice d'aridité. La Météorologie 2, 449-458 (1926).

      (b) Please also cite the R packages used in the analysis. One straightforward approach is the function citation() in R.

      Thank you for the comments. We have cited the references related to the original methods and cited the R packages used in the analysis.

      (2) Several expressions would benefit from clarification:

      (a) Line 182: What has been "referred to as soil moisture"?

      Thank you for the comments. Soil moisture refers to the volumetric water content of the soil, which has been clarified in the revised manuscript as follows:

      Lines 199-200, pages 7: “Soil moisture was the soil volumetric water content.”

      (b) Line 184-185: Do you mean "soil temperature and soil moisture were measured simultaneously with ecosystem CO<sub>2</sub> flux measurements."?

      Thank you for the comments. Yes, we simultaneously measured soil temperature and soil moisture using temperature and moisture probes while measuring ecosystem CO <sub>2</sub> flux measurements. We have made following changes in the revised manuscript as follows:

      Lines 198-199, page 7: “Soil temperature and moisture at a depth of 0-10 cm were measured simultaneously with ecosystem CO <sub>2</sub> flux measurements, using the probes of the LI-8100 system.”

      (c) Line 200: Please clarify what is meant by "the caged plots"?

      Thank you for the comments. The misleading words have been deleted in the revised manuscript.

      (d) Line 242-244: Do you mean "data from non-grazed treatments were excluded when additional treatments were present"?

      Thank you for the comments. Yes, data from non-grazed treatments were excluded when additional treatments were present. We have made following changes in the revised manuscript:

      Lines 268-270, page 9: “(c) Data from non-grazed treatments were excluded when additional treatments (e.g., fertilization, experimental warming, or precipitation manipulation) were present.”

      (e) Line 263: Should this refer to RR<sub>++</sub> instead of RR? Please clarify how RR<sub>++</sub> (or lnRR++) was calculated from the reported response ratios and study weights.

      Thank you for the comments. RR is the dependent variable of the model, representing the response ratio for each observation. RR<sub>++</sub> typically refers to the pooled effect size obtained after all RRs are weighted and subjected to a mixed-effects model, which primarily corresponds to the estimated value of the model intercept β<sub>0</sub>. We have made following changes in the revised manuscript:

      Lines 293-303, page 10: “A linear mixed-effects model, with ‘study’ included as a random factor, was employed to estimate the weighted response ratio (RR<sub>++</sub>) across studies or within a specific group, fitting with restricted maximum likelihood using the ‘lmer’ function in the ‘lme4’ package (Feng et al., 2023).

      where β<sub>0</sub> is the coefficient, π<sub>study</sub> denotes the random effect associated with ‘study’ (accounting for autocorrelation among observations from the same study), and ɛ corresponds to the residual sampling error. We checked the normality of the model residuals using the ‘check_normality’ function in the ‘performance’ package. When the assumption of normality was violated, bootstrapping with 999 iterations was performed using the ‘boot’ package to derive the 95% confidence interval (CI) for each RR<sub>++</sub> (Chen et al., 2021).”

      (3) Line 222: Please list all environmental predictors considered in the analysis.

      Thank you for the comments. The environmental predictors considered in the analysis have been listed in the revised manuscript:

      Lines: 243-245, page 8: “The environmental predictors considered in the analysis included grazing intensity, the response ratio of soil temperature and soil moisture, wetness index, and grazing duration.”

      (4) Line 343: The non-significant response of NEP appears only under MG, while both LG and HG significantly decrease NEP (Fig. 5). It may be helpful to discuss this grazing-intensity-dependent response more explicitly.

      Thank you for the comments. The reason why the non-significant response of NEP appeared only under MG was that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP. This was consistent with the hypothesis of moderate disturbance, which suggested that ecosystem functions remain stable under moderate levels of disturbance. We have made following changes in the revised manuscript:

      Lines 390-393, page 13: “Similarly, the reason why the non-significant response of NEP appeared only under MG may be that moderate grazing affected GPP and ER at similar magnitudes, resulting in limited changes in NEP (Fig. 5). However, the reduction in response of GPP was higher than that of ER in LG and HG, resulting in the decrease in NEP.”

      (5) Line 397: The phrase "global-scale study" typically refers to experiments conducted worldwide. "Studies across global grasslands" may be more precise here.

      Thank you for the comments. We have made following changes in the revised manuscript:

      Lines: 475-477, page 15: “Our meta-analysis has limitations due to the relatively small sample size, which stems from the scarcity of studies across global grasslands exploring the effects of grazing intensity on ecosystem CO<sub>2</sub> fluxes.”

      (6) Line 403: ...have large amounts of grassland for grazing, "but rarely investigated".

      Thank you for the comments. We have made following changes according to the suggestion of the reviewer in the revised manuscript as follows:

      Lines 480-482, pages 15-16: “In addition, future studies could conduct more experiments of ecosystem CO<sub>2</sub> fluxes in South America, Africa and Oceania, which have large amounts of grasslands for grazing, but were rarely investigated.”

      (7) Please remember to cite Figure 8 in the text.

      Thank you for the comments. Figure 8 has been cited in the manuscript as follows:

      Lines 411-414, page 13: “Moderate grazing decreased GPP and NEP under lower WI, but the responses of GPP and NEP to moderate grazing were similar under higher WI in global grasslands (Figs. 6 and 8), indicating that higher wetness offset the response of GPP and NEP to moderate grazing.”

      (8) Figure S1:

      (a) The orange line in panel A appears to represent monthly temperature rather than mean annual temperature.

      (b) Same for the precipitation, the data shown here should be monthly values rather than annual means.

      (c) Consider using a color different from orange for the wetness index in panel C, unless the variable shown is temperature instead.

      Thank you for the comments. We have revised Fig. S1 according to the suggestion of the reviewer as follows:

      Supplementary page 5: “Fig. S1 The monthly total precipitation and average air temperature (A), annual precipitation (B) and wetness index (C) in the study area of the grazing intensity experiment in the typical steppe from 2019 to 2023.”

      (9) Figure S2B: Grazing intensity labels appear in Chinese in the figure. These should be translated into English, and the information on grazing intensity should be provided in the legend or the figure here, as well as in the Method section.

      Thank you for the comments. We have replaced the figures included the Chinese labels. See Fig. S2.

      (10) Figure S6B: LG significantly affects the relationships between wetness index and ER (P<0.05), but a regression line is missing from the panel.

      Thank you for the comments. We have added the regression line between wetness index and ER in Supplementary Fig. S6.

      (11) Figure S7A: "Mean" annual precipitation

      Thank you for the comments. The “Mean” has been revised in Figure S10A.

      References

      Fan, Y., Zhang, X., Wang, J., & Shi, P. (2011). Effect of solar radiation on net ecosystem CO2 exchange of alpine meadow on the Tibetan Plateau.Journal of Geographical Sciences, 21(4), 666-676.

      Jiang, Z., Hu, Z., Lai, D., Han, D., Wang, M., Liu, M., Zhang, M., & Guo, M. (2020). Light grazing facilitates carbon accumulation in subsoil in Chinese grasslands: A meta-analysis. Global Change Biology, 26(12), 7186–7197.

      Li, X., Fu, H., Guo, D., Li, X., & Wan, C. (2010). Partitioning soil respiration and assessing the carbon balance in a Setaria italica (L.) Beauv. Cropland on the Loess Plateau, Northern China. Soil Biology and Biochemistry, 42(2), 337-346.

      Niu, S., Wu, M., Han, Y., Xia, J., Li, L., & Wan, S. (2008). Water‐mediated responses of ecosystem carbon fluxes to climatic change in a temperate steppe. New Phytologist, 177(1), 209-219.

      Rong, Y., Johnson, D. A., Wang, Z., & Zhu, L. (2017). Grazing effects on ecosystem CO2 fluxes regulated by interannual climate fluctuation in a temperate grassland steppe in northern China. Agriculture, Ecosystems & Environment, 237, 194-202.

      Shi, R., Su, P., Zhou, Z., Yang, J., & Ding, X. (2022). Comparison of eddy covariance and automatic chamber‐based methods for measuring carbon flux. Agronomy Journal, 114.

      Wan, L., Liu, G., & Su, X. (2025). Global meta-analysis reveals different grazing management strategies change greenhouse gas emissions and global warming potential in grasslands. Geography and Sustainability, 6(3), 100251.

      Yin, M., Gao, X., Kuang, W., & Tenuta, M. (2023). Soil N<sub>2</sub>O emissions and functional genes in response to grazing grassland with livestock: A meta-analysis. Geoderma, 436, 116538.

      Yu, H., Wang, X., Wu, Y., Wang, C., Yan, R., Xu, D., & Xin, X. (2025). Light grazing tends to enhance ecosystem carbon sequestration and resource use efficiency in a meadow steppe of northern China. Agricultural and Forest Meteorology, 372, 110690.

      Zhang, R., Tian, D., Chen, H. Y. H., Seabloom, E. W., Han, G., Wang, S., Yu, G., Li, Z., & Niu, S. (2022). Biodiversity alleviates the decrease of grassland multifunctionality under grazing disturbance: A global meta-analysis. Global Ecology and Biogeography, 31(1), 155–167.

      Zhou, G., Luo, Q., Chen, Y., Hu, J., He, M., Gao, J., Zhou, L., Liu, H., & Zhou, X. (2019). Interactive effects of grazing and global change factors on soil and ecosystem respiration in grassland ecosystems: A global synthesis. Journal of Applied Ecology, 56(8), 2007–2019.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Choucri and Treiber have reassessed their previous study on TE-gene chimeric transcripts in neural genes in response to Azad et al (2024). Azad and colleagues argued that, contrary to Choucri and Treiber's findings, chimeric TE-mRNAs are relatively infrequent, and they cautioned that further optimization of bioinformatics pipelines is needed to detect TE insertions from RNAseq accurately. In this short response, Choucri and Treiber clearly demonstrate that differences in the tools used between their study and that of Azad et al. likely account for the contrasting results, along with RT-PCR failure in designing primers that would match the chimeric transcript, and the use of different Drosophila lines. The authors emphasize the need for uniform, standardized criteria in such analysis, which would ultimately strengthen and advance the field.

      Strengths:

      The addition of a ratio to compute the number of splice reads specific to the chimeric transcript and compare to the exon-exon splice reads is really interesting because it opens the door to finally quantify the contribution of chimeric TEs to the overall gene expression, although this is not the scope of the present article. The clear dissection of chimeric transcripts, along with the results from Azad et al, allows us to understand the differences between the two studies confidently. Finally, the discussion on Drosophila lines is indeed essential, given that the lines and even individuals have high TE polymorphism.

      Weaknesses:

      I think it is necessary to add more detail to this article, for instance, the differences between TEchim and Tidal could be laid out more precisely.

      We thank the reviewer for this helpful suggestion and agree that a more explicit comparison improves the clarity of the manuscript. Briefly, TIDAL and TEChim are designed to answer different questions. TIDAL is an insertion/deletion caller that was primarily built to determine whether a TE insertion exists at a genomic locus and how it is distributed across strains or populations. It was not designed to resolve splice junctions between exons and TEs. TEChim, by contrast, is purpose-built to detect breakpoint-spanning reads that span exon junctions and putative splice sites within a TE. This difference in design has several concrete consequences in the algorithms used:

      (1) TIDAL clusters reads that support a candidate breakpoint within a window of twice the sequencing read length (e.g. 300nt for 150nt reads). This works well for calling genomic insertions, but it can be too restrictive for a splice junction, which may fall at a variable position within the gene and TE. TEChim does not rely on fixed-window clustering, but instead maps all reads, groups them to individual nt positions within the genome, and then filters for events that were detected in more than one biological replicate.

      (2) TEChim reconstructs long in-silico reads from overlapping paired-end reads, using FLASH, prior to alignment. In-silico paired-end reads are analysed, and full-read merged fragments provide single-nucleotide resolution for a breakpoint. This increases accuracy and confidence in split reads.

      (3) TEChim explicitly intersects candidate breakpoints with annotated exon/intron/UTR. Structures and canonical splice donor- and acceptor sites.

      We have now added a concise summary of the underlying principles of TEChim in the Methods section, and added details to the comparison of TIDAL with TEChim in the discussion. We hope that this addition makes it clearer why the two pipelines produce different results from the same input data.

      Regarding the roo example, one of the caveats of this family, along with others, is the presence of simple repeats. It would be important to show that the simple repeats are not interfering with the read mapping.

      We thank the reviewer for raising this important point. We agree that simple repeats can complicate read mapping and that this requires careful consideration. We have now added to the results the exact locations and lengths of the three known and annotated repeat regions within roo (Domínguez, 2021) and show that the breakpoints we report map more than 4kb away from these regions. In addition, the splice junctions we identify (at positions 5190 and 5462) recur at the same position relative to the roo consensus sequence across multiple genomic insertions, and biological replicates. If these calls were artifacts of copyspecific simple repeats that interfered with read mapping, then we would expect breakpoints to vary between with each insertions local sequence, rather than converge on the same breakpoint. We therefore consider it unlikely that simple repeats account for the observed splice junctions. We have expanded our discussion to make this reasoning clearer.

      Regarding the experiments, if we are looking for a standardized protocol, then we should have a detailed material and methods section, with every experiment, replicate, and PCR temperature clearly defined.

      We thank the reviewer for this suggestion and agree that a more detailed description of the experimental procedures will improve the reproducibility of the study. We have substantially expanded the Materials and Methods section, including the number of biological replicates, primer information, PCR conditions and other methodological details relevant for reproducing the experiments. In addition, we have written a detailed manual for the updated version of TEChim that we used here.

      Finally, and in my opinion, more importantly, the use of RT negative controls on the RT PCRs, along with DNA PCRs to show insertion presence, is mandatory for testing the presence of chimeric genes. Of course, water negative PCR controls are also needed, and unfortunately, absent from Figure 3.

      We thank the reviewer for this helpful suggestion. We have repeated the RNA extraction on 3 new samples, and this time also included minus-RT aliquots for each of the three biological replicates. We have run our PCR for the chimeric transcript between Beadex and opus on all these samples, and in addition on a water control. All these new results are now shown in Figure 3, and confirm our previous conclusions.

      We now also provide results from our DNA testing for the opus insertion in Beadex. We conduct these at regular intervals in our lab to ensure the insertion remains stable in our stock, and mentioned the results in the original version of this manuscript, but we agree that it is important to also show the raw data of this in this study. We use primers at the up- and downstream end of the opus insertion. The downstream pair gave a single band at the predicted size, which we confirmed by Sanger sequencing. This data is now presented in Figure 3 – Figure supplement 1. The PCR around the upstream end resulted in the expected band of 888bp, and two additional bands. Sanger sequencing of the 888bp band produced signals for both Beadex and opus, but we did not get a reliable signal across the precise breakpoint (see Author response image 1). We think this might partly be due to a tandem repeat at the beginning of the opus LTR, which may interfere with Sanger sequencing. Taken together, our data provides strong evidence that the opus insertion is present in our flies.

      Author response image 1.

      Sanger sequencing results of the 888bp band: Two segments of the raw Sanger sequencing trace are shown; the intervening, unmapped section is omitted. The left segment (grey) shows clean, high-confidence signal matching the Bx locus. The right segment (pink) shows the signal falling to near baseline within the opus LTR, so base calls in this region are log-confidence, and the trace does not resolve the breakpoint. Numbers above the sequence indicate position within the raw sequencing read.

      Reviewer #2 (Public review):

      Summary:

      This study by Choucri and Treiber aims to directly address a recent critique regarding the role of transposable elements (TEs) in diversifying the neural transcriptome of Drosophila. The authors seek to demonstrate that TEs are not merely genomic "noise" but are frequently and reliably "exonized" into brain-specific mRNA. By introducing an upgraded computational pipeline, TEChim, and conducting precise experimental validations, the authors set out to show that TE-mediated splicing represents a genuine biological phenomenon that expands the molecular repertoire of the nervous system.

      Strengths:

      The study's primary strength lies in its rigorous technical "forensic" analysis of previous failed replication attempts. The authors convincingly demonstrate that the lack of signal in the opposing study stemmed from a fundamental methodological mismatch: the software used by the critics (TIDAL) is logically incapable of detecting splice sites located within TE sequences. Importantly, the authors complement this computational clarification with definitive experimental evidence through an effective "experimental rescue." By employing correctly designed primers and matching the genetic backgrounds of the fly strains, thereby accounting for genomic polymorphisms, they successfully validated all seven loci that were previously reported as undetectable. This dual-pronged strategy, addressing both algorithmic bias and experimental design, establishes a more robust technical benchmark for the detection and validation of TE-derived exons in neural tissues.

      Weaknesses:

      While the technical rebuttal is highly convincing, the scope of the study remains primarily defensive. As a response to a prior critique, the work focuses on establishing the existence and detectability of chimeric TE-derived transcripts rather than exploring their broader functional consequences. As a result, there is limited new insight into how these TEmodified isoforms influence neural circuit function or organismal behavior.

      We agree with the reviewer that the primary focus of this study is to establish a robust protocol for the detection and validation of TE-derived chimeric transcripts, and to resolve discrepancies raised by Azad et al. We believe this provides an important technical and conceptual framework for future studies investigating the functional impact of TE-driven genetic variation.

      In addition, the detection and validation of these events remain technically demanding, requiring deep sequencing and specialized bioinformatic expertise, which may limit broader adoption by laboratories without dedicated computational resources.

      We agree that the robust detection and validation of TE-derived chimeric transcripts is technically demanding, requiring both high-quality sequencing data and specialised computational analyses. However, we believe that these methodological challenges are justified by the biological insights that can be gained. By providing an updated computational approach together with experimental validation, we hope to facilitate further research into this phenomenon.

      Reviewer #3 (Public review):

      Summary:

      This manuscript by Choucri and Treiber responds to a recent paper by Azad et al., which responds to a paper by Treiber and Wadell (Genome Research, 2020). The controversy relates to the detection of transcripts with transposable elements (TEs) spliced into them in the Drosophila brain.

      Strengths:

      The authors now argue convincingly that these transcripts exist using an improved, updated version of their pipeline. They also validate some of their findings using RT-PCR and explain why Azad et al. failed to detect these transcripts due to methodological errors. Overall, I am convinced that these transcripts exist and that the TE-derived transcripts described by Choucri and Treiber are real.

      Weaknesses:

      The authors should mention that combining PCR-amplified cDNA generation with shortread sequencing is suboptimal for detecting TE-fusion transcripts. Recently, direct longread ONT RNA sequencing, which does not require amplification and spans the entire transcript, has been used to detect similar transcripts in human stem cells and the human brain (PMID: 40848716 & Garza et al, BioRxiv). Had the authors used this technology to validate their findings, there would be no question about these transcripts. If not doing such experiments, then they should at least discuss the possibility and the advantage of the approach.

      We thank the reviewer for this excellent suggestion and agree that long-read RNA sequencing represents a powerful approach to characterise chimeric transcripts, because it can capture full-length transcripts and reduce ambiguity of mapping short reads onto repetitive sequences. We have now expanded the Discussion to highlight the advantages of these technologies and to cite the suggested studies. At present, long-read RNA sequencing remains challenging for Drosophila brain samples, because the total amount of input RNA is limited. As a consequence, amplification is usually required, which itself can introduce artefacts, which we showed previously (Treiber and Waddell, 2017). While we agree that long-read sequencing will be an important approach for future studies, we believe that the combination of computational analysis and targeted experimental validation presented here provides robust evidence for the existence of chimeric TE-gene transcripts.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      To maximize the impact of the work, the authors should consider adding a technical guide that outlines the implementation of the updated TEChim pipeline and provides a standardized protocol for designing chimeric RT-PCR primers. This section should include a clear software workflow and a precise primer design strategy, emphasizing junction spanning probes and the necessity of genomic confirmation, to establish these methods as the definitive technical standard for studying transposable elements in the nervous system.

      We thank the reviewer for this constructive suggestion and want to be upfront about what can and cannot be standardized here. Because TE sequences are repetitive, the specific primer pair that works best for a given locus cannot necessarily be predicted in advance, and some empirical testing of candidate primer pairs is unavoidable. What we can standardise, and now describe explicitly in the Methods section, is, firstly the use of TEChim to identify candidate splicing events, and secondly, the logic behind how we choose primers. Several candidate primer pairs are tested for a given genomic locus, and all resulting bands are confirmed using Sanger sequencing. Together with an expanded TEChim manual on GitHub, we believe this study gives other research groups a clear and reproducible starting point, while remaining honest that, as with most repeat-adjacent primer design, some locus-specific optimization remains necessary.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data.

      The data presented are convincing and support the claims; the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FPexpressing cells in the population) when changing the initial ratio of green:blue FPexpressing cells.

      We thank Reviewer #1 for the Summary and for highlighting the strengths of our work.

      Weaknesses:

      (1) No spatial pattern emerges. There are some isolated colonies that turn on the downstream FPs, but I do not see a pattern, really. Nonetheless, some colonies do differentiate (i.e. they turn on additional FPs).

      The pattern does not arise within single colonies, but when looking at groups of colonies: a green sender is surrounded by a circle of blue-red receivers, which is in turn surrounded by blue-only receivers. This organization can be seen as a collective bullseye pattern (green centre, red annulus, blue background). When multiple green colonies are present in the plate, each gives rise to its own bullseye pattern. We have now clarified this detail in the text:

      Lines 244-245 – “This motif is reminiscent of a bullseye pattern emerging not within a single colony, but in groups of colonies”

      (2) The mathematical model appears somewhat superfluous. While it can clearly reproduce the data, it is not used to make interesting predictions, changing parameters (and not initial conditions) that guide further experimental implementations.

      It is true that we presented the mathematical model primarily as being capable of recapitulating our experimental results. However, the model helped us identify the ideal set of initial conditions (i.e. inducers concentrations) for Figure 6, and to better understand the dynamics of 3O-C6-HSL and 3O-C14-HSL diffusion in our system (Supplementary Movie 6). While changing parameters would be theoretically possible (e.g. diffusion coefficients, protein production rates...) we have limited exploration in this direction, as fine-tuning a single molecular parameter, all other things being equal, is experimentally challenging.

      Since several comments highlighted this weakness, we have now employed our mathematical model to predict patterns theoretically achievable with the sequential differentiation program under substantially different experimental conditions. Please find a detailed answer to this point below (see Reviewer #1 (Recommendations for the authors), point 8).

      Future directions:

      The utility of this differentiation process (e.g. in metabolic engineering or for the study of biofilm formation and antibiotic resistance) will become clearer once the FPs are substituted with functional proteins that exert an effect on the cells.

      We agree with Reviewer #1 that, for the purposes of an application, we would have to replace the fluorescent proteins with functional proteins. In the last paragraph of the discussion, we outline several potential applications of our differentiation system, including division of labor, biocomputation, biosensing, and engineered living materials. However, expanding the system in this direction is beyond the scope of this manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) This sentence is confusing to me: ‘When cells were pre-cultured with 100 nM aTc and 1 mM IPTG, the resulting colonies were almost entirely homogeneously green and blue, respectively.’ It reads as if both inducers were added in the same culture. I would rewrite this as: ‘When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.’

      We adapted the text as suggested:

      Lines 127-129 – “When cells were pre-cultured with either 100 nM aTc or 1 mM IPTG, the resulting colonies were almost entirely homogeneously green or blue, respectively.”

      (2) It would be good to explain why the follow-up experiments are conducted with inducers at the beginning of the differentiation process, given that, in the absence of inducers, there is spontaneous symmetry breaking as noted by the authors: ‘When cells were pre-cultured in absence of inducers, both green and blue colonies grew in the plate (Figure 2b).’ Is it to obtain consistent results across biological replicates?

      We used defined inducer concentrations to control the ratio of senders: receivers. In the absence of inducers, the population consisted of approximately equal proportions of senders and receivers. Under these conditions, nearly all blue colonies were close to a green colony and exhibited high expression of the red reporter. To enable spatial patterning, we therefore required a population strongly biased towards the receiver state. Accordingly, we used low concentrations of IPTG to shift the population to this state (as explained in lines 235-241 of the revised version). For testing and characterization of the system, we spotted senders alongside receivers (e.g. in Fig 7C and some supplementary figures). In these experiments, pre-culturing with high concentrations of aTc or IPTG ensured that the entire population was in the desired state. We have now revised the text to clarify that the choice of inducer concentrations ensured specific population ratios and the reproducibility of the differentiation assay:

      Lines 235-241 – “We therefore selected a condition that would consistently generate a population strongly biased towards the receiver state, with only a few sparse senders to produce the diffusible signal. Having characterized the TS differentiation landscape (Figure 2), we decided to pre-culture cells starting from the green state in presence of 9, 12 and 18 µM IPTG, then plated at a cell density of 500-2000 colonies per plate, in absence of any positional information. This resulted in few sparse green senders densely surrounded by blue receivers, in a highly reproducible manner.”

      (3) In the legend to Figure 4b, please indicate what ‘experimental points’ are (single colonies?)

      We adapted the figure caption as suggested:

      Figure 4b – “Hexagons, circles and triangles represent the average fluorescence intensity of single colonies.”

      (4) This sentence was confusing to me: ‘The time-lapses supported our hypothesis that the differentiation key processes (colony growth, HSL production and detection, expression of the red reporter) happened simultaneously.’ I thought the whole idea behind the work was to have a step-wise differentiation process and that HSL production had to happen first, then its detection and finally the downstream activation of the red FP expression.

      We acknowledge that this sentence was poorly phrased and may have caused confusion. Our multi-step differentiation program indeed operates in a sequential, stepwise manner at the single-cell level: a blue receiver cell can only transition to the red state after detecting 3O-C6-HSL, and red colonies can only transition to the yellow state after detecting 3O-C14-HSL. At the colony level, one might therefore expect colonies to first appear green or blue and only later acquire red fluorescence. However, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period. Experiments at the single-cell level might make the step-wise progression blue - red - yellow more obvious than colony-level time-lapses. We have revised the text to clarify this point:

      Lines 273-280 – “We therefore collected time-lapse videos of the plate assay, first inducing receiver cells with pure 3O-C6-HSL, then spotting sender cells alongside receivers, and finally performing the full differentiation assay (Supplementary Movies 1, 2, 3, 4). Whilst one might expect colonies to first appear green or blue and only later acquire red fluorescence, our time-lapse experiments showed that colonies were already expressing the final fluorescent reporters by the time they became visible. Thus, our observations indicate that, at the colony level, HSL production and detection, induction of the red reporter, and colony growth occur concurrently over the same time period.”

      (5) The authors write: ‘The combination of spatial and temporal information allowed us to develop a mathematical model that could qualitatively recapitulate the patterning properties of the system’. It is strangely put in my opinion. One is ‘allowed’ to develop a mathematical model under other circumstances, too. I suggest rewriting. For example: ‘We developed a mathematical model combining spatial and temporal information to quantitatively ...’.

      We adapted the text as suggested.

      (6) In Figure 6b, the scale bar is missing

      Thank you for noticing this, we have now added the scale bar.

      (7) Figure 7c: Why is the green colony huge compared with the others? The formation of such huge colonies seems to occur sometimes, but not always (it is seen also in Supplementary Figure 12a and in Supplementary Figure 13, but nowhere else; by the way, in Supplementary Figure 12a there is a spatial pattern, with a clearly visible red ring in the otherwise black colony!).

      The huge green colonies occur only in the experiments where 1 µl of green senders were inoculated at specific locations. In most of the differentiation assays, each colony arises from a single cell at unpredictable locations. For Figures 7c, S13a (previously 12a) and S14 (previously S13) we wanted a single green sender colony in the centre of the plate surrounded by blue receivers. To achieve this, as we briefly explained in the methods, we decided to spot 1 µl of culture at OD=1, therefore these green colonies originate from roughly 8×10<sup>5</sup> cells, which justifies their larger size. We have now clarified this detail in the corresponding figure captions:

      Figures 7c, S13a, S14 – “The sender colonies in this figure do not originate from one single cell, but from 1 µl of a culture of sender cells at OD=1, which explains their larger size (see Methods).”

      Concerning the red rings in figure S13a (previously S12a), we hypothesise they are the effect of the leakiness of the pLux promoter. Observing expression of mCherry in the green sender colony was actually the main indication that the sequential differentiation circuit was not working as planned: if mCherry was being expressed in the sender state, then probably cinI was also being expressed, therefore the green sender colony was undesirably producing 3O-C14-HSL. We therefore replaced pLux with pLuxLac, resulting in tighter control over the mCherry-cinI operon in the green sender state. The improvement associated with this modification can be appreciated in Figure 7c, where the green sender colonies do not display any red fluorescence.

      Regarding the specific pattern of this red signal, i.e. a ring, as opposed to an homogeneous signal in the entire colony, we have different hypotheses, but we have not investigated it thoroughly. Since the colony does not originate from a single cell but from a 1 µl inoculum, the effect might partially be ascribed to a ‘coffee ring stain’ phenomenon: upon absorption of the droplet, the outer ring could feature higher cell density compared to the centre, leading to higher red signal. In other cases, we have sometimes seen ring patterns arising in homogeneous colonies due to growth and temperature effects.

      (8) It would be nice to see the model predictions being used to create a system with different properties than the actual one. Can the authors modify parameters such as diffusion of the quorum-sensing molecule or the gene expression response to it, and see how the output would change? Changing the type of quorum-sensing molecule experimentally is doable, as is the addition of, for instance, a delay module in the gene expression module. Nonetheless, I am aware that this would require some time to do, but it would, in my opinion, strengthen the paper.

      We thank Reviewer #1 for the valuable suggestions. We have now employed the mathematical model to test the pattern resulting from various diffusion coefficients combinations and added Supplementary Figure S15 and a paragraph in the main text. Choosing diffusible signals with different diffusion coefficients would indeed allow to generate more complex patterns, including concentric rings where the expression of the red and yellow reporters does not overlap. While we agree that experimentally validating this point would be of considerable interest, we believe that changing diffusion rates is challenging and beyond the scope of a typical three-month revision, as it would most likely require substantial optimization and fine-tuning.

      Concerning the delay module, experimentally it could be implemented by introducing an intermediate step with a transcription factor modulating the expression of the CinI synthase. While we agree it would technically be feasible, we did not implement it due to time constraints. We would expect the effect of a delay module to be the following: production of C14-HSL would be delayed, therefore the front wave of C14-HSL would be consistently lagging behind the front wave of C6-HSL. The distance between the two fronts would be proportional to the delay introduced by the module. We hypothesise this would generate patterns with small yellow circles inside larger red circles. Our simulations already capture these dynamics by varying diffusion coefficients, and given the absence of experimental data to constrain parameters, we decided not to include a hypothetical delay module in our mathematical model.

      See Supplementary Figure S15

      Lines 378-385 – “Our model suggests that varying the diffusion coefficients of the two diffusible signals could generate more complex patterns, particularly concentric rings where the expression of red and yellow reporters does not overlap (Supplementary Figure S15). Interestingly, the largest region of mCitrine reporter expression is predicted when the diffusion of the first signal is slow, while that of the second is fast. This suggests that sufficient local accumulation of the first signal is required to reach the threshold for production of the second signal, which can then spread further away to produce an outer blue-yellow region.”

      Reviewer #2 (Public review):

      In this manuscript, the authors implement a three-step genetic programme in E. coli that converts an initially homogeneous population into spatially structured sender, receiver, and ‘matured’ receiver colonies on agar without externally supplied positional information. They combine a TetR/LacI toggle switch for symmetry breaking, LuxI/LuxR quorum sensing for a paracrine signalling step, and CinI/CinR for an autocrine signalling-like maturation step, and complement the experiments with a mathematical model that qualitatively reproduces pattern formation over a range of initial conditions.

      While the article has many strengths such as a clear conceptual framing using Waddington landscapes, a modular and carefully optimised circuit design, thorough experimental characterisation of the toggle and quorum-sensing modules, integration of spatial modelling with experiments, and generally clear writing and figures, I think it will benefit the article to clarify the definition and stability of ‘differentiated’ states, clarify several quantitative and modelling aspects, better explain how fitted curves and promoter engineering were done, and improve some figure design and wording to avoid ambiguity.

      We thank Reviewer #2 for the summary and for highlighting the strengths of our work. Please find below our detailed responses addressing the weaknesses and gaps you identified.

      Detailed comments below:

      (1) P5-8 / and more generally: A major concern is that producing a reporter output is not, by itself, differentiation. For a state to be credibly called ‘differentiated’, it should be stable (self-maintained) over relevant timescales, ideally in the absence of the inducing context. As written, the manuscript sometimes seems to equate cell type with reporter expression. I strongly suggest adding a short subsection explicitly defining state versus output, and for each claimed state, stating whether it is stable/bistable or unstable/reversible, with evidence. Concretely, the authors should enumerate: a) Togglederived sender versus receiver: stable? under what conditions (inducer ranges, hysteresis window)? b) Paracrine-induced ‘red’ receivers: is this a stable differentiated state, or a context-dependent induction requiring proximity to senders? c) ‘Mature’ (yellow) state: does it persist after removal from the spatial signal field? If not, it should be described as an induced output programme rather than a mature lineage state.

      At present, later sections (and the ‘maturation’ language) risk over-stating what is demonstrated.

      We acknowledge Reviewer #2’s concern regarding the stability and irreversibility of our states and agree this is a limitation in our work. The hysteresis property of the toggle switch is well documented in previous studies (Litcofsky et. al., 2012; Barbier et. al., 2020), so we did not formally re-quantify it. However, the fact that pre-culturing cells without inducer yielded predictable green: blue ratios and very few colonies exhibiting both states indicates that the toggle switch is both irreversible and stable under our experimental conditions. For the second and third steps, the expression of the fluorescent proteins is maintained throughout the duration of the experiment and for several hours to days thereafter. It is indeed established that cells in stationary phase have protein half-lives on the scale of tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025). However, we agree that re-streaking these cells in the absence of sender cells would eventually lead to the cessation of fluorescent protein expression, making the quorum sensing-mediated differentiation steps reversible.

      You can find further details on this point in our reply to Reviewer #3 (Public review), point 1.

      We have now adjusted the wording throughout the manuscript, and in particular we modified the abstract to highlight that our system ‘mimics’ cell differentiation. We have also added sections to explain how red and yellow are reversible cellular programs and not stably differentiated cellular states:

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ONOFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      (2) Figure 2d: It is unclear whether this panel is intended to be qualitative (schematic/ illustrative) or generated from quantitative data. The legend should explicitly state the origin (e.g., representative image, averaged data, simulation output, schematic) and, if quantitative, what was measured, how many replicates, and how the visualisation was constructed.

      We acknowledge the potential confusion generated from this image. While its purpose is to visually illustrate the toggle switch differentiation landscape in 3D, the figure is derived from quantitative data, specifically from the same flow cytometry data used for Figure 2c. Basically, it is the density plot of the (GFP, mCerulean) events recorded for a single replicate (initial state: mixed, induced with 0.003 mM IPTG), but the density is represented with a third dimension (depth) instead of colour intensity (as we did for Figure 2e). We have now adjusted the figure caption to clarify the figure purpose and how it was generated:

      Figure 2d – “Representative 3D energy landscape: valleys represent the green and blue stable states, red line represents the separatrix. The landscape was generated from quantitative flow cytometry data from a single replicate in c (initial state: mixed, induced with 0.003 mM IPTG). We measured GFP and mCerulean fluorescence intensity from 50,000 single cells (see Methods). The depth of each (x,y) point in the landscape corresponds to the number of recorded cells with a specific (GFP, mCerulean) intensity. This specific condition was chosen for illustrative reasons, to show the landscape in an almost symmetrical condition.”

      (3) Figure 2e: The cross-sectional line is described as meant to be comparable, yet the leftmost plot appears to have a different slope from the others. The authors should explain whether this reflects a different scaling/normalisation, a different underlying dataset/condition, or simply a plotting artefact. If these are fitted trends, report the fit function (see also the comment on fitted lines below).

      The observation is correct, the cross-sectional lines do not have the same slope. They are obtained by connecting the local minima of the green and blue population, i.e. the two points with highest density in the density plots. As the mean fluorescence intensity of the two population is not the same across conditions, the resulting lines have different slopes. We chose this strategy instead of using fixed-slope cross-sectional lines to capture the maximum depth of each valley. We have now adjusted the figure caption to clarify how the cross-sectional lines were generated:

      Figure 2e – “Flow cytometry density plots of 5 representative conditions from c (initial state: mixed, inducer concentration indicated below each plot). For each density plot, we computed the coordinates of the local minima in the green and blue populations, then generated an orthogonal plane crossing them. Black lines represent the projections of these planes on the (GFP, mCerulean) plane. [...]”

      (4) Around P7-8: (saddle/separatrix description): When describing the saddle or separatrix between the two valleys, it would be helpful to briefly connect this more directly to a quantitative dynamical-systems perspective: for instance, the intersection of nullclines and how nullcline geometry changes under IPTG/aTc induction. This will make the landscape picture more complete for readers familiar with the original genetic toggle switch work (Garder et al., 2000).

      We thank Reviewer #2 for the suggestion. We have now included the concept of nullclines in our description of the differentiation landscape in the main text:

      Lines 153-159 – “Mathematically, the profile of the toggle switch landscape corresponds to the number of intersections between the nullclines (the curve where the derivative of a given species over time equals zero), which can be one or three (Gardner et. al., 2000). If they intersect only once, there is only one minimum on the potential curve, which corresponds to a single stable state. If they intersect three times, there are two minima and one maximum, which corresponds to two stable states, and the maximum represents the separatrix.”

      Lines 168-171 – “From a quantitative dynamical-systems perspective, the addition of aTc or IPTG modified the geometric shapes of the nullclines, shifting the three solutions. If the inducer concentration is large enough, the nullclines intersect only once, producing a single stable steady state (Gardner et. al., 2000).”

      (5) P9, lines 157-159: The current phrasing (‘in absence of noise, the system would be fully deterministic... in living cells, however, stochastic bursts... change the trajectory’) risks conflating predicting population-level percentages with predicting colony-level trajectories. It would help to clearly separate (i) the ability to predict the overall fraction of ON/OFF (green/blue) colonies from inducer conditions (which is largely deterministic at the population level) from (ii) the intrinsically stochastic choice of state made by any given founder cell and its colony.

      We acknowledge that this wording has generated confusion, but we are not completely sure we understood the suggestion made by Reviewer #2. If we understood correctly, the concern is that single-cell trajectories and population-level ratios are substantially different and should not be confused, and should be investigated and modelled differently.

      To clarify, our initial sentence referred to the toggle switch differentiation landscape, not to the likelihood of a cell or a colony to be in a given state. In absence of noise, the TS landscape is deterministic: one of the two states is always the stronger attractor, for the same initial conditions cells would fall 100% of the times in that valley. In living cells, however, there are sources of noise, which make it possible for two cells starting from the same initial conditions to end up in two different valleys. The stochastic bursts in gene expression can ‘push cells back up’ and beyond the separatrix. Very large noise would allow cells to end up in the green and blue valley regardless of the TS previous state or the inducer concentration.

      The differentiation induced by the toggle switch when the system is initiated close enough to the separatrix is stochastic. The probability of observing a certain ratio green: blue at the population level reflects exactly the likelihood of a single cell falling into one of the two stable states. Our mathematical model does not take into account molecular details (synthesis and degradation/dilution of proteins, affinity of transcription factors for their cognate promoter...) to mimic the stochastic choice at the cell level. Instead, it estimates the TS differentiation landscape from the population-level percentages based on thousands of single cells.

      To avoid confusion, we have removed the initial sentence and replaced it with an explanation of the link between individual cell trajectories and the resulting population-level ratios:

      Lines 159-164 – “When the system is close to the separatrix, cells can progress towards both the green and the blue destiny. The differentiation trajectory of a single cell is largely stochastic and is influenced by bursts of gene expression. Once the TS is locked in one state, the progeny of that cell maintains a memory of that state, hence a colony has the same state as the founder cell. At the population level, the ratio of green and blue colonies reflects the probability of each cell to fall into the green or the blue valley.”

      (6) P11, lines 193-195 (promoter engineering): The main text currently only refers to screening variants and choosing pLux76; I suggest briefly stating in the main text (not only in the supplement) what was changed (for example, promoter box variants, core promoter strength modifications) and what design criteria were used (reduced leakiness, increased dynamic range).

      We adapted the text as suggested:

      Lines 206-211 – we carried out a screening of pLux promoter variants, testing combinations of Lux boxes (G1 [iGem part BBa K1216007] and pLux76 (Grant et. al., 2016) and promoters with different strengths (100% and 54%), in order to identify variants with minimal leakiness and a high fold-change. We identified pLux76 (Grant et. al., 2016) as the regulatory region with the highest fold-change and sufficiently low leakiness among our candidates (Supplementary Figure S4)”

      (7) Use of fitted lines (Figures 2, 4, 5, 7): Wherever fitted curves are overlaid on data, the authors should indicate in the figure legend the explicit form of the fit as well as the fit equation/ parameters. As a reader, it is difficult to interpret what is empirical smoothing versus what is a mechanistic functional form.

      In most cases, with the exception of Figure 2c, the fitted curves have an illustrative purpose and result from empirical smoothing of the data points. Unless otherwise stated, the resulting parameters were not implemented in our mathematical model. For Figure 2c, instead, the experimental data is fitted with a mechanistic functional equation and the calculated parameters were implemented in our mathematical model.

      We have now added the equations and parameters of the lines fitting the experimental points, either directly in the figure caption or in Supplementary Information Tables (Tables VII to XII). In the latter case, the exact table is referenced in the corresponding figure caption.

      (8) P13, lines 232-235: The comparison between induction directly with C6-HSL and induction from sender colonies is qualitative (‘significantly smaller range’). The authors should provide distances (for example, in mm) for the induction range in each case and, if possible, approximate total HSL amounts or concentrations, so that the reader can appreciate the magnitude of the difference.

      We have now calculated the induction range as the distance where half-maximal induction is observed, which allowed us to compare the induction range across conditions. We adapted the text accordingly. We also provide an estimate of the amount of 3O-C6-HSL produced by a green sender colony, based on simulations of our mathematical model, but we highlight that the two conditions are substantially different and care should be used when comparing them: in one case, a fixed amount of C6-HSL is present from t=0 and simply diffuses outwards; in the second case, a source continuously produces C6-HSL that progresses as a wavefront, therefore the concentration profiles and diffusion dynamics are different.

      Lines 220-226 – “The induction range obtained with 100 picomoles of pure 3O-C6-HSL was approximately 6-8 mm (50% of maximal induction at 3.74 mm). The induction range around sender colonies was significantly smaller (50% of maximal induction at 1.8 mm, Figure 4b). Mathematical simulations yielded a similar slope when using 0.25 picomoles of 3O-C6-HSL (data not shown), even though care should be used when comparing diffusion of a fixed amount of inducer with continuous production from a growing source.”

      (9) P13, lines 259-262: The authors model the transition to the stationary phase via a monotonically decreasing sigmoid in time for biosynthetic capacity. What is the rationale or literature basis for this approach to model entry into the stationary phase? The authors should cite prior work and clarify why this form is appropriate here, versus alternatives (nutrient diffusion limitation, logistic growth with resource depletion, etc.).

      In most mathematical models involving pattern formation, entry of cells in stationary phase is simply treated as a step-wise function (Zwietering et al., 1990): the entire population is active (activity = 1) until it suddenly becomes inactive (activity = 0). However, experimental evidence suggests that the decay in activity is better represented by a smooth decreasing function (Gefen et al., 2014). While several frameworks and mathematical equations exist to describe loss of cell viability (for example under scenarios of heat inactivation (Mafart et al., 2002; Van Boekel et al., 2002), or nutrient limitation), we could not find previous works modeling the loss of metabolic activity upon entry in stationary phase. Again, experimental evidence suggest that switches in metabolic state are rather heterogeneous and occur stochastically at the single cell level (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to implement a simple heuristic approximation that captured well our experimental data. We have now added a paragraph in the section Mathematical modelling - Bacterial activity, supported by the appropriate references, to justify our reasoning:

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – In our system, production of diffusible molecules and fluorescent reporters happens on a timescale of several hours, therefore we decided to take into account the entry of cells in stationary phase. Traditionally, bacterial activity has been modelled as a step-wise inactivation function (Zwietering et al., 1990). However, experimental evidence has shown that bacteria support a low constant rate of protein expression even while growth-arrested, suggesting a low decay in bacterial activity (Gefen et al., 2014). Even when considering cell death upon heat inactivation, a Weibull frequency distribution model is preferred over a step-wise viability function (Mafart et al., 2002; Van Boekel et al., 2002). While we could not find mathematical descriptions specifically for entry in stationary phase, several single-cell metabolism studies support gradual, asynchronous state transitions (Nikolic et al., 2013; Kiviet et al., 2014; Van Heerden et al., 2014). We therefore decided to describe the loss of activity function in our system via a monotonically decreasing sigmoid function, that captured well our experimental data.”

      (10) Figure 6c: Are the areas of the plate shown in each column the same field of view across conditions/time, or are these simply representative regions selected per condition (possibly from different plates)? The caption/legend should clarify whether these are matched locations and how images were chosen.

      They are indeed representative regions per each condition. Each condition is a different plate, as the differentiation assays requires plating the culture homogeneously on a fresh plate without inducers. The plates used for imaging were the same used for the quantification in Figure 6b, four images per plate were collected, one representative image was chosen for Figure 6c. We have now adapted the figure caption to clarify how images were collected and chosen:

      Figure 6c – “Representative microscopy images of the spatial patterns generated by cells harbouring the 2-step system, pre-cultured with different inducer concentrations (indicated at the bottom of each column). Each column displays one representative image (from four locations imaged) of the seven plates in b. Rows (from top to bottom): GFP channel, CFP channel, mCherry channel, composite image.”

      (11) Figure 7a: The combination of solid, dashed, and dash-dot arrows/lines is visually hard to read. I suggest replacing the dash-dot line with a fully dotted line or using different colours (if consistent with journal style) to improve readability.

      Thank you for noticing this. We have now replaced the dash-dot line with a simple dotted line in Figures 3c, 7a, S12a (previously S11a) and S13b (previously S12b) to improve readability.

      (12) Figure 7e and similar analyses: The authors should explain in the Methods and/or captions how ‘distance from sender colonies’ is computed when multiple senders exist. Is the distance always measured to the nearest sender, and how are cases handled where a receiver is in the overlapping influence of several senders? This clarification is important for interpreting the fitted curves.

      The calculation of the distance in 4b, 5b and 7e was performed only in cases where sender colonies were sufficiently sparse (i.e., more than 1 cm away) to assume each receiver was under the influence of a single sender. When collecting the microscopy images, we carefully avoided to image fields of view with green senders just outside the edges of the images. Images with multiple senders (e.g., 7c) would require a non-trivial calculation to compute the relative contribution of each sender to the red intensity of each receiver. We have now clarified this detail in the Methods:

      Lines 558-560 – “When collecting microscopy images with senders and receivers, we carefully avoided to image fields of view with green senders just outside the edges of the images.”

      Lines 590-593 – “Calculation of the distance between receiver and sender colonies was performed only for images collected from plates with few sparse (i.e., more than 1 cm away) senders, ensuring that each receiver only sensed the 3O-C6-HSL produced by a single sender colony.”

      Reviewer #2 (Recommendations for the authors):

      (1) P8, ‘upon transformation’: This phrasing is ambiguous and can be misread as referring to DNA transformation. I recommend changing ‘upon transformation’ to ‘after transition’ (or similar) to avoid confusion.

      The phrasing indeed refers to the DNA transformation of circuits into the cells. The two plasmids were co-transformed into MG1655, cells recovered for approximately 1 h, then the bacteria were plated on solid medium in absence of inducers. The resulting colonies, which we refer to as ‘upon transformation’, were a mixture of green and blue (both expressed at low intensity, as highlighted in Supplementary Figure 2). We only used these colonies for Figure 2c, middle row. For all other experiments, we selected colonies that had been pre-differentiated in one of the two states via chemical inducers, and showed strong expression of the respective reporter (Supplementary Figure 2). We have now adapted the figure caption to make this detail more explicit:

      Figure 2c – “For the initial state green and blue, cells were taken from colonies that were homogeneously green or blue, while for the initial state mixed, cells were taken from a colony obtained immediately after transformation of the circuit plasmids into cells.”

      (2) More generally, consider disambiguating terminology around ‘cell type’, ‘state’, and ‘output’, since the current wording occasionally implies stable fate commitment where the data (as presented) may instead support reversible, context-driven induction.

      We went carefully through the text and adapted it appropriately, highlighting which states are stable and which are reversible. See our detailed reply to your point 1 (Public review) above.

      Reviewer #3 (Public review):

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggleswitch-based symmetry breaking with quorum-sensing interactions to generate colonyscale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      However, I think the paper needs to tone down the extent to which the system demonstrates multi-step differentiation or morphogenesis, which is not critical for making the paper valuable. Only the first step of their circuit design (Figure 1), the toggle switch, generates stable alternative states. The latter steps are mainly signal-dependent reporter activation states layered on top of the blue receiver state, rather than true fate transitions. The authors explicitly state that red expression is added without replacing the blue identity, and they also acknowledge that red cells lose their identity upon restreaking unless they remain near sender cells. That substantially weakens the differentiation analogy and makes the Waddington framing too strong.

      We acknowledge that our sequential program does not recapitulate all the features of a differentiation trajectory, and in particular we do recognize the only two stable states are the identities associated with the toggle switch, while red and yellow are outputs indicating activation of a new molecular program. A complete differentiation program would require irreversible cell fate determination, as we mention in the discussion. For the same reason, in our illustrative Waddington landscape (Figure 1) we only represent two valleys, or minima, corresponding to the green and blue states, while the activation of the red and yellow programs does not result in further valleys. We modified the text at various points to clarify this important distinction, and the fact that our system mimics some key steps happening during multicellular differentiation, without claiming that expression of a fluorescent reporter is a stably differentiated state. Notably, we modified the abstract to highlight that our system ‘mimics’ cell differentiation.

      We believe the differentiation analogy and the Waddington framing remain valuable in our work, considering the physical constraints and timescale of our system. While it is true HSL-induced molecular program would eventually deactivate in absence of signal, this would require a significantly longer amount of time than the one we use to observe patterns. Also, upon storage of Petri dishes at 4 °C, the pattern (including red and yellow) remains stable for a few weeks. Finally, our main purpose was to investigate the capacity of a sequential program to generate autonomous spatial patterns, without the need for human intervention. The removal of a colony from the local signal field, e.g. to re-streak it on a fresh plate, represents a strong disturbance of the pattern, not dissimilar from early developmental biology experiments where transplant of tissue portions would sometimes result in fate reprogramming.

      You can find further details on this point in our reply to Reviewer #2 (Public review), point 1.

      See Abstract

      Lines 92-104 – “In the present work, we aimed to address this gap by creating an autonomous multi-step program recapitulating cell differentiation in Escherichia coli. [...] Next, we activated a new molecular program in a subset of receivers that are in close proximity to a sender through intercellular communication, implemented via the quorum sensing (QS) system LuxI-LuxR. Finally, the newly emerged population underwent maturation by producing an autocrine signal via the quorum sensing system CinI-CinR.”

      Lines 171-173 – “Taken together, these observations showed that, owing to the TS bistability, a group of initially undifferentiated cells bifurcated into one of the two possible states, which were stably maintained thanks to the TS hysteresis (Litcofsky et. al., 2012).”

      Lines 255-266 – “It is crucial to distinguish that, while the blue/green identity associated with the toggle switch represents a true bistable state (Gardner et. al., 2000; Litcofsky et. al., 2012; Barbier et. al., 2020), the expression of the red reporter is reversible and dependent on proximity to the 3O-C6-HSL source. Although the QS ON-OFF transition is considerably slower than the OFF-ON activation, reversibility remains a fundamental characteristic of QS (Abraham et. al., 2024). Furthermore, protein half-lives in stationary-phase cells extend over tens of hours (Wiechecki et. al., 2017; Gervais et. al., 2025), allowing for the stable detection of mCherry fluorescence over days. Consequently, mCherry serves as a robust readout for the molecular program’s output; however, a truly differentiated cell state would necessitate irreversible modifications to the gene expression profile. Therefore, in our landscape analogy we represented the receiver valley as a continuum where red intensity decreases as the distance from the sender valley increases, without local minima (Figure 1, 3rd row).”

      Lines 337-342 – “Thanks to the production of an ‘autocrine signal’ that affects only the red cells, this population drifts apart from the blue receiver state, increasing the distance between the two cell types in the differentiation landscape (Figure 1, IV° row). Similarly to the 2nd step, the 3rd step recapitulates a differentiation trajectory via expression of a fluorescent reporter, which in this system is reversible and does not introduce irreversible modifications of the cell state.”

      Lines 428-430 – “In this work, we successfully engineered a multistep program mimicking the differentiation of an initially homogeneous population into three distinct cell types without any external cues, while still achieving fine-tuning of the populations’ ratios through pre-culture conditions.”

      A related concern is that the 3rd step does not introduce a new spatial organizing rule. The authors show that the second signal remains confined to cells already receiving the first signal, and explicitly conclude that it functions only as an autocrine cue rather than a second paracrine layer. As a result, the 3-step system seems more like an added local readout or maturation layer. Overall, the main 2-step outcome is sparse green sender colonies surrounded by red-expressing blue receivers, with distant receivers remaining blue. That is a valid engineered pattern, but it is still a local, threshold-response circuit architecture.

      The comment is correct, we indeed refer to the 3rd step as ‘maturation’ in the manuscript, as the newly emerged red population activates a new molecular program, which cannot be activated in blue receivers that have not been exposed to C6-HSL. While we had initially expected that production of a second diffusible signal could result in signal propagation, therefore generating an outer yellow ring surrounding the red ring, our experiments showed C14-HSL only acted locally, and our mathematical simulations confirmed that the front wave of the second signal was always lagging behind the first one. We have now employed our mathematical model to explore which conditions would support a new patterning rule, and added Supplementary Figure S15. If the diffusion coefficients of C6-HSL and C14-HSL were 10-100 times smaller and 2-5 times larger respectively, a yellow-only area could appear around the red area.

      Regarding the last sentence, we are unsure about the concern raised and which alternatives Reviewer #3 would recommend. We fully agree our genetic program leverages a local, threshold-response circuit architecture, it is indeed a reaction-diffusion system based on quorum sensing signalling. Many natural patterning and morphogenesis systems do rely on local, rather than global, interactions, yet they are capable of forming complex and hierarchical structures.

      The autonomy claim should be toned down and stated more precisely. The plate patterning occurs without externally imposed spatial gradients, which is a strength. However, by design, the overall system behavior depends strongly on pre-culture inducer conditions that set the sender:receiver ratio, and this externally imposed history is central to the final pattern. This property is tied to how the circuit is designed where steps 2 and 3 largely respond to symmetry breaking introduced in step 1, which is dependent on both history and initialization on the plate. In particular, currently the pattern formation process is quite variable (e.g. figure 5), depending on how different colonies flip the toggle switch, and consequently, how many become senders and how many become receivers. It would have been fascinating if they could also demonstrate the differentiation within individual colonies, leading to intra-colony patterns. This aspect should at least be discussed.

      We would like to clarify that, by ‘autonomous’ we refer to the reaction-diffusion system that starts functioning upon the seeding of the cells on the plate. All previous steps (including the cell culturing, dilution and plating) correspond to setting the initial conditions. We agree with the description of Reviewer #3 concerning the features of our system, but we argue that autonomy and dependence on the initial conditions are two different properties. We claim our system is both autonomous and sensitive to initial conditions. Upon initialization of the system on the plate, there is no further intervention, or nudging (e.g. time-dependent light stimuli, addition or removal of inducers at specific times...), therefore the system is autonomous. It is also dependent on the initial conditions, which are set homogeneously for all the cells (no positional information, each cell experiences the exact same conditions). We believe this is a strength of our system, which allows to generate a variety of diverse yet reproducible spatial patterns. Independency from initial conditions would consistently generate the same pattern, regardless of the initial inducer concentration, which might be interesting for certain applications (e.g., ensuring a fixed population ratio, provide robustness to environmental variability...), but it is not universally superior.

      In order to clarify that we only refer to the reaction-diffusion system on the plate as the autonomous component, we have now added a sentence in the main text:

      Lines 132-134 – “In our differentiation assay, cell culturing, dilution and plating correspond to setting the initial conditions for the system. Upon initialisation on the plate, no further intervention was performed, therefore the system evolved in an autonomous fashion.”

      Concerning the intra-colony patterns, that was admittedly our initial interest. We explored conditions that would consistently generate colonies with sectors, for example by inducing cells in one of the two states and then growing them on agar supplemented with the opposite inducer. However, we immediately observed that, due to the close proximity of sender and receiver bacteria, and the rapid diffusion relative to the timescale of gene expression, all blue sectors were also red. This outcome effectively eliminated the distance-dependent nature of the quorum sensing response. In order to take full advantage of the diffusible system, we focused instead on well-separated homogeneous colonies. We have now added Supplementary Figure 5, the corresponding figure caption, and we briefly discuss in the main text the occurrence of intra-colony patterns and why we did not investigate them further:

      See Supplementary Figure S5

      Lines 228-231 – “We then tested the potential of the 2-step differentiation system to generate self-organized spatial patterns. We observed rare motifs arising within the sporadic colonies showing both green and blue sectors. Due to the close proximity of the sender and receiver bacteria within a colony, the blue sectors always showed strong red signal (Supplementary Figure S5).”

      The mathematical model is useful in guiding both the characterization of parts, modules and the overall system. However, the claims around its quantitative predictive power should also be made narrower. The simulations are built from multiple fitted and partly hand-tuned components, including toggle-switch response curves, colony-growth rules, diffusion, reporter-response functions, and activity decline. This supports a calibrated qualitative reconstruction of the observed patterns, but not a strong predictive or mechanistic validation.

      We accepted the suggestion and replaced ‘predicted’ with ‘recapitulated’ or ‘simulated’ at various locations in the main text:

      Lines 105-107 – “Throughout the work, experimental results were used to develop a mathematical model, providing us with insights that guided further experimental efforts.”

      Lines 292-293 – “We developed a mathematical model combining spatial and temporal information to qualitatively recapitulate the patterning properties of the system.”

      Lines 316-318 – “We therefore employed the mathematical model to simulate the patterns generated from the 2-step differentiation system for different initial blue: green ratios. The model suggested a variety of outcomes [...]”

      Other specific points:

      (1) Given the topic of the work, the authors should cite closely relevant studies in programming pattern formation, including: Cao et al, Cell 2016 Collective space-sensing coordinates pattern scaling in engineered bacteria Rajasekaran et al, Cell 2024 A programmable reaction-diffusion system for spatiotemporal cell signaling circuit design Lu et al, BioRxiv 2024 Discovery of interpretable patterning rules by integrating mechanistic modeling and deep learning

      We have added these references to the introduction and discussion.

      (2) The model assumes identical diffusion coefficients for C6-HSL and C14-HSL despite their substantially different molecular sizes and hydrophobicities. This assumption could distort kinetic lag with differential diffusion in explaining the autocrine confinement of the third step. Its impact should at least be explored in the simulations.

      The comment is very appropriate. To the best of our knowledge, no exact values have been published for C6-HSL and C14-HSL diffusion in water, let alone for diffusion in agar. However, estimates in water range between 3 ·10<sup>−6</sup> cm/s and 5 · 10<sup>−6</sup> cm/s at 25°C, depending on the chain length (i.e. a factor of 1.7 at most). In response to your concern, we have now computationally explored the effect of using signals with different diffusion coefficients and their impact on the resulting spatial pattern (Supplementary Figure S15). Varying the diffusion coefficient of the second diffusible signal does not substantially change the pattern: lower diffusion leads to the yellow region being slightly more concentrated around the green sender and more intense. Higher diffusion has the opposite effect, with yellow areas being wider but less intense. Conversely, changing the diffusion coefficient of the first signal leads to significant changes in the final pattern, suggesting that the kinetics of C6-HSL accumulation and dispersal have a strong influence on the production of C14-HSL and therefore of the yellow signal.

      (3) The mCherry response parameters change significantly between the 2-step and 3-step systems. The authors acknowledged this change but did not provide a clear explanation.

      We believe that the elements contributing to the different mCherry response in the 2-step and 3-step systems are: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression. We have now added a paragraph in the section Mathematical modelling Bacterial activity that lists those differences. We have also added Supplementary Figure S20, containing the experimental data that was fitted to obtain the parameters listed in the second row of Table II. Finally, we highlight that for the implementation of our mathematical model we were only interested in the response function to varying 3O-C6-HSL concentrations and not in the exact values of individual parameters.

      Supplementary Material, section A. “Mathematical modelling, subsection 3. mCherry production and 30-C6-HSL sensing – The parameters fitted for mCherry production in the 2-step and 3-step systems are quite different. We hypothesize that the following elements contribute to the difference: a) the replacement of pLux76 with pLuxLac, which leads to tighter regulation and overall reduced expression of the downstream genes; b) the addition of the cinI synthase and a RBS in a monocistronic unit upstream of mCherry, which might justify reduced expression of mCherry; c) circuit-induced cell burden and reduced growth rate for the 3-step system, which leads to a slower response and lower reporter expression.”

      Supplementary Material, section A. “Mathematical modelling, subsection 4. Bacterial activity – For the implementation of our mathematical model, we were not interested in analysing individual parameters values, but rather in the resulting response function to varying 3O-C6-HSL concentrations. Further analysis would be needed to determine whether these parameters are statistically significant, but this is beyond the scope of this present manuscript.”

      (4) The 3-step system is evaluated at only a single condition with no simulation comparison, in contrast to the systematic 11-condition validation of the 2-step system.

      The rationale for not repeating the 11-condition assay with the 3-step system is that the patterns would be substantially the same as Figure 6c (red colonies would also be yellow). Furthermore, the Nikon SMZ25 stereo microscope we used to take images with a large field of view (ideal for the differentiation assay in Figure 6c) would not allow us to discriminate well between green and yellow colonies, making the interpretation of the patterns difficult. However, following up on your comment, we have now used our mathematical model to produce a representative image corresponding to Figure 6d, that we included as Supplementary Figure S16, exploring the effect of varying the green: blue ratio with the 3-step system.

      Reviewer #3 (Recommendations for the authors):

      Please see above. My comments are largely about improving the rigor and clarity of the writing, particularly those related to conceptual claims, as well as some modeling analysis to strengthen their conclusions.

      We thank Reviewer #3 for the suggestions. See our detailed replies to your points above.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study employs a closed-loop, theta-phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats and reports that disrupting theta-timescale coordination impairs performance of challenging aspects of spatial behaviors, while sparing hippocampal replay and spatial coding in hippocampal place cells. The findings are expected to advance theoretical understanding of learning and memory operations and to provide practical implications for the application of similar optogenetic approaches. The experiments were viewed as technically rigorous, but the strength of evidence provided in the current version of the manuscript was viewed as incomplete, mostly due to limited analyses and the descriptions of some of the experimental protocols.

      We thank all reviewers for their overall assessment, thoughtful comments, and suggestions. We have now addressed each of the reviewers’ comments in detail and updated the manuscript on bioRxiv (URL: https://www.biorxiv.org/content/10.1101/2025.09.15.675587v2). In addition, we have shared the raw data, intermediate analysis files, and the complete repository to facilitate replication of the analysis and figures.

      Code repo: github.com/LorenFrankLab/ms_stim_analysis

      Data repo: dandiarchive.org/dandiset/001634

      Docker containers (see GitHub repo for use instructions):

      - Database: https://hub.docker.com/r/samuelbray32/spyglass-db-ms_stim_analysis

      - Python notebooks: https://hub.docker.com/r/samuelbray32/spyglass-hub-ms_stim_analysis

      (1) Novelty and contrast with earlier manipulations:

      We now explicitly contextualize our results with prior pharmacological (Wang et al., 2016; Wang et al., 2015; Koenig et al., 2011; Brandon et al., 2014), systemic (Robbe & Buzsaki 2009; Petersen and Buzsáki 2020), and behavioral (Drieu et al., 2018) manipulations that also assessed some of the physiological features we evaluated. This contrast helps us highlight both the insights and the discrepancies observed in the prior approaches. We also more clearly explain the novelty and importance of our specific approach for temporally and physiologically precise manipulation. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (2) Additional analysis on SWRs during rest:

      Since submitting the manuscript, we have conducted additional analysis on the rate and length of SWRs in the rest box (rSWRs) and found that the rate and length are also indistinguishable between targeted and control animals (effect of manipulation between control and targeted animals; rSWR rate: p=0.45; rSWR length: p=0.94, mixed-effects model). We also find evidence for sequential neural representations (“significant replay”) in the rest box when the encoding was performed in the behavioral arena. Example trajectories and full analysis per animal are shown in the new supplementary figure (Figure S6). These results are consistent with our observations on aSWR rate, length, and content in the behavioral arena. Additionally, based on the reviewer’s recommendation, we have evaluated the fraction of ripples with continuous trajectories during the rest box before the W-Track experience and in the subsequent sleep epochs after the first exposure. We find that with experience on the track, the proportion of continuous replays increases on average in both control and transfected animals, and both groups of animals show an overlapping range of continuous trajectory lengths.

      (3) Theta sequence measurement in the absence of theta:

      We now explicitly explain why our manipulation makes it more appropriate to measure sequential hippocampal representations during locomotion (i.e., theta sequences) without using theta oscillation or an epoch-averaged, relatively large sliding window as a reference. The key insight here is that our manipulation suppresses theta and thus makes it difficult or impossible to accurately identify theta phase. We explain that while theta-phase-based approaches were used in prior work; these prior analyses may have confounded the absence of hippocampal theta sequences during locomotion by the inability to detect theta oscillatory phase reliably. We show that our method of using clusterless Bayesian decoding, in which we estimate the decoded position at every 2ms timestep, is indeed able to capture endogenous hippocampal sequences even without imposing any requirements of aligning to theta oscillations, thus providing an unbiased estimate of the rhythmicity of hippocampal spatial representations.

      (4) Additional analysis on place cell stability and tuning:

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Joshi and colleagues demonstrates that the precise theta-phase timing of spikes is causal for CA1 hippocampal theta sequences during locomotion on a linear track and is necessary for learning the cognitively demanding outbound component of a hippocampus-dependent alternation task (W-maze), independently of replay during immobility. To reach these conclusions, the authors developed a theta-phase-specific, closed-loop manipulation that used optogenetic activation of medial septal parvalbumin (PV) interneurons at the ascending phase of theta during locomotion. This protocol preserved immobility periods, allowing a clean and elegant dissociation from SWR-associated replay.

      The manuscript is well written and was a pleasure to read. The work described is of high quality and introduces several notable advances to the field:

      (a) It extends prior studies that manipulated theta oscillations by examining precise temporal structure (specifically theta sequences) rather than only LFP features.

      (b) The closed-loop manipulation enabled dissociation between deficits in theta sequences during a behavioural task and SWR-associated replay activity.

      (c) As controls, the authors included rats with suboptimal viral transduction or optic-fibre placement, and, within subjects, both stimulation-on (stim-on) and stimulation-off (stim-off) trials. Notably, sequence disruption persisted into stim-off periods within the same session.

      Overall, this is a strong manuscript that will provide valuable insights to the field. I have only minor comments:

      (1) As the authors note, it is striking that both behavioural performance and spike patterns are altered during stim-off trials. They propose that "disruption of theta sequences during the initial experience in an environment is sufficient to have lasting effects," implying that rapid, experience-dependent plasticity is driven by sequential firing. Does this imply that if rats were previously trained on the task, subsequent stim-on and stim-off trials would yield different outcomes, with stim-off trials showing improved performance and intact theta sequences? For example, if the sequence of one-third stim-on, one-third stim-off, one-third stim-on were inverted to off-on-off, would theta sequences be expected to emerge, disappear, and potentially re-emerge? While I am not asking for additional experiments, I think the discussion could be extended in this aspect.

      Alternatively, could the number of stim-off trials (one third of the total) be insufficient to support learning/induce plasticity? In the controls, ~50-100 trials appear necessary to achieve high performance.

      We think it is likely that pretraining would result in a different outcome, although we did not test this possibility. We have modified the discussion to address this point:

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to recover. This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (2) In line with the point above, the authors characterise the behavioural changes induced by MS optogenetic stimulation specifically as a "learning deficit," as rats failed to improve across 300 trials in an initially novel environment (W-maze). While they present this as complementary to prior demonstrations of impaired performance on previously learned tasks (Zutshi et al., 2018; Quirk et al., 2021; Etter et al., 2023; Petersen et al., 2020), an alternative interpretation is a working-memory deficit. This would produce the same behavioural pattern, with reference memory (the less cognitively demanding trials) remaining intact despite stimulation and concomitant changes in theta sequences. This interpretation would also be consistent with work in certain disease models, where reduced synaptic plasticity and working-memory deficits co-occur with preserved place coding despite impaired theta sequences (e.g., Viana da Silva et al., 2024; Donahue et al., 2025).

      We agree that traditionally deficits in alternation tasks have been termed “working memory” but we also note that this may confuse some readers, as memory in these tasks does not engage persistent activity throughout delays in areas like the prefrontal cortex.

      (3) It was not immediately clear whether SWR-associated activity was derived from the interleaved ~15-min rest sessions in a rest box, or from periods of immobility or reward consumption in the maze (aSWR, as in Jadhav et al 2012). Regardless, it would be informative to compare aSWR events within the maze to rest-box SWRs that may occur during more prolonged slow-wave episodes (even if not full sleep). This contrasts with Liu et al. (2024), who analyzed replay during ~1.5-h sleep sessions.

      We thank the reviewer for this comment and suggestion. We will now explicitly mention in the manuscript that we have measured awake sharp wave ripples (aSWRs) on the track during immobility periods. In addition, in line with this and another reviewer’s recommendation, we have included analyses on the proportion of rest SWRs (rSWRs) between control and targeted animals in Supplementary Figure 6, replicating our findings during aSWRs. However, we note that the differences between Liu et al.’s (2024) study and ours. While they waited and analyzed replay during 1.5 hours of sleep sessions, in our study, and in the 1-day w-track learning protocol, sleep sessions are typically shorter (15-20 minutes). That said, control animals in our tasks with an intact hippocampus (previous studies) and intact theta sequences (our study control animals) can learn the task in one day, so a longer replay period is not necessary to learn the task.

      Reviewer #2 (Public review):

      Summary:

      The authors of this study developed a closed-loop optogenetic stimulation system with high temporal precision in rats to examine the effect of medial septum (MS) stimulation on the disruption of hippocampal activity at both behavioral and compressed time scales. They found that this manipulation preserved hippocampus single-cell-level spatial coding but affected theta sequences and performance during a spatial alternation task. The performance deficits were observed during the more cognitively demanding component of the task and even persisted after the stimulation was turned off. However, the effects of this disruption were confined to locomotor periods and did not impact waking rest replay, even during the early phase of stimulation-on. Their conclusion is consistent with previous findings from the Pastalkova lab, where MS disruption (using different methods) affected theta sequences and task performance but spared replay (Wang et al., 2015; Wang et al., 2016). However, it differs from a recent study in which optogenetic disruption of EC inputs during running affected both theta sequences and replay (Liu et al., 2023).

      Strengths:

      The experiments were well designed and controlled, and the results were generally well presented.

      Weaknesses:

      Major concerns are primarily technical but also conceptual. To further increase the impact of this study by contrasting findings from different disruptions, it is necessary to better align the analysis and detection methods.

      We thank the reviewer for their assessment and critical questions. We have addressed each of the comments below. As we note in our responses, our inclusion criteria were based on our analysis approach, in which we aimed to measure the impact of our manipulation where possible for each animal, and ideally at the level of every 20-minute run epoch. This is a strength of our experimental approach, and we will explicitly explain that in a next version of the manuscript.

      Major concerns:

      (1) To show that MS disruption does not affect spatial tuning, the authors computed the KL divergence of tuning curves between stimulation-on and stimulation-off conditions. I have two main questions about this analysis:

      (1.1) The authors seem to impose stringent inclusion criteria requiring a large number of spikes and a strong concentration of tuning curves. These criteria may have selected strongly spatially tuned cells, which are typically more stable and potentially less vulnerable to perturbations. Based on the Figure 2 caption, it seems that fewer than 10% of cells were included in the KL divergence analysis, which is lower than the usual proportion of place cells reported in the literature. What is the rationale for using such strict inclusion criteria? What happens to the cells that are not as strongly tuned but are still identified as significant place cells?

      We thank the reviewers for this question. For the KL divergence analysis, we have imposed a spike-count criterion (100 spikes for each interval type —stimulation-off, stimulation-on, and the stimulus sub-interval) and a coverage criterion (50% HPD of the units’ spatial firing distribution was contained within 40cm on the linear track and 100cm on the w-track). These criteria were chosen to ensure that spatial tuning curves were sufficiently well sampled and localized to allow reliable estimation of KL divergence, which is particularly sensitive to noise arising from low spike counts or diffuse firing. Based on the reviewer’s suggestion, we have relaxed the unit inclusion criteria for KL divergence by relaxing the criteria for the number of spikes (50 spikes) to include more weakly tuned place cells and replicated our results (p=.19, control n=147; targeted n=47).

      Further, we have also evaluated the stability of place field order between stimulation-on and stimulation-off conditions using more standard methods (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). These results are consistent with our observations about place field stability during stimulation-off and stimulation-on conditions (Fig. 2F).

      Modified in results:

      “By contrast, we did not detect any changes in the spatial tuning of putative pyramidal neurons. Specifically, the increased spiking during stimulation-on intervals respected place field boundaries (Supplementary Figure 2B), and the place fields of single cells were not detectably affected by theta disruption (Figure 2E-F, single cell place fields, pooled for G). To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.68, t-test). These results are consistent with previous studies that performed MS manipulations (Zutshi et al. 2018; Etter et al. 2023). Thus, our manipulation protocol provides an opportunity to ask specifically how temporal coding contributes to behavior.”

      (1.2) The KL divergence was computed between stimulation-on and stimulation-off conditions within the same animal group. However, the authors also showed that MS stimulation had lasting effects on theta sequences and performance even during stimulation-off periods. Would that lasting effect also influence spatial tuning? Based on these questions, the authors should perform additional analyses that directly measure spatial tuning quality and compare results across control and experimental groups - for example, spatial information of spikes (Skaggs et al., 1996), tuning stability, field length, and decoding error during running.

      To assess these observations at the population level, we computed the similarity between the place field peaks in control versus targeted animals and could not detect a difference between the two conditions (as in Wang et. al., 2015; Spearman correlation of place field order, control vs targeted, Linear-track, p = .68, t-test, n=6 control and n=4 targeted; W-track p = 0.91, n=36 control, n=18 targeted). Mean decoding error between animals depends on the recording quality. We have reported these values around stimulus times at the choice point in the previous version of the manuscript (Sup. Figure 3G). We also find that the distribution of place field coverage is not statistically different within and across animals (stimulation on versus stimulation off: p = 0.17, control versus targeted: p = 0.9, interaction of stimulation and targeting: p = 0.8, linear mixed effects model).

      (2) The authors compared their results with those from Liu et al. (2023) and proposed that the different outcomes could be explained by different sites of disruption. However, the detection and quantification methods for theta sequences and replay differ substantially between the two studies, emphasizing different aspects of the phenomenon. I am not suggesting that either method is superior, but providing additional analyses using aligned detection methods would better support the authors' interpretations and benefit the field by enabling clearer comparisons across studies. In the current analysis, the power spectrum of the decoded ahead/behind distance only indicates that there is a rhythmic pattern, without specifying the decoding features at different theta phases. Moreover, the continuous non-local representations during ripples could include stationary representations of a location or zigzag representations that do not exhibit a linear sequential trace. Given that, the authors should show averaged decoding results corrected by the animal's actual position within theta cycles and compute a quadrant ratio. For replay analysis, they could use a linear fit (as in Liu et al., 2023) and report the proportion of significant replay events.

      In Liu et al., 2023 study theta sequences were quantified by explicitly segmenting theta cycles into phase quadrants and evaluating the structure of the decoded representations within these phase-defined windows. This approach can be applied in studies that have a stable theta that can provide a temporal reference frame and where there is not enough spatial coverage in the spikes to identify the extent of ahead/behind representations. This analysis is hence not ideal to detect theta sequences in our data, as it is prone to errors due to our disruption of theta oscillatory activity itself. To account for this, we have used a Bayesian clusterless decoding approach in which we can measure the structure of hippocampal sequential representations even without imposing any restrictions on their temporal order. This method allows us to reliably capture theta sequences during locomotion in control animals (Fig. 4B) and their disruption in targeted animals (Fig. 4F).

      Since our experimental paradigm suppresses theta oscillations themselves, we used a clusterless decoding approach (as in Joshi et al., 2023) to obtain an unbiased estimate of rhythmicity for hippocampal spatial representations. Briefly, we estimated the peak of the posterior at every 2ms time step and computed the distance between that value and the actual position of the animal (decode-to-animal distance). We confirmed that, as expected, in control animals, we could detect “theta sequences” as in prior studies without explicitly requiring theta oscillatory cycle windows.

      We have now evaluated the distributions of SWRs that are labeled as continuous and have a trajectory displacement > 10 cm and consistently find continuous replays during both aSWRs and rSWRs (Fig. 5, Fig. S6). We also find that these distributions overlap between control and targeted animals (n=1216 targeted, n=3212 control, p=0.06).

      Author response image 1.

      (3) The finding that theta sequences and performance were impaired even during stimulation-off periods is particularly interesting and warrants deeper exploration. In the Discussion, the authors claim that this may arise from "the rapid plasticity engaged during early learning." However, this explanation does not fully account for the observation. Previous studies have shown that theta sequences can develop very rapidly (Feng et al., Foster lab, 2015; Zhou et al., Dragoi lab, 2025). If the authors hypothesize that rapid plasticity during early stimulation-on disrupts the theta sequence, then the plasticity window must also be short and terminate during the subsequent stimulation-off period. Otherwise, why can't animals redevelop theta sequences during stimulation-off? The authors should conduct additional analyses during the stimulation-off periods of the W-maze task. For example:

      (3.1) What is the spike-theta phase relationship? Do the phases return to normal or remain altered as during stimulation-on?

      We thank the reviewer for this question. We have now looked at theta power on the W Track and find that theta power does not fully recover on the W Track even during stimulation-off periods. We have included this in the results. See Figure S4.

      (3.2) Is there a significant place-field remapping from stimulation-on to stimulation-off? (Supplementary Figure 3F includes only a small subset of cells; what if population vector correlations are computed across all cells, or Bayesian decoding of stimulation-on spikes is performed using stimulation-off tuning curves?)

      We have addressed this question by computing the correlation between the peak of the place fields between stimulation-on and stimulation-off conditions and find that the distributions are largely overlapping between control and targeted animals on both the linear (Spearman correlation of place field peaks between stim-on and stim-off intervals, control vs targeted, p=0.68, t-test, n=6 control, n=4 targeted epochs) and wtrack (Spearman correlation of place field peaks between stim-on and stim-off intervals: p=0.91, t-test, n=36 control, n=18 targeted epochs).

      (3.3) The authors should also discuss why the stimulation-off epochs were not sufficient to support learning, and if the stimulation-off place cell sequences could have supported replay.

      We do not find the aSWR-associated replay to be impacted as a result of our manipulation. We have not conducted a specific experiment to test the impact of longer stimulation-off periods on the formation of place cell sequences, but in response to this and another question from Reviewer 1, we have added the following speculation in the discussion.

      Modified in discussion:

      “Critically, the behavioral effects in targeted animals were seen even though stimulation was off during the middle third of each exposure to the W-track. Consistent with this behavioral result, sequential firing during locomotion (at both the pairwise and population level) was disrupted during stimulation-on periods and remained disrupted in stimulation-off periods, indicating that the 5-6 minutes of stimulation-off trials was not sufficient to allow the system to This surprising result indicates that the disruption of theta sequences during the early experience in a novel environment is sufficient to have lasting effects, potentially by interfering with the rapid plasticity engaged during early learning. In this framework theta sequences may be particularly important for establishing task-relevant structure during the earliest phases of exploration. While we did not explicitly test the effects of pretraining or longer duration of stimulation-off periods, our results raise the possibility that pretraining the animal in the behavioral arena would allow for the development of task-relevant representations, and thereby reduce or eliminate the behavioral impact of theta disruption.”

      (4) Citations and/or discussion of key studies relevant to the current work are missing: Wang et al. in Pastalkova lab 2015-2016 studies for disruption of theta sequence (but not place cell sequence) disrupting learning but not replay, Drieu et al. in Zugaro lab 2018 study on disruption of theta sequence affecting sleep replay, Farooq and Dragoi 2019 for association between a lack of theta sequence and presence of waking rest replay during postnatal development, etc. The authors should discuss what the conceptually new findings in the current study are, given the findings of the previous literature above.

      We thank the reviewer for this question. We have substantially modified the introduction to include this prior work and highlight that our manipulation enabled us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      (5) The assessment of theta sequence is not state-of-the-art:

      (5.1) Detecting the peak of cross-correlograms between neurons (CCG) relates to behavioral timescale CCG, not the theta sequence one; for the theta sequence, the closest to zero local peak should be used instead.

      Here we think we failed to explain our analyses clearly, as we did exactly that analysis. The cross-correlation peak in Fig. 2D is the peak within the theta timescale (+/- 100ms lag), not a slow behavior timescale (e.g. +/- 1s lag). In the revised version of the manuscript, we have improved the explanation so that this confusion does not arise.

      (5.2) How were other methods of detecting theta sequences performing on the stimulation-on/stimulation-off data: Bayesian decoding, firing sequences?

      In the absence of detectable theta oscillations, it is inappropriate to use the commonly used metric of detecting theta sequences. That method also assumes that the theta oscillatory cycle is the correct temporal “reference” for hippocampal theta sequences. To our knowledge, there is no direct evidence for this. Thus, in our manuscript, we have used two approaches to identify hippocampal spatial representations during stimulation-on and stimulation-off periods:

      (1) Clusterless Bayesian decoding approach

      (2) Pairwise correlations between neurons

      Analyses using these approaches provide an unbiased method to detect hippocampal spatial-temporal sequences during locomotion without using theta oscillations as a reference. Indeed, in control animals, we recover the endogenous theta timescale correlation and sequence structure during locomotion. Using the same approach in targeted animals reveals that even though we can measure hippocampal spatial sequential representations in targeted animals, the timing between them is altered.

      (5.3) How was phase precession during stimulation-on/stimulation-off?

      We cannot do this analysis for the theta manipulation condition since there is an unreliable phase estimate in the absence of theta on W Track. Based on the reviewers’ comments, we have now evaluated the theta phase precession on the linear track 10Hz stimulation condition. Phase precession could be observed even when evaluated against the entrained LFP. Further, at a population level, we observed that autocorrelograms followed the entrained 10Hz LFP in the stimulation-on condition compared to the stimulation-off condition (n=41 neurons, solid lines are medians, shaded areas 25/75 percentiles). We have added these results in Supplementary Figure 5.

      (6) It would be important to calculate additional variables in the replay part of the study to compare the quality of replay across the 2 groups:

      (6.1) Proportion of significant replay events out of the detected multiunit events.

      We assume the reviewer is recommending the analysis of significant replays as defined by performing a linear fit on the trajectory. We recognize that there is a fundamental diversity in the replay architecture, as has been shown using clustered and clusterless decoding approaches, and that assuming that only the replays with a linear fit are significant might bias us toward those trajectories. In aSWRs, we have shown that the replay events that are labeled “continuous” are equivalent in proportion between control and targeted animals. In the current version of the manuscript, for rSWRs, we similarly computed the proportion of replay events with a continuous trajectory that traverses at least 10cm on the w-track and found no differences between control and targeted groups. We also found the proportion of continuous rSWRs to increase in targeted animals, similar to control animals. We have included these results in the new Figure S6.

      (6.3) The average extent of trajectory depicted by the significant replay events in the targeted compared to the control, stimulation-on/stimulation-off.

      Here, we show the distributions of continuous replay trajectories (> 10 cm) between control and targeted animals, showing overlapping distributions (p=0.06). See Author response image 1.

      Reviewer #3 (Public review):

      Joshi et al. present an elegant and technically rigorous study examining how the temporal structure of hippocampal spiking during locomotion contributes to spatial learning. Using a closed-loop, theta phase-specific optogenetic manipulation of medial septal parvalbumin-expressing neurons in rats, the authors demonstrate that disrupting theta-timescale coordination impairs performance on the cognitively demanding component outbound trajectory of a spatial alternation task, while sparing hippocampal replay, place coding, and the simpler inbound learning. The work aims to dissociate the role of theta-associated temporal organization during navigation from sharp-wave ripple-associated replay during subsequent rest periods, providing a mechanistic link between theta sequences and learning. The findings have important implications for models of septo-hippocampal coordination and the functional segregation between online (theta) and offline (SWR) network states. That said, there are a few conceptual and methodological issues that need to be addressed.

      We thank the reviewer for the thoughtful comments and suggestions to strengthen our work. In a revised manuscript, we explicitly address prior publications where manipulations (either pharmacological or behavioral) have shown a dissociation between temporal sequence formation, place coding, and replay in the introduction and highlight the specific insights obtained from using our approach. Specifically, our approach (closed-loop theta-phase stimulation during locomotion) provides a level of physiological specificity that enables dissociation of theta-state dynamics from other hippocampal processes. This, in turn, allows us to address a question that has remained unresolved across prior studies: Are hippocampal spatial sequences during locomotion (i.e., theta sequences) necessary for learning a novel hippocampal-dependent task?

      One concern is the overall novelty of this work; the dissociation between online temporal sequence and offline replay events following memory deficits has previously been shown by Wang et al., 2016 elife. While the authors discuss Lui et al., 2023, which demonstrates MEC activation of inhibitory neurons at gamma frequencies during locomotion disrupts theta sequences, subsequent replay and learning (line 65-66), they do not reference Wang et al., 2016 who performed a very similar study with MS pharmacological inactivation, and report large decreases in theta power, attenuated theta frequencies together with behavioural deficits but SWR replay persisted. Given strong similarities in the manipulation and findings, this study should be discussed.

      We agree that this important study should be cited in the introduction, and have now included that in the revised introduction. Importantly in Wang et al., 2016 elife paper, they showed that while replay can exist after periods where theta sequences are disrupted using muscimol inactivation in the medial septum, the rate of replay was higher than in control, leaving open the possibility that the increased replays might contribute to consolidation of non-task memories. We also note that the manipulation was applied for hours, making it impossible to attribute a precise physiological basis for poor behavioral performance.

      Along the same lines, it should be noted that Brandon et al. (2014, Neuron) demonstrated that hippocampal place codes can still form in novel environments despite MS inactivation and loss of theta, indicating that spatial representations can emerge without intact septal drive. Referencing this study would strengthen the discussion of how temporal coordination, rather than spatial coding per se, underlies the learning deficits observed here.

      Thank you. We have referenced the study appropriately in the updated manuscript.

      Our findings, and previous dissociations between precise timescale and place field properties (Petersen and Buzsáki 2020; Liu et al. 2023; Wang, et al., 2016, Brandon et al., 2014) further suggest that different circuits with different time constants are responsible for processing spatial and temporal information in the hippocampal circuit. Spatial/contextual information may arrive from regions with slower timescales (such as the cortex), making them less susceptible to sub-second brief disruptions, while precisely timed inputs from the medial septum coordinate the tightly controlled timing offsets between hippocampal neurons. We hypothesize that learning requires the intersection of these two streams of information in the hippocampal network, and is impaired by the disorganization of the precise temporal templates in which internal plans can be matched to external inputs.

      The conclusion that disrupting "theta microstructure" impairs learning relies on the assumption that the observed behavioral deficits arise from altered temporal coding from within hippocampal CA1 only. However, optogenetic modulation of medial septal PV neurons influences multiple downstream regions (entorhinal cortex, retrosplenial cortex) via widespread GABAergic projections. While the authors do touch on this, their discussion should expand to include the network-level consequences of entorhinal grid-cell disruption and how this could affect temporal coding both online and offline.

      We agree and have expanded our previous discussion to include that possibility.

      Modified in discussion:

      “Understanding precisely why this temporal organization is critical will require more distributed measurements. Notably, MS targets include multiple cortical and subcortical targets (Joshi 2017), and our manipulation may have disrupted precise spike timing throughout these regions. Key amongst these regions include the pre- and para-subiculum, retrosplenial area and the entorhinal cortex, which also receive dense PV projection in addition to the CA3 and DG (Joshi et. al., 2017; Viney et. al., 2018; Salib et al., 2020). Disrupting the spatial code or spike-timing in these regions may contribute to the disruption of sequential activity we have observed. However, we do note that the first response of the stimulation to spiking activity in CA1 is consistent with a strong disinhibitory input to CA3, with spike latencies less than 20 milliseconds. Additionally, monitoring regions beyond the temporal cortex would be informative given the broad coordination between hippocampal theta and other systems (Joshi et al. 2023; Eichenbaum 2017; Buño and Velluti 1977; Berg, Whitmer, and Kleinfeld 2006; Ledberg and Robbe 2011).”

      The finding that replay content, rate, and duration are unchanged is critical to the paper's claim of dissociation. However, the analysis is restricted to immobility on the track. Given evidence for distinct awake vs. sleep replay, confirming that off-track rest and post-session sleep replays are similarly unaffected would confirm the conclusions of the paper. If these data are unavailable, the limitation should be acknowledged explicitly. Moreover, statistical power for detecting subtle differences in replay organization or spatial bias should be added to the supplement (n of events per animal, variability across sessions).

      We thank the reviewer for this suggestion. We have now explicitly evaluated replay properties during rest and indeed confirm that replay rate, length, and content in this manipulation are indistinguishable between targeted and control animals. We have now added statistical power to these claims by reporting the number of events per animal and variability across sessions. As we mentioned above, where possible, we have attempted to replicate each analysis per session (20min session), and all analyses are replicable for each targeted and control animal. Our internal controls are the power of our approach, as there can be significant animal-to-animal variability.

      The exact protocol for optogenetic stimulation is a bit confusing. For the task, the first and final third (66%) of trials were disrupted and were only stimulated when away from the reward well and only when the animal was moving. What proportion of time within "stimulated" trials remained unstimulated? Why were only 66% of trials stimulated?

      We developed a stimulation protocol that included an interval in which stimulation was not applied (stimulation-off) periods to assess the impact of the stimulation at the level of each animal and epoch. This analytical approach gives us the power to study the impact of the stimulation and our experimental approach to the spiking patterns and neural activity observed. We will modify the explanation in the methods. Based on the reviewer’s comment, we have now calculated the proportion of time within the epoch where the laser was on. The laser is ON for ~10% of the total time in the epoch. We implemented a closed-loop algorithm that had both the spatial location of the animal and the theta phase of a reference electrode as online inputs. The trigger was applied when three conditions were met: a spatial inclusion criterion, a speed criterion, and theta phase criteria.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1D & G: should include the frequency band of the filtered trace in the figure caption.

      We have now included the frequency band of the filtered trace (5-11Hz) in the figure caption.

      (2) The observation that hippocampal cells respond to stimulation within ~5-10 ms (Figure 2C), while theta power decreases significantly after 200 ms (Figure 1G), is interesting. Do the authors have any hypothesis explaining this discrepancy?

      Theta is likely best understood as the result of a complex feedback loop between the medial septum and various structures in the hippocampal formation. This would make it robust to individual perturbations. We now discuss this in the text.

      Modified in results:

      “Note, that while we can measure the response to hippocampal cells within 5-10ms, theta power decreases gradually over a period of 200 milliseconds. This is consistent with the view that theta oscillatory activity that can be measured in the hippocampus is a result of a multi-region feedback loop that involves various cortical and subcortical networks, a feature that may make it robust to individual perturbations.”

      (3) Figure 2A: The caption states that "in control rats, endogenous theta sequences are apparent." This is unclear. The purple region marks the ascending phase, but theta sequences are typically defined across peak-to-peak cycles. The shading may distract from this; consider adding dashed lines to indicate theta peaks.

      We thank the reviewers for this suggestion, and we have modified the visualization for this figure per the reviewers’ suggestion.

      (4) Figure 2B (top panel): How are the cells ordered? It appears they are sorted by peak firing time before stimulation (<0). This could misleadingly suggest a sequence before stimulation but not after. Some cells have peak firing in the 0-40 ms range-what is the intended interpretation of this ordering?

      We have modified the spiking by the spiking order after stimulation onset.

      (5) Line 239: Typo - "as for the linear track" should be "as for the W track."

      Modified.

      (6) Lines 279-282: The statement "This suggests a surprising level of preserved representational movement across frequencies ranging from 6 to 12 Hz" is unclear. What does "preserved representational movement" mean? A simpler explanation for the matching of power-spectrum peaks to stimulation frequency could be that stimulation increases firing rates, biasing decoding toward locations with higher mean firing, producing rhythmic fluctuations at the stimulation frequency under Poisson decoding assumptions.

      We have now added examples of the decode to animal distance in the different stimulation conditions to supplement this statement. We will also acknowledge that the change in firing rate might contribute to the observation.

      (7) Line 313: The phrase "while sparing learning on the interleaved trials where a less cognitively demanding choice was required" may be confusing. Although the authors refer to inbound runs, readers may interpret "interleaved trials" as stimulation-off trials.

      We agree and have modified this phrasing.

      (8) The claim that "representations of locations more distant from the animal were preserved during theta disruption" (line 350) is unclear-where is this shown?

      Fig S3G: Distribution of max decode to animal distance within theta cycles

      Consistent with the pairwise analyses (Supplementary Figure 3B-D), the ∼8 Hz peak in the power spectrum of the ahead/behind distance was significantly larger in control animals than in targeted animals across all conditions (Figure 4I; pooled comparison p’s < 10<sup>−4</sup>, hierarchical bootstrap p’s < 0.05). At the same time, the maximal extent of locations represented ahead and/or behind the animal near the choice point, where choices must be made on outbound and inbound trials, did not differ between control and targeted animals (Supplementary Figure 3G). Thus, our findings indicate that the precise timing of non-local representations during theta was disrupted, but the spatial extent was not.

      (9) The title of Figure S1 appears twice.

      Thank you. We have edited it.

      Reviewer #3 (Recommendations for the authors):

      (1) It should be noted that there are several referencing errors that should be addressed. Please check the following:

      (a) Lines 77-81 - a few incorrect references.

      (b) Line 155 - Disruption of hippocampal theta has consistently shown to preserve spatial properties of place cells: Brandon et al., 2014, Koenig et al., 2011.

      Thank you. We have corrected these references.

      (2) Figure 1C- Did the authors stain for colocalization with virus and PV in the MS?

      Yes, our viral construct has eYFP expressed together with channelrhodopsin, and this rat line and viral construct have been previously standardized (Yu et al., 2018, Lepperod et al., 2021). In addition, we conducted three standardization experiments and visually inspected the overlap between eYFP-positive cells and parvalbumin-expressing neurons (86/86 YFP-expressing neurons tested positive for PV). An example of the overlap is now included in Supplementary figure 7.

      (3) Figure 1 - Shows an impressive reduction of theta power. Can Figure 1G be extended to show theta recovery immediately following stimulation?

      Our stimulation protocol restricted the stimulation to periods that were within the spatial and speed inclusion criteria. The period immediately following the stimulation on every trial is reward delivery, during which the animal has already slowed down, and we do not expect high theta power. Based on the reviewer’s suggestion, we have inspected the first few trials after the stimulus is turned off in the linear track and have confirmed that theta power immediately recovers. We have added these additional figures in Sup. Figure S4H.

      Author response image 2.

      (4) Do authors have examples of the same trajectory where temporal coding is intact in baseline and disrupted during stimulation? Does an intact theta sequence ever develop in the target animals?

      Yes, target animals do exhibit intact spatial sequences during the stimulation-off periods on the linear track and w-track (example in a targeted animal during the stimulation-off condition on linear track below). As we show in Figure 4E and Supplementary Figure S4G, on the w-track, while spatial sequences exist, each cycle’s duration is not consistent across the behavioral experience. Thus, on average, power spectrum of the ahead-behind distance is not rhythmic at 8Hz.

      (5) Have the authors computed spike-phase relationships or shown phase-position to evaluate phase precession in individual cells in both the 10 Hz stim vs the phase-specific?

      As discussed above, due to unreliable phase estimates during theta suppression, we have chosen to base our analysis of sequential structure on temporal cross-correlation and decoding analysis. However, in response to the reviewer’s question, we have evaluated phase precession under 10Hz stimulation. Consistent with our overall results for the ahead-behind distance, we find that individual cells phase-precess with the newly entrained theta. Additionally, at a population level, we are able to visualize a clear shift in peak spiking frequency.

      However, we agree that these results do not completely rule out the contribution of additional spikes, and in a revised version of the manuscript, we have included that possibility.

      (6) Did authors perform other types of stimulations that either drove the dominant frequency out of theta range (gamma) or completely desynchronize the system (by stimulating along all different times of theta to perform a phase-specific "scramble")?

      We did not attempt a gamma or scrambling stimulation condition.

      (7) There is overall inconsistency in the formatting of references throughout the text.

      We apologize for these errors and have rectified them in the updated version.

    1. Author response:

      The following is the authors’ response to the original reviews

      Correction: In the process of revising the preprint, we discovered that for the 15SY dataset that a single time point (ZT2) out of the 12 timepoint series was inadvertently combined with temporally adjacent time samples (ZT20, 22, 24). We corrected the accompanying GEO submission (Series GSE145509). With this update, we repeated the rhythm analysis with an updated RAIN algorithm as the original Boot-eJTK could not be run as it was outdated with dependent packages no longer maintained or supported, and some are no longer available through standard package managers. The analysis with the corrected ZT2 sample and did not find any significant changes in the major claims of the papers. One minor change is that we no longer observe significant DR-dependent increases in proteasome gene levels. Nonetheless, we still find DR-dependent cycling of proteasome genes and proteasome module network connectivity consistent with the DR-sensitivity of the proteasome pathway. Figure 4 has been updated to reflect this change. After correcting this issue, we revised the manuscript in response to the reviewer comments.

      eLife Assessment

      This study describes important findings on how a core component of the circadian clock impacts the effect of dietary restriction (DR) on longevity and fecundity in Drosophila, which lead the authors to postulate rhythmic control of proteostasis in the fat body as a critical aspect of DR effects. The evidence presented is still incomplete, not fully supporting the conclusions of the study, as alternative hypotheses/explanations have not yet been systematically explored. The work will nevertheless be of substantial interest to researchers working in circadian and cell biology, metabolism, and aging, with an interesting hypothesis to be explored further.

      We sincerely thank eLife for considering our manuscript and express our gratitude to all three reviewers for their time and constructive comments. While acknowledging that there are alternative hypotheses for some of our findings, which make our evidence incomplete, we appreciate that our manuscript was recognized as important and of substantial interest. We have revised the manuscript to acknowledge that the major findings under light-dark conditions could be attributable to being driven by light rather than the circadian clock. We add new data demonstrating that the far majority of cycling genes in LD in control flies are disrupted in Clk<sup>Jrk</sup> consistent with circadian regulation (see also below). Future investigations are needed to test the alternative hypotheses and explanations. Nevertheless, we believe that the manuscript still represents a meaningful advancement in understanding how molecular circadian clocks in peripheral tissues interact with diet to influence systemic lifespan and aging.

      Public Reviews:

      Reviewer #1 (Public Review):

      The studies by Hwangbo et al. diligently attempt to account for many of the typically neglected dietary and non-dietary factors.

      Strengths:

      - Work addresses many potential artifacts of dietary (e.g., dehydration stress, macronutrient ratios, and protein source) and non-dietary (e.g., leaky expression of S106-GAL4) manipulations-important factors that are too often overlooked.

      - Balanced and complementary behavioral, molecular, and bioinformatic experiments

      - Show necessity of proteostatic subunits in the fat body for DR-mediated longevity. The findings in the current manuscript lay the ground for future studies that test sufficiency of fat body prosβ3 and rpn7, or necessity of other proteostatic genes in other tissues.

      Weaknesses:

      - Could the lack of DR response in clock mutants across dietary concentrations be simply because the clock mutants are better at compensatory feeding adjustments to dietary dilutions? If this were the case, there are two major implications to the authors' conclusions:

      a) The Clk mutants are differently responding to dietary dilutions, not to dietary restriction, per se.

      b) Nutritional intake was unaffected by the dietary manipulations. If the changes in fat body proteostasis and lifespan were due to nourishment, it would be expected that the physiology and lifespan do not change.

      Accurate measurements of food consumption and the resulting protein intake could potentially clarify this critical question.

      We thank the reviewer for their positive feedback and also appreciate their raising the important issue of whether Clk^Jrk mutants may be more effective at compensatory feeding. Xu et al. (2008) reported that overall food consumption in Clk^Jrk flies was indistinguishable from control flies.

      We also directly assessed food intake in ~1 week old flies on three diets (1% SY, 5% SY, 15% SY) over two days (48 hours) using the Con-Ex method (Shell et al. 2018). We observed either no significant changes or relatively modest changes in food consumption between iso31 controls and Clk^Jrk flies that are limited compared to the large differences in caloric content between the diets. While we cannot rule out changes in feeding patterns throughout the lifespan, we believe that minimal differential compensatory feeding in the Clk^Jrk mutants are not sufficient to be the primary cause of the lifespan differences observed across diets. These data are added as new Figure 1-figure supplement 3.

      Reviewer #1 (Recommendations For The Authors):

      Hard to find information:

      - Type of yeast used. Should be addressed in the Methods section at least, instead of having to dig through several paragraphs into the Results section.

      Missing information:

      - Agar type and concentration

      - Mifepristone diets: pipetted on top or mixed into food?

      The relevant information has been updated in the Materials and Methods

      Wrong information:

      - Line #159-160: "Fig. 1 and Fig. S3" should be "Fig. 1 and Fig. S2" and then "(Fig. S3)" added to the end of the sentence.

      This information has been corrected in the revision

      Presentation:

      - Paragraphs are very long.

      - Interactions should be denoted by ×, not * or x.

      - Remove markers from mortality graphs. Having bulky markers reduces perceived differences between curves.

      These changes have been updated in the revised version.

      Reviewer #2 (Public Review):

      Dietary restriction (DR) increases lifespan, an effect that has been consistently observed in several organisms, but we still lack a clear mechanism to explain this phenomenon. In this work, Hwangbo et al. revisited the role of the circadian clock in DR-mediated lifespan effects. They found that the increase in lifespan produced by DR is missing on a clock mutant, a clock dependency that is also observed at the level of nutrient-dependent egg laying. By conducting RNA-seq with an impressive temporal resolution, they showed that DR triggers an increment in the number of cycling genes expressed in the fat body, the fly functional analog of the mammalian liver. Interestingly, from these genes, a group of them are de novo daily expressed genes, meaning that their expression was not rhythmic under the control diet but appear rhythmically expressed under DR. Among those, genes encoding proteasome subunits are enriched. The authors finally showed that adult-specific knockdown of these genes in the fat body prevents the increase in lifespan under DR, further supporting a role of the proteasome in this process. Overall, the conclusions are mostly supported by the evidence presented, and the authors' discussion nicely frame their results with other research in the field.

      Strengths:

      - Many studies have limited their observations of DR on lifespan to a few dietary conditions which makes the reach of some previous conclusions somewhat limited. The dilution strategy that the authors used in this work provides a strong indication that the effect of DR on lifespan relies on clock expression regardless of the conditions used. Furthermore, the inclusion of the egg-laying assay is a good addition to support this hypothesis.

      - Because the strength of the rhythmicity statistics relies heavily on the number of data points collected, the temporal resolution used for the RNA-seq experiments (every 2 hrs per 48hrs) is remarkable. This allows exquisite dissection of the phase of rhythmic genes in different conditions. The dataset produced in this work might be of use to other groups interested in weighting the role of other represented gene clusters in DR.

      We are grateful for the reviewer’s positive feedback regarding the robust experimental design in the manuscript.

      Weaknesses:

      I see only minor flaws in this work, that if addressed, might strengthen the authors' conclusions, particularly:

      - The results of the lifespan assays are quite variable and in some instances contradictory (Fig. S8) across trials, possibly because there are other unaccounted variables we still do not understand. The fecundity assay, in contrast, seems to be a better readout (Fig. 2). Confirming at least the two genes picked for the study (Fig. 5) would be good support for the claim that the proteasome mediates the effects of DR.

      We appreciate the reviewer’s comment of seeing “only minor flaws”. We reiterate that we focused on those results which were replicated across trials, providing confidence in the overall conclusions. Nonetheless, we agree that exploring the proteasome role on the DR effect on fecundity would be intriguing and may complement the lifespan data. We now acknowledge this point in our discussion.

      - According to the model, the acute effect of DR on gene expression is related to CLOCK protein function. However, I am not sure how this link was established. It is tempting to assume that CLOCK upstream is the reason for having an increase in rhythmic genes under DR, but the experiments did not test this. The tests conducted either assessed the role of clk or the effect of an impaired proteasome on DR-dependent extension of lifespan. Thus, it is difficult to assert the authors' claims on the link between CLK and the changes in cycling genes and to the proteasome upon DR.

      We have updated our text and model figure suggesting a direct CLK role in the Discussion.

      Reviewer #2 (Recommendations For The Authors):

      As mentioned before, the experiments, in particular the RNA-seq datasets are excellent. Additionally, the discussion provides a good overview of other relevant papers on DR, and the conclusions are mostly supported by the data. Here I provide a couple of suggestions that I believe might improve this work:…

      - Although the model is simple and understandable (Fig. 6), the inclusion of an overall summary or explanation in the figure legend would be appreciated, especially for readers that are not familiar with the terminology.

      - It might be the formatting while parsing the files but some of the in-text citations are between curly brackets (e.g., lines 80, 92, 93).

      - By definition, and unlike Canton-S or Oregon-R strains, w1118 flies are not wild-type but a genetic control. I believe that reference to this on the figures and text may need correction.

      We updated and/or our corrected each of these in the revised version. For wild-type, we more explicitly define this as wild-type for the relevant genetic locus.

      Reviewer #3 (Public Review):

      In this study, Hwangbo and co-workers investigate the extent to which the well-established life extending effects of DR rely on the molecular circadian clock and how the landscape of clock-controlled gene expression changes in the face of DR within the fat body of the fly, a tissue that performs the functions associate with both the liver and adipose tissue of mammals. The authors evidence that DR extends lifespan in a manner that depends on only one of the two major limbs of the fly's molecular circadian clock, namely the positive limb, that DR produces major changes in the identities of cycling clock output genes, and that genes related to the proteosome represent a major component of DR-induced transcript cycling. Though interesting, these conclusions are not strongly supported by the data and there are two major reasons for this. First, the authors rely on only one loss of function genotype each for the loss of positive and negative limb clock gene function. Second, though they wish to address the "circadian transcriptome" under normal and DR conditions, the authors conduct all their work under strong Light/Dark cycles, making it impossible to address circadian phenomena. These shortcomings are problematic in the extreme, as they leave open obvious alternative explanations for the results and fail to directly determine if the rhythmic expression, they observe are clock controlled or merely driven by the light/dark cycles, which themselves produce major effects on activity, feeding, etc., that may be responsible for differentially driving rhythmic transcripts under normal and DR conditions in the fat bodies.

      Major Weakness One: The use of only genotype each for the loss of positive (Clk^JRK) and negative (Per^01) limb of the circadian represents a major challenge for a central conclusion of the study. Phenotypes caused by the loss of a single clock gene may be due to the loss of circadian timekeeping, or they may represent a pleiotropic effect of the loss of function mutant being used. There are multiple precedents for pleiotropic (non-circadian) effects of clock gene mutants. It is, therefore, possible that the differences in the extent of DR mediated life extension between Clk^JRK and Per^01 may not represent a difference between breaking the positive and negative limbs of the clock but may simply reflect a pleiotropic effect of the dominant negative Clk^JRK. This possibility is acknowledged by the authors (lines 343-344). This could be addressed quite easily by extending the analysis to other loss of function mutants, for example, tim01 for the negative limb and cyc01 for the positive. Given the central focus here on the "circadian transcriptome," leaving open this alternative explanation for Clk's role in DR induced life extension represents a major weakness of the study. Furthermore, given the fact that Clk^JRK appears to be short lived on most of the media tested in the study, is it really surprising or informative that they would display lower life extension under DR?

      We confirmed that the large majority of LD oscillating genes in wild-type controls are disrupted in ClkJrk consistent with circadian clock regulation (Figure 3-figure supplement 1). As noted, we formally acknowledged that the circadian clock mutant alleles used here, and in fact any circadian clock alleles, can have pleiotropic, i.e., non-circadian, clock effects. This would only be partially mitigated by adding more (but also potentially pleiotropic) clock mutant alleles. Very challenging circadian resonance experiments (see Xu et al, 2019) are the gold standard for resolving circadian clock v. non-clock effects which are beyond the scope of this study which we now add to our discussion.

      We also note that foxo mutants are both short-lived and exhibit a robust lifespan extension to dietary restriction and thus the ClkJrk mutant is distinct in this regard. We have added this point to the Discussion.

      Major Weakness Two: The authors have not established that any of cycling transcripts they have detected in the fat body under normal and DR conditions are driven by the circadian clock. This is because: 1.) they have conducted their transcriptomic analysis on cells taken from flies entrained to light dark cycles, which can themselves drive daily changes in expression levels and 2.) they have not shown that the cycling measured on normal diet or DR conditions depends on a functional circadian clock. The "significant reorganization of the circadian transcriptome" is presented as a major conclusion of this study, but the authors have not addressed circadian control of transcription at all here, either by an examination of transcription under free-running conditions and/or in loss of function clock mutants.

      In addition, there is a logical gap in this study. The authors have shown that DR produces less life extension in Clk^JRK mutants than Per^01 or wild-type controls. They then show that DR produces changes in the rhythmic transcriptome when flies are place on DR. The central model presented in Fig. 6 shows/concludes that CLK drives increases in proteome-related transcript rhythms under DR. This conclusion could have been directly tested by asking if the changes in rhythmic gene expression induced by DR are gone the loss of function Clk mutants, or if the transcriptomic landscapes fail to differ between feeding conditions in these mutants.

      In conclusion, the study falls far short of directly testing the ideas it puts forth, greatly limiting its impact and interest.

      As noted above, we also examined the diurnal transcriptome in ClkJrk (at 4 hour resolution) and found that of the 290 genes that were detectably rhythmic in wild-type just 13 were rhythmic in ClkJrk consistent with the notion that oscillations depend on Clk (Figure 3-figure supplement 1). We now add this analysis to the manuscript. Nonetheless, we cannot exclude a role for light and thus have opted to use “diurnal” in place of “circadian” where appropriate for observed rhythms under LD conditions.

      Reviewer #3 (Recommendations For The Authors):

      Line 140 "showed an almost identical response" was a little hard to understand at first. Consider clarifying.

      This has been rephrased for clarity

      The authors claim that Clk mutants are "much longer lived" than wild-type controls on two of the relatively low calorie diets. Figure S3C certainly argues otherwise, and it's not clear how the data in 1C and S2C and warrant the use of "much" here.

      This wording has been rephrased and corrected in the revised version. The low-calorie diet shown in Figure 1-figure supplement 3 contains a higher sucrose concentration (5%) than those used in Figure 1 and Figure 1-figure supplement 2 (1%). This observation suggests that sucrose may play an independent role in the survival of ClkJrk mutants under malnutrition conditions.

      The authors should provide the rationale for the use of a dominant negative form of Clk for their experiments. Would the available amorphic allele be a better choice?

      As ClkJrk is the first described Clk allele and it is probably the most well characterized. As a dominant negative version which is still capable of dimerizing and binding DNA it is less susceptible to compensation by redundant bHLH transcription factors as has been observed for between mouse Clock and NPAS2 (Debruyne et al, 2006).

      It is not clear why the authors have chosen to examine transcriptomes so soon after transfer to DR. Why not wait longer. The authors provide context that changes are already taking place at the early time-point used, but would waiting a bit provide a more robust indication of how DR is changing the fat body?

      We noted in the manuscript that the effects of DR on survival are evident relatively soon (~2d) after a diet shift. We were interested in identifying those changes in daily transcription that would be occurring during that early time span and potentially be a cause rather than an effect of survival changes.

    1. Author response:

      In response to the valuable reviewers’ comments and suggested changes, we are finalising changes to the manuscript, in order to resubmit a revised version, and a document with full author responses, that reflects all the review comments.

      These changes include adding points of clarity, improving accuracy on wording of key messages, adding additional interpretation of the data, including additional data and analyses that reflect open questions raised, and more discussion concerning these unanswered questions, which are subjects of future work.

      In response to specific points we were asked to provisionally address (actions in italics):

      Reviewer #1. We are pleased the reviewer sees the insight these data bring. We indeed think it likely that cilia axoneme orientation is governed by the immediate microenvironment and/or changes to the cytoskeleton and are actively looking to explore this.

      To address the 3 areas of concern:

      (1) Our re-writing of the abstract and the results concerning transcriptomic data seeks to overcome weaknesses in descriptions of these data.

      (2) Further detail is being added on how z-distortion is corrected for, so accuracy is the same in all axis and orientation measurements are robust.

      (3) We will add to the discussion to add our thoughts as to why ciliary and cilia signalling genes are regulated by immobilisation, but that immobilisation does not apparently affect cilia structure.

      Reviewer #2. Thank you for such broadly positive comments, we are pleased the scale and depth of the quantitative analyses comes across, but will make sure that revisions throughout improve the quality of the writing describing these. We are actively exploring means to test ideas for how cilia axoneme become orientated in this way and what the function is. These preliminary ideas will be reflected in the discussion.

      Reviewer #3. Thank you for such a detailed and thoughtful review. They will ensure the data are presented to their very best and we will address the concerns raised. This descriptive study had a hypothesis, and made discoveries which surprised us, we have tried, as you say, to put this in some context of the role of cilia and the role of mechanical forces in GP biology. Most notably we are considering that uncoupled is not the correct term here. To address areas of concern;

      (1) We agree ‘uncoupled’ is not the correct word here. We cannot find a pattern of correlation between centriole position and orientation. However, the two can’t be ‘uncoupled’ and there is no proven independence on a single-cell level (that cilia position and cilia orientation are not in any way mechanistically linked). We will make changes and add more details on what we have considered in this area.

      (2) Similarly, we have not correlated in each cell, cellular orientation and ciliary orientation, though we have made attempts and not yet found a relationship. However, again, this is not the same as one being absent. We might have expected ciliary orientation to change as cell orientation does (through zones or with pertubations) as we have seen in vitro but this remains to be fully explored and is one subject of follow-up work. We are considering column populations and per animal considerations of the data.

      (3) We did take a cautious approach to statistics and specialist advice, but advice was not to overcomplicate things when there are 2 main messages related to centriole position and cilia orientation. Firstly, centriolar position appears random or without preference thus distribution of position on cell, is homogenous. Second, ciliary orientation angle is not a homogenous distribution, as would be expected if random with this number of measurements. We do, and will add comments to this effect, have to mindful of large dataset, but do not think this means we are looking at false discoveries due to number of comparisons. We are considering how better to reflect this. To address these important points we will add a section to the methods and discussion and will endeavour to change results to this end as it is a central point of the manuscript.

      (4) We will add new data concerning IFT88cKO and centriole position and orientation.

      (5) We will add a critique of the immobilisation experiments to ensure the relatively diminished power is clear. We agree force may have set things up initially, we will ensure this is discussed and we will ensure our proposal for why cilia orientation is this way is framed as a hypothesis. We have preliminary data, but this is the subject of an entire new project thus not yet supported by robust experimental evidence so is speculative at this stage and we will ensure this is clear.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The study addresses the organisation of synaptic connections from the medial to the lateral entorhinal cortex. Classic anatomical work has suggested these connections exist, but very little is known about their identity or functional impact. The manuscript argues that these projections are mediated by glutamatergic neurons, providing excitatory input from MEC to all layers of LEC, and by SST+ve interneurons sending inhibitory projections to L1 of LEC. This appears to be the most likely interpretation of the data, although in my opinion, more could be done to rule out the possible impact of the spread of the virus/tracer from the injection site.

      While this concern might seem overly picky, the importance of this level of detail is nicely shown by the authors' previous work clarifying connectivity from postrhinal to entorhinal cortices through careful analysis of similar types of data (Doan et al. 2019). If additional analyses/data can address the concern here, then I think this will be an important set of fundamental results that will influence thinking about circuit mechanisms for spatial cognition and episodic memory. In particular, it will nicely add to an emerging view that MEC and LEC can interact directly, showing that the organisation of these interactions is asymmetric and identifying a potentially interesting long-range inhibitory pathway.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Nilssen et al. presents a comprehensive study of the circuitry linking the medial and lateral entorhinal cortices (MEC and LEC). Using a combination of anatomical tracing, optogenetics, and in vitro electrophysiology, the authors convincingly demonstrate that the MEC sends both glutamatergic and long-range inhibitory SST+ GABAergic projections to the LEC, with distinct laminar and cell-type-specific targeting. Notably, they reveal that SST+ inhibitory projections selectively suppress the activity of layer IIa neurons, whereas excitatory inputs preferentially engage neurons in layers IIb and III, thereby differentially modulating hippocampal-projecting populations.

      Strengths:

      The experiments are carefully executed, the results are compelling, and the conclusions are well supported by the data. This work will be of broad interest to researchers studying memory circuits, cortical inhibition, and the organization of long-range connectivity.

      Weaknesses:

      Although the in vivo relevance of these connections remains to be determined, this is an important and timely contribution to our understanding of entorhinal-hippocampal interactions.

      The request for validation of injection specificity and viral spread, as detailed in the comments and suggestions of the two reviewers has been provided in the revised version. We added supplementary figures 1,2 and 6 as well as an extra insert into the old supplementary figure 5, now supplementary figure 9.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Interpretation of the retrograde labelling experiment in Figure 1A-C is challenging, as the spread of the tracer at the injection site is not shown. It's important to see the full dorsal-ventral extent of the injection site in order to establish that the labelling of neurons in MEC results from projections to LEC and not adjacent areas (including MEC).

      As mentioned in our initial reply, we are fully aware of the risks associated with an incomplete assessment of injection sites and viral spread, so we provide a new Supplementary Fig. 1 showing 6 dorsoventral levels of the case shown in Fig. 1A. The injection site in case of FG often shows a core of damaged tissue with a halo of substantial unspecific fluorescence. Outside of the injection side, one only sees retrogradely labeled somata (often recognizable by FG signal clustered in lysosomes) and dendritic elements. As indicated in the legend, we report some tracer leakage along the needle track in temporal and perirhinal cortex, areas that receive only sparse MEC projections, but there is no apparent spread of the injection into MEC.

      Numbers are also quite low. E.g., Figure 1A is N=1/2 for retrograde labelling experiments.

      The reviewer would be correct if the experiments were meant to analyze MEC projections to LEC in full anatomical detail. This was not our intention (several papers addressed this pathway in detail), we merely aimed to distinguish between glutamatergic and potential GABAergic contributions to this pathway and to establish optimal coordinates in slices to prepare for the electrophysiological recording experiments. Two animals suffice for this purpose and using more would be against the aim of reducing the use of experimental animals as much as possible.

      The rationale here for the use of AAV2-CAG-tdTomato as a retrograde tracer is unclear. My understanding is that this is more effective as an anterograde tracer. Some clarification and validation would be important.

      The reviewer is correct that AAV2 is generally considered an effective anterograde tracer, but tracing the connectivity of entorhinal cortex with AAVs has been proven to be notoriously difficult, in particular retrograde tracing of inputs to layer II. In a neighboring lab in the centre, headed by Edvard and May-Britt Moser, it was established that AAV2 types show very efficient retrograde transport and that is why we decided to use the virus. Also, in our hands the virus showed excellent retrograde transport that served our purpose

      (2) Interpretation of the anterograde experiments in Figure 1D-E would also benefit from showing evidence that the injection sites are restricted to MEC. It should be straightforward to make a supplemental figure showing labelling at all dorsoventral levels.

      More careful analysis of the axon labelling in the dentate gyrus could also help make a case for the selectivity of the injection site for eGFP. In this case, only the intermediate portion of the molecular layer of the DG should be labelled. In the image shown, the labelled band is quite wide, but it's hard to tell if this reflects the plane of section or is because it also includes labelling in the outer molecular layer (which would be indicative of LEC expression).

      We thank the reviewer for these two suggestions, and we have prepared a new Supplementary Fig. 2 in line with this.

      Numbers are also on the low side for these experiments.

      See our response above

      (3) For optogenetic experiments in Figure 2, the selectivity of targeting of AAV to MEC is assessed through the specificity of labelling in the DG. This is great, but it's important to show that this specificity is maintained at all dorsoventral levels.

      Higher resolution images of labelling in LEC could also be helpful. It's hard to tell from the images in 2A if labelling is axonal or is in the soma adjacent to the nuclear NeuN signal (which would indicate a lack of selectivity for MEC).

      We thank the reviewer for these two suggestions and provide a new Supplementary Fig. 6, showing both the details of AAV1 being present only in neuropil in MEC not in somata as well as the specific labeling in the middle molecular layer of DG in detail. Including all dorsoventral levels would not provide additional information in view of the very well-established topographical organization of the entorhinal to dentate projection, reaching approximately 20 -25 % of the full long axis of DG (Van Groen et al., 2003)

      In addition, we have again carefully screened all tissue from the electrophysiological experiments for possible leakage of virus from MEC to LEC. We decided to exclude recordings from one mouse, which had labelling in MEC that was close to the border with LEC. Neuron counts have therefore been adjusted (pages 7-9) and the example recording showing responses to TTX/4-AP exposure in Figure 2B has been exchanged.

      (4) The analysis of excitatory and inhibitory opto-responses in Figure 2 is nice. It may be helpful to report quantification of the rise and decay kinetics of the synaptic currents. They appear much slower for the inhibitory input, which may be functionally important.

      This would indeed be nice to add, but it would not significantly impact or change the main message of our study. Since the lab of the senior author (MPW) has been discontinued and the resources for conducting these analyses are not readily available anymore, we have found it difficult to comply with the reviewer’s request

      (5) More direct evidence for SST axons projecting from MEC to LEC would strengthen the conclusions made. E.g., in experiments where the SST neurons are labelled, is it possible to follow the axons? Do they project as expected from the MEC to the LEC?

      In our view the tracing data provide convincing evidence in support of a direct projection from MEC to LEC by SST neurons, as shown in horizontal brain sections where SST axons labelled in MEC of an SST<sup>Cre</sup> mouse projects within Layer I from the site of origin in MEC to Layer I of MEC (Supplementary Figure 3). Similar visualizations were not possible to obtain in our electrophysiological experiments where semicoronal slices were used. This cutting angle has been shown to be optimal to preserve most of the axon and the dendritic tree of LEC neurons (Tahvildari and Alonso, 2005; Canto and Witter 2012), but does not maintain the projection from MEC to LEC.

      Minor Points:

      (1) "These layers are heavily innervated by medial entorhinal axons (Figure 1F...". I don't see a 1F.

      This has been corrected; should have been Figure 1E.

      (2) Methods should report series resistance values for patch-clamp experiments (range and mean).

      Fully agree and this information has now been added on page 22 of the manuscript:

      Under Voltage clamp: ‘Recordings with series resistance ≤ 25 MΩ were accepted, with an average of 16.2 MΩ for voltage clamp recorded neurons (range, 4.9 – 24.9 MΩ).’

      Under Current clamp: ‘All recordings (series resistance: 18.9 MΩ, 5.0 – 66.7 MΩ; mean, range) were included for analysis.’

      Reviewer #2 (Recommendations for the authors):

      (1) Please specify in the figure or, alternatively, in the figure legend which virus was used in each group shown in Figures 2H and 2I. This is somewhat confusing, since Figure 2E illustrates a specific combination of viruses and mouse lines that only corresponds to part of Figure 2H. While this information is provided in the text, including it directly in the figure would help the reader.

      We thank the reviewer for this excellent suggestion, and we have implemented this in the new version of figure 2.

      (2) In Figure 4, regarding the inputs from PIR, cLEC, and PER to LEC, the inhibitory components recruited by each input were not examined as thoroughly as for the MEC inputs. In fact, some inhibitory interneurons were double-labeled in the GAD67 mice (Figure 4B), which could also influence the responses of LEC neurons, especially for PER inputs. Recordings in Figure 4C appear to have been obtained near the reversal potential for inhibition, which may have prevented the observation of inhibitory effects. The authors could discuss this point in the Results.

      The reviewer is correct and this issue is now addressed in the relevant section in the results (page 11):

      ‘It should be noted, however, that it is possible that inhibitory effects could have been masked in some recordings, due to the resting membrane potential in our recordings being close to the theoretical chloride equilibrium potential. This could be particularly relevant for the inputs from PER, an area where we found LEC-projecting GABAergic neurons (Fig. 4B) and which is known to provide long-distance inhibition to LEC (Pinto et al., 2006; Apergis- Schoute et al., 2007).’

      (2) A diagram summarizing the known connections among MEC, LEC, and the hippocampal formation, highlighting the relevant cell types, layers, and the new connections identified in this study, would be a valuable addition, perhaps as a supplementary figure.

      We appreciate the suggestion, though find a full summary of known connectivity a bit overdone. Instead, we included a new figure 6 that summarizes the main new findings of the paper in the context of LEC projections to the hippocampal formation.

      (3) Although the main focus is on MEC-LEC connectivity, the experiments examining interactions with other cortical areas and converging inputs would benefit from a discussion of how MEC-driven inhibition of LEC might influence those inputs and shape the resulting output to the hippocampus. Including a short paragraph addressing this in the Discussion section would strengthen the manuscript.

      Excellent suggestion although we did speculate briefly in the result section on the possible effect. We have added a short paragraph in the discussion (page 15), reiterating the part in the results (page 13, last paragraph of results). We also briefly discussed the potential functional relevance of the suppression of the pathway from layer IIa to DG-CA3/CA2 versus the facilitation of activity in the pathway from layers IIb/III to CA1 and subiculum (last section of the discussion).

      (4) Lastly, it would be interesting to know what the main source of activation is for the SST long-range LEC projecting neurons. Are these neurons recruited in a feedback manner by the activity of MEC excitatory cells? I realize this question is beyond the scope of the present study, but if the authors have any data or insights related to this point, including a brief discussion would be valuable.

      This is an interesting thought, and we have included a new supplementary figure (supplementary Figure 5) showing data from experiments mapping monosynaptic inputs to MEC SST neurons using rabies virus. Although it was not possible to target only MEC SST neurons that project to LEC, the data show which are the main extrinsic inputs to the population of MEC SST neurons, most likely including those that project to LEC.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This is an important study of critical period plasticity, focused on temperature manipulations, and how different parts of the Drosophila larval motor circuit adapt or maladapt. The work convincingly demonstrates that components of the motor network respond in distinct ways to the heat shock, and the combination of functional, structural, and electrophysiological approaches makes the study of significant interest. The work points to central interneurons as primary drivers of maladaptive changes, while motoneurons and neuromuscular junctions show compensatory or homeostatic adjustments. The study is methodologically rigorous, contributing important insights into critical period biology using a tractable invertebrate model.

      We thank the reviewers for their thoughtful critique and suggestions. We agree with these and, where possible, we have attempted to address these, improving this study. As outlined below, we have revised the manuscript substantively and included additional data and figures.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors examine the impact of heat stress during an embryonic CP in Drosophila, focusing on the larval locomotor network. They show that elevated temperature increases neuronal activity and, when applied during the CP, results in long-term instability of the network, which manifests in prolonged seizure recovery times. At the neuromuscular junction, substantial structural changes occur, including terminal overgrowth and altered receptor composition, yet synaptic transmission remains preserved due to homeostatic regulation. Motoneurons display reduced excitability but receive increased synaptic input from premotor interneurons. These findings suggest that maladaptive instability originates within the central circuitry rather than at the neuromuscular junction, where changes seem to be homeostatically compensated. The study concludes that different network components exhibit distinct and hierarchical responses to CP perturbations, with premotor interneurons setting the tone for downstream adjustments in motoneurons.

      Strengths:

      The work takes advantage of the unique accessibility of the Drosophila system. A major strength of the study is the integration of structural, physiological, and behavioral analyses, which allows the authors to draw a comprehensive picture of how CP perturbations shape the locomotor network. The choice of an ecologically relevant stimulus (heat stress) is particularly convincing, as it links experimental manipulations more closely to natural environmental conditions. The experiments are carefully designed, and the results are robust and consistent with previous findings in the field, while also extending them in new directions.

      Weaknesses:

      The study leaves some uncertainty regarding the experimental design and interpretation. The change from short to prolonged heat shock manipulations raises the possibility that the effects observed may not be confined to the critical period alone - this could be experimentally addressed or simply rephrased in the text.

      We agree that clarity about the experimental paradigm is important and have addressed this as suggested, within text, figures and figure legends: the duration of embryo exposure to 32˚C heat stress is now unambiguously stated and each figure has a graphical illustrations of the heat stress paradigm. For example, experiments represented in Figures 1, 3 (new data) and 8 (new data) used short, defined periods of a few hours of heat stress, aimed to identify specific windows of development that are sensitive to 32˚C heat stress. These also show that behavioural changes result from heat stress experienced during the specific 2-hour window that defines the critical period of the developing central locomotor circuitry, namely from 17-19 hours after egg laying - previously identified by Giachello & Baines (2015). Longer exposure to 32˚C heat stress during embryogenesis result in the same phenotypes when this 2-hour window is included, causing the same level of reduced larval crawling speed and lowered network stability, which manifests in increased seizure recovery times. This is as one might expect from a critical period of nervous system development.

      Where the neuromuscular junction is concerned, where we identified embryonic heat stress causing phenotypes that are evident at late larval stages, a more complex model has emerged. Following suggestions from both reviewers to explore shorter heat stress exposures during embryogenesis, additional experiments (see Figure 3) we identified what might be a critical period for the body wall muscles. This is an earlier window of development, within 13-16 hours after egg laying, which is sensitive to heat stress in terms of the levels of the GluRIIA glutamate receptor subunit that will be expressed in the late larva. This developmental period is characterised by muscles acquiring their electrical properties (Broadie & Bate, 1993), i.e. comparable to the central locomotor network transitioning through its critical period at the time that it becomes active. The neuromuscular junction is composed of both presynaptic motoneurons and postsynaptic muscles, and therefore this composite structure is subject to multiple, sequential critical periods. The characterisation of changes to neuromuscular junction synaptic physiology was carried out using heat stress throughout most of embryogenesis (Figure 4). We think this appropriate from the perspectives of having included all relevant critical periods (muscle and CNS) to explore how this composite structure responds to environmental heat stress; also based on our observations that for each critical period phenotypes are defined by the experience during the critical period and not exacerbated by prolonged heat stress either side.

      In addition, the maladaptive (seizure recovery) and adaptive/homeostatic phenotypes are not always clearly distinguished or highlighted, which makes it harder to appreciate how the different levels of the network plasticity fit together into a single mechanistic framework.

      Following the suggestion, we have tried to clarify the mechanistic framework in a new figure that aims to summarise the model in Figure 9.

      The question of whether phenotypes that result from an embryonic heat stress manipulation are adaptive or maladaptive is difficult to resolve. This is partly due to the nature of critical periods, since perturbations during these developmental windows can cause significant, long-lasting maladaptations that are challenging to reconcile from a perspective of adaptive plasticity. Secondly, in light of the nature of this animal, which has evolved a particularly rapid development and large brood sizes, any deviation from the evolved optimum developmental temperature of 25˚C could constitute a reduction in fitness. We interpret the phenotypes we see along those lines: network instability that results from critical-period perturbations is a manifestation of a sub-optimally tuned network, as is a reduction in larval crawling speed.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a thoughtful and well-executed study of critical period plasticity in the Drosophila larval motor circuit. The authors examined how transient heat, 32 {degree sign}C, during the embryonic stage, altered network properties, showing that premotor interneurons A27h increase excitatory drive onto motoneurons, which respond with a reduction in excitability. At the NMJ, synaptic terminals expand and GluRIIA distribution shifts, yet synaptic transmission remains largely unaffected. Despite these local compensations, the treated larvae display slower crawling and prolonged recovery from seizures, indicating that the network is functionally compromised.

      Strengths:

      (1) One of the major strengths of this study is the elegant dissection of a defined circuit, tracking changes from premotor interneurons through motoneurons to the NMJ. The multimodal approach provides a comprehensive view of how connected elements respond to CP perturbations.

      (2) An interesting finding is that NMJ morphology changes dramatically without corresponding deficits in synaptic transmission, challenging the common assumption that larger boutons necessarily indicate stronger synapses.

      (3) Another intriguing result is that even with two layers of homeostatic compensation, locomotor behavior is still impaired, highlighting the limits of compensation and underscoring the critical role of CP timing.

      (4) Beyond these scientific insights, the study benefits from a well-defined, tractable system and simple experimental manipulations, which together make the results highly interpretable and reproducible.

      Weaknesses:

      There are a few areas where the manuscript could be strengthened.

      (1) Although A27h premotor neurons are well characterized, the claim that they are the causal driver of downstream changes would be strengthened by additional experiments or a clearer discussion of the temporal hierarchy.

      We have tried to clarify the model of the temporal hierarchy (new Figure 9). This is a model and as such will hopefully help us collectively to think about this system and underlying processes, while also inviting this perspective to be challenged. The model we propose is compatible with our observations, namely that the premotor circuitry might change in response to a critical period heat stress (e.g. increasing their synaptic drive onto motoneurons), followed by homeostatic adjustment by the postsynaptic motoneurons (e.g. by reduction of their excitability), thus serving to maintain overall normal motoneuron firing patterns (see Figure 6).

      However, synaptic communication is commonly regulated in both antero- and retrograde directions. Therefore, while compatible with the observations we have made, bi-directional information flow could also be instructive during the CNS critical period.

      (2) While 32 {degree sign}C heat stress is presented as ecologically relevant, it produces maladaptive behavioral outcomes, raising questions about the ecological and mechanistic interpretation of the model. In particular, most experiments, with the exception of Figure 1, used prolonged (24h) heat treatments, which could introduce developmental effects beyond the CP itself. Comparing shorter and longer heat exposures would help clarify the specificity of the CP response.

      We agree. For a detailed response on this point, please response to Reviewer #1 above.

      (3) While there are schematics for experimental procedures, a circuit diagram tracing information flow and indicating where structural and functional changes occur would help readers better understand the findings.

      We have created Figure 9 as a working model.

      (4) Finally, the main paradox of the study, that robust homeostatic compensations occur yet behavior remains impaired, could be explored in more depth in the Discussion.

      We have tried to address this in the discussion.

      Reviewer #3 (Public review):

      Summary:

      During development, neural circuits undergo brief windows of heightened neuronal plasticity (e.g., critical periods) that are thought to set the lifelong functional properties of underlying circuits. These authors, in addition to others within the Drosophila community, previously characterized a critical period in late fly embryonic development, during which alterations to neuronal activity impact late-stage larval crawling behavior. In the current study, the authors use an ethologically-relevant activation paradigm (increased temperature) to boost motor activity during embryogenesis, followed by a series of electrophysiology and imaging-based experiments to explore how 3 distinct levels of the circuit remodel in response to increases in embryonic motor activity. Specifically, they find that each level of the circuit responds differently, with increased excitatory drive from excitatory pre-motor neurons, reduced excitability in motor neurons, and no physiological changes at the NMJ despite dramatic morphological differences. Together, these data suggest that early life experience in the motor neuron drives compensatory changes at each level of the circuit to stabilize overall network output.

      Strengths:

      The study was well-written, and the data presented were clear and an important contribution to the field.

      Weaknesses:

      The sample sizes and what they referred to throughout the distinct studies were unclear. In the legends, the authors should clearly state for each experiment N=X, and if N refers to an NMJ, for example, instead of an individual animal, they should state N=X NMJs per N=X animals. This will help readers better understand the statistical impact of the study.

      This is a good point. For the majority, each data point is derived from a unique specimen, unless explicitly stated otherwise, for NMJ size on muscle DA1 (Figure 3) and for larval crawling data, where each larva was measured up to three time, once per unique 5-minute crawling interval.

      Recommendations for the authors:

      Reviewing Editor Comments:

      In addition to revising the text and making interpretive changes as suggested by the reviewers, we invite you to consider the following:

      (1) Either rephrase the conclusions on the role of the 2h CP and discuss the effects of temperature during embryonic development. Alternatively, to validate the idea of a longer CP window, directly compare the results of a few key experiments using the 2h and 24h heat treatment.

      We have addressed this within the text, as suggested. In the text and figure legends, we have made clear distinctions between exposure to 32˚C heat stress during most of embryogenesis vs a specific developmental window of a few hours. In figures, we have provided diagrams that graphically illustrate the period of heat stress exposure.

      Our ability to experimentally test differences between precise vs broader heat stress periods during embryonic development have been constrained due to the departure of scientists, who were able to carry out electrophysiological recordings (as also explained below). As the next best alternative, we focus on imaging and behavioural analyses. These demonstrated that the developing body wall muscles are sensitive to heat stress during an earlier phase, from 13-16 hours after egg laying, when the body wall muscles become electrically active. It precedes the critical period of the central locomotor circuitry (17-19 hours after egg laying), when neurons in the CNS become electrically and synaptically active (16 hours after egg laying). We think this an exciting additional insight, demonstrating sequential critical periods as different parts of the locomotor network become active: first the body wall muscles, followed by the central circuitry.

      (2) Clarifying the homeostatic responses and shedding light on how they engage with the maladaptive changes described would benefit the study. Furthermore, adding more information about the anatomical and structural changes and how they relate to the intrinsic and synaptic changes would also benefit the study.

      We have tried to address this within the text and with a summary diagram (Figure 8), as suggested.

      Reviewer #1 (Recommendations for the authors):

      (1) It remains unclear whether the authors want to conclude that reduced network stability is not due to changes at the motoneuron level, but rather at the premotor level. Although this idea is mentioned in the results and discussion, it does not appear in the abstract or introduction, which leaves the different findings disconnected. Clarifying and highlighting this conclusion throughout the manuscript would strengthen the narrative.

      We have added additional experiments and changed the manuscript to address this point. These showed that there are distinct phases of embryonic development during which heat stress causes changes to NMJ structure vs to larval behaviour (seizure recovery times/network stability and crawling speed) - additional data in suppl. Fig. 2 and Fig. 7). In the Drosophila embryo, the body wall muscles develop and acquire their electrical properties before central neurons do, and these phases correlate with sensitivity to heat stress.

      (2) In the results section related to Figure 1, the logical link between the CP protocol and the functional assessment of the locomotor network at different temperatures is not sufficiently explained. It is not clear what this assessment is meant to test or demonstrate. A more explicit statement of the rationale and correction of what seems to be a typographical error in the final sentence of the paragraph would help to clarify the authors' intent.

      We have tried to rectify this by changes in the manuscript and to Figure 1, to make the sequence of panels more intuitive.

      (3) In the second results section, the experimental strategy shifts from using a short 2-hour heat shock to a 24-hour manipulation. The reasoning - that short manipulations in different windows yield no phenotype - is understandable, but a 24-hour perturbation may have broader consequences beyond the CP, simply by virtue of its longer duration. Moreover, 24h is roughly the duration of embryonic development at 25C. When at 32C, embryos should develop faster; therefore, is the 24h heat shock extending to L1?

      Yes, the 24 hour heat stress extends into the first few hours of the L1 larval stage.

      In order to validate the use of a longer window, the author should show how it affects the developmental time. Moreover, one should test that a few important observations remain the same with 2h and 24h heat perturbation. Alternatively, one cannot conclude that the phenotype is due to the rather narrow previously defined CP rather than to other effects associated with the overall embryonic developmental time and coordination. This would not make the results less interesting, but it would be important to assess whether the effects can be solely attributed to the 2h CP.

      We have compared the impact of heat stress experience during the majority of embryogenesis, including the CP that had been defined for the central locomotor network (17-19 hours after egg laying) with shorter heat stress manipulations during consecutive phases of embryogenesis until larval hatching. As outlined above in response to point (1) by the Reviewing Editor, reduced stability of the central network and associated reduction in larval crawling occurs when heat stress is experienced during the CP of the central locomotor network (17-19 hours after egg laying). Prolonged heat stress experience for 24 hours leads to indistinguishable outcomes, as long as this 2-hour CP window is included (see Fig. 1 and Fig. 8).

      However, this suggestion by Reviewer #1 led us to identify a second CP for the body wall muscles (see Fig. 3). NMJ overgrowth and changes to the postsynaptic glutamate receptor composition result from earlier heat stress experiences, and those are comparable to the effects caused by 24-hour heat stress exposure when this earlier muscle CP is included.

      Therefore, NMJ development is affected by consecutive CPs, an earlier one linked to body wall muscle development, followed by a later one that impacts the presynaptic motoneurons and their upstream circuitry. Nevertheless, the larval NMJ and behavioural phenotypes that we have identified appear to result from sensitivity to heat stress during these respective CP windows, with no clear evidence of cumulative effects on these phenotypes resulting from longer heat stress exposure during embryogenesis.

      (4) In session 3, the authors note that GluRIIA reductions were most pronounced in proximal regions of the NMJ. However, this is not explicitly quantified in the figures or methods. Including such quantification, or clarifying where it can be found, would make this observation more convincing.

      We have analysed anti-GluRIIA signal intensities in proximal vs distal boutons, comparing different ROI selection processes (e.g. thresholding to a full NMJ/anti-HRP mask and to an anti-GluRIIB mask, which is more selective to postsynaptic sites). Analysis of multiple data sets did not show statistical significance, but instead confirmed that comparable reductions in anti-GluRIIA signal manifest in both proximal and distal boutons, following an embryonic 32C heat stress, relative to controls. We have therefore removed relevant speculative statements.

      (5) In session 4, the authors conclude that motoneurons undergo a decrease in excitability to adjust to greater premotor drive. Is there anatomical evidence for this, such as an increase in input synapses?

      We previously quantified change in excitatory presynaptic synaptic contact number onto aCC motoneuron dendrites in third instar larvae following an embryonic pharmacological activity manipulation: overexcitation of the developing network following introduction of PTX via feeding to gravid females*. No significant structural changes were seen. Although this is a different manipulation of the developing network, all our data to date suggest that heat stress manipulations during the embryonic critical period signal via the same pathways, at least in part due to temperature increases leading to activity increases. Because such a quantification is technically challenging and extremely time-consuming due to the low level of marking individual motoneurons, we did not think it informative or in scope for this project.

      *See Figure 5 in this publication: Hunter I, Coulson B, Pettini T, Davies JJ, Parkin J, Landgraf M, Baines RA. Balance of activity during a critical period tunes a developing network. Elife. 2024 Jan 9;12:RP91599. doi: 10.7554/eLife.91599. PMID: 38193543; PMCID: PMC10945558.

      The interpretation of the optogenetic experiments would also benefit from clarification. If motoneurons are less excitable yet receive more drive, one might expect no net change, rather than the differences observed. Alternatively, could the excitability of the premotor neuron itself have changed, either intrinsically or in relation to Chronos expression? Measuring premotor activity directly during optogenetic activation could help to resolve this ambiguity.

      These are good suggestions. Yes, we think that motoneuron excitability has changed as a result of heat stress - see paper submitted in parallel and published since: Sobrido-Cameán et al., 2025, PLoS Biology. Unfortunately, the team members, who could have carried out this type of analysis had moved on by submission of the manuscript. Therefore, we were no able to experimentally pursue these questions further.

      (6) In session 6, the authors report slower propagation of premotor activity waves after CP heat stress, but the logic of the experiment is not sufficiently explained. How does this finding relate to the enhanced premotor drive described earlier? Only timing is quantified; information about amplitude and wave dynamics would strengthen the interpretation. These results could also be discussed in relation to Figure 1B, where acute heat stress increased motoneuron activity. One interesting possibility could be that CP manipulations might adaptively prepare the larva to function at different temperatures. Experiments testing wave propagation at 32 {degree sign}C (Figure 6) or, conversely, motoneuron activity after CP manipulations (Figure 1B paradigm) would provide valuable evidence for such an adaptive role.

      As per above, unfortunately, the team member who could have carried out this type of analysis had moved on by submission of the manuscript. However, we have tried to address the question of whether there is an adaptive element to the adjustments that result from embryonic heat stress experience. Specifically, we carried out behavioural tests on how larvae respond with changes in crawling speed to acute changes in ambient temperature (new Fig. 7).

      We interpret our findings as follows: that heat stress during embryonic development leads to sub-optimal outcomes with regard to network stability as well as default and maximum crawling speed. The precise causes for this will be difficult to unpick. Behaviourally, when challenging larvae with an acute change in ambient temperature, we saw that slow crawling larvae do respond comparatively normally to a relative increase in ambient temperature by speeding up (in effect an escape response). This demonstrates that embryonic heat stress causes a change in the default crawling speed, while principally maintaining behavioural responses to changes in ambient temperature. It appears that animals that had experienced heat stress during embryonic development, by adjusting their default speed downward, maintain a dynamic response range into a higher temperature range than controls (35C vs 29C, respectively). Potentially, this could be an adaptive outcome to living at higher temperatures, though such an interpretation would require a body of work. 

      Nevertheless, every aspect we have assayed suggests that heat stress experience during the CP leads to sub-optimal outcomes: of network instability, slower default and slower maximal crawling speeds.

      Minor points:

      (1) Abstract: "has suboptimal outcomes;" should be corrected to "has suboptimal outcomes,".

      Corrected

      (2) Abstract: "we find that transient embryonic..." would improve readability with a capitalized "We".

      We are unsure about the sentence this refers to. If this sentence, then we suggest that this could remain as was.

      "Within the central nervous system, we find transient embryonic CP perturbation leads to increased synaptic drive from premotor interneurons to motoneurons..."

      The intention here is to differentiate between changes within the CNS vs at the NMJ.

      (3) Abstract: The sentence "Present the larva ... as an experimental model system..." overstates novelty, as the system has already been established in prior work. Instead, this study could highlight temperature manipulation as an ecologically relevant way to probe CPs.

      We have adopted this suggestion.

      Reviewer #2 (Recommendations for the authors):

      I would recommend:

      (1) Perform additional experimental support and dissuasions for the causal role of the premotor neuron and the network disability.

      Unfortunately, it has not been possible to carry out additional e-phys experimental work due to key people having moved on and now unable to carry out such experiments, and no replacements in sight to do so. Instead, we have focused on other work that we could do, namely to test the effect of different heat stress windows during embryonic development on GluRIIA vs GluRIIB expression at postsynaptic sites. This shows that indeed the effect seen following a 24-hour heat stress is replicated by a much shorter window of heat stress. For NMJs GluRII composition the critical period is different from the critical period of the CNS. This replicates the different developmental timings of maturation: the body wall muscles express ion channels and attain their electrical properties several hours before central neurons. We have provided these additional data as a new supplementary figure to Fig. 2. We have changed the main text to note the caveat of longer heat stress manipulations potentially leading to additional or more exacerbated phenotypes.

      (2) I would suggest that the authors expand the discussion on why two layers of homeostatic adjustment fail to preserve behavior. Is this simply a limit of plasticity?

      This is a difficult aspect to address well, beyond the purely speculative and potentially confusing. We would like to suggest that to do so requires a basic understanding on what pressures neurons/networks respond to (heat-caused over-activation and/or metabolic); and from that perspective to gauge how they adjust to those pressures, i.e. what the adjustments are trying to "achieve". This in turn should inform on whether such adjustments are homeostatic or anti-homeostatic in nature, and whether there are limits beyond which we consider a system "breaking".

      (3) 32C heat was described as "ecologically relevant". However, it produces maladaptive outcomes. The author should consider reframing it as a stressor that reveals CP sensitivity, rather than an adaptive signal.

      This is a good suggestion, and we have implemented changes accordingly.

      (4) It is important to distinguish the effects of transient (2h) vs prolonged heat exposure to confirm the manipulation targeted CP specifically.

      We have tried to make these distinctions clearer within text and figure legends. As per response to (1) above, we generated and analysed additional data to show that changes in GluRIIA are induced during a defined shorter developmental time window, not exacerbated by prolonged heat stress exposure during embryogenesis.

      (5) I would definitely recommend adding a diagram tracking information flow and showing where structural and functional changes occur.

      This is a helpful suggestion, which we have tried to implement with a new figure (Fig. 8).

      Overall, this is a strong and well-written paper that produced some unexpected results, and added a solid model circuit to study CP plasticity at the circuit level.

      Reviewer #3 (Recommendations for the authors):

      I identified one typo: "activity manipulations during the embryonic CP are artificial, We asked to what extent" The "W" of "We" should be lower case.

      Now corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Wang, Zhou et al. investigated coordination between the prefrontal cortex (PFC) and the hippocampus (Hp), during reward delivery, by analyzing beta oscillations. Beta oscillations are associated with various cognitive functions, but their role in coordinating brain networks during learning is still not thoroughly understood. The authors focused on the changes in power, peak frequencies, and coherence of beta oscillations in two regions when rats learn a spatial task over days. Inconsistent with the authors' hypothesis, beta oscillations in those two regions during reward delivery were not coupled in spectral or temporal aspects. They were, however, able to show reverse changes in beta oscillations in PFC and Hp as the animal's performance got better. The authors were also able to show a small subset of cell populations in PFC that are modulated by both beta oscillations in PFC and sharp wave ripples in Hp. A similarly modulated cell population was not observed in Hp. These results are valuable in pointing out distinct periods during a spatial task when two regions modulate their activity independently from each other.

      The authors included a detailed analysis of the data to support their conclusions. However, some clarifications would help their presentation, as well as help readers to have a clear understanding.

      (1) The crucial time point of the analysis is the goal entry. However, it needs a better explanation in the methods or in figures of what a goal entry in their behavioral task means.

      We appreciate Reviewer 1 pointing out this shortcoming and will clarify the description in the revised manuscript. Each goal is located at the end of the arm, and is equipped with a reward delivery unit. The unit has an infrared sensor. The rat breaks the infrared beam when it enters the goal. Figures 1 and 2 have been updated to clearly indicate the time of goal entry. The main text and methods have been updated with the explanation.

      (2) Regarding Figure 2, the authors have mentioned in the methods that PFC tetrodes have targeted both hemispheres. It might be trivial, but a supplementary graph or a paragraph about differences or similarities between contralateral and ipsilateral tetrodes to Hp might help readers.

      We appreciate this suggestion, which has led to an interesting finding. The coherence and burst coordination were similar for ipsi- and contralateral PFC and hippocampus. Interestingly, we found PFC beta activity was more coherent within each PFC hemisphere compared with across hemispheres. This was observed for coherence and burst time. This suggests there is hemispheric localization of beta oscillations. These results are shown in Fig. 2-1.

      (3) The authors have looked at changes in burst properties over days of training. For the coincidence of beta bursts between PFC and Hp, is there a change in the coincidence of bursts depending on the day or performance of the animal?

      This is now reported in Fig. 3-3. After quantifying the proportion of independent and coincident bursts as function of experiment day or performance, we found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (4) Regarding the changes in performance through days as well as variance of the beta burst frequency variance (Figures 3C and 4C); was there a change in the number of the beta bursts as animals learn the task, which might affect variance indirectly?

      The difference in the burst count across days did not explain the results. We performed a permutation test (Fig. 4-4), where we randomly shuffled the day identity to control for the count difference across days. The change in variance remains significant.

      (5) In the behavioral task, within a session, animals needed to alternate between two wells, but the central arm (1) was in the same location. Did the authors alternate the location of well number 1 between days to different arms? It is possible that having well number 1 in the same location through days might have an effect on beta bursts, as they would get more rewards in well number 1?

      The central arm remained the same across days since we needed the animals to learn the alternation task. In our experience, the animal needs a few days to learn the alternation rule when we switch the central arm location. For this experiment, we were interested in the initial learning process, and we kept the central arm constant. Switching the central arm location is a great suggestion for a follow-up experiment where we can understand the effects of reward contingency change on beta bursts.

      (6) The animals did not increase their performance in the F maze as much as they increased it in the Y maze. It would be more helpful to see a comparison between mazes in Figure 5 in terms of beta burst timing. It seems like in Y maze, unrewarded trials have earlier beta bursts in Y maze compared to F maze. Also, is there a difference in beta burst frequencies of rewarded and unrewarded trials?

      We performed the analysis and found burst timing was similar between the two mazes (Fig. 4-2). Bursts on rewarded trials occurred later than those on unrewarded trials (Fig. 4). Interestingly, PFC bursts on rewarded trials were lower in frequency compared with unrewarded trials. CA1 bursts during rewarded and unrewarded trials had similar frequencies (Fig. 4-3).

      (7) For individual cell analysis, the authors recorded from Hp and the behavioral task involved spatial learning. It would be helpful to readers if authors mention about place field properties of the cells they have recorded from. It is known that reward cells firing near reward locations have a higher rate to participate in a sharp wave ripple. Factoring in the place field properties of the cells into the analysis might give a clearer picture of the lack of modulation of HP cells by beta and sharp wave ripples.

      As recommended, we quantified the mean speed, mean distance to goal locations, and spatial information for CA1 cells (Fig. 7 J-L). We found SWR reactivated CA1 cells had higher speed and were spiking further away from the goals compared with non-reactivated CA1 cells. This is consistent with prior work that shows SWR-associated reactivation in dorsal CA1 can correspond to trajectories taken as the animal moves towards goals. In intermediate CA1, the content of reactivations is biased toward place representations closer to goals (Jin et al., 2024). CA1 cells with or without phase locking to beta oscillations had similar spatial firing properties.

      Reviewer #1 (Recommendations for the authors):

      (1) Please make a figure representing what the goal entry means in Figure 1.

      We have updated Fig. 1 to clearly show the definition of goal entry. We also edited the main text and methods to better explain the definition of goal entry.

      (2) For Figure 1-1, please either change the contrast of the pictures, or define the lesioned areas, as it is a bit difficult to see the lesioned parts, especially in PFC.

      We have increased the contrast for Fig. 1-1 for the histology to better show the lesions.

      (3) For Figure 1-2, is it possible to show a beta burst from Hp?

      Yes, we provided two examples of beta bursts from the hippocampus alongside burst examples from PFC in Fig. 1-2.

      (4) Is it possible to make a supplementary table showing the number of tetrodes recorded from each animal per day, plus the number of isolated single cells?

      Yes, we have included tetrode counts in Table 1-2 and cell counts in Tables 5-2 to 5-4.

      Reviewer #2 (Public review):

      (1) When presenting the power spectra for the representative example (Figure 1), it would be appropriate to display a broader frequency band-including delta, theta, and gamma (up to ~100 Hz), rather than only the beta band.

      We agree the extended frequency range provides a better overview of the spectral characteristics during the goal period. We have now included example spectrograms up to 100 Hz to show the spectral content for a wider range of frequencies (Fig. 1-2). Further, we have included additional analyses to compare the spectral characteristics between periods when the animal was moving on the maze or immobile at the goal, for frequencies up to 100 Hz (Fig. 1C-H, Fig. 1-3 and 1-4). We used both Welch’s periodogram (Fig. 1-3) and continuous wavelet transform (Fig. 1-4) to demonstrate our findings on beta oscillations in both regions are robust and consistent.

      What was the rat's locomotor state (e.g., running speed) after entering the reward location, during which the LFPs were recorded?

      Because goal entry is defined as the time the animals break the infrared beam at the goal (response to Reviewer 1), the rat would have come to a stop. We have added the time-aligned speed profile to the spectra and raw data examples in the manuscript (Fig. 1B, Fig. 1-4, Fig. 2A, and Fig. 6A). In addition, we added the quantification of the animal’s speed at the time of beta bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1D).

      If the rats stopped at the goal but still consumed the reward (i.e., exhibited very low running speed), theta rhythms might still occasionally occur, and sharp-wave ripples (SWRs) could be observed during rest.

      We typically find low theta power in the hippocampus after the animal reaches the goal location and as it consumes reward. Reviewer 2 is correct about occasional theta power at the goal. To compare differences in LFP characteristics between maze running and goal locations, we added additional analyses in Fig. 1-3 and 1-4. We did find SWRs during goal periods (Fig. 6) and we quantified SWR properties in an additional analysis in Fig. 6-1.

      Do beta bursts also occur during navigation prior to goal entry? It would be beneficial to display these rhythmic activities continuously across both the navigation and goal entry phases.

      We did not find consistent beta bursts in PFC or CA1 on approach to goal entry. We generated an additional goal entry-aligned spectrogram, and quantification (Fig. 1-4) to show that beta oscillations in both regions increased after goal entry. This was also supported by the Welch’s periodogram method (Fig. 1C-H, Fig. 1-3). Beta oscillations in the hippocampus during locomotion or exploration have been reported (Ahmed & Mehta, 2012; Berke et al., 2008; França et al., 2014; França et al., 2021; Iwasaki et al., 2021; Lansink et al., 2016; Rangel et al., 2015).

      Additionally, given that the hippocampal theta rhythm is typically around 7-8 Hz, while a peak at approximately 15-16 Hz is visible in the power spectra in Figure 1C, the authors should clarify whether the 22 Hz beta activity represents a genuine oscillation rather than a harmonic of the theta rhythm.

      We performed further spectral analysis comparing times when the animal is moving on the maze with times when the animal is immobile at the goal (Fig. 1-3 and 1-4). The results point to the beta frequency oscillations in both regions are unlikely to be harmonics of theta. The beta frequency bands in the spectrogram are independent of the theta band. We were initially concerned about the possibility that the 22 Hz power in CA1 may be a harmonic rather than a standalone oscillation band. If these are harmonics of theta, we should expect to find coincident theta at the time of bursts in the beta frequency. In Fig. 1B, Fig. 1-5, and Fig. 2A, we show examples of the raw LFP traces from CA1. Here, the detected bursts are not accompanied by visible theta-frequency activity. For PFC, we do not always see persistent theta-frequency oscillations like CA1. In PFC, we found beta bursts were frequent and visually identifiable when examining the LFP. We provided examples of the PFC LFP (Fig. 1B, Fig. 1-5, and Fig. 2A). In these cases, we see clear beta frequency oscillations lasting several cycles and these are not accompanied by any visible oscillations in the theta frequency in the LFP trace.

      (2) The authors claim that beta activity is independent between CA1 and PFC, based on the low coherence between these regions. However, it is challenging to discern beta-specific coherence in CA1; instead, coherence appears elevated across a broader frequency band (Figure 2 and Figure 2-1D). An alternative explanation could be that the uncoupled beta between CA1 and PFC results from low local beta coherence within CA1 itself.

      This is a legitimate concern, and we used three methods to characterize coherence and coordination between the two regions. First, we calculated coherence for tetrode pairs for times when the animal was at goals (Fig. 2B), which provides a general estimation of coherence across frequencies but lack any temporal resolution. Second, we calculated burst-aligned coherence (Fig. 2-2), which provides temporal resolution relative to the burst, but the multi-taper method is constrained by the time-frequency resolution trade-off. Third, we quantified the timing between the burst peaks (Fig. 2D), which described the timing differences but the peaks for the bursts may not be symmetric. Each method has its own caveats, but we drew our conclusion from the combination of results from these three analyses, which pointed to similar conclusions.

      Reviewer 2 is correct in pointing out the uniformly high coherence within CA1 across the frequency range we examined. When we inspected the raw LFP across multiple tetrodes in CA1, they were similar to each other (Fig. 2A). This likely reflects the uniformity in the LFP across recording sites in CA1, which is what we saw with coherence values across the frequency range (Fig. 2B). We found that CA1 coherence between tetrode pairs within CA1 was statistically higher than tetrode pairs in PFC across the frequency range (Fig. 2B and C), thus our results are unlikely to be explained by low beta coherence within CA1 itself. The burst-aligned coherence using a multi-taper method also supports this. The coherence values within CA1 at the time of CA1 bursts were ~0.8-0.9.

      (3) In Figure 2-1E-F, visual inspection of the box plots reveals minimal differences between PFC-Ind and PFC-Coin/CA1-Coin conditions, despite reported statistical significance. It may be necessary to verify whether the significance arises from a large sample size.

      We will include the sample sizes in Table 2-1. We repeated the analyses based on average values per day for each animal (Fig. 2-2 E-F). The pattern of significance remains consistent.

      (4) In Figure 3 and Figure 4, although differences in power and frequency appear to change significantly across days, these changes are not easily discernible by visual inspection. It is worth considering whether these variations are related to increased task familiarity over days, potentially accompanied by higher running speeds.

      We agree with Reviewer 2 that familiarity increases across days, and the animal is likely running faster. The analysis for Fig. 3 (previously Fig. 3 and 4) includes only data from periods when the animal was at the goal and was not moving. We added a supplemental figure (Fig. 3-1), which shows that the speed of the animal was below 0.5 cm/s at the time of the analyzed bursts. We used linear mixed-effects models to quantify the relationship between power, frequency and day or behavioral quintile, which accounts for repeated measurements across animals.

      (5) The stronger spiking modulation by local beta oscillations shown in Figure 6 could also be interpreted in the context of uncoupled beta between CA1 and PFC. In this analysis, only spikes occurring during beta bursts should be included, rather than all spikes within a trial. The authors should verify the dataset used and consider including a representative example illustrating beta modulation of single-unit spiking.

      We agree with Reviewer 2 that the stronger modulation to local beta is another piece of evidence indicating uncoupled beta between the two regions. We appreciate this suggestion and have revised Fig. 5 (previously Fig. 6) to include examples illustrating beta modulation for single units. These are spike-phase raster and histograms. We want to clarify that in the revised manuscript, the spikes were only from periods when the animal was at the goal location (5 s after entry) and did not include the running period between goals. Although beta power fluctuates in bursts, our data show there are ongoing beta oscillations throughout this period (Fig. 1-3 and Fig. 1-4), which prompted us to examine the entire period.

      (6) As observed in Figure 7D, CA1 beta bursts continue to occur even after 2.5 seconds following goal entry, when SWRs begin to emerge. Do these oscillations alternate over time, or do they coexist with some form of cross-frequency coupling?

      This is a very helpful suggestion and led to some interesting findings. We performed two additional analyses: 1) burst/SWR cross-correlation on a shorter timescale and 2) SWR-aligned spectrogram for the beta frequency range. PFC beta burst timing and power appear to be anti-correlated with SWRs detected in CA1 (Fig. 6G and K). In contrast CA1 beta bursts were more likely to occur with SWRs (Fig. 6H and L). These results suggest there is temporal coordination between ongoing beta oscillations in PFC and SWRs in the hippocampus during waking. Cortical beta oscillations are reduced during hippocampal SWRs, perhaps to support the transient switch in global cortical states that accompanies SWRs.

      To examine potential cross-frequency coupling between SWRs and beta oscillations, we computed the mean SWR band power (150-250).

      Reviewer #3 (Public review):

      Summary:

      This paper explored the role of beta rhythms in the context of spatial learning and mPFC-hippocampal dynamics. The authors characterized mPFC and hippocampal beta oscillations, examining how their coordination and their spectral profiles related to learning and prefrontal neuronal firing. Rats performed two tasks, a Y-maze and an F-maze, with the F-maze task being more cognitively demanding. Across learning, prefrontal beta oscillation power increased while beta frequency decreased. In contrast, hippocampal beta power and beta frequency decreased. This was particularly the case for the well-performed and well-learned Y-maze paradigm. The authors identified the timing of beta oscillations, revealing an interesting shift in beta burst timing relative to reward entry as learning progressed. They also discovered an interesting population of prefrontal neurons that were tuned to both prefrontal beta and hippocampal sharp-wave ripple events, revealing a spectrum of SWR-excited and SWR-inhibited neurons that were differentially phase locked to prefrontal beta rhythms.

      In sum, the authors set out to examine how beta rhythms and their coordination were related to learning and goal occupancy. The authors identified a set of learning and goal-related correlates at the level of LFP and spike-LFP interactions, but did not report on spike-behavioral correlates.

      Strengths:

      Pairing dual recordings of medial prefrontal cortex (mPFC) and CA1 with learning of spatial memory tasks is a strength of this paper. The authors also discovered an interesting population of prefrontal neurons modulated by both beta and CA1 sharpwave ripple (SWR) events, showing a relationship between SWR-excited and SWR-inhibited neurons and beta oscillation phase.

      Weaknesses:

      Moreover, there is little detail provided about sample sizes and how data sampling is being performed (e.g., rats, sessions, or trials), raising generalizability concerns.

      We appreciate Reviewer 3’s thoughtful suggestions for making our claims convincing. We have included information about sample sizes in the revised manuscript.

      The authors report on a task where rats were performing sub-optimally (F-maze), weakening claims.

      Our experiment was designed to create a scenario in which one task was learned (Y-maze) and another was not (F-maze). This contrast allows us to determine differences in neural correlates of learning versus familiarity. The design produced a learned and not learned task with similar levels of familiarity over 5 days, within the same animal.

      Likewise, it is questionable as to whether mPFC and hippocampus are dually required to perform a no-delay Y-maze task at day 5, where rats are performing near 100%.

      We agree with Reviewer 3 that the mPFC and hippocampus may not be required when the animal reaches stable performance on day 5 (Deceuninck & Kloosterman, 2024). The data we collected spans the full range of early learning (day 1) to proficiency (day 5). We wanted to understand the dynamics of beta across these learning stages, which have not been reported previously.

      Recent studies suggest mPFC and hippocampus are likely to be needed, in some capacity, for learning continuous spatial alternation tasks on a range of maze geometries. Lesions, inactivation or waking activity perturbation of hippocampus or hippocampus and mPFC on the W maze alternation task slowed learning (Jadhav et al., 2012; Kim & Frank, 2009; Maharjan et al., 2018). More recently, optogenetic silencing of mPFC after sharp wave ripples on the Y-maze alternation affected performance when the center arm was switched (den Bakker et al., 2023). The Y and F-mazes in our study both share the continuous alternation rule, where the animal needed to avoid visiting a previously visited location on the outbound choice relative to the center, and always return to the center location.

      Further, the performance characteristics on the outbound and inbound components of our Y task are similar to the W task. We have analyzed the “inbound” and “outbound” performance of the animals on the Y-maze alternation task, and they are similar to the W maze alternation task. The “inbound” or reference location component is learned quickly whereas the “outbound”, alternation component is learned slowly.

      There would be little reason to suspect strong oscillatory coupling when task performance is poor and/or independent of mPFC-HPC communication (Jones and Wilson, 2005) potentially weakening conclusions about independent beta rhythms.

      Although many studies have examined the oscillatory coupling properties at the theta frequency between mPFC-HPC (Hyman et al., 2005; Jones & Wilson, 2005; Siapas et al., 2005), our understanding of beta frequency coordination between the two regions is less established, especially at goal locations. Our work suggests beta frequency coordination at goal locations does not share properties with those of theta frequency coupling between mPFC and HPC, which occurs primarily during movement on the maze. Our first novel finding is that beta oscillations occur at goal locations in mPFC and HPC; our second is that the beta frequency dynamics in these regions are surprisingly distinct. We are not aware of prior work describing these properties at goal locations in spatial navigation tasks, especially their temporal coordination.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusions from this article would be made much stronger if the authors (1) record from rats performing a task known to be dependent on mPFC-HPC communication (e.g. a spatial working memory task) or (2) record from the F-maze in well-trained rats (75-80% performance is common), or (3) show that Y-maze task performance is dependent on mPFC-hippocampal communication (see Maharjan et al., 2018, which used a W-track). It is possible that learning the Y-maze depends on mPFC-hippocampal communication, but that with asymptotic performance, this changes. This would put beta oscillation coupling findings into a nuanced perspective.

      We appreciate the recommendation. The objective of this current manuscript is to present previously unknown properties of beta oscillations in hippocampal-prefrontal cortical networks. We agree that further investigation is required to fully dissect the functional contribution of beta dynamics in these networks. We are in the process of doing that.

      The rule on the Y-maze in our experiment is identical to the continuous alternation rule on the W maze in Maharjan et al., 2018 and Kim and Frank, 2009 (Author response image 1), which showed the PFC and hippocampus are required for normal learning, respectively. The Y and W mazes share the same topology; there is one junction connecting three arms. The W maze has two 90-degree-angle turns which are not choice points. Thus, both maze tasks have one choice point and involve learning an alternation rule. Further the learning properties share similarities. For the Y-maze, the inbound portion (return to center) (Author response image 1) was quicker to learn than the outbound portion (alternation) (Author response image 1). The same pattern is observed in the W maze learning task (Maharjan et al., 2018, Fig. 3A and D, Kim and Frank, Fig. 4C). We agree an experiment is needed to show that Y-maze task learning is dependent on the function of hippocampal-prefrontal cortical networks. The shared topology, rule definition, and behavior profile suggest that the Y and W maze tasks engage similar learning processes.

      Author response image 1.

      W and Y maze alternation tasks share the same rule. Schematics illustrate a comparison of W and Y maze alternation task rules. The W maze has been inverted for visual comparison with the Y maze. Performance grouped by in- or outbound trials. Inbound trials originate from the side goals (2 or 3). Outbound trials originate from the center goal (1). Performance on the inbound trials was higher than that on the outbound trials, which is comparable to previously published alternation tasks on W-shaped mazes.

      (2) Typically, when analyzing LFP profiles, experimenters include running velocity/speed. It should be ruled out whether spectral changes are confounded in any way by speed or time spent in the goal zones.

      We appreciate this suggestion for better conveying our definition of goal period. This was raised by other reviewers. We have now included the speed profiles for the goal-entry-aligned spectrograms (Fig. 1B, 1-4, 2A, and 6A), as well as the speed quantification at the times of bursts (Fig. 3-1, Fig. 4-1) and SWRs (Fig. 6-1 D). We strictly define the goal period after entry into the sensor on the reward delivery device at the end of the arms. This is to ensure the animal is immobile for the period to avoid confounds related to speed. We also ensured the time periods are comparable between the trials we analyzed.

      (3) The authors should describe in each analysis how data are sampled, and for extracellular electrophysiology experiments, cell counts from each rat reported. A table would be ideal. For example, it is unknown if entrainment analysis is performed on data collected from one rat or from all rats.

      We have now added the missing information (Table 1-2, 5-1 to 5-3). We performed analyses using data from all animals.

      (4) There was no profiling of mPFC neurons in terms of their behavioral correlates. The authors should strongly consider examining how individual neurons encode task variables (e.g., trial correctness, reward location...) and can do so using a generalized linear model. Adding an analysis of behavioral correlates could nicely tie into the beta-SWR analyses. For example, are SWR-beta rhythm-modulated neurons also behaviorally modulated?

      This is a very helpful suggestion. We have added analyses on the behavioral correlates of PFC and CA1 neurons. We quantified three metrics: 1) whether PFC and CA1 spiking activity can distinguish goal location based on firing rate or phase preference, 2) firing distance to goal and 3) spatial information (Fig. 5 and 7). We also performed the same analysis based on SWR and beta modulation status. We found PFC cells that were both SWR- and beta-modulated showed the strongest task firing relationship (Fig. 7).

      (5) Figure 2:

      Are these data analyzed from well-trained rats? What is your N (rats/sessions/trials/epochs)?

      These results in Fig. 2 are from all days. We added the breakdown in Table 2-1.

      (6) Figure 3:

      (a) The authors show that on the Y-maze, performance, beta oscillations power, and beta oscillation frequency change over days. However, for the F-maze, performance improves but appears to taper off at 60% and PFC beta frequencies do not change with learning. Do you have rats performing this task well above chance (e.g., 75-80%?), and if so, do beta oscillation frequencies in the mPFC gradually change?

      We did not observe rats performing above chance on the F-maze over 5 days. They do not appear to learn the alternation rule on this task. This is why we used the F-maze as the “non-learner” control. The F-maze task design for this study appears to be difficult to learn and would likely require a much longer training period for the performance to exceed chance. We agree an important follow-up question is whether the effects reported here are generally observed across learning in different tasks. This is a future direction we are actively pursuing.

      One existing data point may partially address the relevance of beta power change with learning. We analyzed the power as a function of performance quintile on the Y-maze (learned) or F-maze (not learned). The power changed with performance quintile on the Y-maze (learned) but not the F-maze (not learned) (Fig. 3D), pointing to the change in power being associated with learning status.

      (b) Does the proportion of beta bursts change with learning? What about the proportion of coherent events?

      We performed this analysis and the results are shown in Fig. 3-3. We found a decrease in the proportion of coincident bursts within CA1, which was specific to the well-performed Y-maze.

      (c) Is the reduction in beta oscillation frequency in the mPFC related to running behavior? This should be ruled out. Same for the hippocampus.

      The reduction is not due to running because we only included bursts when the animal was at the goal and immobile. We added a figure to show speed at the time of bursts (Fig. 3-1).

      (d) Why don't you also show coherence as a function of learning?

      This is a great suggestion. Coherence as a function of day or performance is now shown in Fig. 3-2. Overall, the trends were weak suggesting there was no strong change in coherence over days or as a function of learning. Although some of the linear mixed effects models were statistically significant, the marginal R<sup>2</sup> values (R<sup>2</sup><sub>m</sub>) were very low, indicating the effects of performance or day on coherence were small. Beta frequency coherence within each brain region either remained the same or slightly decreased across days (Fig. 3-2 D and F) or with performance (Fig. 3-2 J). For coherence across regions, we only found a weak but significant increase for the well-performed Y-maze task across days (Fig. 3-2 B).

      (e) The statistics in the caption are great, but maybe consider using a table and a supplemental figure

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (7) Figure 4

      (a) This figure is better suited as supplemental to Figure 3.

      We have combined Fig. 4 with Fig. 3.

      (b) Statistics might be better suited in a table and a supplemental figure.

      We have now included the most relevant model output in the figure. These are the marginal R<sup>2</sup> (R<sup>2</sup><sub>m</sub>) and conditional R<sup>2</sup> (R<sup>2</sup><sub>c</sub>). These describe the contribution of the fixed or fixed and random effects on the model, respectively.

      (8) Figure 5

      (a) The shift of beta rhythms being linked to learning is interesting, albeit confusing, given that error trials were accompanied by even more 'precise' beta rhythm timing. Is it possible that beta rhythms are accompanied by an error signal? Or maybe instead related to running behavior?

      We verified this observation was not due to speed since we only included periods when the animal was at the goal and immobile. We added a new supplemental figure (Fig. 4-1) to report this analysis.

      (b) Changes in beta rhythm timing in the F-maze could be entirely explained by the data in the Y-maze, given that F-maze performance was, in general, very low. To test if beta rhythm shifting is a consequence of learning, the authors should show performance on the F-maze without the Y-maze, and vice versa.

      As suggested, we also analyzed frequency variance data from each maze separately and found learning-related changes were more prominent on the Y-maze (learned) than on the F-maze (not learned) (Fig. 4-4 C). The F-maze data serves as a valuable control within the same animal for the learning-related effects we are reporting.

      We agree that the Y-maze data is consistent with performance-related changes to beta timing. The addition of the F-maze provides a within-animal comparison for a task in which the animal performed less well. We agree an additional control group would provide further evidence to support learning-related changes in beta timing. However, the F-maze will likely take much longer to learn, therefore the additional days of exposure on the task will be a different confound. Thus, we wanted to ensure that we compared a “learned” with a “not learned” task within the same animal, with a comparable exposure period.

      (9) Figure 6

      (a) How many cells are analyzed here? Are they all pyramidal neurons? The authors should add the number of cells analyzed from each rat. A table would suffice.

      We added tables (Table 5-2 to 5-4) to show this information. For our analysis, we included all units.

      (b) The authors should include statistics about recorded units. Peak-to-trough timing and interspike interval are commonly used. Even demonstration of action potentials from individual cells is valuable for proof of concept.

      We have added more information on the cell type classification (Table 5-2) and how we classified the cells (Methods) and also added example waveforms in Fig. 5 (previously Fig. 6).

      (c) Why did the authors stop comparing the F-maze and Y-maze for entrainment analysis?

      The data for each maze was comparable. Per the reviewer’s suggestion, we added the entrainment analysis for each maze separately (Fig. 5-1).

      (10) Figure 7

      (a) The authors should minimally include instantaneous velocity to show that during putative ripple events, the animal is in fact, quiescent.

      We added the speed at the time of ripples in Fig. 6-1 D.

      (b) Figure 7C: The authors should show how time spent in the reward zones varies over days. Is SWR power simply changing due to behavioral occupancy differences (e.g., a statistical sampling problem)? An analysis of the SWR rate over days would be valuable.

      We agree and have added additional SWR properties across days in Fig. 6-1. This includes SWR power (Fig. 6-1 A), rate (Fig. 6-1 B), duration (Fig. 6-1 C) and speed (Fig. 6-1 D). To ensure we are comparing the equivalent time spent at goals, the SWR properties were calculated from the 10 s after goal entry (Fig. 6-1) and we only included goal visits that lasted at least 10 s. This controls for the potential confounds from occupancy differences across days.

      (c) Why did the authors stop comparing the F-maze and Y-maze for SWR analysis?

      The SWR properties were similar in both tasks. We have now added separate analyses for each maze in Fig. 6-1 and 6-2.

      (11) Figure 8

      (a) The authors discovered that mPFC neurons more strongly entrained to SWRs compared to their own beta rhythms, potentially indicating coordination of mPFC and hippocampus (lines 275:276). What about mPFC spiking to CA1 beta?

      We quantified mPFC spiking entrainment to CA1 beta in Fig. 5 O and P. We found mPFC cells were more strongly entrained to the local beta within mPFC compared with CA1 beta.

      (b) Lines 277:278: How did the authors arise at 8.6% being the expected proportion of neurons entrained by both SWRs and beta rhythms?

      We have added a better explanation of how we arrived at the expected proportion. The calculation is based on the joint probability between the proportion of neurons modulated by SWRs (30%) and the proportion of neurons with phase locking to beta rhythms (22%). The expected proportion (6.6%) is calculated by multiplying the two values (30% ´ 22%) under the assumption that the two classes are independent. Deviations from the expected proportions are determined using a Fisher’s exact test.

      We note that in the revised manuscript we reanalyzed the data with more stringent criteria. We now include SWRs with durations greater than 50 ms, rather than 15 ms, and we quantify beta spike phase-locking using only the spikes within the first 5 s after goal entry, for goal visits lasting at least 5 s.

      With these more conservative criteria, the proportion of SWR- and beta-modulated cells (8%) is no longer significantly different from the expected proportion (6.6%, Fisher’s exact test p=0.09). Although this differs from our original significance test, the trend, and importantly, the distinct task correlates for this population (Fig. 7 C-D) still hold. We have updated the revised manuscript with these findings.

      Our original result was: “The subset of PFC cells that are modulated by both SWR and beta (11%) is greater than the expected proportion (8.6%) under the assumption that SWR and beta can modulate the population independently (Fisher exact test, p=0.021), although the size of the difference is small.”

      (c) The discovery that neurons modulated by both beta and SWRs are unique from those simply modulated by beta is really interesting. The authors discovered an interesting relationship among dually entrained mPFC neurons whereby SWR-excited mPFC neurons were entrained to the peak of beta, whereas SWR-inhibited mPFC neurons were entrained to the trough of beta. Then, in lines285:286, the authors write:

      (12) "This relationship was not observed for PFC cells that are modulated by beta but not modulated by SWRs (Fig. 8C, right, Fig. 8-1A)."

      Why would the authors expect this relationship to exist when the mPFC neurons were not SWR modulated? Were the authors referring to something else?

      We should have phrased this more clearly. We edited the text to better convey the expected SWR and beta modulation patterns for the control populations. We wanted to ensure the relationship between SWR modulation direction and beta phase preference was not observed for the cells that were not SWR-modulated.

      Furthermore, what about mPFC neurons that are SWR modulated but not beta modulated? For completeness, the authors should examine these.

      We agree this is an important comparison and have now included this population in the new Fig. 7. As expected, we do not find any relationship between SWR modulation direction and spike preference to beta for the SWR-modulated and not beta-modulated population.

      (13) Lines 294:296: The authors should elaborate on how they obtained an expected proportion of CA1 beta modulated neurons.

      We have now added a description for calculating the expected number of beta-modulated CA1 neurons. These are the neurons that have a Rayleigh test p-value less than 0.05.

      (14) What are these neurons doing to predict task information? Do they at all? Do they differ from beta-only neurons or non-phase-locked neurons?

      This is a great suggestion, and we have added analyses to show the task correlates for the cells. We found brain region-specific differences for task correlates that depended on how the cell was modulated by SWRs or local beta oscillations. This is shown in the updated Fig. 7 and the accompanying supplemental Figs. 7-1 to 7-2.

      For PFC cells, SWR and beta modulation status defined a subpopulation with a strong task structure correlation. There was a positive correlation between the direction of SWR modulation and the spiking distance relative to goals. SWR-excited cells were more active further away from goals, whereas the SWR-inhibited cells were active closer to goals. This correlation was not found for SWR-modulated PFC cells that were not beta-modulated. For PFC cells, the direction of SWR modulation is known to be correlated with movement speed, consistent with the hypothesis that movement-active cells become reactivated during SWRs, and immobility-active cells become suppressed (Jadhav et al., 2016; Yu et al., 2017). We found the same relationship, the direction of SWR modulation was positively correlated with mean spiking speed. Our results show beta modulation marks a subpopulation of PFC cells with stronger task-structure correlates.

      For CA1 cells, we found the expected relationship between SWR modulation and task-structure correlates. SWR-excited CA1 cells spiked further away from the goal and when the animal was moving. This is consistent with the reactivation of trajectory-related spatial firing patterns during movement on the maze. This pattern was observed irrespective of beta modulation status.

      Ahmed, O. J., & Mehta, M. R. (2012). Running speed alters the frequency of hippocampal gamma oscillations. J Neurosci, 32(21), 7373-7383. https://doi.org/10.1523/JNEUROSCI.5110-11.2012

      Berke, J. D., Hetrick, V., Breck, J., & Greene, R. W. (2008). Transient 23-30 Hz oscillations in mouse hippocampus during exploration of novel environments. Hippocampus, 18(5), 519-529. https://doi.org/10.1002/hipo.20435

      Deceuninck, L., & Kloosterman, F. (2024). Disruption of awake sharp-wave ripples does not affect memorization of locations in repeated-acquisition spatial memory tasks. Elife, 13. https://doi.org/10.7554/eLife.84004

      den Bakker, H., Van Dijck, M., Sun, J. J., & Kloosterman, F. (2023). Sharp-wave-ripple associated activity in the medial prefrontal cortex supports spatial rule switching. Cell Rep, 42(8), 112959. https://doi.org/10.1016/j.celrep.2023.112959

      França, A. S., do Nascimento, G. C., Lopes-dos-Santos, V., Muratori, L., Ribeiro, S., Lobão-Soares, B., & Tort, A. B. (2014). Beta2 oscillations (23-30 Hz) in the mouse hippocampus during novel object recognition. Eur J Neurosci, 40(11), 3693-3703. https://doi.org/10.1111/ejn.12739

      França, A. S. C., Borgesius, N. Z., Souza, B. C., & Cohen, M. X. (2021). Beta2 Oscillations in Hippocampal-Cortical Circuits During Novelty Detection. Front Syst Neurosci, 15, 617388. https://doi.org/10.3389/fnsys.2021.617388

      Hyman, J. M., Zilli, E. A., Paley, A. M., & Hasselmo, M. E. (2005). Medial prefrontal cortex cells show dynamic modulation with the hippocampal theta rhythm dependent on behavior. Hippocampus, 15(6), 739-749. https://doi.org/10.1002/hipo.20106

      Iwasaki, S., Sasaki, T., & Ikegaya, Y. (2021). Hippocampal beta oscillations predict mouse object-location associative memory performance. Hippocampus, 31(5), 503-511. https://doi.org/10.1002/hipo.23311

      Jadhav, S. P., Kemere, C., German, P. W., & Frank, L. M. (2012). Awake hippocampal sharp-wave ripples support spatial memory. Science (New York, N.Y.), 336(6087), 1454-1458. https://doi.org/10.1126/science.1217230

      Jadhav, S. P., Rothschild, G., Roumis, D. K., & Frank, L. M. (2016). Coordinated Excitation and Inhibition of Prefrontal Ensembles during Awake Hippocampal Sharp-Wave Ripple Events. Neuron, 90(1), 113-127. https://doi.org/10.1016/j.neuron.2016.02.010

      Jin, S. W., Ha, H. S., & Lee, I. (2024). Selective reactivation of value- and place-dependent information during sharp-wave ripples in the intermediate and dorsal hippocampus. Sci Adv, 10(32), eadn0416. https://doi.org/10.1126/sciadv.adn0416

      Jones, M. W., & Wilson, M. A. (2005). Theta Rhythms Coordinate Hippocampal–Prefrontal Interactions in a Spatial Memory Task. PLoS Biology, 3(12). https://doi.org/10.1371/journal.pbio.0030402

      Kim, S. M., & Frank, L. M. (2009). Hippocampal Lesions Impair Rapid Learning of a Continuous Spatial Alternation Task. PLoS ONE, 4(5). https://doi.org/10.1371/journal.pone.0005494

      Lansink, C. S., Meijer, G. T., Lankelma, J. V., Vinck, M. A., Jackson, J. C., & Pennartz, C. M. (2016). Reward Expectancy Strengthens CA1 Theta and Beta Band Synchronization and Hippocampal-Ventral Striatal Coupling. J Neurosci, 36(41), 10598-10610. https://doi.org/10.1523/JNEUROSCI.0682-16.2016

      Maharjan, D. M., Dai, Y. Y., Glantz, E. H., & Jadhav, S. P. (2018). Disruption of dorsal hippocampal-prefrontal interactions using chemogenetic inactivation impairs spatial learning. Neurobiol Learn Mem, 155, 351-360. https://doi.org/10.1016/j.nlm.2018.08.023

      Rangel, L. M., Chiba, A. A., & Quinn, L. K. (2015). Theta and beta oscillatory dynamics in the dentate gyrus reveal a shift in network processing state during cue encounters. Front Syst Neurosci, 9, 96. https://doi.org/10.3389/fnsys.2015.00096

      Siapas, A. G., Lubenov, E. V., & Wilson, M. A. (2005). Prefrontal Phase Locking to Hippocampal Theta Oscillations. Neuron, 46(1), 141-151. https://doi.org/10.1016/j.neuron.2005.02.028

      Yu, J. Y., Kay, K., Liu, D. F., Grossrubatscher, I., Loback, A., Sosa, M.,…Frank, L. M. (2017). Distinct hippocampal-cortical memory representations for experiences associated with movement versus immobility. Elife, 6, e27621. https://doi.org/10.7554/eLife.27621

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      As a general phenomenon, adaptation of populations to their respective local conditions is well-documented, though not universally. In particular, local adaptation has been amply demonstrated in Arabidopsis thaliana, the focal species of this research, which is naturally highly selfing. Here, the authors report assays designed to evaluate the spatial scale of fitness variation among source populations and sites, as well as temporal variability in fitness expression. Further, they endeavor to identify traits and genomic regions that contribute to the demonstrated variation in fitness.

      Strengths:

      With many (200) inbred accessions drawn from throughout Sweden, the study offers an unusually fine sampling of genetic variation within this much-studied species, and through assays in multiple sites and years, it amply demonstrates the context-dependence of fitness expression. It supports the general phenomenon of local adaptation, with multiple nuances. Other examples exist, but it is of value to have further cases illustrating not only the context-dependence of fitness expression but also the sometimes idiosyncratic nature of fitness variation. I commend the authors on their cautionary language in relation to inferences about the roles of particular genomic regions (e.g.l.140-144; l.227)

      Weaknesses:

      To my mind, the manuscript is written primarily for the Arabidopsis community. This community is certainly large, but there are many evolutionary biologists who could appreciate this work but are not invited to do so. The authors could address the broader evolution community by acknowledging more of the relevant work of others (I've noted a few references in my comments to the authors). At least as important, the authors could make clearer the fact that A. thaliana is (almost) strictly selfing and how this feature of its biology both enables such a study and also limits inferences from it. Further, it seems to me that though I could be wrong, readers would appreciate a more direct, less discursive style of writing, and one that makes the broader import of the focal questions clearer.

      We agree that connecting the paper better to the broader field is desirable, and have tried to do this in the revision, including adding a brief description of A. thaliana that mentions the fact that it is highly selfing. That the availability of inbred lines influences the study design is fairly obvious. As for the inferences, we would argue that our main point—“know your organism”—is universal. For example, our selection experiments and common garden experiments produced diametrically opposite results in terms of fitness, and we attribute this to the importance of seedling establishment. It is hardly surprising that seedling establishment matters for an annual plant that favors disturbed habitats, but this is probably true for both outcrossers and selfers (although it is probable that selfing itself is an adaptation to such habitats), and could well be true for some long-lived trees as well. The Devil is in the details.

      As a reader, I would value seeing estimates of the overall fitness of the accessions in the different conditions, i.e., by combining the survival and fecundity results of the common garden experiments.

      Combining estimates would be possible in the common garden experiments, and would bring us somewhat closer to total fitness estimates, although as noted by another reviewer (and also emphasized by us), the time scale of our experiment is not sufficient to evaluate the trade-off between survival and fecundity. Furthermore, we would still be missing the establishment component of fitness, which we found to be extremely important. Therefore little would be gained by combining the estimates, while at the same time losing resolution to disentangle the fitness components. We thus decided to focus on the individual fitness components and leave a qualitative consideration of their joint effect for the Discussion.

      Reviewer #2 (Public review):

      Summary:

      The goal of this study was to find evidence for local adaptation in survival and fecundity of the model plant Arabidopsis thaliana. The authors grew a large set of Swedish Arabidopsis accessions at four common garden sites in northern and southern Sweden. Accessions were grown from seed in trays, which were laid on the ground at each site in late summer, screened for survival in fall and the following spring, and fecundity was determined from rosette size and seed production in spring. Experiments were complemented by 'selection experiments', in which seeds of the same accessions were sown in plots, and after two years of growth, plants were sampled to determine fitness from genotype frequencies, providing a more comprehensive evaluation of lifetime fitness than can be gleaned from fecundity alone.

      To clarify, fecundity was determined from total plant area using photos of the mature stems, not the rosettes or direct counting of seeds. That said, it is true that our fecundity estimate was well correlated with rosette area. Furthermore, we validate our fecundity estimates by showing they were highly correlated with seed production estimated by measuring and counting siliques on a separate set of plants grown under common garden conditions in one of our sites (Brachi et al. 2022).

      As the main result, southern accessions had higher mortality in northern sites in one of two years, but also suffered more slug damage in southern sites in one year, indicating a potential link between frost tolerance and herbivore resistance. Fecundity of accession was highest when growing close to the 'home' environment, but while accessions from one sand dune population in southern Sweden had among the lowest fecundities overall, they consistently had the highest fitness in the selection experiment. Accessions from this population had large seed size and rapid root growth, which might be related to establishment success when arriving in a new, partially occupied habitat. However, neither trait could fully explain the very high fitness of this population, suggesting the presence of other, unmeasured traits.

      Another clarification: the small set of “beach” accessions that performed well in the selection experiments came from several beaches on the Baltic coast of Skåne, and we also had one such accession from the island of Gotland, 400 km away as the bird flies. Furthermore, we sampled measured seed size in several hundred additional accessions and demonstrated that large seeds are indeed characteristic of this habitat.

      Overall, the authors could provide clear evidence of local adaptation in different traits for some of their experiments, but they also highlight high temporal and spatial variability that makes prediction of microevolutionary change so challenging.

      Strengths:

      A major strength of this study is the highly comprehensive evaluation of different fitness-related traits of Arabidopsis under natural conditions. The evaluation of survival and fecundity in common garden experiments across four sites and two years provides an estimate of variability and consistency of results. The addition of the 'selection experiment' provides an extended view on plant fitness that is both original and interesting, in particular highlighting potential limitations of 'fitness-proxies' such as seed production that don't take into account seedling establishment and competitive exclusion.

      Throughout the study, the authors have gone to impressive depths in exploring their data, and particularly the discovery of 'native volunteers' in selection experiment plots and their statistical treatment is very elegant and has resulted in compelling conclusions. Also, while the authors are careful in the interpretation of their GWAS results, they nonetheless highlight a few interesting gene candidates that may be underlying the observed plant adaptations, and which likely will stimulate further research.

      Overall, the authors provide a rich new resource that is relevant and interesting both in the context of general evolutionary theory as well as more specifically for molecular biology.

      Weaknesses:

      While the repetition of the common garden experiments over two years is certainly better than no repetition (hence its mention also under 'strengths'), the very high variability found between the two years highlights the need for more extensive temporal replication. In this context, two temporal replicates are the bare minimum, and more repeats in time would be necessary to draw any kind of conclusion about the role of 'high mortality' and 'low mortality' years for the microevolution of Arabidopsis. It also seems that the authors missed an opportunity to explore potentially causal variation among years, as they did not attempt to relate winter mortality to actual climatic variables, even though they discuss winter harshness as a potential predictor.

      We agree that two years is insufficient to understand how variation in selective pressures compound over time to generate micro-evolutionary change. The eight-year data in Oakley et al. (2023), which we discuss in the paper, support this. Our results are nonetheless sufficient to demonstrate the idiosyncratic nature of selection. In the revision, we further emphasize that far longer time series would be needed for definitive conclusions on long-term micro-evolutionary change.

      Our short time series is one reason why we do not try to correlate with climate data, as this would amount to doing statistics with four data points (mostly two groups of accession N vs S, with mostly homogenous climates within groups, and two years). Another reason is that we simply have no idea how the kind of meteorological data that are publicly available would affect plant performance on a local scale. Hence we make no claims..

      The low temporal variation also makes the accidental slug herbivory appear somewhat random. Potted plants are notoriously susceptible to slug herbivory, and while it is certainly nice that slug damage predominantly affected one group of accessions, it nonetheless raises the question whether this reflects a 'real' selection pressure that plants commonly face in their respective local environments.

      Characterizing the plants as “potted” is not accurate. A more serious objection is having what is effectively an A. thaliana monoculture. But indeed we have no idea whether slugs exert a significant selective pressure on A. thaliana in Sweden, and we make no claims to that effect. The evidence for selection on glucosinolates by generalist herbivores such as slugs is fairly strong, but the precise agent is not known, and probably varies over time and space. Our results merely demonstrate one possibility.

      The addition of the 'selection experiment' is certainly original and provides valuable additional insights, but again, it seems a bit questionable which natural process really has affected this outcome. While the genetic and statistical analysis of this experiment seems to be state-of-the-art, the experimental design is rather rudimentary compared to more standard selection experiments. Specifically, the authors added seeds from greenhouse-grown mothers to experimental plots and only sampled plants two years later. This means that, potentially, the first very big bottleneck was germination under natural conditions, which may have already excluded many of the accessions before they had a chance to grow. While this certainly is one type of selection, it is not exactly the type of selection that a 2-year selection experiment is set up to measure. Either initially establishing the selection experiment from plants instead of seeds, or genotyping the population over several generations, would have substantially strengthened the conclusions that could be drawn from this experiment.

      We 100% agree that more data would have been beneficial, and hence we do not make any claims about the nature of selection. The selection experiment was an experiment per se, and we were lucky that very large fitness differences turned out to exist. As for initial selection on dormancy “set” incorrectly in greenhouse-grown seeds, we agree that this is a possibility, but we do not think it is likely to explain the data for several reasons. First, why would these differences favor a small set of beach accessions over every other accession in four very different field sites? Second, existing dormancy estimates do not predict fitness in our selection experiments. Third, the same seed batches germinated uniformly in the common-garden experiments with minimal stratification. Fourth, a much simpler explanation—seed size—exists. We clarify this in the revision while retaining our original message that further experiments are needed (and are underway).

      Also, the complete lack of information on population density is a bit problematic. It is not clear if there were other (non-Arabidopsis) plants present in the plots, how many Arabidopsis plants were established, if numbers changed over the year, etc. Given all of these limitations, calling this a 'selection experiment' is in fact somewhat misleading.

      Seeds were introduced into sites that appeared appropriate for A. thaliana, leaving the background community intact. We provided information on sowing density; the density of plants (A. thaliana and other species) that we obtained during the course of the experiments varied considerably between sites, much like in natural populations, although we lack systematic measurements. We provide more information (including photos) in the revision.

      Despite these weaknesses, the authors could achieve their main goals, and despite the somewhat minimal temporal replication, they were lucky to sample two fairly distinct years that provided them with interesting variation, which they could partially explain using the variation among their accessions. Overall, this study will likely make an important contribution to the field of evolutionary biology, and it is another very strong example of how the extensive molecular tools in Arabidopsis can be leveraged to address fundamental questions in evolution and ecology, to an extent that is not (yet) possible in other plant systems.

      Reviewer #3 (Public review):

      Summary:

      The manuscript presents a large common garden experiment across Sweden using solely local germplasm. Additionally, there is a collection of selection experiments that begin investigating the factors shaping fecundity in these populations. This provides an impressive amount of data and analysis investigating the underlying factors involved. Together, this helps support the data showing that fluctuations and interactions are key components determining Arabidopsis fitness and are more broadly applicable across plant and non-plant species.

      Strengths:

      The field trials are well conducted with extensive effort and sampling. Similarly while the genetic analysis is complex it is well conducted and reflects the complexity of dealing with population structure that may be intricately linked to adaptive structure. This has no real solution and the option of presenting results with and without correction is likely the only appropriate option.

      Weaknesses:

      A significant finding from this study was that fecundity is shaped more by yearly fluctuations and their interaction with genotype than it is by the main effect of location or genotype. Another significant finding is that the strength of selection can be quite strong, with nearly 5x ranges across accessions. It should be noted that there are a number of other studies using Arabidopsis in the wild with multiple years and locations that found similar observations beyond the Oakley citation. In general, the context of how these findings relate to existing knowledge in Arabidopsis is a bit underdeveloped.

      We have tried to remedy this in the revision (see also comments by Reviewer #1).

      The effects of the populations across the locations seem to rely on individual tests and PC analysis. It would seem to be possible to incorporate these tests more directly in the linear modeling analysis, and it isn't quite clear why this wasn't conducted.

      We respond to this question below, in Recommendations for the Authors (first item from you).

      I'm a bit puzzled by the discussion on how to find causative loci. This seems to focus solely on GWAS as the solution, with a goal to sequence vast individuals. But the loci that the manuscript discussed were found by a combination of structured mapping populations followed by molecular validation that then informed the GWAS. As such, I'm unsure if the proposed future approach of more sequencing is the best when a more balanced approach integrating diverse methods and population types will be more useful.

      We are puzzled by this comment in return. Our statement about more sequencing (penultimate sentence of discussion) was referring to achieving a better understanding of the history of migration and selection rather than identifying causative loci.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers provide consistent and complementary recommendations that you should consider in a revision. In particular, you should try to better embed your results into the broader literature, both from Arabidopsis and other plant species. Also, try to make your text more accessible to readers outside the special topic.

      Agreed!

      Reviewer #1 (Recommendations for the authors):

      (1) l.545: Here, the text states, importantly, that the accessions were randomized in the greenhouse. However, I have found no mention of this in the referenced Brachi et al. paper. I'm concerned that this be reported accurately and also that, if the accessions were not, in fact, randomized in the greenhouse, the potential for environmentally induced maternal influences to be confounded with genetically based differences be acknowledged, esp. in the case of the seed size difference found for B accessions.

      Accessions were continuously randomized in the greenhouse although we stopped moving them when they flowered to reduce the risk of contamination. All the seeds for all the accessions were produced in the same greenhouse, in the same conditions, in one single planting. We state this clearly in the revised paper.

      As for the seed size variation, note that the measurements we use in the paper were not taken on the seed we used in the experiments, so there is no reason they should be correlated with fitness if fitness was influenced by maternal effects. Moreover, we have plenty of data demonstrating that seed size variation is mostly genetic, and we added a supplementary figure showing that the measurements we used are strongly correlated with those of other experiments, including one done in the field. Finally, we already included a figure demonstrating that beach populations generally have much larger seeds.

      (2) l.593: BLUPs are estimated with statistical uncertainty. I realize that it would be far from straightforward to take into account their sampling variances in the GWAS, but it should be acknowledged that ignoring the uncertainty of the BLUPs has an unknown impact on the findings from the GWAS.

      All phenotypic measures come with error, and we had more replication than most GWAS studies (including essentially all human GWAS). The effect of such error is to reduce heritability and decrease the power of GWAS.

      (3) l.48: As an earlier important reference demonstrating this point: Antonovics, Clay, Schmitt, 1987 Oecologia.

      Indeed: added. Thanks!

      (4) l.49: studied rapid evolution in a natural population[s] - delete 's'

      Done.

      (5) l.502: lead -> led

      Corrected.

      (6) l.53: The present study sought to gain insight into local adaptation in A. thaliana.; Vague, as is the rest of the intro.

      Agreed. As noted above, the intro has been extensively reworked.

      (7) l.533: What is alpha?

      Alpha is the regularization parameter for the snmf algorithm. We added this clarification.

      (8) l.553: Were the accessions also randomized in the field planting?

      Yes, and we now state this.

      (9) l.597: As I've noted in my public comments, the authors have done well to couch their genomic inferences with caveats. In line with such caution, I urge the authors to consider alternate phrasing to 'genetically determined' -> genetically influenced. Also l. 166: 'control of' -> influence on (among other instances).

      No, “genetically determined” was correct on line 597, as this describes the model assumptions. But note that the liability-threshold model is by no means genetically deterministic: it is widely used in epidemiology to model genetic predisposition to diseases caused by the environment. For example, whether you get lung cancer or not primarily depends on luck and on your exposure to pollution (in particular smoking), but there are genes that influence how susceptible you are (to pollution; there are no genes that influence luck). There were some words missing in the sentence; hopefully things are clearer now.

      Re l. 166, "control of” was changed to “effect on”, which is more accurate. Note, however, that there is nothing genetically deterministic here: the context is a variance-partitioning, and the results show very clearly that genetic factors play a minor role relative to environmental ones—and that most of the variance remains unexplained. (Luck?)

      (10) l.595-602: Not clear.

      As noted, some words were missing. Hopefully it is clearer now.

      (11) l.744: Could have been confounded by cryptic native -- missing or extra word? Also next line.

      Corrected.

      (12) l.755-7: More direct phrasing would help readers understand this point.

      Indeed. Sentence straightened out.

      (11) l.129: 'The peak appears to involve a haplotype over 30 kb in length, which is consistent with a history of strong selection on this locus'. I don't understand the logic here. In what way should the length of the haplotype relate to the strength of selection? I would think this relates quite directly to the high degree of inbreeding.

      Strong selection causes rapid allele-frequency change, leaving less time for recombination to break up associated haplotypes, causing increased linkage disequilibrium. This is standard population genetics (e.g., Maynard Smith and Haigh 1974). Inbreeding also causes increased haplotype sharing, but genome-wide. Telling them apart is difficult, hence we used “consistent”. However, Reviewer 3 points to the existence of a segregating inversion in this region, which is an even more likely explanation, and we focus on this in the revision.

      (12) l.132: 'This suggests that the variation for slug damage seen in Figure 4 may be partly mediated by glucosinolate production.' I find the logic unclear here, as well.

      “This” referred to several observations, which is poor English. We rearranged the sentences to make our meaning clearer.

      (13) l.173: 'the accession-effect on fecundity' -> variation among accessions wrt fecundity.

      We rewrote this clunky sentence.

      (14) l.175: PCA unclear. Is this the same PCA described at l. 638? That is the only mention of PCA that I find in the methods, but I don't see that it connects here.

      No this doesn’t refer to the same analysis. In line 175 we talk about a vanilla, textbook PCA: confronted with 200 fecundity estimates in 8 experiments, we used PCR to look for patterns across the 8 dimensions, and found that 3 dimensions captured most of the variation (as shown in the heatmap). We have added a few sentences about this in the methods. Line 638 describes the way images of plants were treated to estimate variation in rosette color.

      (15) l.188: 'reveals what is causing them': Consider rephrasing to avoid language of causality, consistent with care taken elsewhere.

      Well, the context here is very different: we are effectively doing a post hoc analysis, and there is nothing wrong with saying that “the significant value in the chi-square test is caused by an excess of…”, for example. The patterns we see in the PCA are caused by the kinds of patterns we go on to discuss—it captures these patterns, among other things. We changed “reveals” to “helps us understand” to be less biblical.

      (16) l.219: downstream of -> conditional on.

      Yes.

      (17) l.235: Although native plants WERE growing nearby.

      Yes.

      (18) l.264: nearby: make this more explicit, i.e. within x km.

      We changed the sentence to “Although native plants were growing within less than a hundred meters in most cases,...”

      (19) l.285: none of our MAJOR conclusions depend on this.

      We disagree. No conclusion in the paper could be confounded by potential natives.

      (20) l.315: albeit it not as large as B accessions: delete 'it'.

      Corrected.

      (21) l.355: genetic basis of fitness -> genetic contribution to variation in fitness.

      No, this is pretty much exactly Lewontin’s usage—the sentence (and section) is about changes in allele frequencies. The alternative suggested is less precise as we have no fitness estimates, and cannot say anything about variance.

      (22) l.380-2: It seems to me that the authors could articulate a more compelling case.

      Well, what we wrote is the truth. Of course all of this work also fits into a broader context, but those were the specific goals—and we achieved them. There are many papers that ask big questions, but actually answer much more limited ones.

      (23) l.401: latitude vs. latitude???

      Oops. Changed to “north vs. south”, which is hopefully less obscure!

      (24) l.402: 'textbook local adaption,' -- the authors should acknowledge that rigid/simplistic thinking about LA has been recognized as a caricature for quite some time.

      We do. Using “textbook” is meant to convey this: textbooks tend to be full of cartoonish simplifications, and most chapters on local adaptation will have a figure (cartoon or based on real data) showing reaction norms as two crossing lines.

      (25) l.406: could maintain variation AMONG POPULATIONS [right?]

      Could be either, but “among populations” is better in this context.

      (26) l.419: Is there a possibility of a source envt maternal effect? See my comment on l.545.

      As stated in response to the previous comment: In principle yes, but if so, then this environmental maternal effect happens to be strongly correlated with seed size, a trait we know to be genetically controlled and which is also extremely likely to influence seedling establishment. Experiments to confirm these results are underway—meanwhile we would be happy to accept bets against!

      (27) l.434: 'very stable environments dating back to the last glaciation.' REF?

      Ha! That would be Wikipedia references as the precise location of the “Littorina sea” and existence of obvious ancient beachlines now inland have entertained Swedish and Danish school children for generations. But the details are not important: the point is that beaches are vast (by A. thaliana standards) disturbed habitats maintained by the sea, and while they change, they do so on a geological time scale. Of course not all beaches harbor A. thaliana—the beaches of Hanö Bay are geologically unusual for Sweden—but this is beyond the scope of this publication. The sentence has been changed to something less specific making the relevant points.

      (28) l.436: 'It is likely that the existence of S1 and S2 accessions is far more uncertain': unclear wording.

      Yup. We rewrote the paragraph.

      (29) l.446-7: 2nd person is jarring.

      Reworded.

      (30) ll.452-474: This is a welcome acknowledgement of the limits of molecular approaches to elucidating selection. It would be good scholarship to acknowledge earlier authors making such a case. I could suggest Rockman 2012. Evolution; Travisano and Shaw 2013. Evolution; Hoban et al. 2016. Am.Nat., and there are others.

      True; added; thanks!

      (31) l.490: 'might have to run a gauntlet of linkages to genes directly involved in local adaptation': This teleological wording is especially jarring.

      (32) l.494: 'resistance allele to sweep': In concluding this manuscript, it does not seem appropriate to use as an example a single locus case. More broadly, I question the value/effectiveness of this paragraph as a conclusion to this manuscript.

      We agree. The paragraph, gauntlets and all, has been replaced by discussion of why dissecting adaptive traits is hard.

      Reviewer #2 (Recommendations for the authors):

      (1) At the end of the introduction, the authors briefly outline the geographic scale of their study and mention both common garden and experimental evolution plots, but at this point in the manuscript, it is not clear how these differ, and as the methods only come at the very end, it would be better if some more details are provided here. Specifically, it could be mentioned explicitly here that common garden experiments consisted of placing greenhouse-grown plants in pots on the ground (but not burying/planting them), whereas selection experiments were established by sprinkling seeds in 1m 2 plots. This doesn't really become clear anywhere except in the methods, but this is fairly important for the interpretation of results.

      We agree, and have expanded the description of the experiments in the legend to Figure 1, which also links to supplementary photos of the sites. We also note that the differences between the common-garden and the selection experiments should not be exaggerated. In particular, we did not use “greenhouse-grown plants”, but seedlings that had been allowed to establish outdoors in sheltered conditions, and we did not put “pots on the ground” but buried trays with holes in the bottom so that the plants could root in local soil. The main differences between the experiments are guaranteed establishment and lack of competition (from conspecific and other plants). This has also been clarified in the figure legend and in Methods.

      (2) The authors discuss a difference in mortality between the two years of their common garden experiment and suggest that harsher winters could have been the cause of mortality in the north. However, they do not provide meteorological data to support this. Was the 2011-2012 winter harsher than the 2012-2013 winter, and did NM experience harsher conditions?

      The problem is that we do not know what constitutes harshness from the point of view of the plant. Winters differ in many ways: temperature, duration, snow cover, etc. We have changed the relevant paragraph to make clear that our observations are consistent with those of Oakley et al., and that some aspect of winter weather is a plausible explanation.

      (3) The authors find an indication for the involvement of the AOP cluster in overwinter survival, which, among other things controls the accumulation of hydroxyl/alkenyl/methylsulfinyl glucosinolates. Is the chemotype of the accessions in this study known?

      Indeed they are! We had downloaded the data from Katz et al (2021), but the analysis did not make it into the first version of this paper. Thanks for the encouragement. We added a figure showing that their chemotypes are strongly associated with slug damage, supporting a causal relationship.

      (4) L68: 'we added one field site in eastern Skane' - which site does this refer to? It seems odd to mention this before the common garden/selection experiments are mentioned. It should be clear that this is specifically referring to the latter.

      Yes, this was out of place. The site is discussed later, when it becomes relevant.

      (5) Figure 1: It would be useful if common garden and experimental evolution sites used different colors. In contrast, the use of colors for seasons in part C is unnecessary and distracting, as the red and blue colors for fall and winter are very similar to the colors for S1 and B.

      Fixed.

      (6) L90: The authors discuss mortality, but figures show survival. This seems an unnecessary complication for the reader, and the same unit should be used for discussion and presentation.

      Survival is one minus mortality. We trust the readers to be able to do this conversion.

      (7) L99: 'the converse was not true' - it is not immediately clear what this refers to. Referring to the non-linear relationship in panel 3a would make this clearer.

      It refers to the result that the accessions with high mortality in NM did not necessarily have high mortality in NA. We think this is clear from the sentence in question.

      (8) L108: 'S2 accession were also strongly affected' - this can be gleaned from the figures, but it is not immediately obvious as they are relatively complex. Could mean mortality/survival rates be provided to facilitate this?

      We could but, throughout the paper, we have attempted to improve readability by not interrupting the text with numbers or repeating details that are presented in the figures.

      (9) L117: 'however, an indirect association is likely' - what is meant by this?

      We meant that it seems more likely that some accessions are more sensitive to stress regardless of source. This has been clarified.

      (10) Figure 6: What are the two horizontal bars? What are dashed lines? Provide appropriate labels and a figure legend.

      The top is a zoom-in of the bottom and the dashed lines outline the zoomed-in region. Clarified in legend.

      (11) L174: Define 'BLUPS' here.

      Done.

      (12) L198: 'S1 and S2 accessions generally had higher fecundity in 2012-2013 than in 2011-2012' - this is not obvious from Figure 9. Especially for S1 (dark blue presumably), there appears to be no difference visible between years.

      Changed text to note that differences are sometimes small, but that the stated pattern is seen in 13 out of 16 comparisons (which has p = 0.01 using a sign test).

      (13) L401: 'latitude vs. latitude' - what is meant by this? Or is this a mistake?

      Northern vs. southern. This has been clarified. Also noted by Reviewer 1.

      (14) L436: 'the existence of S1 and S2 accessions is much more uncertain' - this is a bit of an odd expression. Could it be rephrased?

      This has been clarified and the paragraph rewritten. Also noted by Reviewer 1.

      (15) L673: Were plots tilled before sowing, or cleared of vegetation? If not, what other vegetation was present at sowing?

      No clearing of vegetation was done except at the SR site, which was a weedy agricultural field. This and other information has been added. No vegetation surveys were made, but we added some more photos to give an idea of what the sites were like.

      Reviewer #3 (Recommendations for the authors):

      (1) In the section on lines 156-190 the large dataset is analyzed by accession in the linear model presented in Figure 6. Then, in Figure 8, the proposed populations appear to be tested individually in each experimental unit, leading to the probability argument in line 193. Is there a reason not to simply have accession nested within population in the model shown in Figure 6 as a way to directly test the between vs within population level components influencing the model? If the accessions are different, then it would be viable to treat this as random rather than fixed. Similarly, the field sites could be parsed into subsets as well.

      Excellent question. We thought a lot about this. Fig. 6 presents a standard ANOVA that simply shows that the overall pattern is what we hoped for: very large effects of site, year, and accession, plus substantial interaction effects. This motivates the exploratory analysis presented in Figs. 7-9, where we focus on how the relative performance of accessions within experiments depended on the fixed effects (year and site) and whether this matched their (arbitrarily but independently defined) group designation—as would be expected under local adaptation.

      We could explicitly add “group” to the model, as suggested, but this would give it a reality we do not think it deserves. Site, year, and accession are very much real, whereas group is the outcome of a somewhat arbitrary clustering of genotypes (i.e., accessions), analogous to race in human genetics, but without the sociological factors that sometimes warrant including race in a model not only because of direct genetic effects. Of course we could try to partition genotype and phenotype into within- and between-population variation using the classical quantitative genetics framework (i.e. Fst/Qst), but then we should have an a priori definition of “population”, which we don’t have. In addition to this fundamental objection, limited experimentation suggests that fitting a far more complex, nested mixed-effects model to our data is difficult in practice.

      Thus we prefer our original approach. We have changed the writing to clarify our logic.

      (2) Similarly, I'm not quite sure that the PC analysis is helping as the section is somewhat difficult to read given that the same impressions from Figure 7 are more explicitly shown in Figures 8 and 9.

      Well, the PCA (Fig. 7) is what led us to the analyses in Figs 8-9, where we interpret the PCs in terms of group behavior. We have tried to clarify our logic in writing.

      (3) Line 87-88 - An honest question, what is considered substantial divergence? The Fst values range from 0.05 to 0.27, suggesting that the divergences range across a spectrum and not all are substantial.

      We removed the sentence. That the divergence is substantial enough to make GWAS difficult becomes clear later (although the extent to which this is due to selection rather than marker divergence is not clear).

      (4) Line 129-130 - The 30kb AOP haplotype identified is likely the inversion associated with the major phenotypic variation identified in Sweden within Katz et al 2021. As such, the local LD structure may represent blocked recombination as much as linked selection.

      Of course; we missed that. Thanks!

      (5) Line 131-132 - I'm not quite sure what the evidence is that the MAM locus is more important.

      The AOP and MAM loci are epistatic, and both have been found to have influence on fitness in Arabidopsis and Brassica ssp in the field, along with sequence signatures of selection for both loci in multiple Brassicaceae. I'm unsure if this statement is supported.

      Indeed. This was a lazy and misleading reference to the fact that the MAM peak in Katz et al explains more of the variance than the AOP peak. It has been consigned to the dustbin of history. Instead we have added a figure showing that the glucosinolate profiles presented in that paper are highly correlated with slug damage in our study. Based on these results, I believe AOP explains more of the variation, but we leave pursuing this for those directly working on these pathways. The correspondence between our studies is certainly a nice confirmation of both.

      (6) Line 131-132 - It should also be noted that the MAM locus is variable in Sweden, albeit having multiple independent haplotypes that convergently create the same phenotype. There is an indication from Gloss 2022 that these haplotypes may create small-effect phenotypic variation.

      Correct again. What we meant to say was that the major polymorphism that is responsible for the highly significant MAM peak does not appear to segregate in Sweden, hence it is not surprising that we do not find an association at this locus either. We now say this. Needless to say, there could still be multiple variants at this locus that we do not have the power to detect.

  2. Sep 2026
    1. Author response:

      General Statements

      We wish to bring to the attention of the editor and reviewers that the title of the manuscript has been modified to more accurately reflect the conclusions of our study, i.e. that PhDEF has a major binding and regulatory role in the petal epidermis.

      Point-by-point description of the revisions

      We thank the three reviewers for their critical reading of our manuscript. We also appreciate that the three reviewers highlighted the conceptual advances and the broad findings that our study brings to the community. They have identified key limitations of our analyses, and we have either provided explanations for our choice, or performed new analyses to circumvent biases. We believe that the manuscript is greatly improved and that our main conclusion, that PhDEF has a major binding and regulatory action in the petal epidermis, is strongly supported by our data.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      The authors previously generated two cell-layer-specific mutants of petunia for the petal identity gene PhDEF. In this study, they profiled differential gene expression in those mutants through single-cell RNA sequencing (scRNA-seq). They found that more genes are highly and specifically expressed in the epidermal cell layer than in mesophyll cells. In addition, they identified cell-layer-specific and -aspecific PhDEF target genes. Using the extensive single-cell transcriptome and layer-specific target identification, the authors concluded that different cell identities affect homeotic regulator PhDEF, thereby influencing transcriptional regulation.

      Major comments:

      This presented work provides comprehensive evidence, that pre-existing cell layer identity (epidermis and mesophyll) modulate transcriptional output of homeotic transcription factor, PhDEF.

      However, a disconnection between PhDEF bindings to genome and transcriptional output undermines the robustness of their conclusion although some binding loci were shown to be correlated with DEG. This may indicate the chromatin state, the existence of interacting partners, and the non-productive binding of PhDEF, suggesting that PhDEF binding alone is not sufficient to predict transcriptional outcomes and additional regulatory mechanisms that shape gene expression in addition to the layer-specific regulatory mechanisms. This disconnection may also be due to the developmental timing. Indeed, it appears authors used different flower stages for ChIP-seq and scRNA-sequencing. In fully differentiated organs, PhDEF binding itself may be no longer transcriptionally productive, and differential gene expression results primarily from the pre-established cell identity rather than directly from the homeotic regulation of PhDEF. Therefore, the main question the authors asked-how homeotic identity works with cell-layer identity and how the homeotic gene, PhDEF, acts in mature organs-was not clearly explained by this study.

      We thank the reviewer for raising this important issue. We agree that the difference in developmental timing between the scRNA-Seq and ChIP-Seq experiments might contribute to the disconnection that this reviewer pointed out.

      First, we want to explain that the reason for performing scRNA-Seq on fully differentiated petals was purely technical, as we were initially aiming to obtain protoplasts at stage 8 (stage used for the ChIP-Seq) but could never retrieve enough of them for proper encapsulation in the 10X Chromium chips. We have now clearly explained this in the manuscript (lines 104-107).

      A hypergeometric test shows that our ChIP-Seq and scRNA-Seq datasets overlap more than by chance (p = 0.000137); however, we agree that differences in developmental stages possibly bring a confounding effect to our conclusions. Therefore, we have decided to add to our manuscript the intersection between ChIP-Seq and bulk RNA-Seq data performed on WT, star and wico flowers at stage 8, that we published previously (Chopy et al., 2024). In that case, both datasets have been obtained with the same exact genetic material and at the same exact developmental stage.

      This intersection confirms that very similar binding profiles are observed for PhDEF target genes, whether they are differentially expressed in the epidermis (star only), in the mesophyll (wico only) or in both layers (star and wico) (Figure 4A). However, star-specific DEGs were more often bound by PhDEF by epidermis+shared binding sites, and wico-specific DEGs more often with mesophyll-specific binding sites, suggesting a weak but significant association between binding and regulatory profiles. Performing similar tests for individual binding categories for DEGs identified by scRNA-Seq also revealed that epidermis-specific DEGs displayed more epidermis+shared binding sites than expected. Since the association between epidermal DEGs and shared+epidermal binding sites is found both in the intersection with RNA-Seq and scRNA-Seq data, we have now stated that “layer-specific binding and transcriptional regulation are partially linked, at least in the epidermis” (line 354).

      In Figure 2, the use of the term "target" is potentially misleading. It sounds like direct target genes (direct binding and differential expression) for PhDEF, but it refers only to DEGs.

      Indeed, the term "target" was referring to both indirect and direct targets of PhDEF. To avoid any possible confusion, we have replaced it by differentially expressed genes (DEGs) throughout the manuscript, when appropriate.

      Lines 496-497: When the authors state, "~ demonstrates for the first time that the regulatory function of homeotic factor is influenced by cell layer identity," it sounds overstated, as prior studies have shown that pre-existing tissue or cell identity can shape transcriptional activity and developmental output.

      We agree that previous studies have shown that cell identity influences transcriptional activity in general. While this might not have been specifically assessed in the context of different cell layers, we have rewritten this sentence accordingly.

      Minor comments:

      In the UMAP presentation, as depicted in Figures 2C, S3, and S5, the cells with zero expression can be colored in light gray (or an inverted color scheme). The purple hue masks the gene expressions of other cells, making it difficult to see the yellow or light green colored cells.

      We have modified all UMAPs depicting gene expression as suggested, in Figures 2C and S5, and replaced UMAPs with DotPlots in Figure S3.

      Reviewer #1 (Significance):

      General assessment: This study is well-designed and technically sound. They utilize single-cell transcriptomics and ChIP-seq by using genetically well-defined genetic materials and layer-specific PhDEF deletion mutants. The analysis showed where PhDEF binds to genomic loci and which genes are differentially expressed in petal epidermis and mesophyll, providing evidence of cell-layer-specific function of homeotic gene in mature organs. Although certain mechanistic aspects were not elucidated, the data from the extensive genome-wide study contributed to drawing their conclusions.

      Advances: This research goes beyond classical models of floral organ identity by showing that homeotic gene function is not uniform in the same floral organ. It represents a conceptual advance in our understanding by integrating cell layer identity into the framework of homeotic gene regulation.

      Audience: This study will be of broad interest to scientists who study transcription networks, cell and organ identity in the context of plant development.

      My field of expertise: Transcriptional regulation by transcription factor, epigenetic regulation of gene expression, plant development

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      This study from Cavallini-Speisser et al. cleverly leverages a tissue layer-specific mutant, single cell and bulk RNA-sequencing, and ChIP-sequencing to decipher tissue layer-specific regulation of petal development in petunia by the PhDEF transcription factor. The authors find common and unique targets of PhDEF between the epidermis and mesophyll and conclude that the activity of transcription factors like PhDEF are influenced by the pre-existing environment in the cell they are expressed in. Understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology that can be elusive outside of highly tractable model systems. As such, I think this study is of high value and has strong potential to expand our understanding of how developmental specificity is mediated by commonly employed transcriptional regulators. However, I think there are some issues with possible over-interpretation and some places where documentation of experimental design and data quality control are lacking. I elaborate on these concerns below.

      Major Comments:

      Line 155: Assigning mesophyll clusters by default without any positive marker genes strikes me as problematic, especially as much of the analysis rests on comparing the transcriptomes of epidermis and mesophyll cells. Can the authors perhaps leverage published scRNA-seq datasets to find potential mesophyll markers, even homologs from other species, to improve confidence in the cluster assignment?

      We thank the reviewer for raising this important point. Our statement of defining mesophyll identity by default was not entirely true (and we have now removed it), since it is supported by GO-enriched terms for cluster markers. For cluster "mesophyll 3", there is a strong enrichment for photosynthesis-related genes; and for cluster "mesophyll 2", there is a strong enrichment for water transport-related genes, both functions being fulfilled by the petal mesophyll. For cluster "mesophyll 1", the most enriched GO term is "glutathione metabolic process" that rather points to stress response. These three clusters also do not express any of the epidermal-identity genes, in contrast to the clusters that we assigned as epidermal. Cluster markers from the mesophyll display the lowest enrichment of all clusters (the best cluster markers have a log2FC between 2.8 and 4.9, in contrast to a log2FC between 8.1 and 9.9 for all other clusters), which is in line with our finding that the mesophyll expresses less specific genes than the epidermis, and suggests a basal identity with a transcriptomic signature that is less clear than in the epidermis. Therefore, it is not entirely trivial to find positive marker genes for the mesophyll with a strong specificity.

      In order to be more transparent about the expression patterns of the genes we selected to assign cluster identity, we modified Figure S3 to include DotPlots of selected photosynthesis-related genes, histone genes, S-phase genes and vasculature genes in the same figure, to compare with DotPlots of epidermal genes and pigmentation genes from Figure 1D, to allow for an informed comparison. Our conclusions remain the same as previously: epidermal clusters are defined based on the specific expression of epidermal genes and/or pigmentation genes. We notice, however, that the cluster "upper limb epidermis" strongly expresses photosynthesis genes, which is likely why it is close in the UMAP space to the "mesophyll 3" cluster. We have no explanation for that, but the extremely high expression of pigmentation genes in this cluster, however, identifies it as epidermal without a doubt. We have also added in Figure S3 the DotPlots of expression levels of homologs of 15 epidermis-enriched and 9 mesophyll-enriched genes from tobacco petal scRNA-Seq published by Kang et al. (doi: 10.1111/nph.17992). This shows that our definition of epidermal and mesophyll clusters and the one from Kang et al. generally overlap.

      Line 228: Through the description and interpretation of the ChIP-seq dataset, the authors use the fact that peaks are more abundant and bigger in the epidermis to conclude that binding of PhDEF is "stronger" in the epidermis. This implies a difference in physical interaction between the TF and the DNA that I don't think can be concluded from the data presented. This could be confounded by biology; if expression of PhDEF is more heterogeneous in the mesophyll than in the epidermis, the peaks from that tissue will be averaged out and appear smaller when in fact the binding is the same strength. This could also be a technical artifact if the ChIP was less efficient in one sample versus another. This conclusion requires reinterpretation. The authors have the power to address this at least partially with the scRNA-seq by measuring PhDEF heterogeneity. I believe assessing ChIP efficiency would have required a spike in control, but perhaps there is a computational way to address this. It is important to discuss these confounding factors in the text.

      This is indeed another important point. We agree that the word "stronger", to describe PhDEF binding in the epidermis, was not appropriate and we have replaced it with "more frequent" which is a more factual interpretation of our results. We also agree that even this interpretation depends on potential ChIP artifacts that we have now evaluated.

      We have used our WT scRNA-Seq data to explore the heterogeneity of PhDEF expression in the mesophyll and the epidermis, as suggested. The barplot in Figure S10B represents the number of cells (y-axis) with given PhDEF RNA counts (x-axis) in the clusters that we assigned as epidermis (left) and mesophyll (right). We have performed this analysis after removing cells that do not express PhDEF at all, which represents 47.7% and 47.4% of epidermal and mesophyll cells, respectively, hence very similar proportions. The distributions of expression of PhDEF in the epidermis and in the mesophyll are within the same ranges, with a slightly higher expression of PhDEF in the mesophyll than in the epidermis. The coefficients of variation (cv) of the two distributions are similar and slightly higher in the epidermis (cv = 0.34 in the mesophyll and cv = 0.36 in the epidermis, p = 0.00106 with Feltz and Miller’s asymptotic test). Therefore, it appears that the expression of PhDEF is actually higher and slightly less variable in the mesophyll than in the epidermis, meaning that it should not result in averaging out the peaks detected. This relies on the assumption that PhDEF protein levels are directly correlated with PhDEF RNA levels, which we have not explored in this study and remains a limitation. We have included this analysis lines 275-278 and Figure S10B.

      Regarding ChIP efficiency, we had run different tests prior to sequencing: first, we have tested different amounts of chromatin, keeping the quantity of antibody unchanged, and tested ChIP enrichment by qPCR on a set of two positive (PhDEF and Pos2) and one negative (Neg1) control binding sites. Second, after selecting the best chromatin quantity, we have performed 4 independent ChIP replicates for each genotype (split between two assays named ChIP-1 and ChIP-2) and measured ChIP efficiency by qPCR. This is depicted in Figure S10A, with the replicates chosen for sequencing highlighted with an orange star.

      We have now explained in greater detail in the Methods our preliminary tests. There is indeed variation of enrichment between replicates, and particularly between assays here (ChIP-1 vs. ChIP-2), which is inherent to the ChIP experiment. It might be particularly prominent in our case due to our custom antibody directed against PhDEF, in contrast to commercial antibodies that are commonly used in ChIP experiments with tagged transgenic lines. However, we see consistently lower enrichment for star as compared to wico and WT, in line with the more frequent epidermal binding of PhDEF. We have followed the ENCODE guidelines for our analysis pipeline, in particular applying the IDR. We have now added other mapping statistics in Table S5 including the FRiP (fraction of reads in peaks) metric, that is in the range of expected values but is lower for star (around 0.5%) than for WT and wico (0.7-1.2 %), again consistent with the more frequent epidermal binding of PhDEF.

      Line 376: The authors risk overinterpreting a lack of differential gene expression detection in their analysis of PhDEF binding profiles. This can be affected by how deeply a library was sequenced or how many cells were analyzed per cell type. A gene might not be found to be DE if low depth or few cells resulted in noise or dropout. Lack of detection does not mean lack of differential regulation so the biological relevance of this portion of the analysis should be interpreted with caution.

      We have added to this new version of the manuscript an intersection between ChIP-Seq and bulk RNA-Seq in WT, star and wico, as bulk RNA-Seq is much more sensitive than scRNA-Seq in detecting lowly expressed genes. We have also added the sentence that bulk RNA-Seq "better captures lowly expressed genes" than scRNA-Seq, line 333. This intersection revealed an association between epidermal DEGs (star-specific DEGs) and the presence of epidermal+shared binding sites for PhDEF. We have modified our conclusions accordingly.

      Line 476: The authors state there is a mismatch in developmental timing between the RNAseq and ChIP datasets. Why is this? This is mentioned briefly in the Discussion, but has the potential to be majorly confounding to the joint interpretation of the ChIP and RNAseq datasets. This experimental design choice should be justified more thoroughly and a consideration of the limitations it brings to data interpretation should be more prominent in the text.

      This concern has also been raised by the first reviewer, and we have now added to our study an intersection between ChIP-Seq and bulk RNA-Seq performed at the same stage. We have also explained the technical reasons for performing scRNA-Seq at a mature stage only (lines 104-107). Indeed, the intersection between bulk RNA-Seq and ChIP-Seq performed at the same stage revealed a significant association between epidermal DEGs (star-specific DEGs) and the presence of epidermal+shared binding sites for PhDEF. Testing for individiual binding categories, we could also detect an enrichment of epidermal+shared binding sites for epidermal-specific DEGs identified by scRNA-Seq. Therefore, we have now stated that “layer-specific binding and transcriptional regulation are partially linked, at least in the epidermis” (line 354).

      Minor Comments:

      Line 229: The authors compare correlations between pseudo-bulked transcriptomes and argue that in the star mutant the epidermis adopts a mesophyll-like identity. The correlation between epidermis and mesophyll in star is 0.97 and the correlations were 0.94 and 0.93 in the other genotypes tested. What is the meaningful cutoff for saying the transcriptomes are similar or not? Is 0.97 so much higher than 0.94 that this conclusion is supported?

      The comparison of pseudo-bulk transcriptomes is a very global and exploratory approach. Given the high number of genes underlying these pseudo-bulk datasets, any difference in the correlation between them is statistically significant, which is not very informative. We agree with this reviewer that the interpretation of these correlation coefficients is somewhat arbitrary. We have simplified this part of the manuscript and have mostly focused on comparing star and wico pseudo-bulk transcriptomes to the WT ones, but not to each other's, which aligns well with our main message of a specific epidermal identity, easily shifting to a mesophyll-identity when PhDEF is missing or not entirely functional. We have also removed Figure 2E to a supplementary figure to give less emphasis to this analysis.

      Line 264: What are the "manually chosen thresholds for differential expression"? Can the authors explain and justify this? There is very little detail on this in the materials and methods and this raises some concerns regarding how a threshold was chosen.

      Seurat gives a default threshold of 0.25 for log2FC, which we found to be very permissive. In order to capture the most informative targets of PhDEF, but still to capture a meaningful number of targets, we empirically decided to increase this threshold to 0.75. On the WT scRNA-Seq dataset, we observed that the layer-specificity factor was also capturing meaningful differences in layer-specific expression (see Figure 1E). We chose a cut-off at 10% since it was the lowest to give a significant difference in the numbers of epidermis- vs. mesophyll-enriched genes in the WT petal. We have added these explanations in the methods.

      Figure 3G: Could this plot be annotated with the classification of peak layer specificity? It is a little difficult for me as the reader to keep up with all the categories in the text, and showing them in the figure might make that easier to follow.

      We have now annotated the Venn diagram in Figure 3G with the classifications of binding profiles.

      Figure 4A: A comparison of only two cell categories should not use scaled expression, as this can over-emphasize small differences in gene expression. Can this be replaced with a dot plot that uses unscaled expression values?

      We thank the reviewer for noticing this issue, we have built a DotPlot with unscaled values and replaced it in Figure 4A, which does not change our conclusions. We have also used unscaled values in Figure S14.

      Figure 4C: Could this be represented more legibly with stacked, space-filled bar charts? As is, this is difficult to read, and might be impossible for someone who is color blind. In addition, the authors claim this analysis shows similar proportions across all categories, but I wonder if that would hold true if they performed an over-representation analysis normalized to the categories shown in "all genes expressed". This could allow them to statistically test whether there are real differences in representation among the categories.

      We have now used a different and color-blind-friendly palette for pie charts of PhDEF binding profiles. Following reviewers' comments, we have analyzed the intersection of bulk RNA-Seq with ChIP-Seq (both performed at the same stage), and performed Chi2 goodness-of-fit tests that indeed support some association between binding profile and regulation profile, although this remains limited.

      Supplemental Figure 2: It's great that the authors include these metrics, but it would be ideal to also include the plots of standard QC metrics for scRNA-seq such as those found here to give a better sense of per cell quality: https://satijalab.org/seurat/articles/pbmc3k_tutorial

      In addition, it is important to include QC metrics for ChIPseq, which I did not find in the supplement. Metrics such as FRiP are important for interpreting ChIP library quality.

      We have now added the standard QC plots (Feature number per cell and RNA counts per cell) to Figure S2, as suggested. Mitochondrial and ribosomal genes are not annotated in the Petunia axillaris nuclear genome that we used, therefore we could not compute mitochondrial or ribosomal gene counts. We have removed cells expressing less than 200 genes, but did not apply any upper thresholds as there were no obvious outliers.

      FastQC reports for scRNA-Seq and ChIP-Seq have been deposited at https://entrepot.recherche.data.gouv.fr/dataverse/PhDEF_Flower_layer, as indicated in the Methods.

      We have also added ChIP metrics in Table S5, including the number of reads, duplicated reads, mapped reads and computed the FriP score. This score ranges between 0.5 % and 1.2 %, which is satisfactory and above the minimum recommended score  by ENCODE of 0.3 %.

      Line 654: What model was used for DESeq2?

      We have used default parameters for DESeq2, ie a negative binomial GLM fitting and Wald significance tests. We have added this information in the Methods.

      Line 752: Can the authors justify why peaks were called separately on input and ChIP samples rather than allowing MACS2 to call peaks in the ChIP sample over input background? That differs from the standard MACS2 pipeline and no explanation for this is provided in the text.

      Our analysis pipeline has indeed been customized, in particular to detect peaks in our positive control PhDEF, for which a binding site of PhDEF in the promoter has been demonstrated experimentally by others in many different species. This binding has a strong experimental support, and we expected PhDEF to bind to its own promoter in the two cell layers. Our ChIP-Seq results show a posteriori that this particular peak is far from being the strongest one over the genome; therefore, we believe it represents a good control to detect binding enrichment for average targets. We have first tried the standard MACS2 pipeline that calculates the enrichment of IP over Input, but we could only detect PhDEF binding for one WT and one wico replicate, although the peaks were visually clear in the two replicates. Our input sample being generally noisy, we explored how separate peak calling between IP and Input would behave (as already done in e.g. Durand et al., 2023, doi: 10.1093/plcell/koad025). We also explored how thresholds for FDR in MACS2, and IDR thresholds for reproducibility between IP samples, would influence peak detection in IP and Input.

      We found that calling peaks on the IP with a relaxed FDR (0.1), then applying the IDR at 0.1, allowed the capture of PhDEF binding to its own promoter in the two wico replicates (but still not in WT, because the peaks were lost after applying the IDR threshold). No peak was detected in the input with these settings, however for other genes we observed spurious peak detection in the input, therefore we decided to lower the FDR thresholds for input peaks to 0.05. Our choices have been made in an attempt to increase specificity at the risk of losing sensibility, and we probably lose true binding events. Considering that this ChIP has been performed on the endogenous PhDEF protein directly, and in chimeric flowers that only express PhDEF in half of the tissue, adapting the ChIP analysis pipeline was a necessary step.

      We have now added a few lines in the Methods (lines 696-706) to explain our rationale.

      Reviewer #2 (Significance):

      General Assessment:

      Strengths: The authors employ a unique and powerful mutant system to explore a fundamental developmental biology question. In addition, the datasets generated will likely be useful to other researchers working in petunia or flower development.

      Limitations: While the mutant system employed here is a creative way to get at tissue-layer specific transcription factor activity, the ChIP samples still include heterogeneous cell types, which may impact the findings presented here.

      Advance: This study uses a unique system to test the function of a transcription factor in distinct cell types. As stated above, understanding when and how a given transcription factor drives expression of unique target genes in various contexts is an important aspect of developmental biology.

      Audience: I believe this work will be of primary interest to the plant development, single cell, and chromatin biology communities. These are specialized, basic research communities.

      Reviewer Expertise: I am a plant developmental biologist who works with multiple modes of cell-type-specific NGS datasets including bulk and single cell RNAseq and ChIPseq among others.

      Reviewer #3 (Evidence, reproducibility and clarity):

      Summary:

      This study addresses a fundamental but underexplored aspect of homeotic gene function: how regulators of cell identity act during late stages of organ development. The authors take advantage of layer-specific mutants of the MADS-box gene PhDEF in Petunia hybrida to dissect the roles of this floral identity regulator in the epidermis and mesophyll. By combining single-cell RNA sequencing with ChIP-seq analyses in wild-type and mutant chimeric petals, the work demonstrates that, although PhDEF is expressed at comparable levels in both petal layers, it binds to and regulates a substantially larger and more layer-enriched set of genes in the epidermis than in the mesophyll. The identification of both layer-specific and shared PhDEF binding sites supports a model in which pre-existing layer identity modulates the regulatory output of homeotic transcription factors.

      Major comments:

      Dissecting PhDEF binding preferences in the epidermis versus the mesophyll using the star and wico mutants is a clever and powerful approach. However, conclusions involving cell identity should be drawn with caution for two reasons. First, cell identities appear to be altered in the mutants: scRNA-seq data suggest that even cells assigned to the same cluster can be molecularly distinct across genotypes. For example, the transcriptomes of wild-type epidermis and mesophyll are highly similar (Pearson correlation R = 0.94), yet both show lower correlation with epidermal cells from star or wico petals. These results raise questions such as are the cells identified as epidermal cells really strictly epidermal in the mutants? Do you need to take into consideration cell composition of the mutants when you do differential expression analysis?Second, PhDEF is expressed in both epidermal and mesophyll clusters in all genotypes, albeit at different levels and in fewer cells in the mutants. As a result, the ChIP-seq profiles should be interpreted as cell type-enriched rather than cell type-specific.

      We fully agree with this reviewer's comments, and alterations in cell identity in the star and wico mutants is indeed a main pitfall for our analysis. The phdef mutation alters the transcriptomic signatures of cells physically located in the epidermis or in the mesophyll, which can result in their artificial clustering with cells located elsewhere in the petal. It might be particularly true for the epidermis, since we see strong depletion in epidermal clusters in the star flowers (and even in the wico flowers), whereas the mesophyll is not much affected in wico flowers. As a result, it is likely that we lose many epidermal cells in star flowers that end up labeled as mesophyll cells, resulting in the under-estimation of DEGs in this layer. We had explored other possible ways to define epidermal or mesophyll cells in our dataset, based for instance on PhDEF or PhGLO1 expression, but since for the majority of cells PhDEF expression is simply not captured, we would have wrongly assigned phdef mutant identity to WT cells. In spite of this limitation, we find more DEGs in the epidermis than in the mesophyll, showing that even an underestimation of DEGs in the epidermis does not affect our main conclusion, which is that PhDEF has a major regulatory action in the epidermis. We have explicitly written this limitation in our manuscript, lines 234-236.

      We agree that ChIP-Seq profiles are rather cell type-enriched than cell type-specific, which we have stated explicitly in the sentence line 308, saying that "differential binding between layers is quantitative". However, for simplicity we prefer to retain the term "specific" since our conclusions are based on the definition of peaks that can either be present or absent in a given genotype, and hence in a given cell layer.

      And there are a few things that need clarification:

      a) Protoplasting and tissue dissection analyses suggest that mesophyll cells constitute more than 80% of the cells in wild-type petals, whereas the scRNA-seq data indicate a substantially lower proportion. Could this discrepancy reflect technical biases in cell recovery or capture efficiency, or issues related to cell identity assignment during clustering and annotation? Notably, the scRNA-seq data from star and wico petals show mesophyll proportions close to 80%. Is this difference due to an increased abundance of mesophyll cells in the mutants, or could it instead reflect differences in transcriptomic separability? In wild-type petals, the epidermal and mesophyll transcriptomes are highly correlated and express similar numbers of genes, with epidermal cells distinguished mainly by higher expression of a subset of genes. This raises the possibility that mesophyll cells in the wild type occupy a more plastic or less differentiated transcriptional state and may therefore be misclassified as epidermal cells, whereas disruption of regulatory mechanisms in the mutants enhances transcriptional divergence and alters cell clustering outcomes.

      Our protoplasting and tissue sections show that the mesophyll should represent 80% of cells in WT tissue, while we estimate it at 70% in our WT scRNA-Seq data based on our cluster assignment (mesophyll = 60% + vasculature = 10%, that we separated from the mesophyll cells but is actually embedded within this tissue). This is not a very strong difference, and considering the multiple steps that protoplasting and cell capture entail, we considered that this was reasonably close to the expected proportions.

      In the star and wico flowers, we assign mesophyll identity to a greater proportion of cells, but as explained above, we believe that cells with altered epidermal identity are easily clustered as mesophyll cells, since they lose their specific transcriptomic signature. This is actually in line with one of the main messages of our article, that petal epidermis transcriptional identity is highly specific.

      b) Cluster 0 appears to show internal heterogeneity, as the expression patterns of KCS3 and LLE2 are largely mutually exclusive. Do these patterns reflect the presence of distinct epidermal cell types within the limb that are currently grouped into a single cluster?

      Indeed, there appears to be some internal heterogeneity within cluster 0. Since our main focus was to compare epidermal and mesophyll clusters, we did not explore further the heterogeneity within epidermal clusters and kept a coarse resolution.

      c) Clusters 0 and 7 both exhibit high expression of pigmentation genes, while cluster 7 additionally shows strong enrichment for cell division genes. Are cell cycle genes the primary features distinguishing these two clusters? If cell cycle effects are regressed out, would cluster 7 merge with cluster 0, potentially yielding a more continuous cell state trajectory and helping to resolve the pattern noted in point (b)?

      To answer one of Reviewer 1's comments, we have added additional DotPlots to better describe our clusters, in Figure S3. Clusters 0 (limb epidermis), 6 (upper limb epidermis) and 7 (replicating cells) exhibit high expression of pigmentation genes (in particular cluster 6), as depicted in Figure 1D. Cluster 7, consisting of only 30 cells, is the only cluster expressing histone genes and S-phase genes. It is possible that these few cells are cluster-6 cells that are replicating, but considering the very low number of cells involved, we have not explored any further their identity and decided to remove them, as it would only marginally affect the conclusions of our analyses. It is indeed surprising that cluster 6 (upper limb epidermis) is quite distinct in the UMAP space to the other epidermal clusters 0 (limb epidermis) and 4 (upper tube epidermis), and we have not observed similar situations in other scRNA-Seq studies. We speculate that this is due to the joint expression of anthocyanin-related and photosynthesis-related genes, which convey a very strong transcriptomic signature to these cells that distinguish them from other epidermal cells.

      d) Pearson correlation is largely driven by highly expressed genes and may therefore be insensitive to changes in cell identity markers. The conclusions that the less clear separation of mesophyll and epidermal cells in star is due to altered cell identity would be more convincing if supported by independent validation, such as in situ hybridization or reporter analyses, to directly visualize molecular alterations in the relevant cell types when comparing wild-type and mutant tissues.

      Following another reviewer's comments, we have now given less emphasis on the Pearson correlation analysis, that was to some extent subjective. Therefore, our conclusion that mesophyll and epidermal cells in star are less separated than in WT has been removed.

      Minor comments:

      (1) For color-coded figure legends (e.g., Fig. 1C), please also include the cluster numbers. This would facilitate interpretation, particularly for readers with reduced color sensitivity.

      We have now used colour-blind-friendly palettes and we have added cluster numbers in Figure 1C.

      (2) For figures containing abbreviations (e.g., Fig. 1D, st. / ca.), please explicitly define all abbreviations in the figure legend.

      We have defined all abbreviations in the figure legends.

      (3) For all UMAP figures, and for figures involving comparisons across clusters, please use consistent color schemes for the same clusters throughout the manuscript.

      We have modified color schemes across the manuscript for colour-blind-friendly palettes, consistently used throughout the manuscript.

      (4) In Fig. S6, the plot showing all genes does not exactly match Fig. 1C, although it appears to represent the same data. Please use the same version of packages, seed values and parameters for all UMAP plots to avoid such discrepancies.

      Figure S6 is the result of integrating with Harmony the WT dataset (as shown in Figure 1C) with WT datasets after removing genes differentially expressed by the protoplasting process. Therefore these UMAPs are a result of integrating different datasets than in Figure 1C, which changes the shape of the UMAPs but cannot be controlled by the seed values, to our knowledge.

      Reviewer #3 (Significance):

      Overall, this work significantly advances our understanding of late homeotic gene function. It establishes compelling evidence for how developmental context constrains transcription factor activity and offers broadly relevant insights for studies of organ patterning. The combination of genetic mosaics with single-cell and chromatin-level analyses represents a powerful and generalizable strategy that will be of interest to both plant developmental biologists and researchers studying transcriptional regulation.

      My expertise: single cell genomics, chromatin biology and plant development.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The uniqueness of this paper is the study of the formation of temporal binding-dependent memories in the cntnap2 mouse, a long-standing mouse model of autism that has been used to test therapeutic modalities.

      Strengths:

      I liked the combination of optical recordings and interventions and the backup of primary observations with control experiments.

      Weaknesses:

      (1) Fiber photometry recordings are too coarse to give salient clues to the underlying mechanism.

      We acknowledge that fiber photometry provides population-level measurements and does not resolve the activity of individual neurons or synaptic mechanisms. Our aim was to identify alterations in the activity of defined neuronal populations during behaviour rather than to delineate the underlying cellular mechanisms. We agree that future studies employing higher-resolution approaches, such as two-photon calcium imaging, in vivo electrophysiology, or single-cell recordings, would provide important mechanistic insights into the circuit changes underlying the observed activity patterns.

      (2) Are perturbed pyramidal cells causally responsible for the altered trace? What can be concluded about the possible role of inhibitory interneurons as potential drivers? The observations focus on abnormal regional activity as observed with fiber photometry and manipulated by optogenetics. The authors should state clearly the limits of their conclusions.

      Our optogenetic experiments demonstrate a causal role for the targeted neuronal population in modulating the observed activity and behavioural phenotype. However, we do not conclude that this population is solely responsible for generating the altered fiber photometry signal, nor do we infer that it represents the exclusive driver of the underlying circuit dysfunction. Rather, our findings demonstrate that manipulating this population is sufficient to alter circuit activity, while acknowledging that the recorded signals likely reflect interactions between multiple neuronal populations. We have revised the manuscript to make these distinctions clearer.

      (3) I found the "trace" nomenclature confusing. "....in which mice are required to memorize the association between a tone (Conditioned Stimulus) and a mild electric foot-shock (Unconditioned Stimulus), separated by a time interval called Trace (Sellami et al., 2017)." It seems that the conceptual model invokes the creation of an [eligibility] trace, characterized by its progressive disappearance over time. It may be a convention in the field or a matter of language, but it seems perverse to use "trace" to label the time interval rather than the entity that is decaying. If this is an accepted convention going back to Howard Eichenbaum, the authors should cite the paper that first introduced the convention.

      We thank the reviewer for this comment. This is indeed an established convention in the classical/Pavlovian conditioning literature rather than terminology specific to our study or to Sellami et al. (2017). The term "trace conditioning" was coined by Pavlov (1927), who used "trace interval" to designate the empty period separating CS offset from US onset, precisely because — as the reviewer intuits — successful conditioning across this gap requires the organism to maintain a memory trace of the CS. In other words, the interval is named for the cognitive/neural entity (the decaying CS trace) that must be sustained across it, not because the interval itself is thought to be a physical or decaying object. So "trace interval" is shorthand for "the interval across which a trace must be maintained," analogous to how "delay conditioning" refers to a paradigm named for a temporal property of the procedure rather than the mechanism per se. We have added a citation to Pavlov (1927) at first use of the term, and clarified the phrasing to make the etymology explicit, as suggested.

      (4) I would advocate for the addition of some discussion points for the authors to consider.

      (a) Is the retention of activity in CA1 related to phenomena at the cellular or subcellular level in CA1 pyramidal cells? I'm thinking of dendritic, delayed, and stochastic CaMKII activation (DDSC) as defined by Yasuda's group or short-term and associative plasticity of calcium dynamics (STAPCD) as delineated by Caya-Bissonette and Beique.

      We thank the reviewer for this insightful suggestion. Our study was designed to investigate network-level dynamics, and the approaches used do not allow us to determine whether the retained activity in CA1 arises from intracellular mechanisms, such as dendritic calcium dynamics or CaMKII-dependent signalling (including mechanisms such as DDSC or STAPCD), recurrent circuit interactions, or a combination of both. We therefore cannot directly assess the contribution of these cellular and subcellular processes. We have now added a paragraph to the Discussion acknowledging that persistent dendritic calcium signalling and CaMKII-dependent plasticity are plausible contributors to sustained CA1 activity and represent an important avenue for future investigation.

      (b) Was the optogenetic intervention ever administered in a delayed fashion, capitalizing on the temporal advantages of optogenetics to probe dynamics?

      This is indeed an important control. While we did not perform this intervention in the current study, this control was part of our seminal study demonstrating the causal role of the dCA1 in temporal binding (Sellami et al., PNAS, 2017, Fig. 1E doi: https://doi.org/10.1073/pnas.161965711). We showed that aged mice had reduced temporal binding capacity, associated with decreased dCA1 activity. We successfully rescued temporal binding capacity in aged mice through ChR2-induced activation of dCA1 pyramidal neurons during the trace interval, but not outside of the trace interval. Given the striking similarity between the temporal binding deficits observed in aged mice and those reported here in Cntnap2 KO mice, we did not repeat this previously established control in the present study. To acknowledge this point, we have also added a statement to the Discussion noting that confirming the temporal specificity of dCA1 optogenetic manipulation in the Cntnap2 KO model will be an important direction for future studies.

      (c) Is the newfound reliance on corticostriatal pathways something more than compensation at the behavioral level? Could it be driven in part by the ASD-related genetic changes?

      We agree that the increased reliance on corticostriatal pathways could reflect both compensatory recruitment at the behavioral level and a direct consequence of ASD-related genetic alterations affecting circuit development and function. Our current data demonstrate a shift in circuit engagement but do not allow us to distinguish whether this represents an adaptive compensation or a primary consequence of Cntnap2 deficiency. However, given that ASD-associated mutations can alter the development, connectivity, and plasticity of corticostriatal circuits, it is possible that the observed changes reflect intrinsic circuit reorganization rather than solely a compensatory behavioral strategy. We have revised the Discussion to acknowledge this possibility and to clarify that altered corticostriatal recruitment may represent a direct consequence of the genetic disruption that subsequently shapes behavioral strategies.

      Reviewer #2 (Public review):

      The authors investigate the contribution of dorsal CA1 hippocampal dysfunction to cognitive impairments in the Cntnap2 knockout mouse model of autism spectrum disorder. Building on previous evidence implicating the hippocampus in episodic and relational memory processes, they combine trace fear conditioning, fiber photometry, optogenetic manipulation, a relational/declarative memory radial maze task, and cFos mapping to test whether altered CA1 function contributes to deficits in temporal binding and memory flexibility.

      The study has several important strengths. First, the work addresses a relatively understudied aspect of autism-related cognition, namely hippocampal-dependent memory processes, whereas much of the literature has focused on social behavior, cortical circuits, or striatal dysfunction. Second, the authors employ multiple complementary approaches that converge on a coherent mechanistic hypothesis. The behavioral data demonstrate a reduced ability of Cntnap2 knockout mice to retain associations across long temporal gaps. Fiber photometry recordings reveal reduced dorsal CA1 activity during conditions that challenge temporal binding, and optogenetic activation of dorsal CA1 neurons during the trace interval is sufficient to rescue memory performance. Together, these findings provide strong support for a causal contribution of dorsal CA1 activity to temporal binding deficits in this model.

      The second major strength of the manuscript is the extension of these findings to a more complex hippocampus-dependent memory task. The radial maze experiments indicate that Cntnap2 knockout mice show impaired memory flexibility and a greater reliance on egocentric learning strategies. The accompanying cFos analyses suggest altered recruitment of hippocampal and striatal networks during learning, providing a systems-level framework that may explain the observed behavioral phenotype.

      Overall, the main conclusions regarding impaired temporal binding and reduced dorsal CA1 engagement are well supported by the data. The optogenetic rescue experiments are particularly compelling because they move beyond correlation and directly test causality. The manuscript therefore makes a meaningful contribution to our understanding of how hippocampal dysfunction may contribute to cognitive abnormalities associated with autism.

      Weaknesses:

      Some conclusions are necessarily more inferential than others. In particular, the interpretation that the observed behavioral phenotype reflects a broader shift from hippocampal-dependent declarative memory toward striatum-dependent procedural learning is supported primarily by cFos activity patterns and behavioral strategy measures. While the data are consistent with this interpretation, they do not directly demonstrate a causal reorganization of memory systems. Similarly, although the findings identify a mechanism in the Cntnap2 model, caution is warranted when extrapolating these conclusions to autism spectrum disorder more broadly; but I believe this caution is addressed in the discussion.

      We agree that our data do not directly demonstrate a causal reorganization of memory systems with cFOS activity patterns. Our intention was to propose that the behavioural strategy together with the brain-wide cFos activation patterns are consistent with a shift in the relative engagement of hippocampal- and striatal-dependent networks. We have revised the manuscript to moderate our interpretation throughout, replacing causal language with wording that reflects an association between the observed behavioural changes and altered recruitment of these memory-related circuits.

      Despite these limitations, the study presents a coherent and well-executed body of work that provides novel mechanistic insight into hippocampal contributions to cognitive dysfunction in a widely used autism model. The findings should be of considerable interest to researchers studying hippocampal function, memory systems, and neurodevelopmental disorders.

      Reviewer #3 (Public review):

      Summary:

      The manuscript evaluated behavioral phenotypes in the Cntnap2 knockout mouse using two behavioral paradigms: trace fear conditioning and a radial maze task. The trace fear conditioning training is normal, but memory generalization is impaired. The inflexibility is suggested to be related to low activity in dCA1 neurons, which can be rescued by ChR2. The radial maze task data suggested a similar conclusion. Brain-wide cFos mapping indicated impairments in the Cntnap2 knockout mouse. The brain-wide cFos mapping does not show direct correlations with Cntnap2, limiting the interpretation of these data in the context of this paper.

      We agree that brain-wide cFos mapping does not directly identify the molecular or cellular mechanisms by which Cntnap2 deficiency alters circuit function. Rather, we use cFos as a functional readout of network recruitment during behaviour. Our interpretation is therefore limited to identifying differences in activity patterns associated with the behavioural phenotype, rather than establishing direct mechanistic links to Cntnap2 function. We have clarified this point in the revised manuscript and moderated the text accordingly.

      Strengths:

      The behavior data are solid.

      Weaknesses:

      The underlying mechanism is not fully investigated.

      Major points:

      (1) The authors should thoroughly check their manuscript as there are many typos in the current version that affect the readability.

      (2) In trace fear conditioning, the tone test impairment can be rescued by ChR2. Have the authors tried rescue experiments with Cntnap2? Rescue experiments in the radial maze task are also essential, either with ChR2 or Cntnap2.

      We agree that rescue experiments in the radial maze would provide additional mechanistic insight by testing whether restoring dCA1 activity optogenetically in Cntnap2 KO mice is sufficient to rescue declarative memory flexibility. However, we have already established the causal role of dCA1 in temporal binding and relational/declarative memory in the radial maze task in our earlier work (Sellami et al., PNAS, 2017, Figure 2E). Indeed, optogenetic inhibition of dCA1 in the radial maze task, specifically during the inter-trial interval, prevented flexible relational/declarative memory formation. In the present study, we confirmed the causal contribution of dCA1 activity to temporal binding in Cntnap2 KO mice in the trace fear conditioning paradigm (Figure 1J-L). Therefore, repeating the optogenetic experiment in the radial maze would provide only limited additional information, while requiring dedicated experimental cohorts and additional validation. Moreover, as the team is based in 2 different institutions and countries (France and Australia), we unfortunately do not have the capacity to perform these experiments.

      Regarding a rescue with the Cntnap2 protein, here, we used the Cntnap2 KO as a well-accepted model of autism spectrum disorder (ASD), rather than understanding the role of this protein in temporal binding capacity and memory formation. The present study was hence designed to determine how dCA1 activity relates to temporal binding and memory performance in ASD, building on the causal evidence established in our previous work. We therefore consider Cntnap2-specific rescue experiments an important follow-up to further test the sufficiency of dCA1 activation, rather than an essential experiment for establishing the conclusions of the present study.

      We have highlighted these points as an important direction for future investigations and revised the manuscript to better distinguish our findings from this additional mechanistic question.

      (3) The quality of the cFos example image in Figure 3 is too low. The authors should also provide example images for the other brain regions in the supplementary data, if possible.

      The low quality of the cFos image in Figure 3 resulted from a formatting issue during figure uploading. This has now been corrected, and a higher-resolution image has been included in the revised manuscript. We have also added representative images for the other brain regions analyzed to the Supplementary Information, as requested.

      (4) The causal link between the brain-wide cFos mapping and the Cntnap2 knockout is weak. How to explain the increase of cFos cell densities in some brain regions, but the decrease in others?

      We agree that cFos mapping cannot explain the mechanisms underlying the regional increases and decreases in activity. We interpret these bidirectional changes as reflecting differential recruitment of distributed brain networks rather than direct effects of Cntnap2 deficiency on individual brain regions, and have clarified this in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Improve clarity of the Figure layout: Figure 1, panel I legend. I think the identification of experimental groups is in the wrong place. It really applies to panels K, N, and L, not to panel I. This is important because the genetic stimulation in panels K, N, and L is the key experiment in this whole figure.

      We agree that the placement of the experimental group identification in the original figure legend could be confusing. We have revised the figure layout and legend so that the description of the experimental groups is now associated with panels K and L, where the optogenetic stimulation experiments are presented, thereby improving the clarity of the figure.

      (2) Fix Figure references. "This indicates impaired retention of trace fear memory and temporal binding ability, relative to WT mice (Figure 2I)....: [there is no Fig. 2I, so I'm guessing you mean Fig. 1I; subsequent Figure references also seem to confuse Fig. 1 and 2].

      We thank the reviewer for identifying these errors. We have carefully checked and corrected all figure citations throughout the manuscript, including the reference to Figure 1I-L in this section, to ensure that each citation now refers to the appropriate panel.

      (3) Missing words? In both T20 and T40 conditions, regardless of genotype, [the response to] tone frequency significantly decreased.

      We thank the reviewer for spotting this omission. The sentence has been corrected to read: "In both T20 and T40 conditions, regardless of genotype, the response to tone frequency significantly decreased."

      (4) Jargon that should be declared at an appropriately early point. R/DM = relational/declarative memory.

      We have now introduced the term as relational/declarative memory (R/DM) at its first appearance in the manuscript and have ensured that the abbreviation is used consistently thereafter.

      (5) Confusing statement? "In the 40s trace-conditioned group, while the frequency of the calcium transients did not differ between WT and Cntnap2 KO mice throughout conditioning (Figure S1G), their amplitude was significantly reduced both during the presentation of the second tone and consistently across the three trace intervals in Cntnap2 KO mice (Figure 1I)."

      We have revised this sentence to improve its clarity and explicitly distinguish between the frequency and amplitude of calcium transients. The revised text now makes clear that, although the frequency of calcium transients did not differ between genotypes throughout conditioning, their amplitude was significantly reduced in Cntnap2 KO mice during the second tone presentation and across the three trace intervals.

      Reviewer #2 (Recommendations for the authors):

      (1) The manuscript would benefit from a clearer distinction between conclusions directly supported by the data and broader interpretations. In particular, statements suggesting a shift from declarative to procedural memory systems could be presented more cautiously, as the cFos analyses provide indirect rather than causal evidence for such reorganization.

      We thank the reviewer for this constructive comment. We have revised the manuscript throughout to more clearly distinguish conclusions that are directly supported by our data from broader interpretations. In particular, statements referring to a shift from declarative to procedural memory systems have been moderated.

      (2) Additional clarification of the behavioral interpretation of the radial maze task would be valuable for readers who are less familiar with this paradigm, particularly regarding its relationship to relational/declarative memory and its distinction from procedural learning strategies. In addition, the results' interpretation in this paradigm is unclear to me. What is the significance of the 20-second delay in this task (why not use 10 sec or 60 sec)? I believe the paradigm tests spatial rather than temporal distant items (in contrast with trace fear conditioning that refers to temporally distant events)? Please explain further the link between the two behavioral paradigms.

      We thank the reviewer for these comments.

      Regarding the relationship between relational/declarative and procedural learning: The R/DM task was designed by our group (Marighetto and colleagues; Mingaud et al., 2007; Sellami et al., 2017, 2018) specifically to dissociate two learning/memory systems that can support the same behavioral output (correct arm choice) but that differ fundamentally in their underlying representations and, critically, in their flexibility.

      During acquisition, an animal can solve each of the three arm-pair discriminations either by forming a flexible, relational representation of the whole spatial configuration (i.e. an allocentric/hippocampus-dependent "cognitive map" strategy, in which the reward's location is encoded relative to distal cues and to the other pairs) or by learning a set of rigid, response-based rules (i.e. an egocentric/striatum-dependent "turn left/turn right" procedural strategy tied to each specific pair).

      Both strategies can produce equivalent accuracy during initial acquisition, which is why acquisition performance alone cannot distinguish them. The flexibility probe (recombining pairs A and B into a novel pair AB, without moving the reward) is the critical dissociation: only an animal that encoded the reward's location relationally, within a broader spatial map, can generalize correctly to this untrained configuration; an animal that relied on a rigid stimulus– response rule for each pair individually would be unable to solve the recombined pair above chance. Performance on pair AB is therefore the read-out of hippocampus-dependent relational/declarative memory, while the degree of left–right lateralization during acquisition (quantified by a lateralization index) is a converging behavioral signature of reliance on the egocentric/procedural system.

      We have clarified this distinction in the Methods and Results to facilitate interpretation of the paradigm.

      Regarding the 20-s delay and its relationship to trace fear conditioning: The reviewer is correct that the R/DM task and TFC differ in the information being integrated: the R/DM task involves spatially distinct discrimination episodes, whereas TFC involves temporally separated stimuli presented in the same context. However, both tasks require information separated by an interval to be integrated into a unified memory representation, a process referred to here as “temporal binding” (Sellami et al., 2017). The 20-s inter-trial interval in the R/DM task was based on our previous work, in which this interval was shown to be sensitive to age-related deficits in temporal binding and relational memory (Sellami et al., 2017). Similarly, our TFC experiments established a longer temporal-binding capacity in young mice (up to 40 s) compared with aged mice (20 s). Thus, although the two paradigms involve different types of information, they share the requirement to maintain and integrate information across a temporal gap. The different intervals used in the two paradigms reflect their distinct task structures and demands.

      (3) Although the authors address potential locomotor confounds, a brief discussion of how hyperactivity and impulsivity might influence behavioral performance in the different paradigms would strengthen the interpretation of the results.

      We thank the reviewer for raising this point. We agree that distinguishing hyperactivity from impulsivity strengthens the interpretation of our behavioral phenotypes, and we have added a discussion paragraph making this distinction explicit. Briefly: in the TFC paradigm, our existing open-field data (Figure S1D) show that while Cntnap2 KO mice travel more distance and move faster, their resting time is unaffected relative to WT; so, the reduced freezing/temporal binding phenotype at the 40 s trace cannot be attributed to a general locomotor confound, since freezing (immobility) was actually comparable between genotypes during acquisition. In the R/DM radial maze task, however, we agree the phenotype is better read as impulsivity than hyperactivity per se: KO mice show shorter decision latencies specifically at the choice point (Figure 2G), with no corresponding genotype effect on their post-choice running speed for correct trials (Figure 2H), indicating that the effect is on deliberation before response rather than on general motor output. We have revised the Discussion to make this distinction, and its implications for interpreting the egocentric/procedural bias, explicit.

      (4) Minor presentation issues:

      - Ensure that all statistical tests, sample sizes, and post hoc comparisons are reported consistently throughout the manuscript and figure legends.

      We have carefully reviewed the statistical reporting throughout the manuscript and figure legends to ensure consistency. We have clarified sample sizes, statistical tests, and post hoc comparisons where required and in the methods.

      - Carefully proofread the manuscript for minor typographical and grammatical errors. For instance, a) several figures numbers indicated in the main text do not correspond to the figures the authors refers to (page 5 Figure 2I, page 6 Figure 2J and 2K, page 7 "data in fig S2"... and so on); b) rephrase (not clear) page 9: "Differences in the relative contribution of individual substructures between trained WT and Cntnap2 KO mice involved dCA1, dCA2, and dCA3, and favored the DMS, IL, and ACC."

      We thank the reviewer for highlighting these issues. We have carefully proofread the manuscript and corrected all figure reference errors and typographical inconsistencies throughout the text. We have also revised the unclear sentence on page 9 to improve clarity and accuracy.

      - Page 9, an additional sentence is needed to explain the significance of the bibliographical reference: "Atypicalities in declarative memory have been reported in ASD, with specific impairments in relating items (Minor et al., 2023)."

      We have revised this section to clarify the significance of the cited study and its relevance to our findings. Specifically, we have expanded the sentence to highlight that impairments in relational processing in ASD may reflect difficulties in integrating individual experiences into coherent memory representations, a process that relies on hippocampal function.

      Reviewer #3 (Recommendations for the authors):

      Major Points:

      As stated in the public review, the authors need to thoroughly check their manuscript before submission. There are too many typos in the current version that affect the readability. For example, but not limited to:

      (1) Page 5, "Figure 2I" should be "Figure 1I"; "Figure 2J" should be "Figure 1J"; "Figure 2K" should be "Figure 1K".

      (2) The figure legends of Figure 1J-L are missing.

      (3) Page 8, "In contrast, Cntnap2 KO mice showed no learning-induced dCA1 activation together with hypoactivation of dCA3 in both naïve and trained groups compared to WT mice (Figure S2A)", should it be dCA2 instead of dCA1? Please double-check the manuscript.

      We thank the reviewer for highlighting these issues. We have thoroughly checked the manuscript and corrected all typos, figure reference errors, and fixed figure legend information. Regarding the point (3), we did mean to describe the lack of training-induced activation of dCA1, but failed to reference back to Figure 3, likely causing the confusion. This has now been clarified. The manuscript has also been carefully proofread to improve clarity and readability.

      Minor points:

      There are two duplicate rows (two wt) in their raw datasheet, Fig3BCD & S2 and Fig3EF. Please correct them in case of further problems.

      WT Naive 178.736 103.821 250.703 185.229 437.575 50.535 266.254 306.187 306.187 WT Naive 178.736 103.821 250.703 185.229 437.575 50.535 266.254 306.187 306.187

      WT Naive 46.87 37.94 15.19 8.57 4.98 12.02 8.88 20.98 2.42 12.77 14.68 14.68 WT Naive 46.87 37.94 15.19 8.57 4.98 12.02 8.88 20.98 2.42 12.77 14.68 14.68

      We thank the reviewer for noticing this mistake. These duplicated rows have been corrected.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This study addresses an important question in insect toxicology by systematically evaluating glycogen phosphorylase as a potential insecticidal target. The authors combine complementary biochemical, molecular, physiological, and structural approaches, including recombinant enzyme characterization, inhibitor assays, RNA interference, metabolite profiling, structural modelling, and measurements of fitness-related traits. This integrative approach provides a comprehensive evaluation of the biological consequences of glycogen phosphorylase suppression. In particular, the biochemical evidence that diflubenzuron does not inhibit glycogen phosphorylase, together with the observation that strong suppression of glycogen phosphorylase produces only transient physiological effects without measurable impacts on development or reproduction, provides strong support for the conclusion that glycogen phosphorylase is unlikely to represent an effective standalone insecticidal target.

      We thank the reviewer for the positive assessment of our study and for recognizing the value of our integrative approach, including recombinant enzyme characterization, RNAi, metabolite profiling, structural modelling, and fitness measurements. We are also grateful for the acknowledgement that our biochemical evidence and phenotypic observations provide strong support for the conclusion that GP is unlikely to be an effective standalone insecticidal target.

      Weaknesses:

      (1) The proposed metabolic compensation mechanism is supported by indirect evidence.

      We agree with the reviewer that our data demonstrate correlation with, rather than direct proof of, increased gluconeogenic flux. We have revised the manuscript throughout to moderate our interpretations and to clearly frame the compensation model as a plausible interpretation supported by multiple lines of indirect evidence, rather than an established mechanism. Specific revisions are detailed below.

      Specific revisions:

      (1) Abstract (Lines 35–39):

      “We provide evidence that insects compensate through a multi-layered metabolic response: upregulation of gluconeogenic enzymes (PEPCK, G-6-Pase), selective upregulation of glycogen branching enzyme (GBE) but not α-amylase, and changes in protein content suggestive of catabolic substrate mobilization”

      (2) Abstract (Lines 41–42):

      “Indicating that the metabolic compensation response is ultimately effective in sustaining development”

      (3) Abstract (Lines 42–44):

      “These findings suggest that GP is functionally non-essential under these conditions, likely through a combination of gluconeogenic compensation and the availability of alternative carbon sources”

      (4) Introduction (Lines 88–90):

      “Our findings reveal that PxGP is functionally non-essential for larval development under the tested conditions, reflecting a previously uncharacterized gluconeogenic compensation mechanism”

      (5) Results – Gene expression (Line 265):

      “Gene expression analysis is consistent with gluconeogenic activation”

      (6) Results – Protein content (Line 278):

      “Changes in protein content suggestive of catabolic substrate mobilization”

      (7) Results – Protein decline (Line 283):

      “This protein decline is consistent with substrate mobilization”

      (8) Results – GBE expression (Line 389):

      “PxGP knockdown selectively upregulates glycogen branching enzyme expression”

      (9) Results – GBE differential response (Lines 400–403):

      “This differential expression pattern—selective GBE upregulation with unchanged α-amylase expression—indicates that the compensatory response to GP suppression involves targeted remodeling of glycogen structure rather than a generalized upregulation of all glycogen-degrading enzymes.”

      (10) Results – Fitness assessment (Lines 411–412):

      “To determine whether the metabolic changes observed following PxGP knockdown are associated with measurable physiological consequences”

      (11) Results – Fitness interpretation (Lines 426–430):

      “GP suppression likely triggers protein catabolism to supply amino acids for gluconeogenesis, contributing to transient weight loss. In the presence of continuous dietary carbohydrate supply, the compensatory response appears sufficiently effective to restore metabolic homeostasis before developmentally critical transitions”

      (12) Discussion – Glycogen accumulation paradox (Line 520-521):

      “a paradox that could be explained by the activation of GP-independent glycogen catabolism via alternative enzymes like...”

      (13) Results – Trehalose and G6P (Lines 299–303, 315–317):

      “Trehalose levels remained stable at 48-72 h but increased substantially by 96 h (7.44-fold elevation, P < 0.05) (Figure 9F). This increase coincided with the significant upregulation of gluconeogenic genes (e.g., PEPCK and G-6-Pase), consistent with the idea that gluconeogenesis contributes to trehalose synthesis and ensures adequate carbohydrate reserves for the upcoming pupation”

      And: “Together, the coordinated behavior of G6P and trehalose is consistent with gluconeogenesis-derived glucose being converted into storage and transport carbohydrates”

      (14) Results – Integrated interpretation (Lines 330–332):

      “Collectively, these metabolite and gene expression data reveal a biphasic metabolic adaptation that provides a coherent explanation for why substantial GP suppression causes no mortality or developmental defects”

      (15) Discussion – Definitive proof (Lines 527–531):

      “At 96 h, trehalose and G6P levels in dsGP-treated larvae were maintained at markedly higher levels than in dsGFP controls, which had declined by this time point, coinciding with 3–4‑fold upregulation of PEPCK and G‑6‑Pase. These changes strongly support, albeit indirectly, de novo glucose synthesis as the primary rescue mechanism.”

      (16) Discussion – Protein catabolism confirmation (Lines 505):

      “provides independent biochemical support for protein catabolism”

      (2) Some mechanistic interpretations extend beyond the data presented.

      We accept this criticism and have systematically revised the manuscript to distinguish more carefully between observed transcriptional/metabolic changes and functional interpretations. We now consistently frame these as correlative evidence consistent with—but not proving—the proposed mechanisms.

      Specific revisions:

      The revisions listed under Weakness 1 above—particularly those modifying language around protein decline (Lines 278, 283), gene expression (Lines 265, 389), and trehalose/G6P changes (Lines 299–303, 314–317)—also directly address this weakness by removing the implication that transcriptional or protein-level changes constitute functional proof of pathway activation. Additionally, we have made the following revisions:

      (17) Additional discussion of GBE/α-amylase limitation (Line 405-408)

      We have added explicit acknowledgement that transcriptional changes alone do not demonstrate functional pathway activation:

      “However, as these observations are limited to the transcript level, further enzymatic activity assays or glycogen structure analyses would be required to determine whether this transcriptional change translates into functional alterations in glycogen mobilization.”

      (18) Fitness section – clarification (Lines 426–430):

      We have revised the interpretation of transient larval weight loss as described above, introducing “likely” and clarifying that the data are consistent with the model rather than conclusively demonstrating it.

      (3) An alternative explanation for the limited phenotype is not fully considered.

      We thank the reviewer for this important point. We agree that the continuous availability of dietary carbohydrates under our experimental conditions represents a plausible alternative explanation for the lack of phenotype. We have added explicit discussion of this alternative interpretation in the Abstract, Introduction, and Discussion sections.

      Specific revisions:

      (19) Abstract (Lines 42–44):

      As shown above, we have revised the abstract to include: “These findings suggest that GP is functionally non-essential under these conditions, likely through a combination of gluconeogenic compensation and the availability of alternative carbon sources.”

      (20) Introduction (Lines 88–90):

      As shown above, we have revised the introduction to include: “under the tested conditions” and to acknowledge that the observed phenotype reflects a combination of factors.

      (21) Discussion – Biphasic response (Lines 531–539):

      We have revised the discussion of the biphasic response to moderate the causal language and acknowledge alternative interpretations:

      “This biphasic response—initial metabolic stress followed by delayed (48-72 h) and robust compensation—creates a temporal buffer by 96 h and is consistent with the absence of overt phenotypes despite substantial GP suppression. This temporal coordination raises the possibility of a developmentally programmed response, potentially involving hormonal regulation, that anticipates energy demands at critical transitions. Critically, this pattern suggests a potential vulnerability window (~72 h) during which combined inhibition of glycogenolysis and gluconeogenesis might overcome compensation, although this remains speculative and would require experimental validation.”

      (22) Discussion – Alternative explanation paragraph (Lines 540–547):

      We have added a following paragraph explicitly addressing the alternative explanation:

      “We also acknowledge that the absence of a severe phenotype may reflect an additional, non‑mutually exclusive explanation: under our experimental conditions, with continuous dietary carbohydrate availability, GP activity may not be rate‑limiting for maintaining glucose homeostasis. The observed transient larval weight loss could thus reflect both active metabolic compensation and the inherently low demand on GP for glucose supply in a feeding larva. Distinguishing between active compensation and the non‑limiting nature of GP will require future studies under nutrient‑restricted conditions or during fasting intervals, where GP's role is likely to become more critical.”

      Additional revisions to moderate language throughout the manuscript

      Beyond the specific revisions addressing each weakness, we have also made the following broader language modifications to ensure consistent and cautious interpretation throughout:

      (23) Changes from “ablation” to “suppression” (Lines 180, 195–197, 357):

      “Ablation” to “suppression” throughout.

      (24) Conclusions – Fundamental principle (Lines 619):

      “our findings uncover a fundamental principle of...”

      (25) Conclusions – Strategic path (Lines 621–623):

      “This metabolic plasticity renders GP non-viable as a standalone insecticidal target but illuminates a potential strategic path forward”

      (26) Discussion – Trehalase inhibitors (Lines 556–558):

      “For example, trehalase inhibitors such as Validamycin A are potent insecticides [44, 45], suggesting that trehalase inhibition may not be subject to the same degree of metabolic compensation”

      (27) Discussion – Metabolic pincer (Line 561):

      “a 'metabolic pincer' attack that might overcome adaptive compensation”

      (28) MM/GBSA and structural predictions (Lines 27–32, 450, 455–457):

      Abstract (Lines 27–32)

      “Molecular docking and MM/GBSA analysis predict that this selectivity reflects differential side-chain engagement: GPI is predicted to occupy the allosteric site at the dimer interface via contacts with seven residues spanning both subunits (ΔG = −34.63 kcal/mol), whereas DFB's difluorobenzoyl moiety is predicted to remain solvent-exposed without productive protein contacts (ΔG = −29.29 kcal/mol).”

      Line 450

      “MM/GBSA analysis indicated”

      Lines 455–457

      “These structural predictions are consistent with established structure–activity relationships of acyl urea compounds [15], in which target selectivity is governed by the side-chain substitution pattern rather than by the shared acyl urea core”

      (29) Scope of conclusion (Lines 471–472):

      “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for diflubenzuron”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major recommendations

      (1) Moderate the interpretation of the proposed compensatory mechanism throughout the manuscript.

      We fully agree with the reviewer's assessment. Throughout the manuscript, we have systematically moderated the language used to describe the compensatory mechanism, ensuring that our interpretations are framed as plausible inferences supported by multiple lines of indirect evidence rather than as established conclusions. The specific revisions are detailed in items (1)–(7), items (10)–(16) of our response to Weakness 1 above, and items (24)–(25) of our response to Weakness 3 above.

      Additional language modifications addressing the reviewer's concern about assertive terminology are detailed in our response to Minor Recommendation 1 below.

      Line 495

      “Our investigation shows that this tolerance reflects a robust”

      Line 498

      “As insects elicited a compensatory gluconeogenic pathway.”

      Line 518

      “Direct metabolite quantification reflected this compensation”

      (2) Clarify the evidence supporting alternative glycogen degradation pathways.

      We thank the reviewer for this valuable suggestion. We have revised the GBE and α-amylase sections to more clearly distinguish between transcriptional changes and functional pathway activation. The revised wording now explicitly acknowledges that our observations are limited to the transcript level and that functional confirmation would require additional experiments. The specific revisions are detailed in items (8)–(9) and item (12) of our response to Weakness 1 above. We have also added a sentence acknowledging that these observations are limited to the transcript level (item 17) of our response to Weakness 2 above.

      (3) Discuss alternative explanations for the limited physiological phenotype.

      We thank the reviewer for this important point. We agree that the continuous availability of dietary carbohydrates under our experimental conditions represents a plausible alternative explanation for the lack of phenotype. We have incorporated this alternative interpretation into the Abstract, Introduction, and Discussion sections, as detailed in items (3)–(4) of our response to Weakness 1 above, and items (21)–(22) of our response to Weakness 3 above.

      We have also revised the subsequent conclusion (Lines 553–554) to align with this alternative interpretation, changing “GP fails as a standalone target due to compensation via gluconeogenesis” to “GP appears to fail as a standalone target, at least in part because of compensation via gluconeogenesis.” This ensures consistency with the preceding discussion of GP’s potentially non‑rate‑limiting role.

      Minor recommendations

      (1) Review the manuscript for statements using terms such as "demonstrates," "confirms," "proves," or "establishes."

      We have systematically reviewed the entire manuscript and replaced overly assertive terms (e.g., “demonstrates,” “confirms,” “proves,” “establishes”) with more cautious language (e.g., “suggests,” “is consistent with,” “provides evidence that,” “indicates”) wherever they refer to the proposed metabolic mechanism. Specific revisions are listed under Major Recommendation 1 above. In addition:

      MM/GBSA and structural predictions (Lines 27–32, 450, 455–457): Abstract (Lines 27–32): Changed “Molecular docking and MM/GBSA analysis reveal that...” to “Molecular docking and MM/GBSA analysis predict that...”

      Results (Line 450): Changed “MM/GBSA analysis confirmed” to “MM/GBSA analysis indicated.”

      Results (Lines 455–457): Changed “These structural data confirm that...” to “These structural predictions are consistent with...”

      Conclusion scope (Lines 471–472): Changed “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for BPUs” to “This provides the first direct biochemical evidence excluding GP as a candidate molecular target for diflubenzuron.”

      Changes from “ablation” to “suppression” (Lines 180, 195–197, 357): Changed “ablation” to “suppression” throughout the manuscript where referring to GP knockdown.

      (2) Ensure that the distinction between changes in transcript abundance, enzyme activity, and metabolic flux is maintained consistently throughout the Results and Discussion.

      We have carefully reviewed the Results and Discussion sections to ensure that we consistently distinguish between transcript abundance (measured by RT-qPCR), enzyme activity (measured by activity assays), and inferred metabolic flux (not directly measured). The revisions listed under Major Recommendations 1 and 2 above directly address this point. Key changes include:

      Consistently using “transcriptional upregulation” or “expression” when referring to qPCR data, rather than “activation” or “pathway activity.”

      Explicitly acknowledging that transcriptional changes do not necessarily reflect metabolic flux.

      Using “is consistent with” rather than “demonstrates” when linking gene expression changes to functional outcomes.

      Additional clarifications and corrections to data presentation

      During the preparation of the revised manuscript, we also made several corrections and clarifications to the data presentation:

      RNAi knockdown efficiency (Line 364–367): We have added a clarification that the knockdown efficiency in the cohort used for enzyme activity and fitness assays (54.66% at 48 h) was lower than that achieved in the dose–response experiment (87.59% at 48 h), reflecting batch-to-batch variation between independently injected cohorts. We have also noted that the enzyme-activity data should be interpreted against the transcript reduction measured in this same cohort.

      Total protein data comparability (Lines 374–380): We have added a note clarifying that the total-protein data shown in Figure 9D and Figure 10–figure supplement 2 derive from independent experimental cohorts processed at different homogenization ratios; absolute protein concentrations are therefore not directly comparable between the two panels. Both datasets nonetheless show a transient decline of approximately 30% in total protein in dsGP-treated larvae within the first 72 h.

      Trehalose and G6P data interpretation (Lines 304–309, 314–315, 345–353): We have revised the description of trehalose and G6P levels at 96 h to clarify that the large fold-differences primarily reflect a pronounced decline in dsGFP control values at this time point, rather than a net increase in absolute metabolite content in dsGP-treated larvae. The revised wording now indicates that GP-suppressed larvae maintain their trehalose and G6P pools at a stage when control larvae are actively depleting them.

      Gene expression recovery kinetics (Lines 251–252): We have corrected the description of PxTre and PxHex expression at 96 h to state that they “returned to, and modestly exceeded, control levels” rather than “returned to near baseline levels.”

      Cell line description update (Lines 644–654): In response to editorial requirements, we have updated the cell line description to include supplier authentication, mycoplasma testing status, and confirmation that the cells were used experimentally within one year of purchase.

      Figure corrections:

      Figure 2 legend: revised to state “maximum inhibition did not exceed 55.07 ± 7.05%” to match the main text.

      Figure 4 legend: corrected normalization description from “control set to 1.0 at each time point” to “calibrated to the 24 h control sample (set to 1.0).”

      Figure 6 legend: corrected reference sample from “L1 set to 1.0” to “Egg set to 1.0.”

      Figure 8D: corrected y-axis label from “Adult emergency (%)” to “Adult emergence (%).”

      Figure 9C: corrected in-figure label from “G-6-P” to “G-6-Pase.”

      Terminology correction: We have standardized the use of “similarity” (rather than “identity”) when referring to sequence comparisons (Lines 190, 1075).

      PROSITE motif name correction (Lines 181–184): We have corrected “phosphatase-pyridoxal phosphate linkage site” to “phosphorylase pyridoxal-phosphate attachment site.”

      Figure 11 statistical description (Lines 1138–1142): We have revised the figure legend to specify that data in (B) and (C) are mean ± SEM of three biological replicates, while data in (D–F) are shown as box plots with sample sizes indicated, and the statistical comparison is dsGP vs. dsGFP (independent samples t‑test). These revisions do not affect any results or conclusions.

      Gene expression peak description (Line 301): Changed “coincided temporally with peak expression of gluconeogenic gene” to “coincided with the significant upregulation of gluconeogenic genes” to avoid implying a single defined peak.

      Minor textual corrections

      We have also corrected the following textual issues:

      Sentence structure (Lines 433–435): We have revised the sentence structure to correct a comma splice and clarify the logical relationship. The original sentence “To provide structural insight into the observed selectivity, GPI potently inhibits PxGP (IC<sub>50</sub> = 2.96 nM) while DFB does not, we performed...” has been revised to “To provide structural insight into the observed selectivity—in which GPI potently inhibits PxGP (IC<sub>50</sub> = 2.96 nM) while DFB does not—we performed...”

      Additional minor corrections: We have corrected a small number of typographical errors and stylistic inconsistencies throughout the manuscript.

      All of these corrections are limited to data presentation, figure labeling, and textual clarity. They do not alter any experimental results, quantitative conclusions, or the overall interpretation of the study.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The rationale behind averaging sentence embeddings across multiple transformer models (with different architectures and training objectives) is unclear. These transformer-based models have different training paradigms and model architectures, which may result in misaligned semantic spaces. The averaging operation may dilute the distinct sentence representations learned by each model, potentially weakening the overall semantic encoding for sentences. Please clarify this choice or cite supporting methodology.

      The reviewer questions the rationale for averaging sentence embeddings across different models. However, our method involves computing correlations separately for each model, then averaging the correlations. We apologize for the confusion. We have clarified this on page 3:

      “Results for the ‘Transformers’ model are computed by computing correlations separately for five different transformer models and then taking a simple average of these correlations. Results for each individual transformer are presented in Supplementary Information Figure S2.”

      (2) All structure-sensitive models discussed incorporate semantics to some extent. Including a purely syntactic baseline, such as a model based on context-free grammar, would help confirm the importance of syntactic structures.

      Following the suggestion, we have implemented two syntactic models and discuss the results on page 10:

      “We also found that purely syntactic models based on constituency parses (see Benepar and CFG) show poor correlations with brain activity (see Supplementary Information Figure S2). Examining the corresponding RSA matrices (see Figure S1), this seems to be due to such models being overly sensitive to syntactic form, and relatively insensitive to which words are assigned to different nodes within the syntactic tree. This is most evident for the edit-distance similarity metric, and to a lesser extent also for the subtree similarity metric. This finding highlights the value of hybrid approaches designed to appropriately balance sensitivity to lexical, syntactic, and compositional information in representing semantic information at the sentence level.”

      (3) In Figure 2, human behavioral judgments show weak correlations with neural data, and even fall below those of computational models, suggesting the behavioral judgments may not reflect the sentence structures in a brain-like way. This discrepancy between behavioral and neural data should be clarified, as it affects the interpretation of the results.

      While the behavioural judgements are made by different participants and involve a different task than the neuroimaging results, nonetheless we agree the difference is surprising and warrants more detailed consideration. We have included a more detailed discussion of this issue on page 11:

      “Our study has several limitations. First, we found a surprisingly low correlation between behavioural ratings and brain activations (see Figure 2). This may be partly explained by differences in task structure. In the behavioural experiment, participants viewed many pairs of related sentences, and were explicitly asked to pay attention to differences in the words of each sentence. In contrast, in the fMRI task, participants read one sentence at a time without an explicit comparison. In addition, we suspect that presentation of so many sentence pairs with highly similar structures may have biased the way in which participants rated sentence similarity. Modifications to the behavioural task to mitigate these aspects may reduce the divergence between behavioural and brain findings.”

      (4) To better contextualize model and neural performance, sentence similarity should be anchored to a notion of semantic "ground truth", such as the matrix shown in Figure 1a. Comparing this reference with human judgments, brain responses, and model similarities would help establish an upper bound.

      While our design matrix served as the basis for constructing a set of stimuli with systematic modifications, we respectfully suggest that it should not be regarded as a ‘semantic ground truth’. Sentence pairs within each category will not have the same degrees of semantic similarity since the words and context differ across sentences in a graded manner. Furthermore, while we anticipated ‘different’ sentence pairs would be less similar than ‘swapped’ sentence pairs, and that within each of the six block diagonals the ‘modified’ or ‘substituted’ sentence pairs would be the most similar, we did not have any prediction about the magnitude of these differences. Our goal was to construct a set of sentence pairs which spanned a range of semantic similarities, and allowed for dissociation between lexical similarity and overall similarity in meaning. The design matrix is not intended to represent a ‘ground truth’ that human judgements or brain representations would be expected to conform with.

      (5) The structure of this paper is confusing. For instance, Figure 5 is cited early but appears much later. Reordering sections and figures would enhance readability.

      We agree that placement of figures was not ideal in the previous draft. We have reworked the manuscript so that all figures appear closer to their mention in the text, and the figure (now Figure 3) appears in the correct order. We have also substantially revised the discussion, and included subheadings to help guide the reader through the various different issues we include.

      (6) While the analysis is broad and comprehensive, it lacks depth in some respects. For instance, it remains unclear what specific insights are gained from comparing across brain regions (e.g., whole brain, language network, and other subregions). Similarly, the results of simple-average and group-average RSA appear quite similar and may not advance the interpretation.

      We included both analyses in line with our preregistration, and also because we believe the fact that two distinct approaches to analyzing the data yield similar results strengthens our conclusions.

      (7) While explaining the grid-like pattern due to sentence length is important, this part feels somewhat disconnected from the central question of this paper (word order). It might be better placed in supplementary material.

      We believe that the grid-like pattern in the RSA results is an important unexpected finding that warrants discussion in the main manuscript.

      Reviewer #3 (Public review):

      (1) The interpretation of findings is nuanced. Although Transformers underperform as brain models on the critical subsets of controlled sentences, a Transformer outperforms all other models when evaluated on the union of all sentences when both word-level content and structure vary. Transformers also yield equivalent or better models of human behavioral data. Thus, although Transformers have demonstrable flaws as human models, which are pinpointed here, in the general case, (some) Transformers are more human-like than the other models considered.

      The reviewer argues that we overstate some of our conclusions, as several transformers achieve higher brain correlations than the hybrid model when computed over all sentence pairs, as well as on the behavioural data. In response, we first note that our primary interest in this paper is on the block-diagonal sentence pairs, as these were specifically designed to interrogate how different models represent sentence structure. The comparison with all sentence pairs is presented for comparison but is not our primary focus on this paper, as also reflected in the pre-registered prediction that our VerbNet-CN hybrid model would show higher brain correlations than transformers over this block diagonal subset.

      Second, we have included a new analysis in the revised manuscript (Figure S9) where we compute brain correlations controlling for the pattern of similarities observed in the primary visual cortex (averaged over participants), as a way to control for visual similarity. This added control substantially reduces the brain correlations of the transformers, such that they all have lower correlations than VerbNet-CN and AMR-smatch even over the set of all sentence pairs. We provide interpretation of this result in the discussion.

      Third, we would like to note one of the disadvantages of transformers as a model of mind or brain representations is that they are largely a ‘black box’ whose workings are poorly understood. One advantage of hybrid models like our simple semantic role model is that they can be much easier to interpret, thereby enabling them to be used to determine which features are most important for brain representations of sentence meaning, and what mechanisms are used to combine individual words into a full sentence. Given their relative simplicity and interpretability, we believe hybrid models have considerable value as scientific tools, even in cases where they achieve comparable correlations to transformers. We have added a short discussion of this issue in the revised manuscript (page 10).

      (2) There may be confounds between the critical sentence structure manipulations and visual representations of sentence stimuli. This is inconvenient because activation in brain regions that process semantics tends to partially correlate with visual cortex representations, and computational models tend to reflect the number of words/tokens/elements in sentences. Although the study commendably controls for confounds associated with sentence length, there could still be residual effects that remain. For instance, the Graph model correlates most strongly with the visual cortex despite these sentence length controls.

      We agree with the reviewer that this is a potential confound. As noted in the previous response, we have implemented a new control analysis in which we directly control for visual similarities as reflected in participant-averaged similarities of primary visual cortex activations in response to all stimuli. These results are shown in Figures S8-S11 in the SI. We show that transformer correlations are reduced much more than graph and hybrid models with this control. Also, we note that the AMR-smatch graph model shows high correlations with other brain regions even after removing correlations with the visual cortex (Figure S10). This indicates that the model represents a range of sentence features, including both superficial visual or length-related features, as well as semantic features that are represented in common with language and other cortical regions.

      (3) Sentence similarity computations are emphasized as the basis for unifying comparative analyses of graph structures and vector data. A strength of this approach is that correlation is not always the ideal similarity metric. However, a weakness is that similarity computations are not unified across models. This has practical consequences here because different similarity metrics applied to the same model produce positive or negative correlations with brain data.

      The reviewer notes that the method for computing similarities differs between the vector-based (mean and transformer) models, and the hybrid and syntax-based models, thereby potentially adding an additional confound to our results. We agree that this is a potential limitation, and our correlations should always be understood as applying to a model paired with a similarity metric. However, we believe that this is mostly unavoidable when comparing different formalisms. In the revised manuscript we have incorporated an entirely new similarity metric for vector-based models (DIEM similarity), as well as an extended discussion of the effect of different similarity metrics for graph and hybrid models.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Objectives of the study and impact of the work:

      The authors of this article primarily aim to reconstruct the evolutionary history of the insect odorant receptor (OR) family, which is responsible for the detection of odorant signals by olfactory neurons. Due to the lack of phylogenetic signal present in the sequences of this multigene family, which evolves very rapidly, phylogenetic analyses have so far never made it possible to precisely retrace how ORs diversified prior to the appearance of present-day insect orders, and what the drivers of this diversification were. For example, one may suspect that the adaptation of ORs to odors emitted by plants constituted a critical step in insect evolution during the "angiosperm terrestrial revolution," which occurred at the end of the Cretaceous, but nothing currently allows this to be asserted.

      There are very nice examples, notably in Drosophilids, derived from comparisons between closely related species and documenting mechanisms of OR adaptation to certain signals. However, what the authors attempt to do in this work is to produce a macroevolutionary analysis at the scale of insects as a whole, based almost exclusively on bioinformatic analyses. To do this, they annotated OR genes in about one hundred insect species and developed pipelines for analyzing sequence similarity, structural similarity, and functional similarity, the latter being estimated through a molecular docking approach. An important feature in the evolution of insect ORs is the emergence of a unique co-receptor, called Orco, which appears to be an OR that has lost the ability to bind odorants. In addition to the largescale bioinformatic analysis, the authors also aim to explore more specifically the factors that favored the emergence of Orco and the selective advantage conferred by the existence of OR-Orco complexes.

      Given the importance of odorant receptors in insect biology and in their adaptation to different environments and lifestyles, retracing their evolutionary history is indeed a major question in evolutionary biology. In principle, this type of work therefore has the potential to become a reference in the field and to provide a basis for significant scientific advances.

      Major strengths and weaknesses:

      The sampling chosen for collecting OR sequences is very impressive, with more than 100 insect families represented, covering most of the major orders. This sampling appears appropriate for the question being addressed. The analysis pipeline used to collect the sequences makes sense, relying on homology-based annotation tools coupled with a structure-based filter. Nevertheless, one can note aberrant numbers of ORs for certain species (much lower than reality), which indicates that the pipeline probably did not function correctly for all genomes. In the absence of a validation step comparing the results with already known OR repertoires, it is difficult to estimate the overall quality of the data. The authors chose to apply a fairly stringent filter on sequence quality (based on predicted 3D structure), which reduces the number from 14,000 to 9,000. This choice seems logical given the subsequent use of these data, but it inevitably leads to data loss. The fact that some OR genes may be missing and that the total number may not be exact for each species is not prohibitive for studying the evolution of the family at a broad scale; however, it calls into question certain results that rely on this total number, such as the correlation between the number of ORs and genome size, lifestyle, and diet.

      We thank the reviewer for raising this important concern. To objectively evaluate how much our strict structural filtering may have reduced OR counts, we collected published OR annotations for species included in our study and compared those values with our structurally intact OR counts (Author response table 1). For most previously annotated species, the difference was within approximately ten ORs, indicating that our counts are generally comparable to published annotations after structural filtering. Diabrotica virgifera virgifera was a clear exception. Previous work, using transcriptomic evidence and manual annotation, reported 193 ORs, including 124 complete ORs and three pseudogenes, whereas our original pipeline retained only nine structurally intact ORs. We found that this discrepancy was mainly caused by insufficient query representation. The previous D. virgifera virgifera annotation was published in 2025, whereas our original annotation was performed in 2022, so our query set likely did not adequately represent this lineage. When we re-annotated this species using new queries, 78 complete ORs were retained after structural filtering. We have updated the D. virgifera virgifera annotation in the revised dataset.

      We also considered genome size as a possible contributor to low automated recovery. Among the 115 species analyzed, D. virgifera virgifera has the third-largest genome, after Locusta migratoria and Thermobia domestica. Large genomes contain many repetitive regions, which can make automated OR annotation more difficult and may reduce the number of recovered intact gene models. The two larger-genome species in our dataset had previously been manually annotated with transcriptomic support, suggesting that automated annotation alone may be less reliable for large, repeat-rich genomes. At the same time, published OR annotations are not always free from error. For example, Propsilocerus akamusi was previously reported to have 17 ORs, whereas our pipeline retained ten complete ORs. After downloading and modeling the 17 published sequences, we found that PaOR1, PaOR5, PaOR6, PaOR9, and PaOR14 differed substantially from canonical OR structures or appeared to be incorrectly annotated (Author response image 1). Truncated genes such as PaOR4, which lacks part of the intramembrane region of TM2-4 but still clearly resembles an OR, would be retained by our strategy; sequences with more severe structural inconsistency would be excluded. Because many previous OR annotations have not been experimentally validated, it is difficult to treat all published OR counts as exact ground truth.

      To test whether possible OR loss affected OR-count-based conclusions, we first removed all potential GR sequences from the dataset. For species with published OR counts, we then repeated the pGLS analyses using the larger of our structurally intact OR count and the published OR count (Fig. S1b-e). The main ecological associations were retained: OR count remained associated with larval diet, adult diet, and habitat, and was not associated with circadian rhythm. This indicates that the main OR count ecological analyses are reasonably robust to moderate underestimation caused by missed ORs or strict structural filtering.

      In contrast, the association between OR count and genome size was no longer supported in this sensitivity analysis (Author response image 2a and b). We therefore removed the analysis and discussion of a negative relationship between genome size and OR count. The original pattern likely reflected the difficulty of annotating ORs in large, repeat-rich genomes rather than a reliable biological relationship.

      The following revision has been added to the revised manuscript (lines 122-126): " We then repeated the phylogenetic generalized least squares (pGLS) analyses using the larger of our structurally intact OR count and the published OR count, and the main ecological associations were retained. This indicates that the main OR-count ecological analyses are reasonably robust to moderate underestimation caused by missed ORs or strict structural filtering."

      These new benchmark and sensitivity analyses directly strengthen the evidence supporting our dataset and its ecological interpretations.

      From the dataset collected, the authors attempted to categorize ORs in several ways, starting with the reconstruction of sequence similarity networks. The approach is interesting, but in the end, the results do not seem to be sufficiently exploited, and it is not obvious what the advantage of this approach is compared with the "classical" phylogenetic approach, which generally fails to reveal homology relationships between ORs from species belonging to different insect orders. Here again, the majority of the clusters identified are "order-specific," and when this is not the case, the authors did not attempt to exploit the results. For example, clusters SeqC26 or SeqC28, which appear to be shared by many insects, are potentially very interesting. It might have been relevant to combine this similarity-based clustering approach with phylogenetic reconstructions within each shared cluster.

      We thank the reviewer for this insightful suggestion. We agree that the previous version did not sufficiently explore the evolutionary information contained in shared sequence communities such as SeqC26 and SeqC28. Following the suggestion, we extracted the sequences from these key shared communities and reconstructed within-community gene trees to search for orthologous groups.

      Because shared SeqCs may contain conserved genes across insect orders, we used the gene trees of SeqC26 and SeqC28 to identify candidate cross-order orthologous groups. In SeqC26, we identified 56 orthologous groups, 11 of which included more than ten species. Among these 11 larger groups, 4 spanned multiple insect orders. One example, OGG54, includes sequences from Orthoptera, Blattodea, Trichoptera, and Hymenoptera; after excluding the possibility of GR contamination, this group includes locust LmigOR5 and LmigOR4. LmigOR5 has been reported to bind geranyl acetone and to be associated with avoidance behavior in the migratory locust (Chang et al. 2023). The functions of the corresponding receptors in other insect orders remain to be tested, and we now present these as candidate conserved modules rather than confirmed functional orthologs. In contrast, we did not identify clear cross-order orthologous groups in SeqC28. This suggests that the ORs in SeqC28 are more similar at the sequence-community level, but do not necessarily reflect traceable orthologous relationships.

      The following revision has been added to the revised manuscript (lines 156-164): "SeqC26 and SeqC28 are large shared clusters that may contain cross-order orthologous relationships. In SeqC26, we identified 56 orthologous groups, four of which spanned multiple insect orders. OGG54 included ORs from Orthoptera, Blattodea, Trichoptera, and Hymenoptera. This group contains the locust receptors LmigOR5 and LmigOR4. LmigOR5 has been reported to bind geranyl acetone and to be associated with avoidance behavior in the migratory locust. The functions of the corresponding receptors in other insect orders remain to be tested, and we now present these as candidate conserved modules rather than confirmed functional orthologs. In contrast, we did not identify clear cross-order orthologous groups in SeqC28."

      The clustering based on structure also leads to the identification of a majority of "orderspecific" clusters, but once again, the clusters shared by several orders are not truly exploited, which does not provide new insight into the evolution of ORs. However, the authors highlight a group of ORs in flies that appear to possess an unusual intracellular region. This is interesting, although it is a result more relevant to OR structure than to their evolution. The function of these ORs in Drosophila melanogaster, if it is known, is not discussed.

      We thank the reviewer for this useful suggestion. Similar to the sequence communities, the structural communities also contain broadly shared groups. In the revised analysis, we focus in particular on StrC23, the only structural community present in 12 insect orders. StrC23 accounts for 46.3% of all structurally intact ORs in our dataset (4124 of 8905 ORs).

      Given the extremely low sequence similarity among insect ORs, the existence of such a broad conserved structural community is important. We found that StrC23 has a significantly larger binding-pocket volume than other ORs and Orco (Fig. S3 d). From a structural perspective, this suggests that StrC23 receptors may have greater potential to accommodate diverse VOCs. We therefore interpret StrC23 not as sequence-level conservation, but as conservation at the level of OR structural evolution: many species appear to retain a large set of ORs with relatively large binding pockets, which may provide a structural basis for broad docking-derived VOC binding potential before lineage-specific OR diversification occurs.

      Following the reviewer’s suggestion, we also examined available functional data for the long-IL3 receptors in StrC17. The long-IL3 orthologous group includes Drosophila melanogaster DmOr13a (droMel44 in our dataset) and Bactrocera dorsalis BdorOR13a (bacDor_18 in our dataset). Previous studies indicate that both receptors respond to 1-octen-3-ol (Kreher et al. 2008; Xu et al. 2023). Earlier work also identified a group of Dipteran fly OR homologs, including DmOr13a and BdorOR13a, as 1octen-3-ol-specific responsive receptors associated with oviposition behavior(Liu et al. 2023). When we modeled these homologous receptors, they also showed an extended IL3 region.

      We therefore propose that the long-IL3 ORs identified here likely correspond to the previously reported fly homologs involved in 1-octen-3-ol responses. However, 1-octen-3-ol has different behavioral functions in different species: it acts as an attractant in blood-feeding mosquitoes and tsetse flies (Hall et al. 1984; Kline et al. 2007), contributes to plant-host localization in parasitoids(Morawo and Fadamiro 2016), functions as an aggregation cue in some beetles(Pierce et al. 1989), and is often associated with oviposition regulation in flies (Kreher et al. 2008; Liu et al. 2023). Mosquito receptors known to detect 1-octen-3-ol, such as AgOR8, AaOR8, TaOR8, and CquiOR118b(Lu et al. 2007; Xu et al. 2015; Dekel et al. 2016; Frunze et al. 2024), do not contain a long IL3 region (Fig. S3 e). This raises the possibility that the long IL3 structure in fly homologs may relate to the oviposition-associated role of 1-octen-3-ol in flies.

      Because OR ligand binding is primarily mediated by the binding pocket, IL3 may not directly determine ligand binding. Instead, it may influence downstream regulation of olfactory responses. Mutational analysis of Orco suggests that IL3 can regulate channel activation and Orco-dependent olfactory responses, indicating that variation in intracellular loops may affect response efficiency or sensitivity(Turner et al. 2014; Bobkov et al. 2021). We now discuss this only as a plausible mechanism requiring future experimental validation.

      The following revision has been added to the revised manuscript (lines 196-202): "StrC23 is the only structural community present in 12 insect orders and accounts for 46.3% of all structurally intact ORs in the dataset. We found that StrC23 has a much larger binding-pocket volume than other ORs and Orco. From a structural perspective, this suggests that StrC23 receptors may have greater potential to accommodate diverse VOCs. This reflects conservation at the level of OR structural evolution: many species appear to retain a relatively large set of ORs with larger binding pockets, which may provide a structural basis for structure-derived volatile organic compound binding potential before species-specific OR diversification."

      The following revision has also been added (lines 408-419): "The long-IL3 OR orthologous group includes Drosophila melanogaster DmOr13a and Bactrocera dorsalis BdorOR13a, both of which respond to 1-octen-3-ol and belong to a group of dipteran fly OR homologs associated with oviposition behavior. Modeling of these homologous ORs showed that they also contain a long IL3 region. We therefore propose that long-IL3 ORs correspond to these fly homologs that recognize 1-octen-3-ol. However, 1-octen-3-ol has different behavioral functions in different species: it acts as an attractant in blood-feeding mosquitoes and tsetse flies, whereas in flies it is usually associated with oviposition regulation. Mosquito receptors known to detect 1-octen-3-ol, such as AgOR8, AaOR8, TaOR8, and CquiOR118b, do not contain a long IL3 region. This suggests that the long IL3 structure in fly homologs may be related to the oviposition-associated role of 1-octen-3-ol in flies, which requires future experimental validation."

      The analysis of structural diversity then leads the authors to focus on the Orco co-receptors, which are characterized by modifications of the binding pocket and the emergence of an extracellular loop that could explain the loss of the ability to bind odorant molecules. This part, which relies on in vitro experiments, is interesting and constitutes the most striking result of the study, which could in itself have been the subject of a separate manuscript. However, the molecular dynamics modelling does not add anything in the way it is conducted (5 ns is too short).

      We thank the reviewer for the positive assessment of our Orco results and for raising this important concern. We agree that 5 ns is insufficient to evaluate the conformational stability or long‑term dynamics of the receptor complex. However, our simulations were designed specifically to examine whether the EL2 β‑sheet could sterically affect the early movement of VOCs near the extracellular entrance of the binding pocket. Importantly, rather than relying on a small number of long trajectories, we performed 60 independent short simulations with the explicit goal of capturing a statistical trend in how VOCs respond to this local steric constraint. Within 5 ns, VOC trajectories were visibly altered by the EL2 β‑sheet. Across the 60 replicates, VOCs reached the binding pocket less frequently in Orco proteins containing this structure, revealing a consistent trend (Fig. 4C and Movie S1). We therefore consider this timescale sufficient to detect the early steric response of small molecules, and the large number of repeats provides statistical confidence in this conclusion.

      The following revision has also been added (lines 688–691): “It should be noted that the molecular dynamics simulations were designed to examine the early trajectories of VOCs near the extracellular entrance of the binding pocket, rather than to assess the conformational stability of the receptor complex. A duration of 5 ns is sufficient for small molecules to respond to local steric constraints.”

      The rest of the manuscript is based on the prediction of OR response spectra using molecular docking. The work that has been carried out is extremely substantial, and the objective of linking clusters based on sequence similarity or 3D structural similarity with functional categories is entirely relevant. Nevertheless, I see two major problems with this in silico functional analysis:

      (1) The docking score threshold used was chosen thoughtfully, which is very good, and according to the calculation performed, should ensure a true positive rate of more than 20%, which is excellent in such a docking analysis. But in the absence of functional validation, this 20% true positive rate is not sufficient to extrapolate OR function as the authors do in the remainder of the manuscript. The risk of error remains too high to compare in such detail the function of ORs from insects with different lifestyles or diets.

      We thank the reviewer for pointing out this important limitation of our computational functional interpretation. We fully agree that, although the threshold calibrated from available experimental data improves enrichment, a true-positive rate of approximately 20% is not sufficient to make deterministic claims about individual OR-VOC pairs. We therefore do not equate docking predictions with true OR response spectra. To prevent any misinterpretation, we have undertaken a major revision of our terminology and framing throughout the manuscript:

      The main purpose of docking in our study is not to predict the real ligand of every OR or to compare the exact functional properties of individual receptors across lifestyles. Because the study includes 115 insect species, thousands of ORs, and a large VOC set, systematic experimental validation of all ORVOC combinations is currently not feasible. Instead, we use docking as a high-throughput and internally comparable theoretical screening method to estimate the potential binding tendencies of OR repertoires toward different functional-group VOCs.

      Under this framework, we focus on repertoire-level trends generated by a uniform structural modeling, docking, and scoring pipeline. The value of this analysis is to provide candidate receptors, VOC categories, and ecological associations for future experimental testing. In the revised manuscript, when the results are based solely on docking analyses, we avoid directly referring to them as OR functions. Instead, we add the prefix “docking-derived” to prevent potential misunderstanding. We also explicitly state that confirming specific OR ligands and behavioral functions will require heterologous expression, electrophysiology, calcium imaging, genetic perturbation, or behavioral assays.

      The following limitation statement has been added to the revised manuscript (lines 285-288): "Although this threshold enriches experimentally responsive OR-VOC pairs, the resulting hit rate is not sufficient to support deterministic inference for individual receptor-ligand pairs. Therefore, subsequent analyses were interpreted at the OR repertoire level rather than as experimentally validated OR response spectra.”; (lines 292-293) “It should be noted that dFunCs represent docking-derived functional communities rather than experimentally confirmed functional classes.”

      (2) The six functional clusters identified are only slightly different from one another, with similar detection of all chemical families except acids and amines (which was expected, given that these families are a priori detected by IRs rather than ORs). This shows that even though the approach is relevant and deserves to be tested, it cannot be used to establish a link between groups/lineages of ORs and response spectra at the scale of insects as a whole. This is reflected in the final analysis by the fact that there is no visible link between sequence or structural clusters and functional clusters. Given the uncertainty surrounding the docking results, the entire subsequent analysis of the relationship between the Binding Breadth Index and ecological variables is highly questionable.

      We thank the reviewer for this key comment. We agree that the differences among dFunCs for any single VOC functional group are often modest. However, the dFunC classification was not based only on whether an OR strongly recognizes one specific functional-group category. Rather, it was based on each OR’s overall docking-derived binding profile across a large VOC set.

      In other words, the differences among dFunCs mainly reflect combinations of relative binding tendencies across multiple VOC categories, rather than a strong difference for one chemical family alone. Therefore, even if the predicted binding level for a single functional group differs only slightly among dFunCs, their multivariate profiles can still form distinct dFunCs.

      We also agree that SeqCs, StrCs, and dFunCs do not show a simple one-to-one correspondence. This is expected because the three classifications are based on different information: sequence similarity, overall structural similarity, and docking-derived binding-profile similarity. At present, there is no direct evidence that insect OR sequence, structure, and potential binding profile maintain a conserved one-to-one relationship across the insect class. We also cannot exclude the possibility that false positives in docking reduce the resolution of the functional classification.

      To avoid overinterpretation, we now refer to FunCs as docking-derived functional communities (dFunC) and define BBI as docking-derived BBI. This metric is intended for macro-scale relative comparisons, not for assigning true ligands or specific ecological functions to individual ORs. Ecological interpretations based on dFunC or BBI have been substantially toned down and are presented as candidate trends for future functional validation.

      The following limitation statement has been added to the revised manuscript: "Ecological associations based on BBI should be interpreted as repertoire-level predicted binding trends, not as direct evidence for lifestyle-specific OR function or behavioral olfactory capacity."

      This reframing ensures our conclusions are appropriately supported by the computational evidence and aligns the manuscript's claims with its methodological strengths.

      Finally, the evolutionary analysis proposed to conclude that the work suffers from an incorrect interpretation: ORs of non-holometabolous insects cannot be considered equivalent to those of species that existed before the Permian-Triassic extinction. The fact that a locust or a cockroach has more narrowly tuned ORs than holometabolous insects does not mean that this was also the case for ancestral insects. To advance this type of conclusion, it would be necessary to conduct a phylogenetic analysis and reconstruct ancestral states, which is not the case here.

      In summary, despite the large number of analyses performed, the authors do not succeed in achieving the stated objective of reconstructing the evolutionary history of insect ORs, and the results obtained do not sufficiently support the conclusions regarding the links between OR repertoires and environment or lifestyle.

      We thank the reviewer for this important point. We agree that a rigorous test of changes in ancestral OR functional spectra before and after the EPME would require ancestral reconstruction within a phylogenetic framework based on reliable orthologous groups. For insect ORs, this is currently limited by rapid sequence divergence, very restricted cross-order and even within-order orthology, and frequent gene duplication, loss, and lineage-specific diversification.

      Therefore, our current data cannot reliably reconstruct complete ancestral OR repertoires for different insect-order nodes, nor can they accurately trace extant ORs back to ancestral OR states before and after the EPME. The analysis is better understood as a comparison of OR repertoire composition among extant insect lineages, rather than as a comparison of ancestral states.

      If major environmental transitions influenced the long-term evolution of insect ORs, such effects may appear either as traceable orthologous changes or as broad differences in extant repertoire composition and binding-profile composition. However, differences observed among extant lineages cannot be directly attributed to the EPME because they may also reflect later lineage-specific expansions, gene losses, and ecological adaptations.

      Accordingly, we have revised the conclusion. We no longer state that ancestral insect OR functional spectra changed before and after the EPME. Instead, we state that comparisons among extant insect repertoires reveal differences in docking-derived binding-profile composition among lineages. These patterns may provide hypotheses for exploring links between major geological/ecological transitions and long-term olfactory evolution, but they require broader phylogenetic sampling, ancestral reconstruction, and functional validation.

      The following revision has been added to the revised manuscript (lines 427-430): "We also detected differences in dFunC composition between extant lineages whose order-level origins fall before and after the EPME. Lineages originating after the EPME showed a higher proportion of broad-tuned ORs, likely influenced by multiple factors."

      We also added the following clarification (lines 444-450): “However, because insect ORs evolve rapidly and cross-order orthologous relationships are limited, the current dataset does not allow reliable reconstruction of complete ancestral OR repertoires at deep insect nodes. The EPME-related comparison should therefore be interpreted as a descriptive comparison among extant lineages, not as a direct test of ancestral OR functional changes before and after the EPME. This hypothesis will require further testing through broader taxon sampling, reliable orthology assignment, ancestral-state reconstruction, and functional assays.”

      This revision directly addresses the reviewer's concern by removing the unsupported causal claim and reframing the analysis within the appropriate, evidence-based scope of our study, thereby strengthening the manuscript's contribution as a source of robust comparative patterns and testable macroevolutionary hypotheses.

      Reviewer #2 (Public review):

      The remarkable evolvability of the olfactory system enables animals to rapidly adapt to dynamic and chemically complex environments. Over the past two decades, substantial effort has been devoted to uncovering the evolutionary principles that drive the diversification of odorant receptors (ORs), yielding key insights into the forces shaping their striking variability in both vertebrates and insects. In this manuscript, Zhang and colleagues analyze the OR repertoires of over 100 insect species, leveraging sequence and structural similarity to infer patterns of gene family evolution within this diverse and ecologically important clade. By integrating sequence-based and structure-based comparisons, their study builds on a compelling and recently emerging line of research made possible by the advent of AlphaFold, which has previously clarified the phylogenetic relationship between insect Ors and the gustatory receptor gene family and revealed the unexpectedly deep evolutionary origins of this ancient structural fold.

      Applying this approach to a large set of ORs derived from species throughout the insect phylogeny, the authors confirm many previously reported patterns of OR evolution. Unfortunately, the way these results are presented lacks clarity in what is already known from previous work in the field versus what is a novel finding based on the analysis of this dataset.

      We thank the reviewer for pointing this out. We agree that the original manuscript did not distinguish previously established findings from the novel results of the present study with sufficient clarity. Rapid OR evolution, lineage-specific expansion, and associations between OR repertoires and ecological traits have been reported in several insect groups. We have revised the manuscript to acknowledge these studies more explicitly and to indicate which results confirm known patterns, which extend them across a broader phylogenetic scale, and which arise from our new analyses.

      Our study contributes more than a confirmation of previous observations. To our knowledge, it provides the first integrated analysis of OR evolution across 115 insect species that combines sequence similarity, structural conservation, and docking-derived binding profiles within a unified macroevolutionary framework. This analysis identifies class-wide repertoire patterns and previously unrecognized structural features, including the conserved structural cluster StrC23, that could not be evaluated using narrower taxonomic datasets.

      The study also provides a mechanistic insight into Orco specialization. We experimentally demonstrate that the distinctive EL2 β-sheet of Orco reduces ligand binding affinity, supporting its contribution to the reduced odorant binding capacity of this conserved coreceptor. In addition, because exhaustive experimental characterization of the large insect OR family is currently impractical, our experimentally calibrated computational pipeline provides a framework for identifying repertoire-level trends, generating testable hypotheses, and prioritizing receptors for future functional studies.

      The reviewer raises several specific examples of the distinction between previous knowledge and novel findings in the comments below. We address each of these points in the corresponding responses. We believe these revisions more clearly position our contribution as testing, extending, and integrating previously reported patterns while also identifying new class-wide structural and functional features.

      It is unclear how complete the odorant receptor sets are. I recommend benchmarking the pipeline by comparing its output to a gold standard and a frequently vetted complete OR set, such as that of Robertson and Wanner 2006 or similar.

      We thank the reviewer for this suggestion. We agree that without benchmarking, readers cannot adequately evaluate the completeness and accuracy of the OR annotation workflow. Following the suggestion, we performed detailed comparisons using three relatively well-curated repertoires: Drosophila melanogaster, Anopheles gambiae, and Bombyx mori. Because the species included in our study did not include Apis mellifera from Robertson and Wanner 2006, we did not use that dataset as the benchmark.

      The benchmark addresses three questions. First, how well does the automated genome-based OR annotation workflow recover full-length ORs reported in previous studies? Second, does the structural filtering step improve the structural quality and reliability of the OR dataset? Third, for candidates in tandem-repeat regions where exon-merging errors may occur, do the candidate ORs share genomic positions with the most similar literature ORs, or are they more likely to represent distinct candidates? For Drosophila melanogaster, the accepted repertoire contains 60 OR genes and 65 protein products including splice variants. At the gene level, the automated workflow recovered most benchmark ORs before structural filtering, but Or47b and Or98b were not annotated. After structural filtering, fragmented or structurally incomplete ORs were removed, leaving 56 ORs that otherwise correspond one-to-one with the benchmark ORs. Two ORs that had been annotated before filtering, DmOr59a and DmOr85e, were removed by the structural filter because DmOr59a was fragmented and DmOr85e lacked TM7 in the predicted structural model (Author response image 3).

      For Anopheles gambiae, the published repertoire includes 79 OR genes. Our automated workflow initially annotated 77 ORs, including 72 published ORs. After removing fragments and five full-length literature ORs, the final structurally filtered set contained 67 ORs. We inspected the five removed full-length ORs and found that AgOr52, AgOr47, and AgOr64 lacked a complete transmembrane region, whereas AgOr58 and AgOr6 contained all transmembrane regions but showed local deletions in some regions (Author response image 3). This indicates that structural filtering removes clearly incomplete structures, but can also be conservative enough to exclude a small number of true ORs with atypical transmembrane-region length variation.

      For Bombyx mori, the published repertoire contains 66 OR genes including two pseudogenes. Our automated workflow initially annotated 103 ORs, including 60 previously reported ORs. After removing fragments and seven literature ORs, the final set contained 64 ORs, of which 53 corresponded to literature annotations. For structurally retained candidates that did not map one-to-one to literature ORs, we checked genomic locations (Table R2). Most were not located on the same chromosome or in the same tandem region as their most similar literature ORs, making exon-merging artifacts less likely. They may therefore represent previously unannotated OR candidates, although transcriptomic evidence or manual annotation will be needed for confirmation.

      Overall, the benchmark shows that the automated workflow recovers most known OR repertoires and that structural filtering produces a high-confidence set suitable for structural comparison and docking. We now explicitly describe the final dataset as a set of structurally intact ORs rather than a complete OR repertoire for every species, and we discuss how strict filtering may underestimate OR counts. We added a benchmark description to the main text (lines 100-104) and included the detailed comparison strategy and results in Supplementary Text S1: "Benchmarking against existing OR repertoires from Drosophila melanogaster, Anopheles gambiae, and Bombyx mori showed that the annotation workflow provides high-confidence annotations suitable for downstream structural comparison and docking (see Supplementary Text S1, Table S1-3)."

      Using their structural clustering approach, the authors identify a structural feature mostly unique to the OR co-receptor ORco, a beta-sheet in EL2, which they functionally show reduces odorant binding affinity - a key aspect of ORco, which does not bind ligands in the ancestral ligand-binding site. This is a particularly strong part of the manuscript, since the authors support their in silico-derived hypothesis with functional data.

      We thank the reviewer for the positive assessment of the Orco-related work.

      Lastly, in an attempt to assess the relationship between sequence identity and structure on one hand and function on the other, the authors perform an in silico structure prediction and chemical docking analysis. As it stands, this part is on the more speculative side since the docking approach has not been verified with available functional datasets.

      We thank the reviewer for this comment and agree that the docking analysis is predictive. Molecular docking cannot be treated as experimental evidence for OR response spectra or as a definitive assignment of ligand specificity for individual OR-VOC pairs.

      In the revised manuscript, we clarify that docking is used as a unified, comparable theoretical framework for macro-scale OR repertoire comparison. It is meant to estimate docking-derived binding potential across thousands of ORs and many VOCs, not to replace heterologous expression, electrophysiology, calcium imaging, or behavioral validation.

      We also clarify that the docking framework was not used without any empirical calibration. We used existing OR-VOC experimental response data to examine the relationship between docking score and experimental hit rate, and we evaluated the ability of docking-derived labels to recover functional tendencies in available datasets. These analyses support the use of docking for repertoire-level enrichment and hypothesis generation, but they do not justify deterministic conclusions for individual receptor-ligand pairs. The revised manuscript now states this limitation explicitly.

      The following revision has been added to the revised manuscript (lines 285-288): "Although this threshold enriches experimentally responsive OR-VOC pairs, the resulting hit rate is not sufficient to support deterministic inference for individual receptor-ligand pairs. Therefore, subsequent analyses were interpreted at the OR repertoire level rather than as experimentally validated OR response spectra."

      Summary of Major Revisions:

      In direct response to the eLife Assessment and reviewer comments, our revisions have systematically strengthened the evidence and clarified the interpretation:

      (1) Added critical validations: Benchmarking of the annotation pipeline and sensitivity analyses for ecological correlations.

      (2) Deepened evolutionary analysis: Phylogenetic exploration of shared clusters and structural characterization of StrC23.

      (3) Clarified the scope and limitations of the computational functional analysis: Adopted “docking-derived” terminology and specified that the results represent hypothesis-generating, repertoire-level comparative trends rather than experimentally validated receptor functions.

      (4) Corrected overinterpretations: Revised conclusions regarding the EPME and holometabolous/non-holometabolous comparisons to be descriptive and hypothesis-generating. '

      (5) Improved scholarly accuracy: Updated framing to properly acknowledge prior work and corrected minor points throughout.

      We believe that these revisions address the concerns underlying the assessment of incomplete evidence and provide stronger, more appropriately qualified support for the manuscript’s main conclusions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The annotation pipeline produces a number of genes that is slightly lower than reality, which is expected given the presence of a fairly stringent filter. However, the number is far too low for certain species, for example, Harmonia or Diabrotica. Whenever possible (for example, for Drosophila, Bombyx, Anopheles, etc.), adding a benchmarking step by comparing the results obtained with already identified repertoires would be a real improvement. Moreover, the dataset contains sequences of gustatory receptors that must be removed in order not to bias the analysis (CO₂ receptors in Drosophila melanogaster and Bombyx mori, for example, based on what I was able to verify). It might therefore be necessary to add a specific step in the pipeline to check for the presence of GRs in the data, since the sequences and 3D structures can be quite similar to those of ORs and therefore difficult to separate.

      We thank the reviewer for this careful observation. To evaluate OR annotation counts, we compared our results with published OR annotations for previously annotated species. Except for Diabrotica virgifera virgifera, most species differed from published counts by fewer than ten ORs, supporting the general reliability of our annotation counts. For D. virgifera virgifera, the very low original count was likely caused by a combination of insufficient query representation and the difficulty of annotating ORs in a large genome. The previous D. virgifera virgifera annotation was published in 2025, whereas our original annotation was performed in 2022, so our query set likely did not adequately represent this lineage. When we re-annotated this species using new queries, 78 complete ORs were retained after structural filtering. We have updated the D. virgifera virgifera annotation in the revised dataset. Large-genome insects over 2 Gb are rare in our dataset, and the sensitivity analysis indicates that this issue does not substantially affect the main ecological associations based on OR count. Regarding gustatory receptor contamination, we found that a small number of sequences labeled as olfactory receptors but actually corresponding to GRs had been retained as queries because of an operational error during OR annotation. To remove potential GRs systematically, we collected 1,418 insect GR genes from the literature and databases and combined them with the OR library to create a receptor reference set. We then compared all annotated OR candidates against this reference set using BLASTP and removed candidates whose best match was a GR. In total, we identified and removed 76 potential GRs distributed across 38 species. These included the Drosophila CO<sup>2</sup> receptor genes Gr63a and Gr21a, as well as Bombyx mori BmGr10 and BmGr9, which the reviewer specifically noted. All downstream analyses have been repeated using the corrected dataset. We are grateful for this careful observation, which substantially improved the rigor of the dataset.

      Both Reviewer 1 and Reviewer 2 requested benchmarking against known OR repertoires, and we agree that this is necessary. We therefore benchmarked the workflow against Drosophila, Anopheles, and Bombyx repertoires, evaluating recovery of known ORs, the effect of structural filtering, and potential exon-merging artifacts in tandem regions. We further added sensitivity analyses using the larger of structurally intact OR counts and published OR counts for species with available literature data. These analyses are now included in the revised Results and supplementary materials.

      The revised text is as follows (lines 100-104): "Benchmarking against existing OR repertoires from Drosophila melanogaster, Anopheles gambiae, and Bombyx mori showed that the annotation workflow provides high-confidence annotations suitable for downstream structural comparison and docking (see Supplementary Text S1, Table S1-3)."

      For the SSN approach, the validation by comparison with the phylogeny of hymenopterans is a good idea, but the result is not as convincing as presented in the manuscript since the cluster SeqC25 does not correspond to a monophyletic group. This, therefore, raises the question of the evolutionary information that can be drawn from this clustering approach. The validation with lepidopteran PRs is also a good idea, but it would be necessary to verify exactly to which PR clades the sequences used for the analysis belong, in order to determine which clade corresponds to each of the identified SeqC clusters.

      We thank the reviewer for this important point. For the Hymenoptera comparison, we agree that SeqC25 does not correspond to a strict monophyletic clade in existing Hymenoptera OR phylogenies. We have clarified that the SSN approach is not intended to replace traditional phylogenetic trees or to reconstruct ancestor-descendant relationships among OR communities.

      SeqC25 should be interpreted as a sequence-similarity community that is strongly expanded in Hymenoptera and partially corresponds to previously defined Hymenoptera OR subfamilies, but not as a strict phylogenetic clade. SeqCs describe modular boundaries in OR sequence-similarity space, whereas phylogenetic trees infer branching relationships under a specific evolutionary model. For insect ORs, where sequence similarity is often very low, duplication and loss are frequent, and tree support is often limited, the two approaches need not produce identical groupings.

      We now state that the evolutionary information provided by SSNs mainly concerns sequence-similarity diversification patterns: highly expanded modules, order-specific or cross-order communities, community boundaries, and differences in repertoire composition among lineages. These results complement phylogenetic analysis, but detailed relationships within close taxa or OR subfamilies still require traditional phylogenetic methods.

      For the Lepidoptera pheromone receptor (PR) validation, we agree that the previous analysis mixed experimentally validated PRs and candidate PRs, making it unclear which PR branch each SeqC corresponded to. We reorganized the PR dataset by separating classic PRs from recently reported new PR branches and mapped these sequences to SeqCs using BLASTP. Classic PRs were assigned almost entirely to SeqC6 (93.9%), whereas the new PR branch mapped to SeqC35 (Fig. S2f). Both SeqC6 and SeqC35 are shared by Trichoptera and Lepidoptera, suggesting that they may correspond to distinct PR-related sequence communities.

      The following revision has been added to the revised manuscript (lines 144-150): "This comparison indicates that SSN-based communities are broadly consistent with established phylogenetic subfamilies, while capturing sequence-similarity modules rather than strictly monophyletic clades. We further evaluated the SSN using curated lepidopteran pheromone receptors (PRs). To distinguish different PR lineages, experimentally validated classic PRs and the recently reported new PR lineage were analyzed separately. Classic PRs were predominantly assigned to SeqC6 (93.9%), whereas the new PR lineage was assigned to SeqC35 (fig. S2F and Data S5), supporting the hypothesis of multiple independent origins of PRs.”

      (3) In Figure 3, the cluster StrC17 appears to be present in both Lepidoptera and Diptera, whereas when looking at Figure S3, it seems to be present only in Diptera. There may be an error somewhere.

      We thank the reviewer for identifying this inconsistency. After rechecking the data, we confirmed that StrC17 contains 47 ORs, of which 46 are from Diptera and one is from Lepidoptera. We have corrected Figure S3. This correction does not affect the main conclusions regarding the long-IL3 ORs, because all downstream analyses were conducted using the correct StrC17 composition.

      (4) The phylogeny presented in Figure S4 does not help clarify the phylogenetic context of the Orco study. The main reason is that it is not correctly rooted: the ORs of M. rhabei should have been used as an outgroup, or the two GRs of D. melanogaster that are present (perhaps by mistake) in this analysis and that clearly diverge strongly from the rest of the sequences. In any case, this phylogenetic analysis does not seem to add much compared with what was presented in the article by Thoma et al. in 2018.

      We thank the reviewer for pointing this out. The original purpose of the tree was to show that early-diverging insect ORs lack clear orthologous relationships, making it difficult to reconstruct ancestral OR states before and after Orco emergence. However, because the figure contributes little to the core conclusions and could introduce confusion, we have removed it from the revised manuscript. The revised text instead directly cites previous phylogenetic studies for the early OR/Orco evolutionary background.

      (5) In Figure 5A, many VOCs appear unclassified when using molecular descriptors. Have you tried classifying VOCs using molecular fingerprints instead of selected molecular descriptors? And can you explain how those 32 descriptors were chosen? Furthermore, it is surprising not to see a "terpenoids" category given the importance of these molecules in insect chemical ecology. One would expect them to cluster together in a chemical space based on molecular descriptors; where are they in Figure 5A?

      We thank the reviewer for this question. We first clarify that unclassified in Figure 5A does not mean that these VOCs were excluded from the analysis. It only means that they were not assigned to one of the nine predefined major functional-group categories. All VOCs were included in molecular descriptor calculation, chemical-space visualization, and downstream docking.

      The 32 molecular descriptors were not arbitrarily selected by us. They come from the optimized odorant metric proposed by Haddad et al. 2008. That study began with 1,664 Dragon molecular descriptors, represented odorants as multidimensional physicochemical vectors, and selected 32 descriptors that best explained similarity in odor-induced neural responses across multiple published datasets. Haddad et al. further showed that this optimized descriptor set performed well across different animals, recording methods, and levels of olfactory-system organization. We therefore adopted these 32 descriptors to represent odor physicochemical space.

      We agree that molecular fingerprints are another useful representation, especially for capturing substructures and scaffold information. We chose molecular descriptors rather than fingerprints because the purpose of Figure 5A is to visualize VOC distribution in a continuous physicochemical odor space, and because the Haddad descriptor set was optimized using neural response data. We have added this rationale to the revised manuscript.

      We also agree that terpenoids are important in insect chemical ecology. In the original classification, VOC categories were based mainly on functional groups, whereas terpenoids are defined by biosynthetic origin and carbon skeleton rather than by a single functional group. Therefore, terpenoid VOCs are distributed across several functional-group categories, such as terpene alcohols, terpene aldehydes, and terpene ketones, and are not expected to form one independent cluster in a two-dimensional odour space.

      Following the reviewer’s suggestion, we added an extra annotation for common terpenoid-related compounds and mapped them onto the odor space shown in Figure 5A. The result shows that terpenoids are distributed across multiple regions and functional-group categories rather than forming a single compact cluster (Fig. S5 c). We therefore retain the major functional-group classification for the main analyses and add terpenoids as an additional annotation, while explaining why they are not treated as a main category parallel to alcohols, aldehydes, ketones, and esters.

      In addition, because the downstream analyses in this study rely on large-scale docking predictions, and because docking itself contains inherent uncertainty, we chose to use relatively basic and chemically explicit major functional-group categories. This classification reduces interpretive instability that could arise from excessive subdivision of VOC classes. By contrast, treating terpenoids as an independent primary category would group together molecules with substantially different functional-group properties, thereby increasing the complexity of interpreting docking-derived binding profiles. Therefore, in the revised manuscript, we retain the major functional groups as the core VOC classification scheme and add terpenoids as an additional annotation.

      The specific revision in the revised manuscript is as follows (lines 268-271): "Although terpenoids are important in insect chemical ecology, they are distributed across multiple regions of odor space. To reduce the docking instability that could result from excessive VOC subdivision, we did not treat terpenoids as a category parallel to the nine basic functional-group categories."

      (6) The methodology used to calculate hit rates with respect to functional data on known ORs is not sufficiently explained. The result could be shown for the different species (drosophila, mosquito, butterfly). In addition, ORs from the locust should also have been used for this comparison.

      We thank the reviewer for this suggestion. In the revised Methods, we now describe the workflow in detail. We first compiled receptor-odorant combinations with clear response relationships from published functional experiments. For each OR-VOC pair, we performed docking using the same pipeline as in the global analysis and extracted the corresponding docking score. We then labeled each pair as response or non-response according to the original experimental data. Next, we grouped pairs into docking-score intervals and calculated, within each interval, the proportion of experimentally positive pairs, which we define as the hit rate. To reduce fluctuation from random sampling, each score interval was sampled three times. Finally, following the hit-rate modeling strategy of Lyu et al. 2019, we fitted the relationship between docking score and hit rate with a Bayesian curve-fitting approach and used the posterior expectation of the parameters to draw the hit-rate curve. The threshold was determined from the docking score at which the hit rate reached a plateau.

      Following the reviewer’s request, we calculated hit-rate curves separately for Drosophila melanogaster, Anopheles gambiae, Lepidoptera, and Locusta migratoria OR data (Fig. S5). Drosophila, Anopheles, and Lepidoptera showed similar trends: the hit-rate plateau was approximately 20%, and the corresponding docking-score threshold was around -8 kcal/mol. This is consistent with the original combined analysis and indicates that the threshold was not driven by a single species dataset. The locust dataset behaved differently. Its experimental positive rate was low, approximately 5.2%, and most reported locust ORs are narrowly tuned. In such a highly imbalanced dataset, docking score was less able to enrich experimental positives to the same level as in the other datasets; most score intervals had hit rates below 10%. When the locust data were combined with Drosophila, Anopheles, and Lepidoptera, the overall plateau decreased to about 18%, and the corresponding threshold shifted to about -11 kcal/mol.

      We consider this difference informative because it shows that OR tuning properties and experimental positive rates can affect the relationship between docking score and hit rate. The locust result suggests that stricter thresholds may be required for systems dominated by narrowly tuned ORs and low positive rates. However, our goal is to establish a unified empirical threshold for large-scale repertoire comparison, not to optimize a separate threshold for every lineage. We therefore retain -8 kcal/mol from the Drosophila, Anopheles, and Lepidoptera datasets as the main threshold, present the locust analysis as a sensitivity test, and note that this threshold may have lower positive-enrichment ability in narrowly tuned OR systems.

      The specific revision in the revised manuscript is as follows (lines 281-285): "In addition, our singletaxon hit-rate sensitivity analysis showed similar trends for Drosophila, Anopheles, and moths. Because the locust OR repertoire had an extremely low positive rate, we did not include the locust functional data in the final threshold assessment (see Supplementary Text S2 for details)."

      Methods section (lines 748-755): "We first compiled published functional datasets containing explicit OR-VOC response relationships. These datasets included ORs from Drosophila melanogaster, Anopheles gambiae, Helicoverpa armigera, Spodoptera littoralis, and Locusta migratoria. For each dataset, OR-VOC combinations were assigned response or non-response labels according to the original experimental results. Each OR-VOC pair was then docked using the same structural modeling, docking, and scoring workflow used in the global OR-VOC analysis. Pairs were grouped into docking-score intervals, and the hit rate for each interval was defined as the proportion of experimentally positive pairs among all pairs in that interval. Each score interval was sampled three times."

      (7) It is not clear how the ROC curves were constructed; they appear to be drawn with very few points. The methodology needs to be explained in more detail here, since this is an important point.

      We thank the reviewer for this suggestion. We have added a detailed description of ROC construction. The ROC analysis was designed to evaluate whether the dFunC labels are directionally consistent with available experimental functional data.

      Specifically, a positive dFunC label for a functional-group VOC indicates that ORs in that docking-derived community are predicted, relative to the all-insect OR background, to bind more VOCs from that functional group. A negative label indicates a lower predicted binding tendency. We mapped Drosophila ORs to the dFunC classification and assigned each OR a positive or negative label for each functional group according to its dFunC. We then used the Drosophila experimental functional matrix to count measured responses of each OR to different functional-group VOCs. For each functional group, we plotted ROC curves and calculated AUC values by comparing the dFunC-predicted labels with experimental response counts. We agree that the curves are based on limited data points for some functional groups. This is because comprehensive OR-VOC functional matrices are still sparse, and many insect ORs have been deorphanized using limited odor panels that do not cover many functional groups.

      The specific revision in the revised manuscript is as follows (lines 794-798): "Drosophila melanogaster ORs were mapped to dFunCs to evaluate the consistency between dFunC labels and experimental response data. For each VOC functional group, ORs were assigned positive or negative labels according to their dFunC labels, and these labels were compared with response counts from the Drosophila functional matrix to generate ROC curves and AUC values."

      (8) What the BBI represents is not very clear, even though the calculation method is presented in detail. This should be clarified in the main text and/or in the figure legends.

      We thank the reviewer for this suggestion. We now define BBI more explicitly in the main text. BBI, or binding breadth index, is a relative repertoire-level metric calculated from docking-derived dFunC labels. A dFunC label indicates whether ORs in a docking-derived binding-profile community show higher or lower predicted binding tendency toward a functional-group VOC category relative to the all-insect OR background. When we calculate BBI for a set of ORs, we summarize the predicted binding tendencies of the dFunCs to which those ORs belong, thereby estimating the relative potential binding breadth of that OR repertoire for the VOC category.

      Thus, BBI should be interpreted only as a relative docking-derived metric. When comparing two OR groups for the same VOC functional group, a higher BBI means that the group is predicted, under our docking-derived framework, to have broader potential binding breadth. It does not directly represent experimentally confirmed olfactory breadth, odor perception, or behavioral function. We have revised the terminology accordingly to docking-derived potential binding breadth.

      The following revision has been added to the revised manuscript (lines 319-321): "BBI is a docking-based, repertoire-level relative metric used to summarize the potential binding breadth of a set of ORs toward a given functional-group VOC category."

      (9) Insects with saprophagous larvae in your dataset are almost exclusively flies. Therefore, the conclusions concerning the saprophagous diet could just as well result from phylogenetic constraints rather than from an adaptation to the diet.

      We thank the reviewer for pointing this out. We agree that, because the currently available saprophagous larval species are mainly concentrated in Diptera, this result may be influenced by lineage effects. We have therefore toned down the interpretation and present it as a candidate association. Future inclusion of more saprophagous species from non-dipteran lineages will allow this relationship to be tested more rigorously.

      The following revision has been added to the revised manuscript (lines 493-496): "This interpretation should be treated with caution, as the saprophagous larval species currently available in our dataset are mainly concentrated in Diptera and may therefore reflect lineage effects. Future inclusion of saprophagous species from broader insect lineages will help test this association more rigorously."

      Reviewer #2 (Recommendations for the authors):

      Main feedback

      (1) This study presents an interesting approach to unify the exceptionally large OR multigene family in a single macroevolutionary framework. While this is a non-trivial task, given the low sequence similarity of ORs, the inferences broadly overlap with previous findings, suggesting that the approach works. Unfortunately, the line between previously identified patterns of OR biology and evolution and new insights derived from this study is somewhat blurry throughout the text at the moment and needs to be sharpened. As it reads, the manuscript is prone to overselling the results.

      A few examples:

      (a) Lines 37-41 - contrary to the claim in this sentence, previous work has described the principles of OR function, structure, and evolution to an extent that would warrant concluding that the fundamental logic of insect olfaction is actually comparatively well understood.

      We thank the reviewer for this comment. We agree that previous studies have already established important principles of insect OR biology, including OR-Orco complex structure, ion-channel function, ligand recognition, and rapid OR evolution. Our intention was not to imply that the basic logic of insect olfaction remains unknown, but to emphasize that an insect-class-scale framework integrating OR sequence divergence, structural variation, and docking-derived binding-profile diversity has been lacking. We have revised the Introduction to clarify this point and to better acknowledge the contribution of previous studies.

      We therefore revised this sentence in the revised manuscript as follows (lines 38-42): "However, a classwide integrative framework linking OR sequence divergence, structural variation, and docking-derived binding-profile diversity is still lacking. Such a framework is needed to distinguish broadly conserved patterns from lineage-specific features and to place species-level OR diversification in a broader macroevolutionary context."

      (b) Lines 47-48 - work in several systems has emphasized the "evolutionary interplay between receptor family diversification and macroecological adaptation", both in species with highly specialized ecologies (plant host specialists, slave making ants -> these are even cited in the manuscript) and in a broader macroevolutionary context (e.g. bee diet breadth, Singh et al. 2025), which does not really justify the claim that OR evolution with respect to ecology is a profound enigma. The present study indeed presents the first one combining data throughout the entire insect tree of life, which is an impressive feat. The fact that inferences drawn from smaller datasets with smaller phylogenetic breadth are supported is a useful finding. I recommend emphasizing this.

      We agree that the relationship between OR repertoire evolution and ecological adaptation has been investigated in multiple insect systems and should not be described as an unresolved “profound enigma.” The novelty of our study is to test, integrate, and extend these previously reported patterns across 115 insect species under a unified sequence-structure-docking profile framework. We have revised the text to emphasize this class-wide extension rather than overstating the unknowns in the field. The following revision has been added to the revised manuscript (lines 46-47): "Although several studies have linked OR repertoire evolution to ecological adaptation in specific insect lineages, a unified class-wide comparison remains lacking."

      (c) Lines 347 - 348 - what is meant by the statement that the manuscript 'revealed the mechanism of insect olfactory perception underlying macroenvironmental influences' is unclear to me. What mechanism is referred to here? What do the authors mean by 'macroenvironmental influences'? It reads as if the authors claim to have shown how olfactory perception works in insects, which is not an accurate statement.

      We thank the reviewer for pointing out that this statement was unclear. We did not intend to claim that our study reveals the mechanism of insect olfactory perception itself. Rather, our results identify class-scale associations between OR repertoire variation, ecological traits, and docking-derived binding profile composition. We have revised the Discussion.

      In the revised manuscript, this sentence has been changed to the following more accurate wording (lines 399-402): "Our results reveal class-level associations between OR repertoire variation and ecological traits and identify differences in docking-derived binding-profile composition among extant insect lineages that may be relevant to long-term macroevolutionary transitions."

      (d) Lines 348-350: I do not agree with the statement that the authors "describe the evolutionary process through which Orco emerged". The data that the beta-sheet in EL2 impedes ligand binding in non-orco ORs is compelling and suggests an intriguing contribution to why Orco lack ligand-binding function, but is not sufficient to comprehensively describe the evolutionary process of Orco evolution from ancestral ORs. This would require more sequences of early-branching insect species. Accordingly, this statement should be toned down.

      We agree that the original wording overstated the extent to which our data describe Orco emergence. Our results identify structural features, including the EL2 beta-sheet and specialized binding-pocket properties, that may have contributed to the loss of ligand-binding function during Orco specialization. However, a complete reconstruction of Orco origin will require broader sampling of early-diverging insect lineages and more reliable ancestral-state reconstruction. We have revised the relevant statements accordingly.

      We therefore toned down this statement in the revised manuscript and described it more accurately as follows (lines 402-405): "We further examined Orco specialization within the insect 'conserved chassisdiversified sensors' model and identified structural features, including the EL2 beta-sheet and specialized binding-pocket properties, that may have contributed to the loss of ligand-binding function during Orco evolution." We also revised another sentence as follows (lines 522-525): "Overall, we systematically explored the relationships among sequence, structure, and docking-derived function in insect ORs, identified structural features potentially associated with Orco specialization, and provided a framework for testing how OR repertoire evolution may relate to macroenvironmental and lifestyle variation."

      It is understandable that the authors want to emphasize the value of their work, but this can't be achieved by inaccurately portraying the state of the field. The framing should be adjusted throughout the manuscript.

      We thank the reviewer for this important suggestion. We agree that the manuscript should emphasize its contribution without overstating gaps in the field. We have therefore revised the framing throughout the manuscript to better acknowledge previous advances in insect OR structure, function, evolution, and ecological adaptation. The revised text presents our main contribution as an insect-classscale integration and extension of these findings using a unified sequence, structure, and docking-derived binding-profile framework.

      (2) It is unclear how complete the odorant receptor sets are, because the annotation pipeline used is not benchmarked. The high similarity of ORs often located in clusters of up to 50 genes on the genome represents a major obstacle for purely automated annotation, leading to artifacts such as the erroneous joining of exons across genes. The lack of manual verification in this study, combined with stringent filtering of annotations that do not meet structural criteria, does likely lead to a set of sequences that present a correct fold, but at the same time, the approach may discard misannotated genes and cannot distinguish between correct gene models and artifacts where exons of more than one gene could have been merged. Please benchmark the pipeline by comparing its output to a gold standard and frequently vetted complete OR set, such as that of Robertson and Wanner 2006 or similar. That would allow assessing the presented work better and reveal how complete the OR counts are. This is important since the authors use OR counts to derive conclusions about the evolutionary dynamics of the gene family.

      We thank the reviewer for raising this concern. Because this point overlaps with the benchmark issue raised in the public review, we provide the detailed benchmark results in our response above and in the revised supplementary text. Briefly, we compared the annotation and structural-filtering pipeline with curated OR repertoires from Drosophila melanogaster, Anopheles gambiae, and Bombyx mori. The results show that the workflow recovers most known ORs, while the structural filter removes fragmented or structurally incomplete models. We also added sensitivity analyses using published OR counts where available, which supported the main OR-count ecological associations.

      (3) The in silico functional docking analysis is intriguing and bears the potential to produce testable hypotheses on the functional evolution of ORs. Docking can produce a wide range of results, including artifacts, and thus should be benchmarked with existing datasets derived from functional experiments (there are several large datasets available). While the authors have used previous functional data to constrain their model, they should benchmark it by running it on a set of ORs with known ligand-binding profiles to convincingly show that their predictions are an accurate assessment of the actual functional properties of ORs.

      We thank the reviewer for this suggestion. We agree that docking-derived results should be evaluated using ORs with known ligand-response data. In this study, however, the key prediction used for downstream analyses is the dFunC functional label derived from each OR’s overall docking-derived binding profile, rather than the exact ligand assignment of each OR-VOC pair. Therefore, we evaluated the dFunC labels using the Drosophila melanogaster experimental OR response matrix.

      Drosophila ORs were mapped to dFunC communities, and for each VOC functional group, ORs were assigned positive or negative labels according to the predicted binding tendency of their dFunC. These labels were compared with experimentally measured response counts from the Drosophila OR functional matrix. ROC curves and AUC values were calculated for each VOC functional group. Most functional-group labels achieved AUC values above 0.75, indicating that the dFunC classification is broadly consistent with the Drosophila experimental response matrix.

      We did not use individual OR-VOC pair accuracy as the primary validation criterion because the hit rate analysis showed that the experimental positive rate in the enriched docking-score range is approximately 23%. This indicates that docking is more appropriate for evaluating dFunC-level binding profile tendencies than for exact ligand assignment. We have revised the manuscript to clarify this validation logic. The functional labels of dFunCs should be further tested with experimental response data from more insect species and can be updated as prediction methods improve.

      The following details have been added to the revised manuscript (lines 794-798): " Drosophila melanogaster ORs were mapped to dFunCs to evaluate the consistency between dFunC labels and experimental response data. For each VOC functional group, ORs were assigned positive or negative labels according to their dFunC labels, and these labels were compared with response counts from the Drosophila functional matrix to generate ROC curves and AUC values."

      (4) The authors conclude that holometabolous and non-holometabolous insects exhibit distinct OR differentiation patterns, stating that sequence diversity "increases progressively with increasing degree of insect order divergence, suggesting a gradual accumulation of sequence variation during insect evolution" (Lines 332-333). If the pattern is indeed gradual, is the phylogenetic splitting of holometabolous vs. non-holometabolous insects not arbitrary? Couldn't it be split at any point along the phylogeny and lead to a similar difference?

      We thank the reviewer for pointing out this important logic issue. We agree that, if OR sequence variation accumulates gradually with phylogenetic distance, the difference between holometabolous and non-holometabolous insects should not be interpreted as two completely separate OR differentiation modes.

      Our result is better understood as a comparison of extant deep insect lineages. In this comparison, holometabolous insects show higher OR sequence-community diversity than non-holometabolous insects. This pattern is still informative because it suggests that, during long-term evolution, OR repertoires in holometabolous lineages have accumulated broader sequence-community diversity. This may be related to lineage-specific OR duplication and loss, ecological diversification, life-history differences, and differences in repertoire size.

      We have revised the manuscript accordingly. We no longer describe the result as two distinct OR differentiation patterns. Instead, we describe it as a difference in OR sequence-community diversity between extant holometabolous and non-holometabolous lineages.

      The following revision has been added to the revised manuscript (lines 370-371): "Additionally, extant holometabolous insects showed higher OR sequence-community diversity than non-holometabolous

      insects."

      (5) Further, if I understand correctly, the assessment of the differentiation patterns seems to rest on total counts of ORs in 'SeqCs', quantified in Figure S9. These counts, however, are not normalized and thus seem influenced by the total number of sequenced genomes belonging to a specific clade in the dataset. These have a heavy bias towards holometabolous species. If this is true, the analyses should be repeated based on normalized counts. OR is this assessment based on the number of 'SeqCs' across the phylogeny? If yes, does this correlate with OR numbers? Please clarify.

      We thank the reviewer for pointing this out. We agree that unnormalized SeqC counts can be affected by unequal sampling, because holometabolous insects include more sampled species and more ORs in our dataset.

      To address this issue, we added normalized analyses. First, we calculated SeqC diversity at the species level (Fig. S9a). Second, because OR number was significantly correlated with SeqC count (pGLS p-value = 0.001035), we performed equal-OR random sampling (Fig. S9 b). Specifically, we randomly sampled the same number of ORs from holometabolous and non-holometabolous insects, 100 ORs per group, calculated how many SeqCs were covered by the sampled ORs, and repeated this procedure 100 times to obtain the distribution of SeqC diversity under equal OR sampling depth. These analyses showed that holometabolous insects still exhibited higher OR sequence-community diversity after controlling for OR sampling depth. We have revised the manuscript to clarify the calculation and now interpret the result as a normalized difference in sequence-community diversity, rather than as an uncorrected difference in total SeqC number.

      We added the following normalized analysis to the revised manuscript (lines 371-379): "Because total SeqC counts may be affected by unequal species sampling and OR numbers, we added normalized analyses. First, we calculated SeqC numbers at the species level and found that holometabolous insects contained more SeqCs on average. Second, because OR number was significantly correlated with SeqC count (P=0.001035), we performed equal-OR random sampling between holometabolous and nonholometabolous insects. Under the same OR sampling depth, holometabolous insects still covered more SeqCs, indicating higher OR sequence-community diversity after controlling for OR number. Therefore, this result supports a genuine difference in SeqC diversity rather than a difference driven by unequal species sampling or OR number. "

      (6) Phylogenetic differences in OR evolutionary patterns have been described previously for Paleoptera compared to Neoptera. Are the authors simply picking up on these patterns? Is the holometabolous vs non-holometabolous hypothesis supported when only Neoptera are taken into account?

      We thank the reviewer for raising this point. Differences in OR evolutionary patterns between Paleoptera and Neoptera have already been described in previous studies, so it was necessary to test whether our pattern simply reflected this older phylogenetic division.

      Following the suggestion, we restricted the comparison to Neoptera and compared Holometabola with non-holometabolous Neoptera. After excluding Paleoptera and earlier-diverging lineages, Holometabola still showed higher OR sequence-community diversity (Fig. S10). Additional normalized analyses indicate that this trend is not explained only by unequal species sampling or OR counts (Fig S9a-b).

      Therefore, our result is not simply a rediscovery of the known Paleoptera versus Neoptera difference. Instead, within Neoptera we observe an additional difference in SeqC diversity between Holometabola and non-holometabolous Neoptera.

      We added the following comparison of neopteran insects to the revised manuscript (lines 379-382): "The higher SeqC diversity observed in Holometabola was retained when the comparison was restricted to Neoptera, indicating that this pattern is not simply a restatement of previously reported Paleoptera-Neoptera differences."

      Minor comments

      (1) Line 10: "all the insect orders were found to contain fully functional OR repertoires". It is unclear what the authors mean by this. What is a fully functional OR repertoire? Is it the number of receptors? The number of functional receptors? Given this lack of clarity and the fact that the manuscript does not present functional data on Or binding profiles, this part should be altered.

      We have revised this sentence. The intended meaning was that all sampled insect orders include ORs assigned to all six docking-derived functional communities.

      (2) Lines 17-18: I am not convinced the presented data explain the adaptive relationship between insect olfactory potential and diverse ecological environments. It suggests that few aspects of insect ecology (larval diet and terrestriality) correlate with receptor number and predicted ligand binding breadth. This should be toned down.

      We have toned down the statement. The sentence has been revised as follows: " Our findings provide a class-level framework for investigating insect OR evolution and generate testable hypotheses about how receptor repertoire diversification may relate to ecological adaptation."

      (3) Line 44: the low amino acid sequence identity among ORs has been described before. Please add appropriate references.

      We have added the requested references.

      (4) Lines 52 - 55: The 1:3 stoichiometry is not yet confirmed in vivo - indeed, other work suggests a 2:2 stoichiometry as a possibility. Please adjust this sentence to reflect this uncertainty.

      We have revised the sentence to reflect this uncertainty (lines 51-56): Unlike vertebrate G protein-coupled ORs, insect ORs function as ligand-gated ion channels by assembling with the conserved co-receptor Orco, with recent cryo-EM structures supporting a 1:3 OR–Orco heterotetrameric model. However, the exact in vivo stoichiometry remains to be fully resolved, and alternative arrangements such as 2:2 may also be possible.

      (5) Lines 94-95: Please cite studies that previously annotated ORs.

      We have added the relevant citations.

      (6) Lines 99-100: It is well established that hymenopterans have the largest sets of ORs among insects. Please add appropriate references.

      We have added the relevant citations.

      (7) Lines 101-102: It has been shown before that Odonata have very few ORs. However, there are other insects with even smaller sets of ORs. Please add appropriate references.

      We have clarified that the statement refers only to our intact OR dataset. The sentence has been revised as follows: "In our OR dataset, Odonata had the fewest intact ORs, with a mean of 4 (N = 3)."

      (8) Lines 109-111: This sentence appears to be illogical. Please revise.

      We re-evaluated analyses that depend on OR counts. Because the association between genome size and OR count was sensitive to annotation completeness, structural filtering, and a few extreme species, we removed this result and the related discussion from the revised manuscript.

      (9) Line 113: What is ecological behavior? Do you mean ecology?

      We have replaced ecological behavior with ecological traits.

      (10) Lines 148 - 150: This has been described previously. Please add appropriate references.

      We have added the appropriate citation.

      (11) Lines 161-162: This has been previously described. Please add appropriate references.

      We have added the appropriate citation.

      (12) Lines 199-201: Could it be that a lack of orthologs is a result of sampling bias? Only very few genomes of early-branching lineages have been sequenced.

      We thank the reviewer for this helpful comment. We agree that the apparent lack of direct orthologs among early-diverging insect ORs may be influenced by limited genome sampling from these lineages. We have revised the manuscript to clarify that, with the currently available data, we cannot reliably determine direct orthologous relationships among early-diverging insect ORs or reconstruct ancestral OR states. We now state that broader sampling of high-quality genomes and transcriptomes from early-diverging insect lineages will be needed to test this question more rigorously.

      The revised text is as follows (lines 218-220): "With the currently available early-diverging insect genomes, we could not reliably identify direct orthologous relationships among early ORs, which limits the reconstruction of ancestral OR states."

      (13) Line 204-205: Do you mean TdomOR1-8 may be ancestral to Orco? This is not possible, given that Orco and TdomOR1-8 co-exist in one genome. But previous phylogenetic analyses of multiple silverfish genomes suggest that Orco is closely related to TdomOR1-8 and similar receptors in other species. Please revise.

      We thank the reviewer for pointing this out. We did not mean that TdomOR1-8 are the ancestors of TdomOrco. Because currently available OR data from early-diverging insects are limited, direct orthologous relationships cannot be reliably identified and the ancestral OR state of Zygentoma cannot be reconstructed with confidence. Therefore, we can only search among extant zygentoman ORs for comparators that are closely related to the Orco clade.

      The current gene tree shows that TdomOR1-8 are closely related to TdomOrco. Thus, they can serve as structural references for comparing Orco-related receptors and for understanding structural changes that may have been involved in Orco specialization.

      In the revised manuscript, we have changed the relevant sentence to (lines 223-225): “Therefore, we speculate that Orco recruitment in the ancestral zygentoman lineage may have been accompanied by expansion of a set of candidate ORs, and that TdomOR1-8 represent extant retained members of this receptor set.”

      (14) Lines 387-388, this was already known. Please cite accordingly

      We have added the appropriate citation.

      (15) Line 389: Please include a reference for the finding that early ORs have a homologous tetrameric configuration.

      We have added the appropriate citation.

      (16) Lines 403-407: This hypothesis has been previously put forward by del Marmol and colleagues. Add reference.

      We have added the appropriate citation.

      Author response table 1.

      Number of OR genes in previously annotated species.

      Author response table 2.

      ORs in this study that did not show one-to-one correspondence with published Bombyx mori ORs.

      Author response image 1.

      Structurally erroneous ORs in Propsilocerus akamusi.

      Author response image 2.

      Sensitivity analysis of OR counts. (a-b) Correlations between updated species-level OR counts and genome size.

      Author response image 3.

      OR protein structures removed by structural filtering from the automated annotations of Drosophila melanogaster and Anopheles gambiae.

      Author response image 4.

      Normalized analysis of SeqC diversity between holometabolous and non-holometabolous insects within Neoptera at the species level (a) and OR-count level (b). Welch's t-test was used. ***p < 0.001.

      Reference

      Bobkov YV, Walker Iii WB, Cattaneo AM. 2021. Altered functional properties of the codling moth Orco mutagenized in the intracellular loop-3. Sci Rep 11: 3893.

      Carey AF, Wang G, Su CY, Zwiebel LJ, Carlson JR. 2010. Odorant reception in the malaria mosquito Anopheles gambiae. Nature 464: 66-71.

      Chang H, Unni AP, Tom MT, Cao Q, Liu Y, Wang G, Llorca LC, Brase S, Bucks S, Weniger K et al. 2023. Odorant detection in a locust exhibits unusually low redundancy. Curr Biol 33: 5427-5438 e5425.

      Dekel A, Pitts RJ, Yakir E, Bohbot JD. 2016. Evolutionarily conserved odorant receptor function questions ecological context of octenol role in mosquitoes. Sci Rep 6: 37330.

      Frunze O, Lee D, Lee S, Kwon HW. 2024. A single mutation in the mosquito (Aedes aegypti) olfactory receptor 8 causes loss of function to 1-octen-3-ol. Insect Biochem Mol Biol 167: 104069.

      Guo S, Kim J. 2007. Molecular evolution of Drosophila odorant receptor genes. Mol Biol Evol 24: 11981207.

      Hall DR, Beevor PS, Cork A, Nesbitt BF, Vale GA. 1984. 1-Octen-3-ol. International Journal of Tropical Insect Science 5: 335-339.

      Kline DL, Allan SA, Bernier UR, Welch CH. 2007. Evaluation of the enantiomers of 1-octen-3-ol and 1octyn-3-ol as attractants for mosquitoes associated with a freshwater swamp in Florida, U.S.A. Med Vet Entomol 21: 323-331.

      Kreher SA, Mathew D, Kim J, Carlson JR. 2008. Translation of sensory input into behavioral output via an olfactory system. Neuron 59: 110-124.

      Liu WB, Li HM, Wang GR, Cao HQ, Wang B. 2023. Conserved Odorant Receptor, EcorOR4, Mediates Attraction of Mated Female Eupeodes corollae to 1-Octen-3-ol. J Agric Food Chem 71: 1837-1844.

      Lu T, Qiu YT, Wang G, Kwon JY, Rutzler M, Kwon HW, Pitts RJ, van Loon JJ, Takken W, Carlson JR et al. 2007. Odor coding in the maxillary palp of the malaria vector mosquito Anopheles gambiae. Curr Biol 17: 1533-1544.

      Morawo T, Fadamiro H. 2016. Identification of Key Plant-Associated Volatiles Emitted by Heliothis virescens Larvae that Attract the Parasitoid, Microplitis croceipes: Implications for Parasitoid Perception of Odor Blends. J Chem Ecol 42: 1112-1121.

      Paddock KJ, Corcoran JA. 2025. Life-stage dependent behavior mimics chemosensory repertoire diversity in a belowground, specialist herbivore. G3 (Bethesda) 15.

      Pierce A, Pierce H, Borden J, Oehlschlager C. 1989. Production Dynamics of Cucujolide Pheromones and Identification of I-Octen-3-o1 as a New Aggregation Pheromone for Oryzaephilus surinamensis and O. mercator (Coleoptera: Cucujidae). Environmental Entomology 18: 747-755.

      Qiu L, Tao S, He H, Ding W, Li Y. 2018. Transcriptomics reveal the molecular underpinnings of chemosensory proteins in Chlorops oryzae. BMC Genomics 19: 890.

      Rondoni G, Roman A, Meslin C, Montagne N, Conti E, Jacquin-Joly E. 2021. Antennal Transcriptome Analysis and Identification of Candidate Chemosensory Genes of the Harlequin Ladybird Beetle, Harmonia axyridis (Pallas) (Coleoptera: Coccinellidae). Insects 12.

      Tanaka K, Uda Y, Ono Y, Nakagawa T, Suwa M, Yamaoka R, Touhara K. 2009. Highly selective tuning of a silkworm olfactory receptor to a key mulberry leaf volatile. Curr Biol 19: 881-890.

      Tian Z, Sun L, Li Y, Quan L, Zhang H, Yan W, Yue Q, Qiu G. 2018. Antennal transcriptome analysis of the chemosensory gene families in Carposina sasakii (Lepidoptera: Carposinidae). BMC Genomics 19: 544.

      Turner RM, Derryberry SL, Kumar BN, Brittain T, Zwiebel LJ, Newcomb RD, Christie DL. 2014. Mutational analysis of cysteine residues of the insect odorant co-receptor (Orco) from Drosophila melanogaster reveals differential effects on agonist- and odorant-tuning receptor-dependent activation. J Biol Chem 289: 31837-31845.

      Wang Q, Smid HM, Dicke M, Haverkamp A. 2024. The olfactory system of Pieris brassicae caterpillars: from receptors to glomeruli. Insect Sci 31: 469-488.

      Xu L, Jiang HB, Yu JL, Pan D, Tao Y, Lei Q, Chen Y, Liu Z, Wang JJ. 2023. Two odorant receptors regulate 1-octen-3-ol induced oviposition behavior in the oriental fruit fly. Commun Biol 6: 176.

      Xu P, Zhu F, Buss GK, Leal WS. 2015. 1-Octen-3-ol - the attractant that repels. F1000Res 4: 156.

      Xu Q, Wu Z, Zeng X, An X. 2020. Identification and Expression Profiling of Chemosensory Genes in Hermetia illucens via a Transcriptomic Analysis. Front Physiol 11: 720.

      Yan C, Sun X, Cao W, Li R, Zhao C, Sun Z, Liu W, Pan L. 2020. Identification and expression pattern of chemosensory genes in the transcriptome of Propsilocerus akamusi. PeerJ 8: e9584.

      Zhang S, Zhang Z, Wang H, Kong X. 2014. Antennal transcriptome analysis and comparison of olfactory genes in two sympatric defoliators, Dendrolimus houi and Dendrolimus kikuchii (Lepidoptera: Lasiocampidae). Insect Biochem Mol Biol 52: 69-81.

      Zhang Y, Wang B, Zhou Y, Liao M, Sheng C, Cao H, Gao Q. 2023. Identification and characterization of odorant receptors in Plutella xylostella antenna response to 2,3-dimethyl-6-(1-hydroxy)-pyrazine. Pestic Biochem Physiol 194: 105523.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer # 1 (Public review):

      (1) Structure and Presentation of Results

      • I recommend reordering the visual-cue experiments to progress from simpler conditions (no cues) to more complex ones (cue-conflict). This would improve narrative logic and accessibility for non-specialist readers. The authors have chosen not to implement this suggestion, which I respect, but my recommendation stands.

      Thank you for this suggestion. We understand your point that presenting the experiments from simpler to more complex conditions may seem more intuitive. However, we have kept the original order because it better reflects the logic of the study itself. Our work first asked whether fall armyworms, like the Bogong moth, use a magnetic compass that is integrated with visual cues. Only after establishing this behavioral feature did we go on to test whether visual cues are required to maintain magnetic orientation. To make this reasoning clearer to readers, we have explicitly stated in the Introduction that magnetic orientation in the Bogong moth depends on the integration of visual cues, which provides clearer context for the experimental design.

      (2) Ecological Interpretation

      • The authors should expand their discussion on how the highly simplified, static cue setup translates to natural migratory conditions, where landmarks are dynamic, transient, or absent. Specifically, further consideration is needed on how the compass might function when landmarks shift position, become obscured, or are replaced by celestial cues. Additionally, the discussion would benefit from a more consolidated section with concrete suggestions for future experiments involving transient, multiple, or more naturalistic visual cues. This point was addressed partially in one paragraph of the Discussion, which reads as follows:

      "In nature, they are likely to encounter a range of luminance-gradient visual cues, including relatively stable celestial cues as well as transient or shifting local features encountered en route. Although such natural cues differ from our simplified laboratory stimulus, they may represent intermittently sampled visual inputs that can be optimally integrated with magnetic information, with the congruency between visual and magnetic cues likely playing a key role in maintaining a stable compass response. Whether the cues are static or changing, brief periods without them may still allow the subsequent recovery of a stable long-distance orientation strategy. Determining which types of natural visual cues support the magnetic-visual compass, and how they interact with magnetic information, including how their momentary alignment or angular relationship is integrated and how such visual cue-magnetic field interactions may require time to influence orientation, together with elucidating the genetic and ecological bases of multimodal orientation, will be important objectives for future research." While this paragraph is informative, the wording remains lengthy, somewhat unclear, and vague. Shorter, clearer statements would improve readability and impact. For example:

      • How could moths maintain direction during periods when only the magnetic field is present and visual landmarks are absent?

      • Could celestial cues (e.g., stars) compensate, and what happens if these are also obscured?

      • What role does saliency play when multiple visual landmarks are present simultaneously?

      • How might a complex skyline without salient landmarks affect orientation?

      Including simple, concise sentences that pose concrete open questions and suggest experimental designs would strengthen the discussion without creating space issues. In my view, a comprehensive discussion of how the simplified, static cue setup relates to natural migratory conditions-where landmarks are dynamic, transient, or absent-would add significant value to the paper.

      Thank you for this constructive and insightful comment. You correctly point out that our articulation of the ecological relevance of the simplified, static cue setup was not sufficiently clear. We also agree that the original wording in the Discussion remained overly general. In the revised Discussion, we updated the manuscript to incorporate recently published findings on the use of light–dark gradients for orientation in fall armyworms. However, we explicitly note that it remains unclear whether fall armyworms can exploit naturally occurring luminance gradients, such as those generated by the moon, for orientation under natural conditions. We further emphasize that during natural migration the visual environment is dynamic, with celestial cues available intermittently and local visual features changing continuously during flight. In this context, we outline several key unresolved questions, including whether celestial cues can compensate when local landmarks are absent; how multiple visual cues are weighted and integrated with geomagnetic information; how transient visual cues (like moving clouds or changing illumination) influence orientation; and how luminance gradients that are common in natural nocturnal environments interact with the geomagnetic field to support orientation. For each of these issues, we briefly suggest experimental approaches to guide future research.

      (3) Methodological Details and Reproducibility

      • The lack of luminance level measurements should be explicitly highlighted.

      Thank you for your helpful suggestion. You are right that luminance level is an important experimental parameter. We have stated this information in the Methods section under Behavioral apparatus: “The ambient light level in the experimental environment was measured to be below 1 lux using a Testo 540 lux meter (Testo SE & Co. KGaA, Titisee-Neustadt, Germany). Further work is still required to compare the illuminance used in this study with that under natural conditions, which are inherently variable.” This point is also clarified in the legend of Figure S3 in the supplementary material.

      • The authors chose not to adjust figure legends by replacing "magnetic South" with "magnetic North." While I believe this would be more conventional and preferable, this is ultimately a minor stylistic issue.

      Thank you very much for your suggestion. We understand your point and agree that using “magnetic North” would be more conventional. However, because our experiments focus on the orientation behavior of the autumn population, magnetic South is aligned with the landmark direction representing the potential migratory direction, which we believe makes the figures more intuitive for readers. We therefore consider this a minor stylistic issue.

      (4) Conceptual Framing and Discussion

      • Although the authors made a good attempt to explain the limitations of using an artificial visual cue, I believe there is room or a more explicit argument. For example, it could be stated clearly that this species is unlikely to encounter a situation in nature where a single, highly salient landmark coincides with its migratory direction. Therefore, how these findings translate to real migratory contexts remains an open question. A sentence or two making this point directly would strengthen the discussion.

      Thank you for your helpful suggestion. We now address this point explicitly in the Discussion, noting that fall armyworms are unlikely to experience a natural visual environment dominated by a single, static, and highly salient landmark coinciding with their migratory direction. Consequently, how these findings translate to real migratory contexts remains an open question.

      (5) Technical and Open-Science Points

      • Sharing the R code openly (e.g., via GitHub) should be seriously considered. The code does not need to be perfectly formatted, but making it available would be highly beneficial from an open-science perspective.

      Thank you for the suggestion. We agree that making code openly available is valuable from an open-science perspective. The MMRT script used in this study is Moore’s Modified Rayleigh Test, available from the original publication by Massy et al. (2021; https://doi.org/10.1098/rspb.2021.1805). In the previous version, we only cited this reference in the Materials and Methods section; we have now added a direct link to the script to improve clarity and accessibility. We have also provided a public link to the data-recording scripts used in the Flash Flight Simulator (https://doi.org/10.17632/6jkvpybswd.1). This repository additionally includes a map-based optical flow script that was not used in the present study but is shared for completeness.

      Reviewer #1 (Recommendations for the authors):

      • LL. 133-137 (end of paragraph starting with "The fall armyworm is a migratory crop pest native to the Americas"): Suggest splitting into shorter, clearer sentences. The limitations of this method could be better articulated here and elaborated in the Discussion.

      Thank you for this suggestion. We have revised this paragraph by splitting it into shorter, clearer sentences and by articulating the limitations of this method more explicitly. These limitations are further elaborated in the Discussion.

      • LL. 181-185 (end of paragraph starting with "To examine if fall armyworms integrate geomagnetic and visual cues for seasonal migratory orientation"): It would be helpful to state explicitly that season-specific headings have been confirmed in the lab using a flight simulator, but destination regions remain unknown without further tracking experiments.

      Thank you for this helpful suggestion. We have now clarified in the revised manuscript that season-specific orientation headings have been confirmed in the laboratory using a flight simulator, while the actual migratory destination regions remain unclear in the absence of tracking experiments.

      • LL. 230-234 (start of paragraph "Our previous research showed that fall armyworms reared under artificially simulated fall conditions…"): Clarify which migratory season is being referenced.

      Thank you for this helpful suggestion. We have clarified in the text that the migratory season referenced here is the autumn migratory season. In addition, we have added information in the Methods to specify the actual calendar season during which the insects were reared under the simulated conditions.

      • LL. 270-272 (middle of Fig. 2 caption): Suggest explicitly mentioning that for this population, the seasonally appropriate direction is southbound in autumn and northbound in spring, as this may not be clear to non-specialists.

      Thank you for this helpful suggestion. We have now explicitly stated the seasonally appropriate migratory directions for this population, indicating southbound migration in autumn and northbound migration in spring, to improve clarity for non-specialist readers.

      • LL. 421 (middle of paragraph starting with "We also considered the limitations of the Rayleigh test…"): Add that the groups lacking visual cues exhibited "lower directedness as per lower vector length (r)" in addition to lower flight stability.

      Thank you for this helpful suggestion. We further note that the conclusions drawn from the flight stability analysis are consistent with those based on individual r-value analyses.

      • LL. 499-501 ("unlike some vertebrates that can rely solely on magnetic information (Mouritsen, 2018)"): This point is slightly downplayed. It should be emphasized that nearly all tested vertebrates and invertebrates (e.g., birds, mole rats, fish, frogs, and other insects) demonstrate a magnetic compass without requiring visual landmarks. Moths are the only tested invertebrates so far that show landmark-magnetic field dependency for their magnetic compass to be manifested in a behavioural orientation response in Flight Simulator.

      Thank you for this important comment. We agree that this point represents a key synthesis in the Discussion, as it concerns how our findings relate to, and differ from, magnetic orientation demonstrated in other animal groups. We have therefore expanded the Discussion to note that studies have shown that some animals can exhibit directional preferences in simplified visual environments solely in response to changes in the magnetic field, and we now cite representative examples from birds and mole rats. At the same time, we also acknowledge important methodological and phenotypic differences among taxa. In particular, moths’ magnetic orientation has been assessed using a flight simulator, a setup in which stable directional behavior must be actively maintained during continuous movement. This is an important difference from orientation assays in birds during take-off or in terrestrial mammals such as mole rats. Moreover, whether birds and other animals rely on visual input to detect or calibrate magnetic information under certain conditions remains an open question. We therefore emphasize here both the phenotypic differences observed across experimental systems and the methodological considerations.

      • LL. 560-565 (paragraph starting with "Our flight simulator system (Dreyer et al., 2021) …"): Suggest clarifying what the Flash flight simulator system is and how it differs from the Mouritsen-Frost flight simulator.

      Thank you for this suggestion. We have added a brief clarification of the Flash flight simulator and how it differs from the Mouritsen–Frost system.

      • LL. 605-608 ("Spectral measurements …"): Explicitly mention that total illuminance was not measured and that further work is required to compare the illuminance used with natural conditions which of course vary.

      Thank you for this helpful suggestion. We agree that total illuminance is an important factor. We have now added a statement noting that the ambient light level in the experimental environment was measured to be below 1 lux using a Testo 540 lux meter, and we further acknowledge that additional work is required to compare the illuminance used in this study with that under naturally variable conditions.

      • LL. 628-641 (end of paragraph starting with "Electromagnetic noise at the experimental site ... "): Explain why this matters for interpreting behavioural responses. Highlight that although conditions were somewhat magnetically noisy which based on the past work may disrupt magnetic compass as it was shown in birds (eg Engels et al. 2014 Nature), the observed magnetic response under certain conditions indicates that the magnetic sense remained functional when landmark and magnetic field were aligned. This way you can pre-empt this criticism of your magnetic conditions being not ideal and noise on the left handside of the spectrum measured (which is not uncommon).

      Thank you for this helpful suggestion. We have now cited Engels et al. (2014, Nature) in this section and expanded the text to explain why electromagnetic noise at the experimental site is relevant for interpreting the behavioural responses. We also clarify the rationale for measuring electromagnetic noise and discuss the observed low-frequency (“left-hand side”) noise in the spectrum.

      • Fig. 51: Suggest adapting Y-axes and using violin or box plots (e.g., panels A/B starting from 30 up to 50, etc.).

      Thank you for this helpful suggestion. We have revised Fig. 5 accordingly by adapting the Y-axis scaling and replacing the original plots with box plots, as suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      (1) Davis and co-authors used many mouse models to investigate mechanisms that regulate the contractility of mouse popliteal collecting vessels, primarily chronotropy. Many of the mechanisms studied were previously shown to regulate pressure-induced constriction in small arteries. The authors use prior literature from the vasculature as a framework to test similar concepts in lymphatic vessels. The mouse models used provide evidence for and against the involvement of multiple proteins in regulating chronotropy and other contractile properties in lymphatic vessels. They propose that mechano-activation of GNAQ/GNA11-coupled GPCRs generates IP3, which induces Ca2+ release through IP3R1 and drives depolarization through the activation of ANO1 Cl- channels. Major concerns include the author's major conclusion that GNAQ/GNA11-coupled GPCRs contribute to chronotropy. This conclusion is not supported by the data presented.

      Although we have not yet identified a specific GPCR in LMCs that mediates pressure-induced chronotropy, our experiments have established that Gq/11-coupled GPCRs are critical for this response. The almost identical phenotypes of GNAQ/GNA11 double KO vessels with ANO1 KO and IP<sub>3</sub>R1 KO vessels point to an upstream, mechanosensitive GPCR that regulates ANO1-regulated pacemaking activity through calcium release by IP<sub>3</sub>R1. New experiments show that the highly specific GNAQ/GNA11 inhibitor YM254890 decreases contraction frequency to the same degree as, or even stronger than, that observed in Gq/11 double KO mice (Fig. 22). Additional data assessing the responses of inguinal-axillary lymphatic vessels (IALVs) from G11<sup>-/-</sup> and Gq/11 double KO further confirm our findings and conclusions (Suppl. Fig. 1). To address your concern about identifying a specific GPCR, we have performed new scRNAseq analyses to identify and rank the most prominently expressed GPCRs in LMCs (Fig. 21) along with new experiments to test the most likely GPCR candidates that emerged from that analysis (Fig. 22). Please see our specific responses below and the additional data that we now provide.

      While our results still fall short of identifying a specific Gq/11-coupled GPCR involved in this LMC mechanosensitive response, they rule out all of the major mechanosensitive ion channels previously implicated in this process (Figs. 4-15), rule out G12/13-coupled GPCRs (Fig.20) and strongly support a critical role for the canonical Gq/11-coupled mechanism (Fig. 18 and Suppl. Fig. 1). New data to eliminate contributions of the 7 most consistently expressed LMC GPCRs, along with the most likely mechanosensory GPCR in VSMCs (AT<sub>1</sub>R) (Fig. 22), out of the 136 GPCRs detected in our LMC scRNAseq dataset, raise interesting GPCR candidates that require future testing. The remaining list of 11 strongly or moderately expressed candidates includes 3 frizzled GPCRs that typically do not signal through G proteins, an olfactory GPCR, and several adhesion GPCRs (aGPCRs) that encode their own agonist (containing the Stachel peptide sequence). Some of these adhesion GPCRs recruit and signal through Gq/11, and aGPCRs have been implicated in mechanosensation in other cell types (PMID: 25937282, PMID: 40395494). However, pharmacological tools to specifically test most of these aGPCRs are currently lacking. Another possibility, of course, is that the key LMC protein may be a GPCR with a very low expression level.

      At present, the only way to test most of those candidates would be to generate new mice with smooth muscle KO of the specific GPCRs, as Offermann’s group has done in a few other cases (PMID: 40623996, PMID: 38597112, PMID: 39920161), but we are without the resources to embark on such an expedition. We are therefore unable to resolve this last issue with currently available resources and/or technology but now acknowledge this as a shortcoming of our study and a future direction. We have adjusted our title to be more congruent with the data presented in the revised manuscript.

      Strengths:

      (2) One major strength of the study lies in the vast number of mouse knockout models that were used to test the importance of ion channels and G protein signaling pathways in the regulation of lymphatic vessel contractility. In this regard, the study is a valiant effort. The authors achieved several objectives to find that ANO1 and IP3R1 regulate chronotropy, and many other potential proteins do not regulate chronotropy. This study will have a major impact on the field if additional support for G proteins is provided.

      Please see our specific responses below and the additional data that we now provide to strengthen support for a GNAQ/GNA11 mechanism.

      Weaknesses:

      (3) Major conclusions concerning the involvement of G proteins are drawn from the global Gna11 knockout mouse models. This conclusion is weak. Global Gna11 knockout mice are highly likely to have a multifactorial phenotype that could create significant differences in the data.

      We agree that global GNAQ/11 DKO mice likely have a multifactorial phenotype and that, IF our hypothesis involved testing some aspect of in vivo function with multiple cell types or organs, this would complicate our conclusions. However, because our approach focuses on ex vivo testing of the F-P relationship of isolated lymphatic vessels, a response mediated specifically by LMCs (i.e., independent of other circulating factors, endothelium, neural control, etc.), our results are NOT likely to be influenced by such extrinsic factors. Stefan Offermanns and colleagues have used similar mice (single floxed gene + Cre + global KO) and ex vivo assays of arteries in multiple peer-reviewed studies (e.g. PMID 31549965), so there is clear precedent for our approach. An alternative approach to the global GNAQ/11 DKO mice would be to use Myh11Cre; GNAQ<sup>f/f</sup>; GNA11<sup>f/f</sup> double floxed mice, but then those would likely suffer from the potential limitation of competing recombination between the two floxed genes, as we and others have encountered this problem in previous studies using a single Cre with multiple floxed alleles (PMID: 29336844, PMID: 36868507).

      (4) Control experiments need to be performed on vessels from the global knockout mice if these major conclusions are to be made.

      Control experiments were conducted on popliteal vessels from Myh11Cre; GNAQ<sup>f/f</sup>; GNA11<sup>+/+</sup> mice (uninduced littermates) and from GNAQ<sup>f/f</sup>; GNA11<sup>+/+</sup> mice (lacking Myh11Cre), the combination of which had a normal F-P relationship (Fig. 18). We have also conducted experiments on Myh11Cre; GNAQ<sup>f/f</sup>; GNA11<sup>-/-</sup> mice that were not induced with tamoxifen (GNA11<sup>-/-</sup>). It is not clear if the reviewer is suggesting that we include GNAQ<sup>-/-</sup> mice, but we did not have access to those mice. Gnaq<sup>-/-</sup> mice exhibit significant developmental abnormalities (PMID: 22723772), although according to Offermanns and colleagues (PMID: 9687499), only GNAQ homozygous deficient mice displayed “obvious phenotypic defects”.

      (5) Similarly, pharmacological tools or alternative approaches to manipulate G proteins should be used to support the data from these mouse models to draw these major conclusions.

      To address this comment, we have performed new experiments on WT mice using the potent GNAQ/11 inhibitor YM254890 (at 100 nM, the concentration used in many previous publications by other groups), which had nearly the same effect on abrogating the F-P relationship as induced Myh11Cre; GNAQ<sup>f/f</sup>; GNA11<sup>-/-</sup> mice (see Fig. 22). The specificity of YM-254890 for the Gq family has recently been verified by excellent work out of David Yule’s lab (PMID: 40563241).

      (6) The Gnaq smKO mice are the most specific G protein model studied here. However, there is no phenotype. Do not discuss trends in the data. If the data are not significant, conclude so. If more experiments are required to reach significance, provide more data in the manuscript.

      With respect to the Gq smKO data:

      (1) We have added more data to the control group.

      (2) We have now listed the p value (0.0819 in Fig. 18) for the Gq/11 controls vs Gq smKO, instead of simply designating it “n.s.” with a cut-off of p<0.05; the readers can now judge for themselves how close the values come to being significant.

      (3) Although the F-P slope does not reach significance (Fig. 18B), the frequency is significantly different between the Gq/11 controls and Gq smKO at 3 pressures (see the frequency graph in Fig. 18C), so significant differences ARE indeed detected in Gq smKO vessels.

      (4) The F-P slope for the DKO is lower than that of the G11 KO and it appears the two knockouts are cumulative, suggesting that Gnaq deletion is contributing to the lower values of the Gq/11 DKO. As further support for this conclusion, we have performed new experiments on IALVs (in part also to address concern #8 below) and that analysis is shown in Suppl. Fig. 1. We include those data as a Supplemental Figure because the contraction data in all the other figures are for popliteal lymphatic vessels. Note that the F-P analysis is not appropriate for IALVs because the control vessels typically show only a 2-2.5x change in frequency over the entire pressure range and frequency essentially reaches a plateau at P = 2 cm H<sub>2</sub>O. The frequency plots for control, G11<sup>-/-</sup> and Gq/11 DKO IALV vessels show the same trend as for popliteal vessels (significantly lower than control vessels at all pressures, with only a single exception), but in this case the DKO vessels have significantly lower frequencies than G11<sup>-/-</sup> vessels at the lower pressures (1-2-3 cm H<sub>2</sub>O), further supporting the results from popliteal vessels.

      (7) The conclusions repeatedly refer to a signaling pathway wherein the upstream component is GPCRs, which activate G proteins. While this may be the case, no GPCRs were identified here, and the involvement of G proteins is questionable, as the authors outline in lines 693-695 and noted above. The conclusions should be tempered, including in the abstract, unless additional experiments are performed to support the involvement of G proteins. Perhaps then the authors may be able to infer that GPCRs are involved.

      Please see our response to your comment #1. We have now included additional experiments using the Gq/11 inhibitor YM254890 (at a concentration of 100 nM that is reputed to be “selective” for Gq/11, according to the literature (PMID: 40563241). The blunting of the F-P relationship in the presence of that compound is comparable to, and perhaps even more significant than, the blunting observed for the DKO (the summary analysis is shown for P= 3 cm H<sub>2</sub>O in the middle of Fig. 22, in a format that matches the other analyses in that figure, but the frequencies were significantly attenuated over the entire lower pressure range). These results strengthen support for a ligand-independent mechanotransduction mechanism through one or more GPCRs of the Gq/11 family. However, because 1) there are at least 136 possible specific GPCR candidates and 2) our new specific tests of the 7 most abundantly expressed GPCRs were negative, pinning down the exact GPCR(s) involved will require much more work. Given the amount of data already presented in this study, and the possibility that the unidentified GPCR is an orphan receptor or other GPCR with very low expression, we think it is appropriate to defer that line of experimentation to a future study. We now clearly state that this is a limitation of the present work and a future direction.

      (8) Line 318. The point regarding the choice to use popliteal vessels versus IALVs will be unclear to the uninitiated, particularly as the authors previously used IALVs. Including additional justification in the text and/or data from IALVs in Figure 1, which compares IALVs to popliteal vessels, would better explain the logic.

      We now include additional justification in the text by citing Zawieja 2018 and the similarities to Ano1 KO and IP<sub>3</sub>R1 KO findings in IALVs in our other papers. Note also the very close agreement between the Gq/11 data for popliteal vessels vs IALVs in Fig. 18 C and Suppl. Fig. 1).

      (9) The conclusions drawn for TRPC6 and TRPC3 are less convincing. Germline global knockout mice, which are known to undergo compensation, were used, and high data variability is apparent. Using TRPC3 and TRPC6 blockers in the mouse models studied in Figure 4 would strengthen the arguments made regarding these proteins.

      Our primary reason for using the global KO mice is that pharmacological inhibitors of TRP channels, particularly TRPC6, are quite non-specific. We did not have access to Trpc6 floxed mice but our ex vivo assay focuses a specific aspect of LMC function in isolation from extrinsic influences (see answer to comment #3 above), and of the 18 cell subtypes detected in these vessels by scRNAseq, Trpc6 is only detectable in LMCs (Fig. 5 panel A). With regard to consequences of global knockouts and possible compensation, that is why we generated the Trpc6/Trpc3 DKO mice, as upregulation of Trpc3 is a known compensatory mechanism for Trpc6 deletion (PMID: 16055711); to our knowledge only one other group has taken the trouble to test this (PMID: 16055711) with respect to vascular smooth muscle function.

      (10) Did you perform power analysis to ensure that experimental numbers were sufficient to conclude that no statistical difference exists between datasets? If not, this needs to be done. For example, data shown in Figure 5C for tone and 6C for frequency and tone appear to be significantly different, but are concluded not to be so.

      With respect to Fig. 5, there are significant differences in contraction frequency between the TrpC6/C3 DKO mice and their SV129 controls at 5 of 7 pressures, but the frequencies of the KO mice were actually HIGHER—the opposite of the result expected if TRPC6 or TRPC3 were critical for pacemaking. With respect to the effects of TRPC6 KO on tone, there were no significant differences in tone between TRPC6<sup>-/</sup>- vessels and their SV129 controls, but there were indeed significant differences between the TrpC6/C3 DKO mice and controls at 4 of 7 pressures, but with tone actually being HIGHER than normal (opposite to what would be predicted if those channels were critical for the development of myogenic tone). We point out how that observation is at odds with the conclusions of at least one study on arterial myogenic tone (PMID: 11861411), but that topic is not the focus of our study.

      (11) At the end of each result section, a concluding statement is made regarding the effects on pressure-induced chronotropy. In many cases, there are additional effects of manipulating protein expression on other contractile properties. One example is for TRPC3 and TRPC6 (lines 414-416), but others are TRPV4, TRPV3, ENaC, Kir, Cav3.1/3.2, etc. Some interpretation is in the Discussion, but the concluding statements at the end of each result section should be expanded to summarize what the authors think the other significant differences in the data represent.

      Thank you for allowing us the flexibility to address this at the end of each relevant section in the Results. We have now expanded several of these sections as you suggest.

      (12) Kv7.4 channels. You state you have data (not shown) with linopiridine and XE991. Why not show those results here to support the experiments with the Kcnq4 smKO mice? Otherwise, I suggest you remove the statement from the unpublished data.

      We have removed the comments about the unpublished data with other Kv7.4 inhibitors.

      (13) Figure 13A. Kcnj2 is modestly expressed in LECs, but very little is present in LMCs. This likely underlies the effect of barium. If you remove the endothelium, does the effect of barium disappear? While this is not the major focus of the study, the effects of barium are dramatic, and it should be made clear whether this is due to inhibition of Kir channels in smooth muscle or endothelial cells.

      Thank you for making this very good point. We cannot rule out an effect on LECs in this context and denuding popliteal LVs is nearly impossible without compromising their contractile function because of their small size and the presence of intraluminal valves. The valves prevent easy passage of an air bubble (the opening is only 1/4 the area of the lumen) and attempts at mechanical removal result in extreme valve damage and/or residual LECs at valve sites. If there is an influence of inhibiting Kir channels in LECs, it is not via electrical communication between LECs and LMCs, which is minimal or non-existent as we have shown in previous studies (PMID: 30143234, PMID: 28994159, PMID: 30355030). Thus, a potential LEC influence would be mediated by release/production of vasoactive factors, which is a point that we now discuss.

      (14) Figure 18C tone. Several values for losartan look different but are not labelled as such. Please clarify and discuss if different.

      All the significantly different values were labeled as such, but that panel is no longer in the revised paper as the AT<sub>1</sub>R frequency data have been included in a different form (Fig. 22).

      (15) The manuscript should include raw data traces in figures that show the major pathways that you conclude regulate chronotropy.

      Thank you for this suggestion. Three new figures have now been added to illustrate the blunting of F-P relationship in Ano1 smKO (Fig. 3), IP<sub>3</sub>R1 smKO (Fig. 17) and GNAQ/11 DKO vessels (Fig. 19).

      Reviewer #2 (Public review):

      Summary:

      In this study, Davis et al. embarked on the quest for the molecular elements responsible for the regulation of lymphatic phasic contractile activity in response to variation of transmural pressure, a mechanism (termed pressure-induced lymphatic chronotropy by the authors) critical for drainage of interstitial fluid from the tissue and transport of lymph back to the blood circulation. Their aim was to investigate the mechanism(s) involved in the pressure-induced regulation of lymphatic pumping, and test whether activation of cation channels, shown in other systems to play mechanosensitive roles are directly at play, and/or whether mechano-activation of GNAQ/GNA11-coupled GPCRs is necessary to generate second messengers to activate those channels, as it has been suggested for the regulation of myogenic tone in arteries. To achieve their goal, the authors used their well-described, highly reliable protocols of mouse lymphatic vessel isolation, pressure myography, and data acquisition to obtain frequency-pressure relationships and other contractile function parameters from transgenic mice where specific channels or molecular elements of interest have been ablated. They combined these data with scRNAseq analysis of these gene targets to determine their respective role and levels of expression in lymphatic muscle cells. Their conclusion is that none of the exhaustive list of tested ion channels was critical, except ANO1 Cl channels, part of the contractile pacemaker mechanism, but that transmural pressure activates GNAQ/GNA11-coupled GPCRs, which generate IP3 to induce SR Ca2+ release through IP3R1 and activate ANO1-mediated depolarization.

      Strengths:

      The manuscript's strengths reside primarily in very robust, clean, and unequivocal pressure myography data and analysis. The research team is mastering these techniques they developed more than a decade ago and have implemented in mouse lymphatics to study their contractile properties, with consistent and convincing outcomes. They also provide data from an impressive list of transgenic mice in order to determine the role of the targeted gene in pressure-induced lymphatic chronotropy, relying on pharmacological small molecule inhibitors only when necessary. Finally, the use of scRNAseq analysis they gathered from previously published datasets brings novelty with respect to the expression of the genes of interest in all populations of cells comprising the lymphatic vessels, but more critically, to validate or contrast the potential impact of genetic alteration of the given gene on the ability of lymphatic muscles to respond to a change in pressure.

      Weaknesses:

      (1) The main weakness may reside in the fact that while the authors provide a convincing demonstration that GNAQ/GNA11 are involved in the regulation of the F-P relationship, they give little evidence of the involvement of "upstream" receptors. Indeed, inhibition of AT1R, shown to be involved in myogenic regulation of arteries (a phenomenon the authors rightfully compare to pressure-induced lymphatic chronotropy), didn't lead to a similar effect (decrease in F-P) in lymphatic vessels. Arguably, other GPCRs might be involved in lymphatic vessels, but as such information is not provided in the manuscript, the author's conclusions should be dampened. More in-depth discussion would be required. In fact, it can be argued that the discussion is very restricted with respect to the amount of data and information the manuscript provides.

      To address these valid concerns, we performed a detailed scRNA-seq analysis of the GPCRs expressed in mouse LMCs. Approximately 136 GPCRs were identified as being expressed in over 0.5% of LMCs, the top 20 of which are expressed at moderate levels in a substantial percentage (>36%) of LMCs. We have added a new Figure showing bubble plots of these 20 GPCRs in the various cell populations of IALVs, ranked according to their expression level (Fig. 21); we also show their expression in the other cell types. We then experimentally tested 7 of the most likely candidates for which inhibitors were available by determining whether blocking the GPCR would alter pacemaking by lowering the contraction frequency at P=3 cm H<sub>2</sub>O. This was a streamlined assay compared to measuring the complete F-P relationship but produced similar results; to demonstrate this point, we also performed the streamlined analysis for the ANO1 inhibitor, Ano1 smKO, IP<sub>3</sub>R1 smKO and Gq/11 DKO, all of which showed significantly reduced frequencies, compared to their respective controls at P=3 cm H<sub>2</sub>O (see Fig. 22, black symbols). The alignment of statistical significance for the latter results with the corresponding F-P analyses shown in Figs. 2, 16 and 18) demonstrates that this assay is sufficiently powerful to detect a significant effect on pacemaking. Unfortunately, the results for each of the GPCRs tested were negative (Fig. 22, red symbols). We conclude that none of those particular GPCRs (NPY1R, ET<sub>A</sub>R, NPR3, TBA<sub>2</sub>XR, S1PR1, AT1<sub>A</sub>R, AT1<sub>B</sub>R), or two additional GPCRs (HT-2R, HRH1) that were tested in previous studies (PMID: 12770929, 28453392), are critical for pacemaking or the F-P relationship in LMCs. Of the remaining 11 GPCRs, several are adhesion receptors (Adgrl1, Adgrl2, Adgrl3, Adga2), two are orphan receptors (Gprc5b, Olfr1033), one codes for the GABA-B1 receptor (Gabbr1, with unknown function in the lymphatic system), two code for proteins in the Wnt signaling pathway (Fzd4, Fzd2), one codes for a protein in the Hedgehog signaling pathway (Smo) and one encodes a beta-amyloid binding protein (Tm2d1). The possible roles of these GPCRs, and the other 100+ GPCRs expressed at much lower levels remain to be tested in future studies. At any rate, systematic investigation of these GPCRs is well beyond the scope of this study, and in many cases not possible due the lack of known soluble ligands/antagonists. We have significantly expanded the Discussion with regard to this topic and stated that our inability to identify a specific Gq/11-coupled mechanosensory GPCR in LMCs is a limitation of our study.

      (2) Overall, the authors convincingly achieved their aim by performing an impressive number of technically challenging experiments, leading to solid datasets. While these support their main conclusions, a more elaborate discussion might be required to refine them.

      We hope this limitation has been tempered by the addition of the new data and the expanded Discussion section mentioned above.

      (3) This study is likely to have an important impact on the field as it provides some answers to the lingering question of how lymphatic vessels regulate their contractile activity to variation in transmural pressure and certainly proposes an experimental means to further explore and address that question.

      Thank you.

      Reviewer #3 (Public review):

      In this manuscript, Davis and colleagues aimed to identify the molecular sensors and signaling cascade that enable collecting lymphatic vessels to increase their spontaneous contraction frequency in response to intraluminal pressure (pressure-induced chronotropy). They tested whether the process is similar to blood vessel myogenic constriction by relying on cation channels (TRPC6, TRPM4, PKD2, PIEZO1, etc.) or instead require the activation of G-protein-coupled receptors (presumably mechanosensitive GNAQ/GNA11-coupled receptors), using ex vivo pressure myography of mouse popliteal lymphatics, smooth muscle-specific conditional knockouts, quantitative PCR validation, and single-cell RNA sequencing for target prioritization. The authors convincingly demonstrate that pressure-induced chronotropy does not require the cation channels implicated in arterial myogenic tone but is blunted by deletion of GNAQ/GNA11 or IP3 receptor 1, supporting a model of GPCR > IP3 > Ca2+ release > Cl⁻ channel activation > depolarization. The core conclusion is robust. The work redefines lymphatic pacemaking as G-protein-coupled receptor-dependent mechanotransduction, distinct from arterial mechanisms, and provides a genetically validated toolkit that is useful for studying lymphatic function and dysfunction.

      Strengths:

      (1) The data are of high quality and highly sensitive functional readouts

      (2) The systematic genetic targeting is a major strength that overcomes pharmacological artifacts

      (3) Careful quantitative analyses of frequency-pressure slopes

      Weaknesses:

      (1) The use of inguinal-axillary vessels for single-cell RNA sequencing rather than the popliteal segment studied functionally.

      Agreed. The need to use IALVs derives from the much smaller amount of tissue needed for scRNAseq analysis of individual popliteal vessels (and 3-5x longer time required for cleaning) before pooling them. We have tried to clearly state this limitation. However, functional comparisons between popliteal and IALV function are quite similar in almost all aspects that we have studied. In the majority of our protocols, preliminary experiments using IALVs were performed side by side with popliteal vessels (for Ncx1 smKO, Trpc6<sup>-/</sup>-, Trpc6/3 DKO, G<sub>12/13</sub> DKO, Piezo1 smKO, Trmp4 smKO) without detection of a significant difference in phenotype. In addition, we have performed new experiments using IALVs from G11<sup>-/-</sup> and GNAQ/GNA11 DKO mice to support the major positive findings drawn from the popliteal vessel studies (compare Fig. 18 with Suppl. Fig. 1).

      (2) No direct testing of the specific G-protein-coupled receptor involved.

      Agreed. As described in our response to the other reviewers (see point 1 to Reviewer 1 and point 1 to Reviewer 2), to address these concerns we have added a new figure (Fig. 21) ranking the top-expressed GPCRs in LMCs and have experimentally tested several of the most likely candidates (Fig. 22). Unfortunately, our results were still negative, allowing us to rule out those particular GPCRs but not to pinpoint a specific GPCR. The possible roles of other 100+ GPCRs in LMCs remain to be investigated in future studies. We have adjusted our title to better reflect our results.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors state multiple times that several cation channels have been shown to contribute to vascular tone. Arterial smooth muscle cell ANO1 channels also contribute to arterial myogenic tone (PMID: 22872152). An ANO1-dependent mechanism is proposed here in lymphatic vessels, and the same ANO1 mouse models are studied. PMID: 22872152 is not discussed or cited and should be.

      Now cited.

      (2) Lines 88-97. It is important to note that the Bulley paper studied tamoxifen-inducible smooth muscle cell-specific knockout mice, whereas the Sharif-Naeini study used germline SM22-driven PKD1 knockout mice. The mice used in the Sharif-Naeini paper were subsequently found to have an ADPKD-like phenotype due to non-specific knockout of PKD1 in other cell types (PMID: 20856231). That should also be noted. The authors also state "PKD1 channel" in this paragraph when PKD1 does not form a channel.

      Corrected and clarified. A new paper on PKD1 by Jaggar’s group is now cited in that context.

      (3) Line 119. Perhaps you mean PKD1/2 channels?

      Corrected.

      (4) Line 312. It is unusual that data from Figure 2A is described before that in Figure 1. Can you reorder the figures to fit the text?

      Corrected.

      (5) Line 130. Can you provide references that show that ANO1 is not intrinsically mechanosensitive?

      Added.

      (6) Line 160. I believe the Pkd2f/f mice are available from the Maryland PKD Core and were not generated by the Jaggar lab.

      Corrected.

      (7) Line 319 needs a reference.

      Added.

      (8) The color choice for the symbols in many figures, particularly Figures 2, 3, and 4, makes it difficult to know which condition is which.

      Corrected.

      (9) I may have missed it, but the Figure 1A legend seems to need a definition for every cell type abbreviation.

      It is shown in panel A (Figure 2).

      (10) In several places, the authors state that they deleted gene expression in LMCs, when the mouse model deletes expression from all smooth muscle cells. Please refine wording.

      Clarified.

      (11) Line 664-666. Sharif-Naeini did not propose that PKD2 channels contribute to myogenic tone. That was Bulley et al.

      Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) The role of TRPV4: It is interesting to note that despite a quasi-non-expression of the TRPV4 gene in LMC, KO this gene in these cells results in an increase in contraction frequency and a strong decrease in amplitude at low pressures, suggesting the channel is expressed and active in LMCs. It is not too clear how macrophages could be indirectly involved in this process by releasing products. Are they also sensitive to pressure? How? And why should they be more involved than any other cell type present in the vessel wall (i.e, LECs)?

      Macrophages definitely release products that affect LMC activity (PMID: 38826322) and they may be sensitive to pressure but that has not yet been demonstrated. However, as shown in the paper cited, selective KO of Trpv4 in macrophages does not abrogate the F-P relationship in lymphatic vessels.

      (2) The role of AT1R in pressure-induced chronotropy: The lack of effect of AT1R inhibition on F-P doesn't mean other GPCRs could play a similar role in lymphatics. A more in-depth discussion could be provided. Furthermore, could G-proteins be activated by pressure independently of coupled receptors?

      We have extensively expanded our discussion of possible mechanosensitive GPCRs in LMCs. We have also cited at least one paper providing evidence for G-proteins being activated by pressure independently of coupled receptors (PMID 28148497), although we believe our added assessment of LMC GPCR expression reveals novel targets to explore for mechanotransduction in LMCs.

      (3) Figure 1: Please explain the apparent discrepancy between the described protocol and the experiment displayed in A. Please explain and confirm that the experiments presented in the study followed the protocol described in the Methods section. Add text to Methods.

      The example does indeed follow the stated protocol in the “METHODS section” for the sequence and duration of pressure steps. Note that we stated 2 min was typical at each pressure level but for some of the lower pressures 4-5 min were required to get a sufficient number of contractions, as in this case. The diameter and pressure traces at P=8, 10 cm H<sub>2</sub>O are not shown in the example so that we could expand the time axis for better resolution.

      It can also be inferred from the trace that the F-P curve was built from frequencies obtained when pressure was lowered and increased. Are values at a given pressure identical if recorded when pressures were lowered or increased? Does it influence the relationship?

      There is indeed some variability in the response to stepping pressure up vs down because of rate-sensitive effects, as described in one of our previous studies (PMID: 19001046). In the protocol we implemented here, the FREQ at P=3 cm H<sub>2</sub>O upon step up from 0.5 cm H<sub>2</sub>O was usually slightly higher than the initial FREQ at P=3 cm H<sub>2</sub>O. Those two values were averaged together and this is now stated.

      The FREQ graph in B displays values up to 10 cm H2O, which are not shown in A.

      The data at P=8, 10 cm H<sub>2</sub>O were obtained for that same vessel. The complete range is shown to illustrate that the relationship reaches a plateau above 5 cm H<sub>2</sub>O.

      (4) Figure 1: Is there a need to repeat/duplicate the FREQ graph in B and D?

      It has now been replaced with Ejection Fraction.

      (5) Figure 19: With respect to point #2, and whether G-proteins are involved/important in the process, given the current state of understanding, that part of the illustration might need to be amended.

      The figure has been revised.

      Reviewer #3 (Recommendations for the authors):

      Below are some suggestions for the authors to consider:

      (1) Test specific GPCRs pharmacologically (e.g., AT1R) on wild-type vessels to determine if pressure-induced chronotropy is blunted.

      Done, Figure 22.

      The authors could also rank Gq/11 receptors based on the expression in LMCs (e.g., qPCR using FACS-sorted cells).

      Done, Figure 21.

      (2) Residual chronotropy in ANO1 knockout mice: Quantify the remaining frequency-pressure slope in Myh11-CreERᵀ²;Ano1ᶠ/ᶠ vessels and test whether it is sensitive to low-chloride buffer or chloride channel blockers to probe an alternative Cl⁻ conductance.

      The role of chloride in lymphatic pacemaking is well appreciated and the reviewer offers some intriguing experiments for future experimentation. The use of low Cl<sup>-</sup> solutions was part of the fundamental discovery of lymphatic muscle pacemaking and the role of a calcium-activated chloride channel as the basis for spontaneous transient depolarizations (STDs) by Dirk Van Helden in 1993 who reported that “Lymphatic STDs reversed at potentials near -35 mV (range -40 to -25 mV; n = 4)” and were “suppressed by exposure to low chloride (replaced with sodium isethionate) solution [(> 5 minutes)]”, although more details were not provided. The effect of low Cl<sup>-</sup> and lymphatic muscle excitability was followed up in a 2008 paper by von der Weid et al., which showed that STD frequency and amplitude increased acutely (within 2-3 min) after the bath solution was changed to a solution containing 10% of the initial Cl<sup>-</sup> concentration (replaced with methane sulphonate). This result likely suggests that the acute reduction of extracellular Cl<sup>-</sup> increased the reversal potential for Cl<sup>-</sup> and that within 2-3 minutes the intracellular Cl- content had not quite run down. Work by Boedtkjer’s group (Mohanakumar et al. 2018) also demonstrated that human lymphatic pacemaking was also heavily dependent on the bath Cl<sup>-</sup> concentration. Using wire myography preparations, they showed that human thoracic rings and mesenteric vessels ceased spontaneous contractions when Cl<sup>-</sup> was removed (replaced with aspartate), and that contractions returned upon washout with a normal Cl<sup>-</sup> containing solution. In that study, the authors noted that there was some variability in the timing required for contractions to cease with the mesenteric LVs. In some cases, an acute large constriction immediately followed the exchange with Cl<sup>-</sup> free solution, which would also point to an acute depolarizing/stimulatory effect of a rapid resetting of E<sub>Cl</sub> to a more positive value when external Cl<sup>-</sup> was removed acutely removed. A limitation to the approach is the use of wire myographs as opposed to pressurized vessels, as the wire method does not allow the vessels to be stretched in a physiological manner. Similarly, diastolic depolarization was not typically observed by Van Helden and von der Weid, as they used pinned-out tissue sections (required for recording STDs from short vessel sections), which highlights the necessity of using a physiological stretch stimulus.

      On the face of it, the complete cessation of contractions reported in the Boedtkjer paper is at odds with our previous and present results using Ani9 and Ano1 smKO and IP<sub>3</sub>R1 smKO vessels, which retain some baseline level of pacemaking, albeit largely pressure-independent. Thus, the hypothesis that a separate Cl<sup>-</sup> channel may be active and important holds some merit. However, it is worth noting that we do not expect this basal pacemaker drive to be calcium-dependent, as IP<sub>3</sub>R1 calcium oscillations are maintained. IP<sub>3</sub>R1 smKO vessels, which largely lack subcellular calcium transients in diastole, also have a lower-than-expected contraction frequency compared to Ano1 smKO vessels, although this is in part due to an elongation of the action potential plateau. We also do not think the residual activity is simply due to the inability of inducible Cres to drive complete recombination, as we performed experiments with Ani9-treated Ano1 smKO vessels (Fig. 2B) and did not observe a significant additional reduction in the concentrations at which maximal inhibition is assumed, 2-10 mM (see also Harlow et al. 2025). We do not see a robust presence of Ano2 (TMEM16b) and while Ano6 (TMEM16) is expressed, it is a calcium-activated scramblase with permeability to both cations and Cl<sup>-</sup> (Ye et al 2019).

      We also have personal experience with Cl<sup>-</sup> substitution (replaced to 10 mM by aspartate), originally obtained as part of our work identifying Ano1 as the primary calcium-activated chloride channel, although the inability to ascertain the intracellular Cl<sup>-</sup> at any point in time while the bath exchange is occurring significantly limited the interpretability of that data. During review of that manuscript, we were asked to remove the Cl<sup>-</sup> substitution experiments but spontaneous contractions still persisted in the majority of the vessels (n=6). Notably this was not a complete removal of Cl<sup>-</sup> as was done in Mohanakumar et al. 2018. However, the complete removal of Cl<sup>-</sup> is a much more significant intervention than the inhibition of a single channel, likely altering HCO<sub>3</sub><sup>-</sup> and Na<sup>+</sup> flux through the associated Cl<sup>-</sup> antiporters and symporters. How Cl<sup>-</sup> removal affects LMC intracellular pH regulation via impaired bicarbonate transport also remains unknown. Nor is it well appreciated how Cl<sup>-</sup> removal affects mitochondrial function and ROS production in LMCs, which regulate potentially significant confounding pathways related to lymphatic muscle pacemaking. Without control of these significant confounding circumstances, the hypothesis that the baseline depolarization is still a Cl<sup>-</sup> channel is far from a certainty.

      In Author response image 1, we provide the reviewer with the expression data for other members of the Anoctamin family, the potential Anoctamin-regulating CLCA family, and other documented chloride channels in LMCs (clusters 5-6).

      Author response image 1.

      (3) I understand this would be a huge undertaking, but aligning scRNAseq with functional vessels using popliteal LMCs or inguinal-axillary LMCs could help rule out potential tissue-specific differences.

      We understand the reviewer’s concerns and, owing to the small amount of cells that can be recovered from popliteal vessels, we instead now provide new data using pressure myography of inguinal axillary vessels (IALVs) for the data with positive findings (Gq/11 inhibition) that support our findings in popliteal vessels (Suppl. Fig. 1). This points to a fundamental role for Gaq, Ga11, and possibly Ga14, in signaling upstream of IP<sub>3</sub>R1 mobilization and Ano1 activation. We are aware of the length of this manuscript (now further increased in order to respond to the reviewers’ suggestions) and, given the negative data for the vast majority of putative mechanosensory pathways, we have omitted some of the parallel data collected in IALVs. In the majority of our protocols, preliminary experiments using IALVs were tested side by side with popliteal vessels (from Ncx1 smKO, Trpc6<sup>-/-</sup>, Trpc3/6<sup>-/-</sup>, G12/13 DKO, Piezo1smKO, Trmp4 smKO mice) without detection of a significant difference in phenotype.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This study provides a valuable genome-centric characterization of microbial communities across deep sediment cores from a Spartina patens salt marsh. The study provides claims on the metabolic capabilities of the deep sediment microbiome as well as on a burial microbial assembly process and functional complementarity at depth. However, some of these claims remain incomplete and would benefit from further supporting evidence. Overall, this work will be of interest to microbial ecologists working on wetlands.

      We appreciate all of the efforts of the reviewers and editors. Additional analysis and extensive edits were made to the original manuscript to address the comments of both reviewers and editor. We believe this effort has significantly improved the manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Vineis et al. examined the structure and functional potential of microbial communities along a vertical sediment profile of a salt marsh, using a genome-centric metagenomic approach. They attempted to test whether (1) the microbial communities within dynamic upper layers contain genomes with diverse functional potential, (2) the energy limited deeper sediments contain microbial consortia assembled to metabolise complex carbon, and (3) microbial compositional changes in the low energy sediments mirror the burial processes observed in marine environments with similar energetic limitations. Results revealed a core microbial consortia that contains a collective metabolic potential for complex carbon and aromatics degradation, suggesting putative syntrophic interactions. Besides, the recovery of MAGs assembled independently from multiple depths in the same core and the consistent relative abundance structure of MAGs within co-occurrence network modules together suggest burial process as a likely mechanism for microbial assembly.

      Strengths:

      (1) Two long sediment cores (down to 240 cm deep) were collected in this study, allowing investigation of the less well characterised subsurface microbiome in salt marsh.

      (2) A genome-centric metagenomic approach was employed here, which provides information on both the structure and functional potential of the salt marsh sediment microbiome, which is not possible in commonly performed 16S rRNA-based surveys.

      Weaknesses:

      (1) In both the abstract and conclusion, the authors claimed that results from this study provide a "mechanistic understanding" of the assembly and distribution of the microbial communities in salt marsh sediment (P2, L31 and P35, L645-649). However, both claims are speculative and not supported by solid evidence. Firstly, the genomic data presented in this study and supplementary physical properties of sediments in the broader area are not enough to make a solid claim (that appears in the title) on microbial assembly being governed by a burial process. Alternative explanations include residual bioturbation, slow porewater advection, etc. Therefore, this remains an interesting hypothesis unless additional evidence is provided to rule out the alternative explanations. Similarly, the claim on the detailed syntrophic interactions among members within a co-occurrence network module (e.g. P36, L649-652) is purely speculative and warrants functional validation experiments to prove.

      The reviewer makes two points in section 1 of their review, and we have addressed each point as follows: 1) We have removed the term “mechanistic understanding” from the manuscript and instead focused on the co-occurring group of microbes and the evidence for overall community and population structure with depth. Statistical analyses were added to the manuscript including a test for the influence of depth on the overall community structure (Fig. 3) and sample-specific single nucleotide variants (SNVs) within a collection of MAGs that were abundant and prevalent throughout the subsurface (Fig. 7). The text now contains additional text in the discussion (lines 702-720 and 843-845) that addresses alternative explanations of our observations including the potential for advection and residual bioturbation. 2) We agree that to prove syntrophy, functional validation through experimentation is required and we have added a statement to this effect on line 858-864.

      (2) A major aim of this work was to study complex carbon degradation. However, neither CAZymes, the first-line carbon degradation enzymes, nor peptidases, which can be important contributors to carbon degradation at depth, was examined here. METABOLIC, which the authors used for functional annotation of MAGs, by default generates peptidases outputs and can be easily integrated here.

      We expanded our analysis to include CAZymes and peptidases including an analysis of the number of peptidase pathways among our network modules (Table S3, Fig. S6), and a supplemental table of glycoside hydrolase and polysaccharide lyase (Table S2). Methods regarding the CAZyme and peptidase analysis are included on lines 282-296, results on 490-509, and 581-588.

      (3) No geochemical data is available to provide context for the genomic analysis here. Without such information, readers cannot even tell whether the surface sediment samples were oxic or anoxic. A reference to a PhD thesis is provided (P6, L126) but it would be most helpful to extract relevant data from there and provide as a supplementary table.

      Oxygen concentrations were not measured as part of that study. Although oxygen concentrations are heterogeneous, due to bioturbation, and radial oxygen loss from roots in marsh systems similar to ours they are generally below detection at the shallowest depth examined in our study. We have added additional citations relevant to this point on lines 747-753.

      (4) A single metagenomic binning tool, CONCOCT, was used in this study, which very likely has resulted in a limited number of MAGs recovered. More (high-quality) MAGs are expected with the use of additional binners and a bin consolidation procedure.

      We agree that multiple binning tools can potentially identify more bins, contigs can be mistakenly binned when using purely automated processes. To sidestep the chance of including erroneous contigs in our collection, we chose to invest a large amount of time and effort required to bin manually. CONCOCT was used as an initial guide to identify some of the most easily reconstructed MAGs but all contigs within MAGs were subjected to visual inspection of the sequencing coverage profile over several samples, sequence composition congruency, and real-time completion and contamination estimation as contigs were manually added and/or subtracted from the collection in each bin. This approach is not more widely applied because of the expertise and time required to manually reconstruct MAGs from each sample individually. Many researchers find that human involvement is often required to improve the accuracy of automated binning tools, which has given rise to several tools in addition to Anvi’o, BinaRena, ICoVeR, and ggKBase. To elaborate on this procedure, we made a video tutorial of our approach to binning manual-binning-approach. An example of the manual bin refining process can be found here (merenlab-MAG-refinement)

      (5) Several terminologies are misleading here. Firstly, the term "co-occurring" or "co-located" microbes or MAGs (e.g. P1, L19 and P31, L537) can be misleading as it could imply a close spatial relationship. However, co-occurrence networks rely on correlations of (relative) abundance and show statistical associations instead of direct spatial or physical relationships. I would suggest alternative names such as co-abundant or statistically associated microbes.

      We have added text to address the important point raised by the reviewer regarding the term “cooccurrence”. We added additional language in the methods section to define co-occurrence and to make clear that the term should not be interpreted to infer a direct spatial or physical relationship (lines 121-123). However, the term “cooccurrence” is commonly used to describe the significant correlations identified molecular ecological networks and we think the introduction of additional terminology such as “statistically associated” could lead to additional confusion.

      Secondly, the term "persistent conversion of soil organic carbon" (P36, L654) in the conclusion is also misleading as it implies an active process, which cannot be tested without metatranscriptomics or metaproteomics data.

      We agree and thus we added metatranscriptomic data to the manuscript to assess the activity of the microbes in the subnetwork of commonly co-occurring MAGs within the subsurface. Our results indicate that the pathways for complex carbon are both present and active at the time of sampling. The methods section describing this additional analysis is included on lines 155-168 and 314-324 and we created a new figure to communicate the present and active MAGs in the Bathyarchaeia BA1 subnetwork (Fig. 7). Even with the additional metatranscriptome data, we agree that the term “persistent conversion of soil organic carbon” is still not supported by our data because we sampled at a single timepoint and we have removed this text from the manuscript.

      (6) Based on a NMDS plot of KEGG IDs (Figure 4B), the authors claimed that the functional potential among MAGs in modules 1, 2 and 7 was very similar (P18, L346). However, the dispersions of modules 1 and 2 were just too large. A proper statistical test, such as PERMANOVA, should be used to support the claim.

      We replaced the analysis of the KEGG modules with a more detailed characterization of the functions identified in the MAGs within each of the modules (Fig. 5) (including the CAZy and peptidase analysis described above) (Table S2 and S4). This revised approach provides more resolution on the higher-order functions that are differentially detected among the modules. This allows for greater clarity on the functions that are specific to the MAGs within each module.

      (7) Genome-scale metabolic networks was analysed using Metag2Metabo (M2M) and results were discussed in detail (P26, L453-466). However, the source data should be provided in a supplementary table to show what metabolites are producible by which MAGs.

      The M2M analysis generates a map of thousands of genes and reactions for each MAG which is stored in an XML-formatted file. Files for each MAG are available upon request so the metabolic pathways can be explored in Fluxer (https://fluxer.umbc.edu/) or Escher (https://escher.github.io/). The metabolites produced through complimentary reactions within the Bathyarchaeia BA1 subnetwork are included in Table S6 and the complimentary gene content for the Benzoyl-CoA pathway is shown in Table S7.

      Reviewer #2 (Public review):

      This work provides a detailed metabolic reconstruction of sediment microbiomes along a depth profile in a Spartina patens salt marsh in Massachusetts, USA. Using a combination of genome reconstruction, co-occurrence network analysis, and metabolic profiling, the authors describe the metabolic potential of co-occurring microbial consortia in understudied deep sediments.

      Major strengths of this study include the detailed metagenomic characterization of the understudied deep marsh sediments. The authors recovered genomes representing a substantial portion of the deep sediment microbiome (up to ~60%) and provided an initial explanation of pathways related to the potential for organic carbon decomposition in this environment. Of particular interest is the capability of the deep sediment microbiome to process aromatic organic compounds, highlighting the need for a collaborative consortium to carry out their decomposition. Improved understanding of the microbial transformation of deep sediment organic carbon in blue carbon ecosystems is vital to better understand the fate of this large carbon pool in the face of climate change.

      However, I have a few concerns in the interpretation of the results, and in the case of the surface sediments there is a lack of strong evidence in my opinion.

      (1) A stronger ecological interpretation is needed regarding the meaning of the co-occurrence network analysis. The authors correctly note that their analysis identifies groups of co-occurring genomes, which may indicate shared niche space, not necessarily interspecific ecological interactions (as the authors imply for instance in lines 423-425). When performing network analysis using samples from the entire sediment profile (0-240 cm), they identified consortia that co-vary in relative abundance along the depth gradient most likely because of shared environmental filtering forces, such as changes in redox potential and sediment chemistry. Supplementary Figure S4 showing that different modules have distinct abundance distributions along the sediment profile supports this idea. Being that the case, I would like the authors to define the ecological significance of the "connector hub". Is it merely taxa that is prevalent in the whole sediment profile? Since the modules are physically separated (in different sediment depth layers), they are not really interacting between each other. As it stands, it is not clear why the authors decide to study connector hubs in greater detail, along with their subnetworks.

      These are all very good points raised by the reviewer and we have taken the following steps to clarify the role of connector nodes. Additional text regarding the environmental role of a connector node is included on lines 273-278, 440-442. We added additional text regarding our decision to study the Bathyarchaeia BA1 MAGs that were classified as connectors in greater detail (lines 514-534).

      (2) I question if the lack of network modules in the surface sediment is really a consequence of non-significant interspecific ecological interactions and not the result of methodological biases. The low MAG recovery and thus short read recruitment in surface-level metagenomes may hinder the ability of the authors to identify co-varying microorganisms in the surface sediment. The high diversity of the surface sediment prevents proper assembly of the surface microbiome. I would also argue that as redox potential declines sharply in salt marsh sediments just below the root surface, the microbial community in the first few centimeter's changes rapidly and is significantly different from the more stable deep sediment microbiome. Due to the sampling design, the study has less representation of the surface layer (only 0-30 cm, while the cores extend down to 240 cm). Grouping sediment microbiomes by depth based on similarity in their sequence space (e.g., Mash) or functional profile (e.g., KEGG annotation) before performing network analysis could help to better infer ecological relationships within the distinct ecological niches of the marsh sediment profile, rather than performing a single network analysis of all samples combined.

      We agree with rationale that the low MAG recovery from surface sediments inhibits our ability to build a reliable network that would capture interactions within this region of the sediment profile. While the reviewer’s idea of grouping samples based on functional similarity or depth is and interesting idea, this decreases the number of samples used for estimating co-occurrence. Because we had few MAGs derived from surface sediments, it is unlikely that this recommended approach would yield additional connections. Additional analysis of surface sediments is certainly warranted in order to tease apart the interactions and their relevance to the biogeochemistry of the sediment. The reviewer comments may be useful to others attempting to identifying niche spaces and connections in the surface. See lines 775-783.

      (3) Normalizing the relative abundance of MAGs by dividing by the total reads mapping to a particular sample can be misleading due to differences in recruitment levels across samples (and depths). A better approach would be to normalize by metagenome library size, or preferably by genome equivalents (e.g., using MicrobeCensus) or a similar approach.

      We decided to normalize to the number of reads mapped to our non-redundant collection of MAGs instead of to the total number of sequences in each sample for several reasons. 1.) We do not know what the proportion of unmapped reads represents and could include virus and eukaryotic life. Comparing the amount of microbial DNA to the eukaryotic fraction can be misleading because changes in the relative abundance from these sources would skew the microbial relative abundance. In human microbiome studies, subtraction of human reads is often carried out prior to calculating relative abundance, but we lack any such reference in our system. We added text to explain this in the methods (lines 219-232). 2.) We are not comparing the total abundance of MAGs across samples or comparing individual genes across samples, so tools such as MicrobeCensus would not offer an improved analysis. Additionally, many of these tools rely on reference genomes within the workflow which is challenging in these sediments due to the many novel taxa. There are challenges with any relative abundance calculation and clear explanation of our read mapping strategy and Fig. S2 allows the reader to see the potential bias.

      Recommendations for the authors:

      Reviewing Editor Comments:

      It appears that the main issue is that some of the claims in the paper are not adequately justified. You can find the detailed assessment of the two reviewers in their reviews, and their recommendations are listed below. We encourage you to examine these comments and incorporate the recommendations in your manuscript. In particular, after consulting the reviewers, we encourage you to prioritize the following:

      (1) Addressing alternative hypotheses.

      (2) Improving MAG binning and assembly.

      (3) Revising statistical analyses.

      (4) Including CAZymes and peptidases in the analysis.

      (5) Revising the terminology and interpretation (e.g. regarding co-occurrence).

      Reviewer #1 (Recommendations for the authors):

      (1) Title: consider shortening it.

      We agree and shortened the title

      (2) P3, L49-52: incomplete sentence.

      The corrected sentence in on lines 45-47.

      (3) P3, L54: consider adding a citation for "blue carbon stocks" for a broad readership.

      We added a citation. Line 54.

      (4) P7, panel A: consider adding coordinate axes; panel B: would be helpful to add more details on the tools (e.g. CONCOCT for binning), key parameters (e.g. dereplication at 95% ANI), and/or key statistics used (e.g. the number of high-quality and medium-quality MAGs).

      We added additional details to figure 1.

      (5) P8, L131 (and L148): define MAG at its first use. Are they bins with >50% completeness and <10 contamination MAGs?

      MAGs are first defined on line 24. Bins are a minimum of 50% complete and maximum 10% redundant (lines 200-205).

      (6) P9, L168: the default setting for dRep uses -comp 75, which won't result in MAGs with a completeness <75%; however, some MAGs in Table S1 have a completeness <75%.

      The methods on this point are clarified on lines 191-206.

      (7) P10, L173: by the "75% threshold", do you mean for completeness/completion?

      We have clarified that sentence to specify “completion” on line 203.

      (8) P10, L175: provide the version number for GTDB used.

      The version is now provided in the text on line 205.

      (9) P10, L185: information in the GitHub repository is incomplete; e.g. the code for running CONCOCT binning is missing.

      The GitHub is now up to date including a link to a video explaining how binning and MAG refining were conducted for one of the samples.

      (10) P14, L268: how many MAGs are novel at each taxonomic ranks? Table S1 does not include genus level data.

      We have included additional text regarding the novelty of MAGs on line 370-375. Table S1 now has genus-level classification.

      (11) P14, L271, 274, and 276: please provide the exact numbers and/or percentages.

      The percentages are now reported on lines 372-375.

      (12) P14, L272: however, Figure 2 only shows phylum data.

      We have reworked this paragraph and corrected this error (lines 384-406)

      (13) P14, L281: deep instead of shallow?

      The paragraph was rewritten and this point now appears on lines 392-394.

      (14) P16, L298-302: would be helpful to provide a supplementary table showing these values.

      These values were added to table S1.

      (15) P16, L314: aren't they 30% and 50%, respectively?

      This error was corrected. See lines 378-379.

      (16) P18, L339: five Dehalococcoidales and no GIF9 MAGs?

      We revised this section and the details are now included on lines 438-440.

      (17) P21, L385: module 5 instead of 4?

      This section was revised and the corrected details can be found on lines 460-462.

      (18) P21, L386: Figure S4 is irrelevant here.

      This section was revised and the corrected details can be found on lines 460-461.

      (19) P22, L394: Figure 5 instead?

      This section was removed during revision

      (20) P24, L413: to me, these MAGs were present along the entire core instead.

      The section regarding “bimodal” MAGs was removed during revision.

      (21) P24, L411: there are multiple typos in the scientific names in Figure S6.

      We have removed this supplemental.

      (22) P25, L430: Fig. 6A instead?

      The correction of this error is on line 525 and references figure S7A.

      (23) P25, L440: module 1?

      The GIF9 classification was replaced with the family-level classification AB-539-J10 and we have corrected any error assigning these MAGs to the incorrect module. The original sentence containing this error was revised, but the section including these results can be found on lines 514-534.

      (24) P28, L501: these pathways were present in not even half of the MAGs.

      The details of selenate and arsenate are now presented in figure 5.

      (25) P30, L508: also provide a key for the size of nodes.

      The original figure 6 is now figure S7 and was revised to include a key for genome size in the network.

      (26) P31, L553: please provide a supplementary table showing the metabolic potential of each MAG to support this claim.

      We created figure 5 to address this comment in addition to table S9. The associated text can be found on lines 739-745.

      (27) P32, L564: I would suggest changing the word "consistency" to "correlated pattern" for clarity.

      We are trying to describe the correlated pattern as being consistent across multiple samples. See lines 730-732 for the revised use of this term.

      (28) P32, L567: also Bacteroidales.

      We have clarified the genomic content of MAGs containing benzoyl-CoA reductase subunits and production of benzoyl-CoA from phenol on lines 630-635

      (29) P32, L575: please tone down this statement; some MAGs are capable of both (Figure 6C).

      We have toned down the statement and discussion of the metabolic handoffs and syntrophy. See lines 801-817.

      (30) P33, L582-585: aromatics are not products of fermentation so their decomposition doesn't "follow" the initial fermentation.

      We revised the statement on lines 815-817.

      Reviewer #2 (Recommendations for the authors):

      (1) Performing co-assemblies could help improve the recovery of microorganisms from the surface sediment layer.

      We have responded to this comment in response to reviewer 1, point #4

      We agree that multiple binning tools can potentially identify more bins, contigs can be mistakenly binned when using purely automated processes. To sidestep the chance of including erroneous contigs in our collection, we chose to invest a large amount of time and effort required to bin manually. CONCOCT was used as an initial guide to identify some of the most easily reconstructed MAGs but all contigs within MAGs were subjected to visual inspection of the sequencing coverage profile over several samples, sequence composition congruency, and real-time completion and contamination estimation as contigs were manually added and/or subtracted from the collection in each bin. This approach is not more widely applied because of the expertise and time required to manually reconstruct MAGs from each sample individually. Many researchers find that human involvement is often required to improve the accuracy of automated binning tools, which has given rise to several tools in addition to Anvi’o, BinaRena, ICoVeR, and ggKBase. To elaborate on this procedure, we made a video tutorial of our approach to binning manual-binning-approach. An example of the manual bin refining process can be found here (merenlab-MAG-refinement)

      (2) Performing diversity analysis of metabolic pathways in surface vs deep sediments could help prove the idea of more versatility in the surface marsh sediment.

      To address this comment, we created a figure detailing the proportion of MAGs containing central microbial functions among the modules with distinct depth distributions (Fig. 5), analyzed the number of complete pathways (Fig. 4, Fig. S5) and differences in CAZyme and peptidase genes (Fig. S6).

      (3) To strengthen the idea that burial is key in deep marine sediments, looking at the abundance of mobility structures (i.e., flagella) along the sediment profile could help make a case.

      This is an interesting idea, however, many of the same genes used for motility are also used to form biofilms. This analysis could be an interesting aspect of these sediments that warrants an independent study.

      (4) Lines 60-63, but oxygen leaked from macrophytes could also exacerbate microbial respiration rates.

      This point is clarified on lines 59-64.

      (5) Line 49-52: Check this sentence. It does not read properly. "2this" -> "this".

      We corrected this error on lines 45-47

      (6) Line 173: I assume >75% is only for completion, not contamination, right? I would rephrase to make this clear.

      This point is clarified on lines 201-206.

      (7) Line 179, what is Bacteria_71? Citation is needed.

      A citation has been added for this single-copy gene collection on line 209.

      (8) Line 201: When performing log transformations, how are 0s, which are highly abundant in microbiome data, handled?

      The details of matrix transformation are clarified on line 252-254. However, unlike amplicon data (ASVs and OTUs), the read recruitment data contained very few zeros in the dataset.

      (9) Line 417: The Pseudolabrys MAG most likely utilizes the dsr and apr genes in reverse, as is typical for other Alphaproteobacteria, not for sulfate reduction. Additionally, sat can also function in assimilatory sulfate reduction.

      Yes, we agree with the reviewer’s assessment of Pseudolabrys. However, during revision, the section regarding MAGs with bimodal distribution was removed.

      (10) Line 414 uppercase mags.

      During revision, the section containing this error was removed.

      (11) Line 605: "examined".

      This error was corrected on line 682.

      (12) Line 610: Are independent MAGs from adjacent sediment layers more similar to each other than those from layers that are farther apart? Demonstrating such a pattern could provide stronger evidence that burial acts as a major force shaping the assembly of the deep sediment microbiome.

      We conducted a community-level analysis based on MAG relative abundance (Fig. 3) and a population analysis of SNVs for six abundant MAGs to test for an effect of depth (Fig. 8). While we no longer focus our analysis on a mechanistic explanation of the depth-dependent patterns, with the exception of lines 769-772. As requested by the reviewers we discuss the observed patterns in light of several other potential mechanisms (lines 838-845)

      (13) While acknowledging that selection pressure following burial is a major driver of microbial community assembly in marine sediments, I would ask whether other biogeomorphic forces could also influence the assembly of sediment microbiomes over millennial timescales, such as bioturbation or sediment redistribution within the salt marsh environment.

      Based on our community and population-level analysis, we suggest that bioturbation and sediment redistribution are likely to shape the upper 40 cm. Beyond this depth, we did not find evidence for major restructuring events. This topic is discussed on lines 830-845, but the formation of distinct relative abundance peaks for the modules primarily found in deeper sediments remains an outstanding question (lines 772-774).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Torpor can be induced by chemogenetic activation of the medial preoptic area. This activation leads to protection from myocardial infarction in an isolated heart preparation despite normalization of the ambient temperature, thus, in principle, uncoupling hypothermia from torpor-induced neuroprotection. Putative pathways of protection are suggested by proteomic studies.

      Strengths:

      (1) Elegant strategy for inducing torpor in rats.

      (2) Appropriate controls for verifying the neuron transducer.

      (3) Cardiac protection is significant and appears independent of hypothermia.

      (4) Interesting omic strategy to begin to find established and novel pathways mediating organ autonomous torpor-induced protection.

      We thank the reviewer for their positive feedback and review of our study.

      Weaknesses:

      (1) The study would benefit from using inhibitory chemogenetics of the same neurons to demonstrate that this might make cardiac response to ischaemia worse.

      This is an interesting experiment proposal and is something we will consider for the future. It is not immediately clear to us whether the hypothesis would be that inhibiting these neurons would make cardiac ischaemia worse, or that it would be similar to the control groups as we would not be inducing synthetic torpor.

      (2) Infecting an area of the brain not known to be involved in torpor would be a useful control.

      We chose to infect neurons in the same brain region without the presence of the HM3DGq DREADD construct as our control. This is a well-established control for this type of experiment.

      (3) In vivo cardio protection seems essential as the validation of the strategy requires support that is in the intact animal.

      We agree that an in-vivo model is important to establish whether the same cardioprotection occurs in the intact animal and whether it is observed when synthetic torpor is induced following an ischaemic injury. Our data suggests that synthetic torpor pre-conditions the cardiac tissue, prior to isolation of the heart, as the protection remains post removal of the heart and without any ongoing in-vivo systemic signal/drive. This provides some evidence that this may be used as a pre-conditioning model for organ protection.

      (4) The assumption that the positive effects of torpor are mediated via a phosphoproteomic change rather than a translational or transcriptional control mechanism is not established.

      Cardiac tissue for the proteomics experiment was collected following 90 minutes of synthetic torpor. This time point was selected for several reasons, including: in-vivo body temperature and heart rate reach their nadir at this time point, and we collected the hearts for the Langendorff ischaemia-reperfusion injury data set at 90 minutes and wanted to keep this consistent.

      While we tested for both total protein and phosphoprotein changes, the only significant differences occurred in the phosphoprotein levels. At 90 minutes, the predominant mechanisms for cellular responses are phosphorylation changes – somewhat too early for changes in transcription or translation. In future, it would be of interest to look at later time points to determine whether changes have occurred at the transcription/translation level.

      (5) A 40 percent reduction in infarct size may work for genetically identical rats with no co-morbidities, but is unlikely to be significant enough to weather the variability that emerges in humans because of these differences and more. The question is not what the mechanism is, but how do we make it more robust? Overall, this is at best a preliminary data set that requires more experiments to deliver on its immense promise.

      We agree with the reviewer that these are preliminary findings, as this is the first demonstration of a torpor-like state in a species that does not naturally enter torpor is protective. However, understanding the mechanisms by which the cardioprotection occurs is going to be important if we want to replicate this state in humans and to understand whether this is specific to cardiac tissue or if it extends to multiple different organs. Furthermore, by understanding the mechanisms, this may reveal avenues that would allow us to enhance these protective effects and make the method more robust.

      Reviewer #2 (Public review):

      Summary:

      Elley and colleagues induced a synthetic torpor-like state in rats (a non-hibernating species) by chemogenetically activating neurons in the medial preoptic area of the hypothalamus. They show that this state substantially reduced cardiac infarct size in an ex vivo ischaemia-reperfusion model. They further report that protection persisted when ambient temperature was raised to prevent hypothermia, and used exploratory phosphoproteomics to identify candidate cardioprotective signaling pathways.

      Strengths:

      This is the first demonstration that a torpor-like state is cardioprotective in a species that does not naturally enter torpor, which meaningfully advances the potential clinical utility of synthetic torpor. The experimental design is logical, and the controls are generally appropriate. The characterisation of the responsible neuronal population using ISH against QPLOT markers adds mechanistic depth and supports the cross-species conservation argument. The phosphoproteomic analysis, though exploratory, generates plausible and biologically coherent hypotheses grounded in the hibernation literature.

      We thank the reviewer for their positive assessment of our work.

      Weaknesses:

      The primary weakness is that the central conclusion - that hypothermia is not necessary for cardioprotection - exceeds the evidence. The thermoneutral groups were not demonstrably normothermic (36.4 vs 37.05{degree sign}C, p=0.44 with n=6), core temperature telemetry was absent in the majority of control animals contributing to the infarct endpoint, and the decisive test, i.e., a correlation between individual nadir temperature and infarct size, was never performed. Additional weaknesses include the absence of sex-stratified analysis despite known estrogenic contributions to torpor

      The core body temperature in the thermoneutral synthetic torpor animals was ~ 0.6°C cooler than the control animals housed at room temperature, which was not significantly different. A 0.6°C body temperature reduction is akin that the changes observed around a sleep-wake cycle and we consider it to be biologically implausible that this degree of temperature reduction in the ‘thermoneutral synthetic torpor’ group is sufficient to account for the observed cardioprotection. We have since added a graph and correlation analysis of individual surface temperature against infarct size in the synthetic torpor animals (see supplementary figure 5). This demonstrates no correlation between nadir temperature during synthetic torpor and infarct size, supporting the hypothesis that the protective effect is not driven by cooling.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Elley and colleagues describes experiments on the effects of synthetic torpor on ex vivo heart ischaemia. The key aspect of the study was the use of viral-vector mediated manipulation of the hypothalamic medial preoptic area (MPA) in rats. They used AAV-CaMKIIa-hM3D(Gq). The authors report that chemogenetic activation of the MPA prior to an ex vivo heart ischaemia-reperfusion insult induces cardio protection against infarct size that is independent of prior in vivo hypothermia. Phosphoproteomic analysis of cardiac tissue suggested changes in cell survival and death pathways.

      Strengths:

      This study has important strengths. The idea is novel. The experimental design is appropriately rationalized and fascinating. The manuscript is written and presented concisely.

      We thank the reviewer for their positive evaluation of our paper.

      Weaknesses:

      The study has important weaknesses in the experimental design and validation of the model.

      (1) The study is based on the use of a DREADD-designed viral vector (AAV-CaMKIIa-hM3D(Gq) -mCherry) that is activated by 2 mg/kg IP injection of CNO. The rationale is to putatively activate the MPA. The authors show no evidence for chemogenetic activation of neurons in the MPA. This could be done using a variety of different approaches, even phosphoproteomics.

      Chemogenetics is a very well-established technique and we demonstrate histological data showing expression of the DREADD in the MPA region and large physiological changes in body temperature, heart rate and oxygen consumption that we feel is sufficient to demonstrate that these neurons are activated. The field of chemogenetics is sufficiently established that it is not common practice to demonstrate neuronal activation in the context of large induced physiological responses.

      (2) The stereotaxic injections are difficult to precisely and locally place, particularly bilaterally. Figure 2F is only a schematic. It would be better to show actual low magnification brain sections (bregma +0.12 to -0.48) from a representative rat to show the placement of the AAV.

      Figure 2F is a schematic produced from imaging data, showing the regions within the MPA that expressed the AAV across all animals. We mapped the areas of viral transfection in each animal and then mapped this onto the brain atlas demonstrating which regions were consistently transduced across all animals. We have added a supplementary figure (supplementary figure 2) showing overlaid sections from 6 animals that underwent detailed vector mapping, which shows the full extent of vector transduction in these animals.

      (3) The control rats were injected with AAV-CaMKIIa-EGFP. Why was EGFP used instead of mCherry for the control?

      We prioritised using the same serotype and promoter for our experiments and would not anticipate the colour of fluorophore to have an effect on experimental results.

      (4) Ideally, a mutant non-activatable variant of AAV-CaMKIIa-hM3D(Gq) should have been used for a better control.

      We thank the reviewer for this suggestion, however, to the best of our knowledge, a viral vector supplier Addgene, does not stock this variant of virus, nor have we seen one reported in the literature.

      (5) The authors should comment on whether there is any neurotoxicity in the MPA associated with the forced AAV expression of hM3D-Gq.

      This is not something we have assessed, however, administration of AAV into the CNS is not associated with a robust inflammatory response or neurotoxicity [1]. We used a viral titre of 1.7x10<sup>13</sup> viral genome copies per ml and injected 200 nL at each POA site, which is within reported guidelines by other users of the Addgene construct.

      (1) Mastakov MY, Baer K, Symes CW, Leichtlein CB, Kotin RM, During MJ. Immunological aspects of recombinant adeno-associated virus delivery to the mammalian brain. J Virol. 2002;76(16):8446-8454. doi:10.1128/jvi.76.16.8446-8454.2002

      (6) Is there any inflammatory pathology seen in the MPA with AAV transduction?

      This is not something that we assessed.

      (7) There are no experiments to show that the systemic torpor is specifically associated with the MPA region. Experiments should be done with injections of AAV-CaMKIIa-hM3D(Gq)-mCherry placed in other brain regions, for example, the nearby nucleus accumbens.

      The medial preoptic area is widely believed to be key for triggering torpor. Data from our lab and other groups has demonstrated that neurons within the preoptic area are sufficient to trigger torpor entry in mice. However, we concede that it is possible that other brain regions form part of the same circuit and that activation of the circuit at this / these points could also trigger a synthetic torpor bout. Should this be the case, our key message would remain the same: animals that do not naturally enter torpor can be induced into a similar state through targeted CNS activation, and this state is cardioprotective.

      (8) The mapping of the distribution of neurons responsible for synthetic torpor is not mechanistic enough and is not directly to the point. While excitatory and inhibitory markers are examined, a more interesting and deeper approach would have been to use glutamate receptor antagonists to manipulate the torpor response.

      We thank the reviewer for this interesting suggestion. This type of experiment is not commonly performed and would be extremely technically challenging. A more common approach would be to genetically access glutamatergic neurons in the region and silence them with chemo or optogenetics.

      (9) The ischaemia and reperfusion aspects of the Langendorff method need to be clarified. The isolated hearts are already ischaemic after their removal from the rat. The reperfusion aspect is caused by reflow of blood to generate oxidative stress, but in the ex vivo model, is there really reperfusion injury?

      The Langendorff model is a well-established ex-vivo technique where the excised heart is perfused retrogradely through the aorta. This allows oxygenated buffer that also contains metabolic fuel to be pumped through the aorta, forcing the aortic valve closed and diverts the buffer directly into the coronary arteries to sustain the organ outside of the body. By temporarily stopping buffer perfusion then reinstating it as we did here, we can mimic a full ischaemia-reperfusion injury. It has been commonly used since the late 19th century to study heart physiology and pathophysiology including ischaemia-reperfusion injury.

      We agree with the reviewer that the hearts may have experienced a degree of ischaemia when being removed, unfortunately that is one of the confounds associated with ex-vivo preparations such as the Langendorff model. However, it should be noted that both the control and synthetic torpor hearts were exposed to the same conditions following removal from the animal. To minimise ischaemic damage, all hearts were placed in ice-cold buffer immediately following excision. Preparations were excluded from our data if a) cannulation took longer than 3 minutes, b) more than 2 cannulation attempts were made and c) if the heart rate was slower than 200 bpm during the baseline 30-minute recording. This criteria is in accordance with other laboratories and is frequently reported in the literature.

      (10) The authors show that whole animal oxygen consumption is reduced in the torpor state. The measurement is crude and most likely reflects the inactivity of the animal's skeletal muscle in the torpor state. A more relevant and direct experiment would be to do oxygen consumption (or Seahorse) assays on extracts of the isolated hearts.

      We did not simultaneously measure oxygen consumption alongside animal activity so we cannot determine whether muscle inactivity is a contributing factor in this data. However, previous work on mice has shown that VO<sub>2</sub> to be decreased by 15% during the light phase (mouse inactivity/sleep period) compared to the dark phase (their active phase) [1]. Whereas during fasting-induced torpor, a ~45% decrease in VO<sub>2</sub> has been reported [2], which is consistent with what we observed during synthetic torpor in rats (~39% decrease in VO<sub>2</sub>).

      Furthermore, we simultaneously recorded activity of our rats while recording core body temperature and heart rate before and during synthetic torpor. We have since updated the manuscript to include this data as a supplementary figure (supplementary figure 6), showing that there is no significant difference between the activity of animals in synthetic torpor compared to controls (a curious contrast with natural torpor).

      We also agree with the reviewer regarding the seahorse experimental suggestion and plan to use the seahorse assay to measure oxygen consumption in isolated cardiomyocytes and cardiac tissue punches in the future.

      (1) Nie Y, Gavin TP, Kuang S. Measurement of Resting Energy Metabolism in Mice Using Oxymax Open Circuit Indirect Calorimeter. Bio Protoc. 2015;5(18):e1602. doi:10.21769/bioprotoc.1602

      (2) Hrvatin S, Sun S, Wilcox OF, et al. Neurons that regulate mouse torpor. Nature. 2020;583(7814):115-121. doi:10.1038/s41586-020-2387-5

      (11) The authors report that the synthetic torpor induces bradycardia. There is no follow-up on this important observation. The MPA-heart connection is not analyzed. (A) Is the link through cardiovascular centers in the brainstem? (B) Is the torpor-induced bradycardia mediated through increased parasympathetic or decreased sympathetic autonomic tone? Pharmacological experiments could also be done.

      We thank the reviewer for these suggestions and agree that these will be important follow-up experiments to understand how the cardioprotection is mediated. We currently have a grant under review that will perform these important experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments or analysis:

      (1) Chemogenetic inhibition of the MPO.

      This is an interesting experiment proposal, we chose to focus on the activation of these neuronal populations to establish whether activation would be sufficient to drive synthetic torpor in a species that does not naturally enter torpor. Inhibiting these neurons would be an interesting avenue to explore in the future.

      (2) In vivo validation of torpor-induced protection.

      We plan to use an in-vivo cardiac ischaemia-reperfusion injury model in future experiments and this forms a work package in a grant we currently have under review.

      (3) A search for the dose and timing of torpor that provides the most robust protection.

      Using the in-vivo cardiac ischaemia-reperfusion injury model described above, in future we aim to determine whether synthetic torpor is protective following the ischaemic episode, which would better model a clinical scenario such as when a patient presents with a heart attack.

      (4) Use of inhibitors of transcription and translation as a strategy to understand what are the nodes of regulation that mediate neuroprotection.

      An interesting suggestion. We might consider this in future studies but feel it is beyond the scope of this manuscript.

      (5) Testing in males, females, and rats of different genetic backgrounds.

      This data set includes both male and female rats, however we did not power our experiments to determine sex differences. In future, we will plan to power our studies to allow determination of sex differences, and to use rats from different genetic backgrounds.

      (6) Use of chemical or molecular tools to assess the relevance of phosphoproteomic data.

      Our future experimental plans, which are currently under review in a grant application include the use of inhibitors and activators of the kinases we identified in our phosphoproteomic data, to determine whether these can modulate the infarct size in the ischaemia-reperfusion injury model and modulate the cardioprotective effect from synthetic torpor.

      Reviewer #2 (Recommendations for the authors):

      (1) Clarifying core body temperature measurement:

      The authors report post-CNO core body temperature (mean, 90-100 minute window) for four groups: room temperature MPAGq (31.64 {plus minus} 1.02{degree sign}C, n=10), room temperature MPAEGFP (37.05 {plus minus} 0.43{degree sign}C, n=6), thermoneutral MPAGq (36.4 {plus minus} 0.73{degree sign}C, n=6), and thermoneutral MPAEGFP (37.05 {plus minus} 0.43{degree sign}C, n=6). Please clarify the provenance of the MPAEGFP temperature data, as the room temperature and thermoneutral MPAEGFP groups report identical means to two decimal places (37.05{degree sign}C) with identical standard deviations ({plus minus}0.43{degree sign}C). It should also be noted in the limitations that continuous core temperature telemetry was available in only a subset of animals - specifically, 10 of 12 room temperature MPAGq animals and 6 of 16 room temperature MPAEGFP animals that contributed to the primary Langendorff infarct endpoint. It is correct that this means that for the majority of room temperature MPAEGFP control animals (10 of 16), no core temperature measurement of any kind was paired with their infarct data? If so, this should be stated as a limitation. Additionally, pre-CNO baseline temperatures are not reported numerically for any group, and individual nadir temperatures across the 90-minute induction period are not provided, meaning the full thermal exposure of individual animals cannot be characterized from the data as presented.

      We thank the reviewer for these suggestions, and we have since corrected the in-text error for the MPA<sup>EGFP</sup> thermoneutral and room temperature data. While we did not have the in-vivo core body temperature in all of our animals, we did record surface body temperature if core body temperature was absent and we have since plotted this surface body temperature data against the infarct size and performed a correlational analysis. We found no significant correlation between the two. We have since also updated our supplementary figure 1 to include the pre-CNO baseline temperature and heart rates for all animals in the synthetic torpor and control groups and those exposed to room temperature or thermoneutral environments.

      (2) Conclusions exceed evidence:

      While the demonstration that chemogenetic MPA activation under thermoneutral conditions is associated with reduced infarct size is an interesting and valuable finding, the conclusion, as stated, that "the hypothermic component of torpor is not necessary for the protective effects of synthetic torpor" (lines 303-304), goes beyond what the evidence can support. The cardioprotective effect observed in the thermoneutral condition may indeed reflect the contribution of other (non-temperature) factors - but the involvement of a temperature effect cannot be ruled out on the basis of the data presented. The thermoneutral MPAGq group, although not significantly different from the thermoneutral MPAEGFP control in core temperature (36.4 {plus minus} 0.73{degree sign}C vs. 37.05 {plus minus} 0.43{degree sign}C, p=0.44), had a mean core temperature that was in fact 0.65{degree sign}C lower than controls, and with n=6 in each group the non-significance of this difference reflects limited statistical power to detect equivalence as much as true normothermia. Given the limitations in telemetric coverage described above, a non-significant difference between group means is not sufficient to conclude that hypothermia is categorically uninvolved in mediating the cardioprotective effect. The appropriate test would be to demonstrate an absence of correlation between individual core body temperature nadir and individual infarct size across all animals in both the room temperature and thermoneutral conditions - a continuous analysis that the current telemetry coverage may not fully support, but that would directly address the question. We would therefore ask the authors either to provide such a correlation analysis if the paired data permit it, or to temper their conclusion accordingly - stating that factors other than hypothermia are sufficient for cardioprotection, rather than that hypothermia is not involved, which is the stronger and less well-supported claim.

      We agree with these suggestions and have adjusted our conclusions within the text. We have also performed a correlational analysis between the surface temperature of the animals and the infarct area, which we have included in the supplementary figures. We found no significant correlation between the two, which further supports our conclusion that other factors are likely to contribute to the cardioprotection observed, rather than the hypothermia. We feel that it is biologically implausible to suppose that the 0.6°C drop in body temperature observed in the thermoneutral synthetic torpor group is sufficient to recapitulate the cardioprotection observed.

      (3) Use of the word torpor: The statement:

      "this cardioprotection induced by synthetic torpor persisted in the absence of hypothermia"

      Presupposes that the thermoneutral condition is still synthetic torpor, just without one feature. But the authors have no basis for that framing - they have not demonstrated that the thermoneutral condition preserves the metabolic suppression component, and they have in fact abolished the temperature component that is part of the very definition they invoke. A more accurate statement would be:

      "Chemogenetic activation of MPA neurons under thermoneutral conditions, which prevents hypothermia and whose effect on oxygen consumption was not assessed, still produced a reduction in infarct size."

      We thank the reviewer for this suggestion, and we have amended the text in our manuscript to reflect this.

      (4) Lack of data addressing possible sex differences. Given the known role of estrogen in contributing to torpor, a more thorough analysis of sex differences beyond underpowered stratification is needed. In the absence of this analysis, this should be noted as a limitation.

      The preoptic area of the hypothalamus is enriched with estrogen receptors and previous work by our group has demonstrated that the estrous cycle modulates fasting-induced torpor. Our current work did not aim to investigate sex differences and is not powered to do this. As a result, we have now added this limitation to our discussion.

      (5) QPLOT marker characterization is incomplete. The QPLOT classification requires co-expression of markers including Qrfp and Tacr3, but the authors only tested three of the five (Ptger3, Lepr, Opn5). Ptger3 was expressed in 100% of transduced cells, but the question is whether Ptger3 expression alone is sufficient evidence of QPLOT identity without testing Qrfp and Tacr3. The conclusion that these neurons are "analogous" to mouse QPLOT neurons is reasonable but somewhat circular, given the marker selection. The assertion that 29% expressed "all three QPLOT markers" should be contextualized as three of five.

      We have updated our manuscript to state that our assessment is for three of the five QPLOT neuron markers.

      Reviewer #3 (Recommendations for the authors):

      Other comments:

      (12) The figure legends should state the group sizes (animal numbers) for the graphical data.

      We have added this information to the figure legends.

      (13) The control and experimental groups do not appear to be balanced. Were any animals removed from the treatment groups?

      Some MPA HM3DGq injections failed and the animals did not enter synthetic torpor following CNO injection. These were excluded from the study (the success rate for inducing synthetic torpor was 80%). Some animals were excluded from the ischaemia-reperfusion study, due to time from heart excision to cannulation on the Langendorff apparatus being more than 2 minutes, more than 2 cannulation attempts or a heart rate of less than 200 bpm at the end of the equilibration period. 8 hearts were excluded from the heart infarction data set based on these criteria. We have updated the methods section of our manuscript to emphasise this.

      (14) With small group sizes, it is better to show the data set variation with standard deviations rather than SEM.

      We thank the reviewer for noting our mistake and have corrected it and/or reported individual data points.

      (15) The precise P value for the infarct size reduction with torpor should be stated rather than p<0.05.

      We have updated the manuscript and the figure legends to include the precise P values.

      (16) Is there translational relevance of this work using an in vivo model of bona fide heart ischaemia-reperfusion injury?

      We absolutely plan to complete those experiments, but feel they are beyond the scope of this initial reporting of the principle that synthetic torpor in the rat is cardioprotective.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This work by Pyne and Pandey et al. addresses DNase X (DNase1L1) activity at the macrophage phagocytic cup, using an innovative imaging approach that couples visualization of cup formation to spatially resolve DNA degradation. The methodology is technically sound, and the central finding that DNA digestion begins prior to phagolysosomal maturation is considered well supported, though some mechanistic claims may benefit from further evidence and more cautious framing. Overall, the study is solid and provides a valuable framework for investigating early events at the phagocytic cup that may shape responses to pathogens and inflammatory disease.

      Thanks for the overall review.

      Public Reviews:

      Reviewer #1 (Public review):

      Pyne and Pandey et al. report the observation of early DNA degradation at the phagocytic cup during macrophage engulfment. Using an elegant experimental system that combines actin staining to visualise cup formation with direct monitoring of DNA degradation, the authors identify rapid recruitment of the membrane-bound nuclease DNase X (DNase1L1) to nascent phagocytic cups. This recruitment occurs within minutes of cup formation, is independent of DNA presence at the substrate, and appears to originate from intracellular membrane structures rather than from the extracellular environment. The results support the conclusion that DNase X activity is present at the phagocytic cup and that DNA digestion can begin prior to phagolysosomal maturation.

      The study is technically strong. The experimental system is clean, specific, and allows precise spatial and temporal detection of DNA degradation. The imaging-based approaches are carefully executed and enable convincing visualisation of DNase X recruitment and activity. The use of an alternative substrate beyond the primary SNS system strengthens the core observation, and the data broadly support the authors' central claim.

      Thanks for the positive comments.

      However, several limitations temper the physiological interpretation. The system relies largely on short, free DNA substrates, leaving open how efficiently DNase X processes more complex or physiologically relevant DNA structures, such as nucleosome-bound DNA or neutrophil extracellular traps (NETs). It remains unclear whether DNase X deficiency would alter macrophage responses to larger nucleic acid structures, influence engulfment efficiency, or modify downstream inflammatory signalling pathways such as TLR9 or STING activation.

      Thanks for the comments. It is a constructive suggestion to test how DNase X in phagocytic cups (PCs) may degrade long DNA or chromosomal DNA, as the current experiment setup mainly used 18bp DNA as the DNase sensor construct. Now we included new data of DNase response of macrophage to plasmid DNA (extracted from E. coli) immobilized on microbead surface. The result show that macrophage also degrade the plasmid DNA, suggesting that the DNase in PC is rather versatile and can degrade long DNA. This result is included as fig. S5.

      It would be certainly interesting to explore the role of DNase X in phagocytic efficiency and downstream inflammatory pathways. While these could be the future research directions, the goal of current manuscript is to demonstrate a new DNase activity in PCs and calibrate its several core features (temporal and spatial dynamics and response to biofilm, etc.). Expanding the topic further will require substantial resource and effort, and may make the current manuscript bloated with data. Therefore, we didn’t pursue the study of DNase of PCs in the context of phagocytosis efficiency and downstream pathways.

      Moreover, the experimental setup prevents full phagocytic cup closure, potentially prolonging DNase activity compared with physiological phagocytosis, which typically proceeds rapidly to cargo internalisation. For example, the peak signal observed in Figure 5 occurs approximately 90 minutes after phagocytic cup formation, a time point at which many phagocytic cups would be expected to have already closed under physiological conditions.

      Additional work using fully engulfed cargo in more physiological contexts would clarify whether early DNase X activity meaningfully contributes to overall DNA clearance kinetics.

      Real-time imaging of SNS signal and F-actin in macrophages, shown in Fig. 1I in the manuscript, suggests that DNase activity in PCs appear within one minute after PC formation. We do resonate with the reviewer’s concern about the experiment platform used in this work, namely, the microbeads are immobilized on the glass surface, preventing the natural closure of PCs. This artificial design may create biased observation of DNase activity during phagocytosis.

      In light of this concern, we developed a new assay by using free SNS-coated microbeads (non-immobilized). These beads can be completely internalized by macrophages through phagocytosis. During experiments, we searched for the PCs just starting to form over free SNS-coated beads. Despite the rare occurrence of these events (phagocytosis process has a rather short time windows), we did find the evidence that PCs already exhibit DNase activity before their closures, as shown in fig. S8.

      Accordingly, we added one paragraph to the main text, “Because surface-immobilized microbeads prevent PC closure, they likely prolong PC formation and may alter the observed temporal dynamics of DNase activity within the PC. To determine whether the PC exhibits DNase activity before closure, we prepared SNS-coated free microbeads and fed them to adherent macrophages. The results showed that PCs indeed exhibited DNase activity before cup closure in response to free microbeads (fig. S8), confirming that DNase activity is initiated rapidly during PC formation.”

      Mechanistically, the signal that triggers DNase X recruitment remains unresolved. Although actin rearrangement was excluded as the primary driver, the upstream cues that direct DNase X-containing membrane structures to the forming cup are not yet defined.

      We agree that the biochemical signals responsible for initiating DNase X recruitment remain unresolved. Despite our efforts to identify such signals by exposing macrophages to a variety of immunogenic stimuli, we consistently observed DNase X recruitment to phagocytic cups under all conditions tested. Based on these findings, we conclude that DNase X recruitment to phagocytic cups is constitutive rather than stimulus-dependent.

      Nevertheless, even if this recruitment is constitutive, it is likely that an intracellular biochemical cue initiates or regulates the recruitment process. Unfortunately, we have not yet been able to identify this upstream signal. We have acknowledged this limitation and included this consideration in the Discussion as an important direction for future investigation.

      In previous studies, we also examined DNase activity at phagocytic cups following inhibition of the cGAS-STING and TLR9 pathways, two major cellular DNA-sensing mechanisms. However, inhibition of either pathway did not alter DNase activity at phagocytic cups. As these experiments yielded negative results and did not provide mechanistic insight into DNase X recruitment, we did not include them in the current manuscript.

      In the broader context, early DNase X activity at the phagocytic cup could represent an additional safeguard against inflammatory signalling by limiting extracellular or surface-associated DNA before phagolysosomal degradation by DNase II. This mechanism may be particularly relevant in settings where DNA fragmentation before engulfment is incomplete, such as necroptosis or NET formation. Determining whether DNase X deficiency exacerbates inflammatory responses, alters DNA clearance efficiency in vivo, or contributes to immune pathology will be critical for establishing its physiological and disease relevance.

      Thanks for the constructive comments which provide valuable suggestion for the future work.

      Overall, this is a compelling study that introduces a novel concept of pre-phagolysosomal DNA digestion. The conclusions are well supported within the in vitro system used, but further investigation using diverse DNA substrates and physiologically relevant models will be required to fully define the impact of this mechanism on immune regulation and disease.

      We appreciate the insightful comments and agree that in vivo experiments would provide valuable information for future studies. The current manuscript focuses on a cellular function identified using an in vitro experimental system. The results suggest that DNaseX may play important physiological roles in immune defense and extracellular DNA clearance, highlighting the need for further investigation using animal models or other in vivo systems. These are exciting directions that we are actively planning to pursue in our future research.

      We are not able to incorporate in vivo studies into the current revision, as such experiments would require the establishment of fundamentally different experimental models, the acquisition of new technical expertise, and substantial additional financial support. These efforts are beyond the scope and timeline of the present manuscript revision. We believe that the current in vitro findings provide a foundation for these future in vivo investigations.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an elegant and innovative imaging approach to visualize DNase activity at the interface between macrophages and extracellular substrates. The platform is technically strong and enables the study of localized DNA degradation with high spatial resolution. The work is of clear interest and provides a useful framework to investigate how immune cells process extracellular DNA. However, several aspects of the mechanistic interpretation and conceptual framing would benefit from clarification.

      Strengths:

      (1) The study introduces a creative and well-designed imaging platform that allows visualization of localized DNase activity at cell-substrate interfaces.

      (2) The approach is technically robust and represents a valuable tool that could be broadly useful to the field.

      (3) The experiments are thoughtfully designed and address an important question regarding how immune cells interact with extracellular DNA.

      (4) The work opens interesting avenues for studying DNA processing in contexts such as infection and inflammation.

      Thanks for the positive comments.

      Weaknesses:

      While the experimental approach is strong, several key conclusions rely on interpretations that would benefit from further clarification:

      (1) First, the conclusion that DNaseX is recruited to phagocytic cups from the "cytoplasm" appears conceptually imprecise. Given that DNaseX is a membrane-anchored protein, it is unlikely to exist as a freely soluble cytoplasmic pool. A more plausible interpretation is that DNaseX is supplied from intracellular membrane compartments. This interpretation would also be more consistent with the data showing dependence on a membrane anchor.

      We thank the reviewer for this insightful comment. We agree that, as a GPI-anchored protein, DNaseX is likely transported from the endoplasmic reticulum and Golgi apparatus to the plasma membrane via secretory vesicles carrying the protein. The original subtitle “DNaseX in PCs is recruited from the cytoplasm, not from the plasma membrane” in the manuscript, is misleading, as it may imply that DNaseX is a cytoplasmic soluble protein. Our original point is to state that DNaseX is recruited intracellularly, not from the adjacent cell membrane domain. To avoid the ambiguity, we have revised that subtitle to “DNaseX in PCs is recruited intracellularly, not from the plasma membrane”. In the revision, we also emphasized the that DNaseX is likely delivered by the vesicles, which were observed in Fig. S13.

      (2) Second, the interpretation that actin polymerization is not required for DNaseX recruitment raises concerns. Phagocytic cup formation is known to depend strongly on actin dynamics, and it is therefore unclear whether the structures observed under actin inhibition represent fully formed functional cups or partial cell-substrate contacts. This distinction is important for interpreting recruitment versus activity, particularly since enzymatic activity is reduced under these conditions.

      We appreciate the reviewer’s comment and agree with this point. The data presented in the manuscript demonstrate the relationships, but not causalities, among F-actin signal, DNaseX signal (detected by immunostaining), and SNS signal (reflecting DNase activity) in PCs. Specifically, our results show a strong correlation between the F-actin and SNS signals, whereas the correlation between the F-actin and DNaseX signals is weak. These correlations do not warrant the causality relations.

      Experimentally establishing causality is challenging because F-actin is an essential structural component of the phagocytic cup and cannot be completely eliminated from PCs. To ensure rigorous interpretation of the data, we have revised the manuscript to present the findings more cautiously and avoid implying a causal relationship where it has not been experimentally demonstrated.

      We revised the section title: “Actin polymerization is dispensable for DNaseX recruitment but essential for DNase activity in the PCs” to “Actin polymerization is correlated with DNase activity in PCs, but not with DNaseX recruitment”. We also revise the abstract and the main text correspondingly to reflect this point.

      (3) Third, the identification of DNaseX as the main nuclease responsible for the observed activity is not fully resolved. The conclusions rely primarily on gene silencing and staining approaches, but the specificity of these strategies relative to other nucleases is not addressed. It therefore remains possible that additional enzymes contribute to the observed activity.

      The reviewer raises an important point. We acknowledge that gene silencing and immunostaining may not be entirely specific, as they could potentially affect or detect other members of the DNase family that share structural similarities with DNaseX. Indeed, the official name of DNaseX, Deoxyribonuclease-1-like 1, suggests a close structural relationship with DNase I.

      Our conclusion regarding the specificity of DNaseX is based on the combined evidence from multiple independent approaches rather than any single experiment. In addition to gene silencing and immunostaining, we demonstrated that cleavage of the GPI anchor with PI-PLC markedly reduced or completely abolished DNase activity in PCs. This experiment provides particularly strong evidence because DNaseX is the only known membrane-bound DNase that is anchored to the plasma membrane via a GPI linker.

      Taken together, the complementary results from GPI-anchor cleavage, gene silencing, and immunostaining provide convergent evidence supporting the conclusion that DNaseX is the primary enzyme responsible for the DNase activity observed in PCs. We therefore believe that the combined data provide sufficient specificity to support our conclusion.

      (4) Finally, the interpretation of the biofilm experiments may be overstated. While the data clearly show localized DNA degradation in contact with macrophages, it is not fully established that this process depends specifically on phagocytic cup structures. An alternative explanation is that membrane-associated DNase activity more generally mediates this effect. In addition, the physiological relevance of this mechanism would benefit from further discussion.

      We agree with the reviewer on this point. Although our data demonstrate that macrophages degrade eDNA in biofilms through direct physical contact, they do not establish that phagocytic cups are the primary structures responsible for mediating this process.

      We attempted to obtain direct evidence by simultaneously imaging F-actin structures in macrophages and eDNA degradation within biofilms, with the goal of visualizing phagocytic cup formation around eDNA filaments. However, this proved to be technically challenging. Unlike a flat glass surface, biofilms possess substantial thickness and an uneven three-dimensional topology, making it difficult to achieve high-resolution imaging of F-actin structures with sufficient clarity. In addition, eDNA filaments, unlike microparticles, may not induce the formation of canonical F-actin-rich phagocytic cups. Consequently, we were unable to obtain convincing imaging data demonstrating that phagocytic cups are the structures responsible for eDNA degradation within biofilms.

      To ensure a rigorous interpretation of our findings, we have revised the manuscript to avoid overstating this conclusion and to more accurately reflect the limitations of the current data. We added this paragraph in the revision: “While the results showed that macrophages degrade eDNA in biofilms through physical contact, it has not been confirmed that this degradation is mediated by DNaseX within the PC, as eDNA filaments are not expected to induce the formation of a typical PC structure. An alternative possibility is that DNaseX localized on the cell membrane, including within the plaque-like clusters (fig. S13), comes into direct contact with eDNA and mediates its degradation through physical contact.”

      Overall, the study is technically strong and introduces a valuable methodology, but several central conclusions are only partially supported by the current data and would benefit from more cautious interpretation and clearer conceptual framing.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The data are clearly presented, and the manuscript is well written, making it a pleasure to read.

      Reviewer #2 (Recommendations for the authors):

      (1) Clarify the description of DNaseX origin by replacing "cytoplasmic" with more precise terminology (e.g., intracellular membrane compartments), and revise the interpretation accordingly.

      Revision has been made according to the reviewer’s suggestion. Please see the point-to-point response.

      (2) Provide a clearer definition of what constitutes a phagocytic cup under actin inhibition conditions, and discuss whether the observed structures represent fully functional compartments.

      Revision has been made according to the reviewer’s suggestion. Please see the point-to-point response.

      (3) Consider including additional controls or discussion addressing the specificity of the gene silencing approach, particularly in relation to other nucleases. If feasible, assess the potential contribution or localization of other nucleases (immunofluoresce for DNase I, for example), or explicitly acknowledge this as a limitation.

      Revision has been made according to the reviewer’s suggestion. Please see the point-to-point response.

      (4) Refine the interpretation of the biofilm experiments to distinguish between phagocytic cup-specific mechanisms and more general membrane-associated activity. Expand the discussion on the physiological relevance of the findings, particularly in the context of complex environments such as biofilms or extracellular DNA structures in vivo.

      Revision has been made according to the reviewer’s suggestion. Please see the point-to-point response.

      (5) Minor: Improve consistency in terminology throughout the manuscript (e.g., membrane-associated vs cytoplasmic localization) to avoid conceptual confusion.

      We revised all “membrane-associated” to “membrane-bound” for the consistency of terminology. The confusion due to the use of “cytoplasmic” has been rectified. Please see the point-to-point response.

    1. Author response:

      We sincerely thank the reviewers and editors for their thorough and constructive comments. Our goal was to study characteristics of excitatory, inhibitory, and neuromodulatory components of voluntary motor commands in participants with multiple sclerosis. To do so, we drew on a novel reverse-engineering framework designed to estimate characteristics of these components from a selection of features extracted from motoneuron discharge patterns (Chardon et al., 2024).

      We appreciate the reviewers’ recognition of the considerable size and richness of our dataset of motor unit firing patterns recorded from the tibialis anterior and soleus of 89 participants with MS and 34 neurologically intact controls, the rigor of our statistical analyses, the novelty and importance of our descriptive findings, and the potential value of our dataset and analytical approach to other researchers.

      We address reviewer comments and discuss our planned revisions to the paper below.

      Concerns about physiological inference of findings from the reverse engineering framework. Reviewer 1, for example:

      The principal limitation concerns the physiological interpretation assigned to these discharge patterns. The excitation, inhibition, and neuromodulation variables are not directly measured physiological inputs. They are composite variables derived from several features of motor unit discharge, with each feature weighted according to relationships identified in simulations reported previously by Chardon and colleagues. Thus, the step from observed discharge behavior to specific underlying synaptic mechanisms is necessarily model-dependent. The distinction between these two levels of inference is important for interpreting the main conclusions of the study.

      The issue is especially relevant because several different physiological processes can plausibly influence the same discharge features. … The interpretation of inhibition is a particularly clear example of this general inverse problem.

      We acknowledge all three reviewers’ concerns that, as with any computational model, assumptions of the Chardon reverse engineering framework place limits on the physiological interpretation of our findings.

      However, two key strengths of the reverse engineering framework warrant further emphasis when considering the extent to which model assumptions limit physiological inference: (1) the extensive experimental foundation and physiological realism of the motoneuron model, and (2) the demonstrated ability to address the problem of non-uniqueness when predicting characteristics of synaptic and neuromodulatory inputs from discharge patterns of model motoneurons. We will discuss these strengths and clarify the physiological properties that are modeled.

      (1) Physiological foundation of the motoneuron model

      We first highlight the extensive experimental foundation underlying the motoneuron model. The Heckman laboratory has devoted considerable effort to the development of a physiologically realistic motoneuron model. The motoneuron model is grounded in over 20 years of in situ voltage-clamp experiments that characterized motoneuron intrinsic properties and their responses to excitatory, inhibitory, and neuromodulatory inputs (e.g., Lee & Heckman, 1998a, 1998b, 1999; Kuo et al., 2003; Hyngstrom et al., 2008). The resulting model motoneurons used for simulations were described by the authors as “designed to closely recreate behaviors documented in our extensive database of current and voltage clamp studies in motoneurons within animal preparations” (Chardon et al., 2024).

      Moreover, the motoneuron model did not originate with the reverse-engineering framework used in the present study. Earlier versions were refined and evaluated across multiple prior publications (Kim et al., 2009; Kim & Jones, 2012; Powers et al., 2012; Powers & Heckman, 2015, 2017; Beauchamp et al., 2023) and optimized to balance realism with computational speed.

      While the motoneuron model’s extensive experimental foundation does not eliminate the inherent limitations of model-based inference, it constrains the model substantially and reduces the extent to which model predictions depend on arbitrary or purely theoretical assumptions.

      (2) Addressing non-uniqueness in the inverse problem

      We next discuss whether characteristics of motoneuronal inputs can be inferred with reasonable accuracy from motoneuron discharge patterns. This is the inverse problem for which non-uniqueness is a fundamental challenge.

      Non-uniqueness must be considered when making physiological inferences from motor unit discharge. Many motoneuron discharge features, when examined on their own, are sensitive to changes in more than one type of motoneuronal input. For example, discharge hysteresis is affected by both the level of neuromodulation and the pattern of inhibition relative to excitation. Reviewers were concerned that the reverse engineering framework also suffered from non-unique solutions when predicting motoneuronal inputs from discharge patterns, suggesting as an example that changes to motoneuron firing patterns resulting from a change in the pattern of inhibition could also result from changes to the level of neuromodulation.

      However, we must distinguish the general non-uniqueness problem from the performance of the reverse engineering framework specifically. Non-uniqueness is not an unaddressed weakness of the reverse-engineering framework. Rather, it is the central problem that the framework was explicitly designed to address.

      The ability of the reverse-engineering framework to address and substantially reduce non-uniqueness was explicitly evaluated in Chardon et al., (2024) using simulated, known inputs to a population of model motoneurons. Characteristics of motoneuronal inputs (i.e., the distribution of excitatory inputs across motoneurons of different size, the pattern of inhibition relative to excitation, and the level of neuromodulation) were systematically varied in all combinations, generating a “library” of simulated spike trains from each model motoneuron. Seven reverse engineering features were calculated on the simulated discharge patterns and used to train linear and non-linear machine learning models. When applied to test data, the machine learning models predicted the known inputs with substantial accuracy. 

      The key finding of Chardon et al. is that the discriminatory power of the reverse engineering framework comes from the collection of discharge features used in the inverse problem. Because individual features can be sensitive to more than one type of motoneuronal input, considering the features as an ensemble drastically reduces the inverse problem’s solution space. This allows the known inputs to the simulated model motoneurons to be predicted with substantial accuracy.

      (3) Modeled and unmodeled parameters

      Having discussed the unique strengths of the reverse engineering framework, we also acknowledge its limitations, as do the authors of the approach. There are aspects of motoneuronal inputs and intrinsic properties that were held constant or not explicitly represented in the Chardon et al. (2024) simulations. Therefore, application of the current framework to human motor unit data introduces uncertainty regarding the extent to which variation in unmodeled parameters may influence discharge patterns and, by extension, the inferred characteristics of excitatory, inhibitory, and neuromodulatory inputs.

      However, we consider this uncertainty in the context of the physiological information gained by applying the framework to human data. Specifically, the framework allows us to interpret motor unit discharge patterns recorded in vivo using relationships established through systematic manipulation of physiologically grounded model parameters.

      Thus, despite its limitations, the framework provides substantially greater physiological insight than can be obtained by interpreting individual motor unit discharge features in isolation and represents an important step toward identifying the combinations of motoneuronal inputs that contribute to typical and pathological motor unit discharge in humans. Further variation in these parameters could be incorporated in future iterations of the framework, as is possible with more complex models like that of Mousa and Elbasiouny (2026).

      With regard to the inclusion of specific parameters, we discuss two points:

      (A) It is necessary to clarify that some physiological properties identified in the reviewer comments as unmodeled were, in fact, represented in the motoneuron model and simulations used in Chardon et al. (2024).

      Perhaps the most crucial is the presence of persistent inward currents (PICs), whose realistic representation is a defining feature of the motoneuron model due to its four dendritic compartments with varied Ca<sup>2+</sup> channel densities (Powers & Heckman, 2017). In fact, in Chardon et al. (2024), the PIC-induced non-linearities in simulated firing patterns were not merely present in the model, but were a key factor contributing to the success of the reverse engineering framework.

      Contributions from afferent inputs were also considered within the synaptic inputs to the model. The simulated distributions of excitatory input and patterns of inhibitory input both encompass potential contributions from afferent sources in addition to descending and spinal sources (Binder et al., 2002; Johnson et al., 2017; Chardon et al., 2024). Further, the opposing influence of inhibitory input from any source on facilitation of PICs is incorporated in the simulations.

      (B) We agree with the reviewers that it is important to consider potential between-group differences in intrinsic membrane properties that affect PIC behavior in addition to neuromodulatory input. This is a point that we did not adequately discuss in the initial version of the paper. However, we see this primarily as a lack of precision in our discussion of the physiology rather than a deficit related to unmodeled parameters in the Chardon et al. model.

      In the Chardon et al. (2024) simulations, the level of neuromodulatory input was varied with the intent of scaling PIC amplitude accordingly. This was done by changing the density of dendritic PIC channels to simulate how the number of channels activated scales with the amount of neuromodulatory input (i.e., serotonin and norepinephrine) present. Thus, the reverse-engineering features that predicted the simulated level of neuromodulatory input in the Chardon study more directly predicted the resulting variation in simulated PIC amplitude. Because other intrinsic membrane properties (e.g., NaV conductance) were held constant, changes in PIC amplitude could be attributed specifically to the simulated level of neuromodulatory input.

      In contrast, variation in PIC amplitude estimated from our experimental data cannot be attributed to differences in neuromodulatory input alone. We must also consider other factors that affect PIC amplitude that are unknown in our participants, including differences in intrinsic membrane properties. Importantly, several discharge features incorporated into the reverse engineering framework, including braceheight and delta-F, are commonly used to estimate PIC amplitude in human motor unit recordings, independently of the Chardon et al. (2024) framework (e.g., Mesquita et al., 2024).

      Thus, to the extent that brace height, delta-F, and the “neuromodulation” composite variable reflect PIC amplitude in our human data, our findings reflect variation in the physiological factors that determine PIC amplitude. These include neuromodulatory input and intrinsic membrane properties, as well as, for delta-F in isolation, the pattern of inhibition. In our revised version of the manuscript, we will incorporate discussion of these additional factors in addition to our current discussion of neuromodulatory input and the evidence for its alteration in MS.

      Finally, while identifying the specific motoneuron inputs and properties that contribute to altered PIC amplitude among participants with MS is an important goal for future work, uncertainty regarding those contributors does not preclude the potential functional significance of our finding that estimated PIC amplitude can be abnormally high or low among participants with MS. Because PIC amplitude directly influences motoneuron excitability, abnormally high or low PIC amplitudes could contribute substantially to motor deficits that emerge in MS and their variation across patients.

      Calculation of composite variables using mutual information scores. Reviewer 3, for example:

      Mutual information measures how much a feature tells you about a parameter. It does not carry the direction of that association, and it does not become a regression coefficient by having a sign attached to it. Since the direction of every conclusion in the paper depends on those associated signs, the reader has no way of judging how faithfully the composite scores track the physiological quantities they are named after. I do not understand why this was not done with simulations. … The simulations have known inputs, so the composites could be computed on the simulated discharge patterns and their accuracy reported directly.

      A related difficulty is that the three composite scores are not independent of one another. The hysteresis measure contributes to all three of them and several other features contribute to two. Finding abnormality in all three components of the motor command may therefore reflect a single underlying signal expressed three times over. The correlations between the composites are not reported, and without them the reader cannot tell which of these two readings is correct.

      We agree that mutual information scores are non-directional and that they are not regression coefficients. We did not use them as regression coefficients; rather, we used them as an informed way to create a weighted average of the reverse engineering features that, individually, were most informative about each type of input. Signs were assigned to each feature to reflect the direction of the relationship between the feature and each type of input. We took great care to base the assigned signs on published literature, and we will update the paper to more explicitly link each assigned sign to the supporting literature.

      Thank you for the useful suggestion to calculate our composite scores from the simulated features from Chardon et al. and compare them with the known input values. We obtained these data from the authors and found that two of the signs in the excitation composite (for torque at recruitment and duration) needed to be updated based on the Chardon data. After correcting these signs, the neuromodulation and excitation composite variables were both highly correlated with their respective known inputs (r > 0.85). Correction of the two signs in the excitation composite did not change the general pattern of our results. The inhibition composite variable was moderately correlated with the known input (r = 0.50), consistent with the original Chardon et al. study showing that machine learning predictions of inhibition using non-linear regression greatly outperformed those using linear regression. We will include these validation results in the revised manuscript, as they provide a direct assessment of how well the composite variables reflect the known inputs.

      We agree that it is important to demonstrate the independence of the composite variables, and it was an oversight on our part not to include that information. The correlations among the composite variables were very low, ranging from r = 0.006 to r = 0.18, with none reaching statistical significance. As discussed, many of the discharge features are sensitive to more than one type of input, which is why it is difficult to interpret them physiologically in isolation. However, the different types of input affect the features in different ways. For example, push-pull/reciprocal inhibition increases both delta-F and rate attenuation slope compared with uniform inhibition. In contrast, an increase in neuromodulation increases delta-F but decreases rate attenuation slope. The signs within the composite variable calculations reflect these differences in the relationships between each feature and the inputs. Therefore, the shared constituent features do not necessarily result in strongly correlated composite variables, as confirmed by the very low correlations observed in our data.

      We are now collaborating with the authors of Chardon et al. (2024) to explore application of their reverse engineering framework (and/or its subsequent updates currently under development) to our MS and control data to supplement our mutual information-based composite variable analyses.

      Comparison with stroke and spinal cord injury populations, Reviewer 3.

      Second, the comparison with stroke and with spinal cord injury, which carries much of the novelty of the paper, is asserted rather than demonstrated. A good deal of recent motor unit work in spinal cord injury and in stroke is also omitted, which is an important weakness, and I would encourage the authors to engage with it directly rather than treat those populations as a settled contrast. The abstract and the discussion state that the variability seen here is fundamentally different from the consistency seen in those populations, but no stroke or spinal cord injury data are presented, and no quantitative comparison with published values is offered. This matters because the individual patterns illustrated in the paper, that is, reduced peak discharge rate, compressed rate modulation, a narrowed recruitment range, synchronisation between units, and continued firing after the end of the task, are all well-described features of spastic paresis of other causes.

      Although our discussion of the literature closely reflects the quantitative results from those populations, we agree that directly presenting those values would better support the comparison. We therefore will include quantitative comparisons with published values in these populations, to the extent that they are available, in the revised paper.

      We are uncertain which additional studies the reviewer has in mind when stating that a “good deal of recent motor unit work in stroke and spinal cord injury is omitted, which is an important weakness.” We cited the published papers most directly relevant to our specific line of inquiry. Nonetheless, we will conduct an additional review of recent literature and update our references and discussion to address any relevant omissions in the paper revision.

      Finally, we agree that the motor unit firing characteristics mentioned above can be found in spinal cord injury and stroke populations, although their prevalence and expression differ between the populations. However, the presence of those characteristics in some of our participants with MS does not undermine the primary distinction we intended to make between our findings and those reported in stroke and spinal cord injury. One novel aspect of our findings is the marked between-participant heterogeneity in motor unit firing patterns within the MS group, as shown in Figure 3. In particular, among participants with MS, alterations in multiple discharge features were not just more variable than controls, but in some cases, they deviated from controls in opposite directions. We will revise the manuscript to make clear that it is this heterogeneity, rather than the presence of any individual discharge characteristic, that distinguishes the patterns observed in our MS sample from the predominantly group-mean difference patterns typically reported in stroke and spinal cord injury. We will support this comparison quantitatively where published data permit.

      Further, the motor unit characteristics listed by the reviewer were not our only finding. Another novel aspect of our findings is that a substantial number of participants with MS demonstrated decreases in estimated PIC amplitude. In contrast, recent work in spinal cord injury and stroke reported increases in estimated PIC amplitude during voluntary drive in both groups (Hassan, 2021; Benedetto et al., 2026).

      Alternative explanations to increased variability in MS. Reviewers 3, for example:

      … the central finding of greater variability between patients has plausible alternative explanations that have not been excluded.

      In the revised manuscript, we will address the comments from Reviewers 1 and 3 regarding potential alternative contributors to the increased between-participant variability observed in MS, including motor unit yield, decomposition quality, task performance, anti-spastic medications, and other factors raised by reviewers.

      Group differences in the constituent features of the composite variables. Reviewer 3:

      The feature carrying by far the largest weight in the neuromodulation score shows no group difference at all, and depending on which version of the hysteresis measure entered the composite, either one or none of its six constituent features differs between groups.

      There are two main points to clarify. First, a statistically significant group-mean difference in a composite variable does not require the individual constituent features to also have statistically significant group mean differences. As discussed, many of the features are sensitive to more than one type of motoneuronal input. Further, Chardon et al. (2024) demonstrated that prediction of motoneuronal inputs improved as more features were included.

      Second, the absence of a group-mean difference in an individual discharge feature or composite variable does not imply that the feature or composite variable is unchanged among participants with MS. This is particularly important given that many of our MS distributions included values that deviated in opposite directions from controls, which can result in little or no difference in the group mean despite substantial differences at the individual level. For this reason, our analyses compared MS and control distributions not only in terms of central tendency but also their spread and shape.

      Relationships between motor unit discharge features, composite variables, and clinical measures. Reviewer 3, for example:

      Moreover, it would have been valuable to see more relationships reported between the clinical measures and motor unit behaviour, in particular disability, walking speed, lesion location and medication.

      We agree that exploring these relationships is a logical next step and will be especially valuable for understanding the heterogeneity of discharge features and composite variables among participants with MS. The primary goal of the present study was to characterize this heterogeneity, whereas determining how specific clinical characteristics relate to that heterogeneity represents a substantial additional question. We are currently preparing a follow-up paper focused specifically on these relationships so that we have sufficient space to present and discuss them thoroughly. As mentioned above, however, in the paper revision, we will include more information about data from participants who were and were not taking anti-spastic medications.

      In our revision, we will also use more cautious terminology when discussing physiological inference vs. motor unit phenotypes, where appropriate, revise the denominator used in calculating the weighted averages, add a formal test of multimodality, add more raw motor unit firing traces, and address the reviewers’ other outstanding minor suggestions.

      Beauchamp JA, Pearcey GEP, Khurram OU, Chardon M, Wang YC, Powers RK, Dewald JPA & Heckman C (2023). A geometric approach to quantifying the neuromodulatory effects of persistent inward currents on individual motor unit discharge patterns. J Neural Eng 20, 016034.

      Benedetto A, Jenz S, Farley M, Heit B, Sangari S, Beauchamp JA, McPherson L, Heckman C, Perez M & Pearcey G (2026). Muscle-specific motor unit firing characteristics in elbow flexors and extensors after cervical spinal cord injury. ; DOI: 10.64898/2026.06.03.729825. Available at: http://biorxiv.org/lookup/doi/10.64898/2026.06.03.729825 [Accessed June 9, 2026].

      Binder MD, Heckman CJ & Powers RK (2002). Relative Strengths and Distributions of Different Sources of Synaptic Input to the Motoneurone Pool. In Sensorimotor Control of Movement and Posture, ed. Gandevia SC, Proske U & Stuart DG, pp. 207–212. Springer US, Boston, MA. Available at: https://doi.org/10.1007/978-1-4615-0713-0_25.

      Chardon MK, Wang YC, Garcia M, Besler E, Beauchamp JA, D’Mello M, Powers RK & Heckman CJ (2024). Supercomputer framework for reverse engineering firing patterns of neuron populations to identify their synaptic inputs. eLife. Available at: https://elifesciences.org/articles/90624 [Accessed November 2, 2024].

      Hassan AS (2021). Changes in Motor Unit Firing Patterns as a Function of Age, Muscle, and Following a Unilateral Brain Injury: Ionotropic and Metabotropic Effects (Ph.D. thesis). Northwestern University, United States -- Illinois. Available at: http://libproxy.wustl.edu/login?url=https://www.proquest.com/dissertations-theses/changes-motor-unit-firing-patterns-as-function/docview/2572564921/se-2?accountid=15159.

      Hyngstrom AS, Johnson MD & Heckman CJ (2008). Summation of Excitatory and Inhibitory Synaptic Inputs by Motoneurons With Highly Active Dendrites. J Neurophysiol 99, 1643–1652.

      Johnson MD, Thompson CK, Tysseling VM, Powers RK & Heckman CJ (2017). The potential for understanding the synaptic organization of human motor commands via the firing patterns of motoneurons. J Neurophysiol 118, 520–531.

      Kim H & Jones KE (2012). The retrograde frequency response of passive dendritic trees constrains the nonlinear firing behaviour of a reduced neuron model. PloS One 7, e43654.

      Kim H, Major LA & Jones KE (2009). Derivation of cable parameters for a reduced model that retains asymmetric voltage attenuation of reconstructed spinal motor neuron dendrites. J Comput Neurosci 27, 321–336.

      Kuo JJ, Lee RH, Johnson MD, Heckman HM & Heckman CJ (2003). Active Dendritic Integration of Inhibitory Synaptic Inputs In Vivo. J Neurophysiol 90, 3617–3624.

      Lee RH & Heckman CJ (1998a). Bistability in Spinal Motoneurons In Vivo: Systematic Variations in Persistent Inward Currents. J Neurophysiol 80, 583–593.

      Lee RH & Heckman CJ (1998b). Bistability in Spinal Motoneurons In Vivo: Systematic Variations in Rhythmic Firing Patterns. J Neurophysiol 80, 572–582.

      Lee RH & Heckman CJ (1999). Enhancement of Bistability in Spinal Motoneurons In Vivo by the Noradrenergic α<sub>1</sub> Agonist Methoxamine. J Neurophysiol 81, 2164–2174.

      Mesquita RNO, Taylor JL, Heckman CJ, Trajano GS & Blazevich AJ (2024). Persistent inward currents in human motoneurons: emerging evidence and future directions. J Neurophysiol 132, 1278–1301.

      Mousa MH & Elbasiouny SM (2026). A Multi-Scale, High-Fidelity Computational Model of the Mouse Triceps Surae Motor Pool.

      Powers RK, ElBasiouny SM, Rymer WZ & Heckman CJ (2012). Contribution of intrinsic properties and synaptic inputs to motoneuron discharge patterns: a simulation study. J Neurophysiol 107, 808–823.

      Powers RK & Heckman CJ (2015). Contribution of intrinsic motoneuron properties to discharge hysteresis and its estimation based on paired motor unit recordings: a simulation study. J Neurophysiol 114, 184–198.

      Powers RK & Heckman CJ (2017). Synaptic control of the shape of the motoneuron pool input-output function. J Neurophysiol 117, 1171–1184.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) miRNA-based regulation, by definition, requires the miRNA and its target mRNA to be present in the same cell type. CCL5 is expressed in many cell types, making it impossible to propose any miRNA–mRNA interaction based solely on tissue-level expression changes. In vivo cell-type specificity data would substantially strengthen the claims.

      We fully agree with this critique, and we consider this the most critical point to address. To directly examine cell-type-specific expression in vivo, we performed FISH combined with immunofluorescence for GFAP (astrocytes) or IBA1 (microglia) in the peri-infarct region of MCAO mice at D3 (Figure 1E–G). Our results show that Ccl5 mRNA is expressed in both GFAP-positive astrocytes and IBA1-positive microglia, with comparable co-localization rates across both cell populations. In contrast, miR-324-5p showed a significantly higher positive rate in astrocytes than in microglia (Figure 1G). Given that miR-324-5p is more abundantly expressed in astrocytes, the regulatory capacity of the miR-324-5p/CCL5 axis is predicted to be more pronounced in this cell type. These findings provide direct in vivo evidence supporting an astrocyte-predominant miR-324-5p/CCL5 regulatory interaction in the peri-infarct region. We acknowledge that the current data do not constitute cell-type-specific in vivo manipulation, and we discuss this limitation and propose future directions (including viral or transgenic approaches) in the revised Discussion.

      (2) The authors treat an extensive area of ipsilesional cortex uniformly as "IP." Astrocytic and microglial responses to localized injuries such as stroke are highly location-dependent and undoubtedly change dramatically within this area. Data cannot be interpreted without confirmation that samples were collected at identical, defined distances from the injury. Similarly, it is difficult to interpret the Sholl and spine data without knowing where within the large IP region these neurons were found.

      We thank the reviewer for identifying this important concern. Upon review, we recognized that part of the tissue samples used for qPCR, ELISA and Western blot in the original submission were collected from an excessively broad region of the ipsilateral cortex, which likely introduced heterogeneity into the data. We have re-collected these samples specifically from the peri-infarct zone, defined as the 1–2 mm cortical rim immediately surrounding the visibly pale infarct core, beginning from the second and third coronal slices from the most rostral aspect of the cerebral cortex. The qPCR, ELISA and Western blot data have been updated accordingly (Figures 1C–D, 2A, 3A, 6A), and the sampling definition has been specified in the Methods. We also confirmed that all immunofluorescence and Golgi staining analyses were performed within this same peri-infarct zone, where cells retain intact morphology and show the most informative between-group differences. This sampling region has now been defined consistently across all in vivo analyses in the revised manuscript.

      (3) Astrocytes are notoriously prone to dramatic change in serum-containing culture. The shaking-based culture system makes it difficult to conclude much about the role of astrocytes in the CCL5 pathway, particularly without cell-type-specific validation in vivo.

      We acknowledge this limitation. We attempted to implement the immunopanning protocol described by Barres et al. (doi: 10.1016/j.neuron.2011.07.022) to obtain a more purified astrocyte culture. However, several essential reagents — including sodium selenite, putrescine, and N-acetyl-L-cysteine — could not be procured due to import and purchasing restrictions in our region. The use of a commercially available O4 antibody substitute (clone O4, R&D, MAB1326) in place of O4 hybridoma supernatant, combined with the absence of these chemicals, likely contributed to the very low astrocyte yields and poor cell viability observed across multiple independent attempts. We therefore retained the shaking-based isolation and purification method for the present study.

      To characterize the composition of our primary cortical astrocyte cultures, we have included immunofluorescence data from co-labeling of GFAP with Tuj1, Olig2, and IBA1 at P0 and P1 in Supplementary Figure S4, confirming that GFAP-positive cells comprised approximately 88% of total cells at P1. We have also added a paragraph to the Discussion acknowledging that more refined culture systems, as well as cell-type-specific in vivo manipulation of CCL5 and miR-324-5p via viral or transgenic approaches, would further consolidate the conclusions of the present study.

      (4) Missing methodological information, including infarct size measurements, TUNEL staining, and statistical testing.

      We apologize for these omissions. Detailed descriptions of infarct volume quantification (including the edema-correction formula), TUNEL staining procedures, NeuN/TUNEL co-labeling, and all statistical tests have been added to the Methods section.

      (5) The TTC figures appear unusual, with infarct edges resembling overlapping stars rather than natural smooth boundaries. It is unclear whether infarct volume measurements accounted for edema, and no quantification protocol is described.

      We apologize for the confusion. The unusual appearance of the infarct edges in the original TTC figures resulted from dotted-line annotations we had added to highlight the infarct boundaries; these have now been removed to present the unmodified TTC images (Figure 2B, 3B). Infarct volume was corrected for edema-induced hemispheric swelling using the formula: [(contralateral hemisphere volume − ipsilateral non-infarcted volume) / contralateral hemisphere volume] × 100%. This formula and the complete quantification protocol have been added to the Methods section.

      (6) Repeated t-tests between subgroups are used instead of the more appropriate ANOVA, making it difficult to have confidence in the results.

      We agree. We identified that t-tests had been applied inappropriately in the original qPCR and ELISA analyses. All qPCR and ELISA data have been re-analyzed using two-way ANOVA with Tukey's post-hoc test, as appropriate for datasets with multiple groups and time points. We have also reviewed all other figures and corrected any inappropriate use of t-tests.

      Reviewer #2 (Public review):

      (1) The temporal and spatial expression patterns of miR-324-5p do not match those of CCL5, especially at D1 and D3. Despite the inverse relationship between miR-324-5p and CCL5 being apparent only at D7 after MCAO, what was the purpose of administering miR-324-5p agomir (or antagomir) at D1 post-MCAO? If the connection cannot be clearly established, the conclusion reached at the end will be difficult to accept.

      We thank the reviewer for identifying this critical issue. Upon re-examination, we recognized that the original qPCR and ELISA samples had been collected from an excessively broad region of the ipsilateral cortex, which likely introduced heterogeneity into the expression data and obscured the true temporal dynamics in the viable peri-infarct tissue. We have re-collected samples specifically from the peri-infarct zone — defined as the 1–2 mm cortical rim immediately surrounding the infarct core — and updated the data accordingly (Figures 1C–D, 2A, 3A).

      The updated data reveal a clearer and more consistent temporal pattern. Ccl5 mRNA levels in the IP region are significantly elevated as early as D1 compared with sham controls, and continue to increase progressively through D7. Regarding miR-324-5p, although IP region levels at D1 do not yet differ significantly from sham controls, they are already significantly lower than in the contralateral CP region at this early time point. This ipsilateral-versus-contralateral difference at D1 indicates that miR-324-5p downregulation begins in the acute phase following stroke, even before it reaches statistical significance relative to the sham baseline. From D3 onwards, miR-324-5p levels in the IP region are significantly reduced relative to both sham and CP groups, coinciding with the period of sustained and progressive CCL5 upregulation. Taken together, these updated findings support an early and progressive inverse relationship between miR-324-5p and CCL5 in the peri-infarct cortex following MCAO, consistent with our previously published finding that miR-324-5p suppresses astrocytic CCL5 expression (Sun et al., Cell Death Dis., 2019), and functionally validated by the ELISA data showing that miR-324-5p agomir injection significantly reduces CCL5 protein concentrations in the IP region at D3 and D7 (Figure 3A).

      Regarding the rationale for administering miR-324-5p agomir at D1: this timing was chosen to model a clinically realistic therapeutic scenario targeting the early post-stroke period. MicroRNA agomir/antagomir interventions typically require 2–10 days to achieve peak target gene modulation in the mouse brain; administration at D1 therefore ensures that meaningful miR-324-5p-mediated suppression of CCL5 is achieved during the critical acute-to-subacute transition period, as confirmed by the ELISA results at D3 (Figure 3A).

      (2) Would administering miR-342-5p or anti-CCL5 at later time points (e.g., after D3) reduce infarct size or improve functional recovery? If this is not the case, the effect of CCL5 on neuronal cell damage must occur within a very short time after MCAO. Additionally, if the increased CCL5 expression is due to the downregulation of miR-342-5p, its impact would likely be less significant.

      We acknowledge that the current study did not include experimental groups with delayed administration, and we recognize this as a limitation.

      However, several points inform our interpretation. First, as described in our response to Comment 1 above, the updated data demonstrate that miR-324-5p downregulation in the IP region is already detectable relative to the contralateral CP region at D1 and progresses further through D3 and D7, indicating that the miR-324-5p/CCL5 regulatory axis is engaged from the acute phase of stroke, providing a biological basis for early intervention. Second, in experimental stroke models, the infarct core is largely established within the first 24–72 h following vessel occlusion, with the majority of ischemic neuronal death occurring during this window. CCL5, as a pro-inflammatory mediator, is expected to amplify immune cell recruitment and inflammatory cascades most consequentially during this early period, making D1 administration mechanistically rational for limiting neuronal loss. Third, the superior early behavioral outcomes in the CCL5 antibody group relative to the miR-324-5p agomir group (Figures 2D–E, 3D–E) support the value of early CCL5 suppression. As the antibody acts immediately while the agomir requires time for post-transcriptional regulation, this difference highlights that timely agomir delivery is essential to achieve effective CCL5 suppression during the critical early window.

      (3) The study would benefit from the exploration of potential translational applications.

      We thank the reviewer for this constructive suggestion. We are planning to investigate whether astrocyte-derived extracellular vesicles engineered to overexpress miR-324-5p can enhance neurological recovery after stroke. Extracellular vesicles offer several translational advantages: they can traverse the blood-brain barrier, provide a stable and biocompatible vehicle for miRNA delivery, and may be less immunogenic than viral approaches. This would leverage the neuroprotective regulatory mechanism identified in the present study while offering a clinically viable delivery strategy.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In vivo cell-type specificity data for CCL5/miR-324-5p would substantially strengthen the manuscript.

      We have performed FISH combined with GFAP and IBA1 immunofluorescence in the peri-infarct region of MCAO mice at D3 to characterize the cell-type-specific in situ expression of Ccl5 mRNA and miR-324-5p (Figure 1E–G).

      (2) Cell-type-specific in vivo manipulation of CCL5/miR-324-5p (e.g., virally or transgenically) would further substantiate the conclusions. A more compelling astrocytic culture model is needed.

      We fully agree that cell-type-specific in vivo manipulation would represent a major advance. As described in our response to Comment 3 above, we were unable to successfully implement immunopanning in the current study. The FISH data provide in vivo evidence supporting the astrocyte-enriched expression of miR-324-5p in the peri-infarct region. We have added a Discussion paragraph explicitly identifying viral or transgenic astrocyte-specific manipulation of CCL5 and miR-324-5p in vivo as a critical next step to validate and extend the conclusions of this study.

      (3) The measurement shown in Figure 4D is unclear. A more informative measure might be TUNEL/DAPI, with additional cell-type-specific markers to identify what cells are dying in what proportions.

      We agree with this suggestion and have revised the quantification accordingly. In the updated manuscript, apoptotic cell death is reported as the proportion of TUNEL-positive cells among total DAPI-positive nuclei (TUNEL/DAPI). As a complementary cell-type-specific measure, we quantified the proportion of NeuN-positive neurons among total DAPI-positive nuclei (NeuN/DAPI) to specifically assess neuronal survival within the co-culture system. These two measures together provide a clear and interpretable readout of both overall cell death and neuronal viability under each experimental condition. The revised quantification is presented in Figure 4C–E and Figure 5B–D.

      (4) qPCR and ELISA data should be normalized to internal controls, sham values should be presented, and ANOVA (or appropriate non-parametric tests) should be used.

      All qPCR data are now normalized to Gapdh (for mRNA) or U6 snRNA (for miRNA). Sham group values are presented in all relevant figures. Two-way ANOVA with Tukey's post-hoc test is now used for all qPCR and ELISA comparisons. The updated statistical approach is summarized in the Statistical Analysis section of the Methods.

      (5) The description "within 24 hrs" for the timing of CCL5 in vivo manipulation is ambiguous.

      We have revised the description to "at 24 h post-MCAO" throughout the Methods and Results sections to specify the precise time point of intervention.

      (6) Only some statistical comparisons are shown in Figure 2E, which inaccurately implies that the other groups are not different.

      We have updated Figure 2E and Figure 3E to include all statistically significant pairwise comparisons, ensuring that the significance markers accurately represent the complete set of statistical relationships among all groups.

      (7) "Activation" is not the appropriate term for astrocytes in pathological contexts; A1/A2 terminology should be removed.

      We thank the reviewer for this important correction. In line with the consensus recommendations by Escartin et al. (Nat Neurosci, 2021), we have replaced all instances of "astrocyte activation" in pathological contexts with "astrocyte reactivity" or "reactive astrogliosis" throughout the manuscript, including the Abstract, Results, and Discussion. All references to A1 and A2 subtypes have been removed from the Discussion.

      (8) There are typographical errors, and the repeated use of "Besides" is awkward.

      We have carefully proofread the entire manuscript to correct typographical errors. All instances of "Besides" used as a sentence-opening connector have been replaced with contextually appropriate alternatives, such as "Furthermore," "Moreover," "In addition," or "Additionally."

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Hsiung et al. investigated whether the effects of autophagy gene knockdown on the lifespan of long-lived C. elegans mutants depend on experimental conditions. The authors first compiled published data on autophagy-dependent lifespan regulation in daf-2 and wild-type backgrounds, highlighting that prior results are notably inconsistent and likely context-dependent. They then systematically tested the lifespan effects of RNAi knockdown of six autophagy genes (atg-2, atg4.1, atg-9, atg-13, atg-18, and bec-1) in wild-type (N2), daf-2 (reduced insulin/IGF-1 signalling), and glp-1 (germlineless) animals, while varying temperature, daf-2 allele, FUDR concentration, and bacterial infection status.

      The key findings are as follows. In wild-type animals, lifespan suppression by most autophagy gene knockdowns was more pronounced at 20˚C than at 25˚C, where little or no effect was observed. In daf-2 mutants, stronger lifespan suppression was seen in the weaker daf-2(e1368) allele at 20˚C, but not in the stronger daf-2(e1370) allele, and effects were largely absent at 25˚C. In glp-1 mutants, four of six gene knockdowns suppressed lifespan to a greater extent than in N2, though again in a temperature-dependent manner. FUDR at a high concentration (800 µM) abolished the life-shortening effects of most knockdowns and, in the case of atg-9 and atg-13, led to lifespan extension. Kanamycin treatment to eliminate bacterial proliferation did not fully account for the lifespan effects, suggesting that increased susceptibility to infection is not the primary mechanism. The authors also tested the programmed aging hypothesis that autophagy promotes lifespan reduction through biomass repurposing, but found no changes in vitellogenin levels upon knockdown of any of the six genes.

      Altogether, among all genes tested, atg-18 knockdown produced the strongest and most consistent lifespan suppression across nearly all conditions, including both daf-2 and glp-1 backgrounds. The authors probed whether atg-18 acts through the FOXO transcription factor DAF-16 by examining dauer formation and ftn-1 expression, but found no evidence for this, suggesting a DAF-16-independent mechanism.

      We thank the reviewer for their meticulous review of the manuscript. Responding to their comments (particularly those not in the public review) has improved it substantially.

      Strengths:

      The primary strength of this work lies in its systematic and comprehensive approach to dissecting how experimental variables influence the outcome of autophagy-lifespan epistasis tests. The compilation of prior data alongside the authors' own multi-condition dataset is a genuinely useful resource for the field. The study raises a timely and important point about condition selection bias, which is relevant not only to autophagy research but to C. elegans aging studies more broadly. The finding that atg-18 behaves distinctly from other autophagy genes across all conditions is noteworthy and opens avenues for future mechanistic work.

      Weaknesses:

      Despite its breadth, the study has several weaknesses that limit the strength of some conclusions.

      (1) Variability in control lifespan data. The N2 lifespan values under ostensibly identical conditions (e.g., GFP RNAi at 20˚C) differ substantially across experiments (compare Tables S2, S5, S6, S7, and S9). Since N2 serves as the baseline for calculating whether the effect is greater in long-lived mutants via Cox proportional hazard (CPH) analysis, this variability in controls directly affects the reliability of those comparisons.

      Such inter-trial variability in N2 lifespan is not unusual in studies of C. elegans aging, ostensibly identical conditions notwithstanding, and its causes are unknown. A careful C. elegans lifespan study from 2017 compared results of tests performed independently at 3 sites under similar conditions. This showed that inter-trial variation occurred mainly at each site over time, rather than between sites (Lucanic et al., 2017). We note that the data for our study was gathered over a 5-year period. One possibility is that this long time frame may have contributed to the variability seen. 

      Regarding statistical tests (whether the log-rank test to assess differences between pairs of populations, or the CPH test to assess differences between effects of treatments under two conditions): comparisons were made either of data within individual trials, or for pooled data. Thus, valid comparisons were made. We did not compare controls from one trial with treatments from another, which would have yielded misleading results. 

      (2) Limited biological replication. Most experiments were performed with only two biological replicates. In several cases, the two replicates yield contradictory outcomes: one showing significant lifespan suppression and the other showing no effect or even extension. The authors combine these into cumulative datasets for analysis, which, while not incorrect in principle, may obscure genuine irreproducibility. Given that the central message of the paper concerns variability and condition dependence, additional replication would have substantially strengthened confidence in the reported results.

      This is a natural issue to raise. In tests of an effect of a given treatment on C. elegans lifespan, a minimum of 3 trials is standard, and in this lab as a rule we follow this convention. However, after careful consideration during the design stages we opted not to do so for this particular study. Our reasons are set out in the manuscript as follows.

      “A methodological note: for tests of effects of a given intervention on C. elegans lifespan an often-applied standard is to include 3 biological replicates. This is true of several recent studies where the effect of knockdown of a single atg gene on daf-2 longevity was studied (Minnerly et al., 2017; Wilhelm et al., 2017; Yang et al., 2024). However, given that the present condition dependence study effectively performs this test in 18 different ways, involving RNAi of 6 atg genes, 2 daf-2 mutants and 2 temperatures, N = 2 biological replicates were judged to be sufficient to draw robust conclusions; similarly, an earlier study of RNAi 14 atg genes under two conditions used 2-3 biological replicates (Hashimoto et al., 2009); for an overview of N sizes in previous studies, see Table S1.”

      While this approach has yielded robust broad conclusions (e.g. that atg gene RNAi generally does not suppress daf-2(e1370) Age), it is true that for any one given treatment (say, effects of atg-9 RNAi on daf-2(e1370) Age) one may not draw conclusions with a high degree of confidence, and we do not do so. We therefore argue that in a study of this nature, as for instance in a whole genome RNAi screen for lifespan effects, it is reasonable and expedient to drop below the 3 replicates minimum standard; here we agree with the Nishida lab’s similar judgement, and from that study too robust conclusions may be drawn.

      (3) Low sample sizes in individual trials. A number of lifespan assays were conducted with only 40-50 worms per replicate, and in some cases, as few as 30. Such sample sizes are below the standard commonly used in the C. elegans aging field and are likely to contribute to the variability observed.

      Please see our response to point 2, which in essence responds to this concern.

      (4) RNAi efficacy measured only in N2 at 20˚C. The authors demonstrated that atg-2 and atg-4.1 RNAi did not significantly reduce target mRNA levels, which may explain their weaker lifespan effects. However, these same RNAi treatments significantly affected lifespan in several other conditions (e.g., daf-2(e1368) at 20˚C, glp-1 at 20˚C and 25˚C, and N2 with 15 µM FUDR). Measuring RNAi efficacy across different genetic backgrounds and conditions would be needed to properly interpret these variable results.

      The study would indeed be strengthened by inclusion of target atg mRNA measurements under all of the various conditions tested. However, this would have required a very large number of qPCR tests to be run; we note that in previous assessments of atg RNAi effects on lifespan (listed in Table S1 and Table S6), such tests were rarely performed. Regarding RNAi effects on daf-2 mutants, we note in the text the following: “While mRNA levels after RNAi under the various other conditions tested were not assayed, reduced IIS (including daf2(e1370)) has been shown to intensify the RNAi response (Wang and Ruvkun, 2004), thus lack of effect on lifespan in daf-2(e1370) is unlikely to reflect suppression of mRNA knockdown.” Here we have at least assessed, using N2, the most important issue relating to RNAi efficacy: the differential effects of different RNAi feeding clones on atg mRNA levels.

      (5) Incomplete mechanistic exploration. The investigation of why atg-18 knockdown has uniquely strong effects was limited to DAF-16. Given published evidence that atg-18 may regulate HLH-30/TFEB, a master transcriptional regulator of autophagy and lysosomal biogenesis, testing whether atg-18 specifically affects HLH-30 nuclear localisation or activity could have provided valuable mechanistic insight and would distinguish atg-18 from the other genes tested.

      We would have readily investigated this. However we learned of the interactions between atg-18 and hlh-30 only in Nov 2025, when one of us (David Gems) bumped into a member of Evandro Fan’s research group at an aging meeting at the Crick Institute in London. Their findings were very interesting for us, as they offered a possible explanation for the seeming idiosyncrasy of atg-18 RNAi effects on lifespan. This subject is currently under investigation by the Fan lab at the University of Oslo, who recently posted a preprint describing the work (Schmauck-Medina et al., 2026).

      Reviewer #2 (Public review):

      Summary:

      This study examines how genes involved in cellular recycling (autophagy) influence lifespan under different experimental conditions. The findings help clarify why previous studies have reported conflicting results about whether blocking autophagy shortens or extends lifespan. The work will be of interest to researchers studying aging and cellular stress responses, particularly those using model organisms.

      We thank the reviewer for their helpful remarks. Responding to their comments (particularly those not in the public review) has enabled us to improve it.

      Strengths:

      The findings are valuable, as they help resolve inconsistencies within a specific subfield of aging research. The evidence presented is solid, as the data broadly support the primary claims of the study. In addition, the discussion is thorough and thoughtfully integrates the findings within the broader context of the field.

      Weaknesses:

      Additional functional validation would further strengthen the conclusions.

      We very much agree. Our original plan for this study was to include autophagic flux assays under different conditions. However, this line of investigation led us to a careful reassessment of reporter-based approaches to measuring autophagic flux in C. elegans, and attempts to improve them. This includes development of an automated, AI-based quantitative image analysis pipeline to improve reproducibility and data interpretation across studies. This investigation is still ongoing, and we are currently preparing a separate manuscript focused specifically on methodological clarity and quantitative assessment of autophagic flux.

      Recommendations for the authors:

      Reviewing Editor Comments:

      To increase the evidence provided by the authors, they should at least address the comments from Reviewer 1 regarding the experimental inconsistencies. We acknowledge that the additional experiments suggested by Reviewer 2 regarding the use of C. elegans mutants and additional methods to assess autophagic flux would likely be a lot of additional work. However, adding results from such experiments would, of course, make the evidence more compelling.

      Reviewer #1 (Recommendations for the authors):

      Writing and presentation

      (1) The abstract discusses results for daf-2 in detail but does not mention the glp-1 findings. Given that a substantial portion of the study addresses glp-1 longevity, including a summary of those results in the abstract would better represent the scope of the work.

      Agreed. glp-1 is now referred to in the abstract.

      (2) In the abstract, the sentence regarding FUDR effects is placed between statements about daf2, while it is referring to N2 lifespans, which may give the impression that the FUDR results were obtained in a daf-2 background. Consider restructuring this section for clarity.

      Agreed. To improve clarity this now reads as follows. “In wild-type C. elegans, FUDR at a high concentration caused knockdown of several atg genes to increase lifespan”

      (3) The definition of "robust" used in Figure 4E could be misleading to readers. For instance, bec-1 and atg-4.1 knockdowns in glp-1 at 20˚C are classified as robust, yet the actual percent suppression is modest (~8% and ~3%, respectively) and non-consistent in individual replicates. The "robust" designation arises because these knockdowns slightly increased N2 lifespan. Clarifying the definition in the figure legend or text would help readers interpret this correctly.

      For glp-1 at 20˚C robust suppression is seen with atg-2, atg-18 and bec-1 RNAi, not atg-4.1 RNAi. But regarding bec-1: yes, the suppression is modest. bec-1 RNAi is an unusual case insofar as it meets the <30% definition partly because it caused an increase in N2 lifespan. Under the circumstances, arguably, it makes little sense to view it as an example of robust suppression, and Figure 4E has been altered accordingly, with a note added to the legend as follows. “Note that the bec-1 RNAi effect on glp-1 at 20˚C is not classified as robust here even though it reduces lifespan to within <30% of the mean lifespan of N2 under bec-1 RNAi, since the fact that it does so partly reflects an increase in N2 lifespan, rather than a robust life-shortening effect on glp-1.”

      To try to improve clarity we have altered the definition of “robust” to read as follows. “R, robust suppression, i.e. knockdown reduces the extended lifespan of daf-2 or glp-1 to within <30% of the mean lifespan of N2 under the same RNAi. This designation (“robust”) indicates a high degree of suppression of the mutant longevity phenotype (see Figure 2, Figure 5 and Fig. S2).”

      (4) On page 12, the text should read "9/30 suppresses robustly" (currently appears to contain a numerical error).

      Fixed. This now reads “In 8/30 the RNAi effect was robust, i.e. the mutant longevity was largely suppressed.” (Now 8/30 since bec-1/glp-1/20˚C is no longer viewed as robust suppression).

      (5) In Supplementary Sheet 8, the lifespan data from Hashimoto et al. are presented in a different format than the data from other studies. Standardizing the presentation would improve readability.

      This is Supplementary Table 1. The inconsistencies have been ironed out.

      Data and calculations:

      (6) In Supplementary Table S3, the ΔΔCt values for atg-9 appear to be incorrect. Please verify and correct.

      We thank the reviewer for highlighting this error. The ΔΔCt values for atg-9 in Supplementary Table S3 were incorrectly entered; the Fold Change values had been mistakenly placed in that column. This error has now been corrected. Please note that analyses and conclusions reported in the manuscript were based on the correct values.

      (7) In Supplementary Table S4, the standard deviation values do not match my independent calculations. Please double-check these values.

      This is correct: there was an error in the standard deviation (SD) values. We thank the reviewer for their diligence. We have updated both the SD and SEM (standard error of mean) columns in Table S4. The mean ΔΔCt values remain unchanged, as do the conclusions from analyses and statistical tests.

      (8) In Supplementary Table S7 (kanamycin experiment), there appear to be several errors in the percent change calculations. Additionally, the statement "In the absence of Kan, atg-13 RNAi caused a slight reduction in lifespan" is not supported by the combined data, which actually shows a slight increase. Given that the reported changes are subtle but statistically significant, it would be prudent to re-verify the p-value calculations as well before drawing conclusions from this experiment.

      The calculation for trial 1 atg-13 (-Kan) as a percentage of control (L4440 Kan) has been corrected so that the mean lifespan of the knockdown is divided by the mean lifespan of the control. In the previous version, the ratio was inadvertently calculated in the opposite orientation. All other values remain unchanged.

      Responding further to this point, to strengthen the data here we have also conducted an additional trial, and Figure 3D and Table S7 have been updated accordingly. The more robust data still supports the conclusion that E. coli infection does not mask a life extending effect of atg-13 RNAi. However, in the new, summed data, atg-13 RNAi on no Kan does not shorten lifespan at all, in contrast to our previous trials (conducted several years earlier), but consistent with several other instances of variability in the study. Moreover, the modest life-shortening effect of atg-13 RNAi on Kan (-6.7%) is now statistically significant. The manuscript has been updated accordingly.

      Experimental interpretation

      (9) On page 9, the authors report testing N2 and daf-2(e1370) lifespan at 15˚C and 20˚C, but only the daf-2 results are discussed in the text. The N2 results at 15˚C appear in Table S5 but are never addressed. Notably, bec-1 knockdown significantly suppressed N2 lifespan at 20˚C in Table S2 but appears to significantly extend it at both 15˚C and 20˚C in Table S5. These discrepancies should be discussed.

      The issue of inter-trial variability is discussed in our response to point 1 in the public review. More specifically: here it may be significant that the trials listed in Table S5 were performed several years after those in Table S2. The discrepancy is now noted and discussed as follows. “In these trials bec-1 RNAi also modestly increased N2 lifespan at both temperatures (Table S5), surprisingly given that in previous trials (performed several years earlier) bec-1 RNAi shortened N2 lifespan (Table S2). The reason for this discrepancy is unknown.”

      (10) Regarding the CPH analysis of bec-1 in daf-2(e1368) (Figure 1), the authors state that bec-1 knockdown does not have a significantly greater effect in daf-2(e1368) relative to N2, and then note that this is consistent with the earlier observation by Hansen et al. that bec-1 shortens daf-2 lifespan without affecting N2. However, in the cumulative dataset, bec-1 does significantly suppress N2 lifespan. A more precise statement here would prevent readers from drawing an incorrect conclusion.

      The point here is that our study and the Hansen et al study both point to atg RNAi suppression of daf-2 Age being limited to class 1 mutants, not that there are no effects on N2. To try to improve clarity we have rephrased as follows. “These findings are broadly consistent with the earlier observation that bec-1 and vps-34 RNAi shortened the lifespan of the daf-2(mu150) class 1 mutant but not of N2 at 20˚C (CPH analysis not performed) (Hansen et al., 2008).”

      (11) The reference to Hashimoto et al. (2009) on page 10 states that atg-9 and atg-13 RNAi increased lifespan, but that study does not include data for atg-13. The lifespan extension reported by Hashimoto et al. was for atg-7, atg-9, bec-1, and unc-51. Please correct this citation.

      Done. It now reads “where atg-9 (and also atg-7, bec-1 and unc-51) RNAi increased lifespan”.

      (12) In a previous publication from this group, atg-2 and atg-13 knockdown with 15 µM FUDR led to significant lifespan extension, whereas in the current study, the same treatments significantly suppressed lifespan. Although this discrepancy is briefly mentioned in the Discussion, a more thorough discussion of possible explanations would strengthen the manuscript's value as a reference dataset for future studies.

      A more detailed discussion has been added, as follows. “Regarding the causes of variability between results of ostensibly identical tests performed under ostensibly identical conditions: one clue is provided by a study comparing results of lifespan assays performed across three sites under similar conditions. This revealed that inter-trial variation occurred mainly at each site over time, rather than between sites (Lucanic et al., 2017). One possibility is that this reflects batch variation in media components, such as the BactoPeptone constituent of nematode growth medium (Petrascheck, 2014).”

      (13) The N2 mean lifespan on GFP RNAi with 0 µM FUDR at 20˚C is approximately 15 days, whereas the N2 lifespan at 20˚C in Table S2 is approximately 20 days. While inter-experiment variability is expected, a difference of this magnitude warrants acknowledgement, as it could influence the interpretation of subsequent comparisons.

      We have now acknowledged this in the legend to Figure 3, as follows. “We note that in (A) the lifespan of the gfp RNAi control is somewhat lower than in other experiments (mean 14.96 days, Table S6); see Discussion for consideration of possible reasons for inter-trial variability.”

      (14) Regarding the glp-1 experiments (page 11 and Table S9), the N2 lifespan values in these experiments differ from earlier N2 results at 20˚C for several knockdowns (e.g., bec-1 knockdown appears to increase N2 lifespan in Figure 4). Additionally, the glp-1 lifespan results are not consistent between the two replicates for most genes except atg-2 and atg-18. The developmental shift (raised at 25˚C, then moved to 20˚C to obtain the glp-1 phenotype) could plausibly account for some of this variation compared to animals raised continuously at 20˚C. If so, this should be explicitly discussed.

      Agreed. The following has been added to the Figure 4 legend. “That bec-1 RNAi increases N2 lifespan in (A) (+16.7%, p < 0.0001) but not (B) could imply an interaction with temperature during development, or merely variability of atg RNAi effects (see Discussion).”

      (15) At 25˚C, N2 lifespan shows significant suppression upon atg-13 and atg-18 knockdown in Table S9, while these same effects were non-significant in Table S2. These and other interexperiment discrepancies should be noted and, if there are identifiable experimental differences, those should be specified.

      This discrepancy has now been noted on page 9, immediately after the description of the data in Table S2, as follows. “(although in later tests at 25˚C, life-shortening effects of atg-13 and atg-18 RNAi were seen; Table S9)”

      (16) If autophagy is already regulated by heat stress at 25˚C, this could explain the diminished effects of autophagy gene knockdown at higher temperatures. Measuring autophagy gene expression by qPCR across different temperatures and genetic backgrounds could provide useful mechanistic insight.

      Agreed. However, more informative will be to measure effects of temperature and genotype on autophagy more directly, using fluorescent reporters of autophagic flux. We are addressing this as part of an ongoing study using improved and fully validated autophagic flux measurement methodologies (please see our response to reviewer 2, public review). 

      (17) The authors report that atg-18 knockdown upregulates other autophagy genes (supplementary data). This is intriguing given that atg-18 shows the strongest phenotype. Whether this reflects a compensatory mechanism and why it does not rescue the lifespan suppression deserves further discussion.

      On reflection we decided to remove from the manuscript the data relating to effects of atg-18 RNAi on mRNA levels of other atg genes due to concerns about data quality. 

      (18) Regarding the FUDR and infection hypothesis: the logic that reduced bacterial infection upon FUDR treatment explains the loss of lifespan suppression is reasonable, but it does not account for why atg-13 knockdown actively extends lifespan in the presence of FUDR. This point could benefit from further discussion. In the kanamycin experiment, they further see that life-shortening effects of atg RNAi are not solely attributable to infection, but the question of lifespan extension remains unanswered.

      Good point. We have addressed this as follows. “We also conclude the increase in lifespan upon atg-13 RNAi in the presence of 800 μM FUDR (Figure 3C) is not attributable to suppression of E. coli infection, but rather to some other, unidentified mechanism.”

      (19) The timing of when lifespan assays are initiated relative to other experimental treatments is another potential source of variability (as seen from earlier reported data) that could be acknowledged as a consideration for future studies.

      Good point. We have added the following to the section of the discussion about tackling condition dependency issues. “Another factor to take into account is the apparent tendency of results of C. elegans lifespan assays to vary over time (Lucanic et al., 2017).”

      Reviewer #2 (Recommendations for the authors):

      Major Comments:

      (1) Assessment of autophagic activity

      The authors demonstrate by qPCR that feeding RNAi reduces mRNA levels of autophagy-related genes. However, reduced transcript levels do not necessarily confirm functional inhibition of autophagy. Incorporating established assays of autophagic flux, such as Western blot analysis of lipidated ATG-8 (LGG-1/ATG-8-II) or validated fluorescence-based reporters, would substantially strengthen the mechanistic conclusions.

      We agree with the reviewer that direct assessment of autophagic flux would provide additional functional insight. We are currently performing complementary assays to directly assess autophagic flux using reporter-based approaches, with a focus on standardizing reporter-based measurements and developing a quantitative image analysis pipeline to improve reproducibility and interpretation across studies. These analyses will be presented in a separate manuscript focused specifically on methodological clarity and quantitative assessment of autophagic flux.

      (2) Use of genetic mutants

      Several C. elegans loss-of-function mutants for autophagy genes are available. Validation of key findings using selected genetic mutants, where feasible, would provide complementary evidence and enhance confidence in the RNAi-based results.

      In principle is this a good idea. In practice many loss-of-function mutants in core autophagy genes exhibit developmental defects or impaired viability, which can confound interpretation in aging studies. As our aim was to examine the effects of autophagy gene perturbation specifically during adulthood, RNAi provided a practical approach that allowed post-developmental knockdown while minimising disruption of normal development.

      Minor Comments:

      (1) Context-dependent effects of autophagy

      The findings are conceptually consistent with prior work demonstrating dual roles of autophagy in C. elegans survival during starvation, where physiological levels promote survival but insufficient or excessive autophagy contributes to mortality (Kang et al., Genes & Development, 2007). Including a discussion of this study would help frame the present results within a broader biological context.

      Good idea to cite this study, and we have now done so in the introduction, as follows. “It is by now clear that autophagy can enhance as well as inhibit the development of pathologies in C. elegans, including senescent ones (Kang et al., 2007).”

      (2) Relevance beyond C. elegans

      It would be helpful for the authors to clarify whether similar context-dependent effects of autophagy on lifespan have been reported in other organisms. Briefly referencing comparable findings in additional model systems would broaden the relevance and impact of the study.

      This is a good idea, but our search for similar cases in other model organisms failed to identify clear examples. Perhaps more to the point here is that context dependent effects and, perhaps, condition selection bias, are a serious issue in scientific research in general. To emphasize this, the following has been added as the last line of the discussion. “More widely, condition dependency and conditional selection bias risk diminishing the reliability of research findings in many scientific disciplines.”

      Hashimoto, Y., Ookuma, S. and Nishida, E., 2009. Lifespan extension by suppression of autophagy genes in Caenorhabditis elegans. Genes Cells. 14, 717-726.

      Lucanic, M., Plummer, W., Chen, E., Harke, J., Foulger, A., Onken, B., Coleman-Hulbert, A., Dumas, K., Guo, S., Johnson, E., Bhaumik, D., Xue, J., Crist, A., Presley, M., Harinath, G., Sedore, C., Chamoli, M., Kamat, S., Chen, M., Angeli, S., Chang, C., Willis, J., Edgar, D., Royal, M., Chao, E., Patel, S., Garrett, T., Ibanez-Ventoso, C., Hope, J., Kish, J., Guo, M., Lithgow, G., Driscoll, M. and Phillips, P., 2017. Impact of genetic background and experimental reproducibility on identifying chemical compounds with robust longevity effects. Nat Commun. 8, 14256.

      Minnerly, J., Zhang, J., Parker, T., Kaul, T. and Jia, K., 2017. The cell non-autonomous function of ATG-18 is essential for neuroendocrine regulation of Caenorhabditis elegans lifespan. PLoS Genet. 13, e1006764.

      Schmauck-Medina, T., Anisimov, A., Meyer, D.H., Hu, Y., Huang, Z., Wu, Y., Taylor, S., Takla, M., MacArthur, M.R., Mitchell, S.J., Ai, R., Simonsen, A., Jensen, V., Labbadia, J., Shen, H.-M., Hansen, M., Schumacher, B., Rubinsztein, D., Lautrup, S., Lu, G. and Fang, E.F., 2026. ATG-18/WIPI2 drives longevity in an HLH-30/TFEB-dependent manner. bioRxiv.

      Wang, D. and Ruvkun, G., 2004. Regulation of Caenorhabditis elegans RNA interference by the daf-2 insulin stress and longevity signaling pathway. Cold Spring Harb Symp Quant Biol. 69, 429-31.

      Wilhelm, T., Byrne, J., Medina, R., Geisinger, J., Hajduskova, M., Tursun, B. and Richly, H., 2017. Neuronal inhibition of the autophagy nucleation complex extends life span in postreproductive C. elegans Genes and Development. 31, 1561–1572.

      Yang, Y., Arnold, M.L., Lange, C.M., Sun, L.H., Broussalian, M., Doroodian, S., Ebata, H., Choy, E.H., Poon, K., Moreno, T.M., Singh, A., Driscoll, M., Kumsta, C. and Hansen, M., 2024. Autophagy protein ATG-16.2 and its WD40 domain mediate the beneficial effects of inhibiting early-acting autophagy genes in C. elegans neurons. Nat Aging. 4, 198-212.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting manuscript by Kirk and colleagues describing a highly valuable knock-down system that leverages CRISPRi in order to further elucidate the role of the Kruppel-Like Factor (KLF) transcription factor family in regulating the maturation of postnatal cortical projection neurons. The authors firstly use RNA-Seq and ATAC-Seq data in order to identify the KLF TF family as a potential regulator of cortical neuron maturation in the postnatal brain and subsequently knock down four KLF family members; KLF9, KL13, KLF6 and KLF7, in order to ascertain the functions of specific KLF genes in the developing cortex. The described CRISPRi knock down strategy is highly robust and penetrant as evidenced by a KD efficiency > 95% (assessed by both qPCR and single molecule FISH) and demonstrates that KLF6 and KLF7 play an activating role in driving the expression of target genes relating to axonal growth whereas KLF9 and 13 play a repressive role that inhibits the expression of overlapping gene targets. Together, the authors propose a model where the KLF TF family acts as a regulatory "switch" from activation to repression in the postnatal cortex as a mechanism to control a shift in projection neuron function from axonal growth to circuit refinement. The findings and conclusions of the manuscript offer a valuable contribution to the field of postnatal cortical development and further our understanding of the regulatory mechanisms that govern neuron maturation.

      The conclusions of this manuscript are generally supported by the data, but some aspects of the data collection and analysis require some further clarification. Specifically:

      (1) The authors comprehensively assess the molecular effects of KLF TF knock-down, however, the authors do not deeply address the cellular effects of these knock-downs. The authors conclude that knockdown of KLF6/7 and KLF9/13 cause downregulation and upregulation, respectively, of a common set of genes involved in cytoskeletal or axon regulation such as Tubb2 and Dpysl3. How is the morphology of the cells affected by these knockdowns? For example, does KLF9/13 knockdown cause neurite/axonal outgrowth? The authors should perform some basic experiments to assess changes in cell morphology following KLF TF KD. This is the one key point that needs addressing, in my opinion.

      We appreciate this comment and agree that the cellular effects of KLF activator and/or repressor KD are not addressed by the experiments in this manuscript. However, the effects of KLF9, KLF13, and KLF9/13 knockdown on neurite outgrowth have been previously examined in vitro (Avci et al., 2012; Avila-Mendoza et al., 2020) and in vivo (Apara et al., 2017). The collective findings of these papers (enhanced neurite outgrowth and axon regeneration following KLF repressor KD) align with the predicted outcome of upregulating the set of cytoskeletal remodeling genes identified as putative KLF family targets in this manuscript. Furthermore, the in vitro results demonstrate that the partial redundancy between KLF9 and KLF13 we identified at the transcriptional level is relevant for their roles in repressing neurite outgrowth. Our results suggest that these previous findings are likely to hold true in cortical neurons in vivo while offering a molecular explanation for these effects at the level of gene expression. Other groups have examined the effect of KLF6 or KLF7 overexpression on corticospinal axon regeneration in vivo and found that these transcription factors can individually promote axon regrowth after injury (Blackmore et al., 2012; Wang et al., 2018). Similarly, this earlier work did not identify transcriptional mechanisms underlying the observed effect so our results also offer a plausible set of targets that could mediate the link between KLF activator overexpression and axon regeneration.

      (2) The authors identify 374 DEGs in P10 Klf6/7 KD neurons and 115 DEGs at P20 (figure 6B). Have the authors looked to see what proportion of these DEGs are upregulated in the KLF9/13 KDs in order to get a more global understanding of the degree of overlap in the genes regulated by the KLF family members? [MOU2] Along similar lines, the authors later indicate that there are 144 shared targets between the KLF activator and repressor pairs (Figure 7C). What percentage does this represent of the total number of DEGs between the KLF pairs. This could further illustrate the degree to which the KLF pairs regulate the same set of genes. If it is already indicated in the manuscript, it should be made a bit more clear to the reader.

      We thank the reviewer for the suggestion to include more comprehensive quantitative measures of the degree of overlap between targets of KLF activator and repressor pairs. We have included the exact number and percentage of P10 and P20 KLF6/7 targets that are differentially expressed (adj. p-value <= 0.05) in KLF9/13 KD neurons and vice versa in the text in the section associated with Figure 7, which is devoted to describing this class of overlapping targets. Furthermore, figures 7D and S7.2B both show how the full set of KLF6/7 targets are affected by Klf9/13 KD and vice versa. Collectively, this demonstrates that both KLF activator and repressor targets trend towards opposite regulation by the opposing pair.

      (3) Figures 5B and 6D2 are very interesting as they relate the changes in gene expression over time in neurons from P2 to P30 to the functions of KLF9/13 and KLF6/7, respectively. I would be curious to see how these two forms of analyses overlap with one another. For example, in Figure 6D2, where would the KLF9/13 upregulated genes fall on the plot shown in Figure 6D2? And would those overlapping genes fit a similar correlation?

      We agree with the reviewer that this is a powerful way to visually demonstrate the overlap between KLF6/7 and KLF9/13 targets in a more unbiased way and within a developmental context. The suggested analysis has been included in the supplement to Figure 7 (Fig S7.3).

      (4) Figure 7E shows expression levels of shared KLF TF targets in control or KD conditions. Interestingly, the expression of Tubb2b, shows higher expression in ScrGFP P10 when compared to KLF9/13 P20, suggesting that derepression of KLF9/13 does not fully restore the expression level of Tubb2b seen at P10. This may suggest that other repressive regulators may be involved in the downregulation of Tubb2b from P10 to P20[MOU4] . Can the authors further comment on this, perhaps in the discussion, and speculate if there are other regulatory factors at play that may be controlling some of the shared targets by KLF6/7 and KLF9/13?

      We thank the reviewer for this keen observation. We have included a comment on this within the text associated with Figure 7. While other regulatory factors are plausible, this is easily explained by low expression of Klf6 and Klf7 in the P20 cortex, which cannot drive transcription of Tubb2b and other targets to the same level as what is observed at P10 when the KLF activators are still relatively abundant, even after repression by Klf9/13 is removed. Within this framework, overexpression of KLF activators on a KLF repressor background would be the only way to restore expression of Tubb2b to its P10 expression levels. This is wholly compatible with our model of ‘push-pull’ regulation of KLF targets outlined in Figure 9. The finding could also be affected by the addition of repressive histone marks by the KLF9/13-associated SID complex in early neonatal development that may persist following the knockdown of these repressors, preventing complete restoration of neonatal expression patterns.

      Reviewer #2 (Public review):

      Summary:

      Kirk et al. use RNA-Seq and CRISPRi to provide evidence that KLF family transcription factors regulate postnatal neuronal maturation of pyramidal neurons. The genetic programs regulating postnatal neuronal maturation are not well understood. The authors first analyzed chromatin accessibility and gene expression data from layer 4 and 6 pyramidal neurons and found that KLF TFs are predicted regulators of postnatal neuronal maturation. They then use CRISPRi knockdown and find that KLF activators first activate genes and then this is followed by KLF repressors repressing genes. Interestingly, some genes, such as those with cytoskeletal functions, are shared targets of KLF activators and repressors.

      Strengths:

      The study is well-executed and the paper is well-written. A major strength of this study is the application of state-of-the-art transgenic approaches. The CRISPRi approach used to knock down multiple KLFs is compelling. The genomic data generated appears to be high quality and is carefully analyzed. The presented findings provide important insights into the genetic programs that regulate postnatal maturation in cortical pyramidal neurons. The discovery that KLF family activators/repressors regulate gene expression changes during this critical step of neuronal development fills an important gap in the field.

      Weaknesses:

      A limitation of the current study is that the functional importance of KLF for postnatal neuronal maturation is unclear. Although the authors find that KLFs regulate some of the gene expression changes during postnatal neuronal maturation, it is still unclear whether such gene expression changes mediate the postnatal changes in morphology and physiology. While beyond the scope of the current study, future studies should investigate the contributions of KLFs on postnatal morphological and physiological changes.

      We thank the reviewer for their helpful comments on this manuscript. We agree that the effects of KLF knockdown identified in this study are primarily descriptive, but – as noted - a detailed analysis of the morphological and physiological consequences of KLF knockdown are beyond the scope of this paper. However, we believe that a more mechanistic model of KLF function during neuronal maturation can be obtained by considering our findings on the bidirectional regulation of core cytoskeletal genes by KLF activators and repressors alongside published in vivo and in vitro data on the opposing roles of KLF family members on axon outgrowth/regrowth (Apara et al., 2017; Avila-Mendoza et al., 2020; Blackmore et al., 2012; Moore et al., 2009; Wang et al., 2018). Thus, we offer a set of transcriptional targets that likely mediate these opposing effects and a developmental context within which they might operate.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript "Multiplexed CRISPRi Reveals a Transcriptional Switch Between KLF Activators and Repressors in the Maturing Neocortex", Kirk and colleagues seek to dissect the developmentally regulated pan-neuronal gene programs that control the postnatal maturation of cortical neurons. For this, the authors analyzed newly generated and existing RNA-seq and ATAC-seq of Layer 4 and Layer 6 cortical pyramidal neurons at postnatal day 2 (P2) and day 30 (P30), and identified thousands of shared developmentally regulated genes and genomic (promoter) regions, including genes involved in axon growth (tend to be downregulated) and synaptic function (tend to be upregulated). Motif enrichment analysis of promoters of differentially regulated genes revealed a strong presence of KLF/Sp family binding motifs, pointing to Krüppel-Like Factors (KLFs) as key transcriptional regulators of cortical maturation. Expression profiling showed a developmental switch from activating KLFs (Klf6, Klf7) expressed neonatally to repressive KLFs (Klf9, Klf13) upregulated during maturation. Using an elegant in vivo multiplexed CRISPR interference (CRISPRi) system, the authors achieved efficient, cell-type-specific knockdown of these TFs and showed that Klf9 and Klf13 repress a set of genes that includes cytoskeletal regulators such as Tubb2b, Dpysl3, and Rac3. Conversely, Klf6 and Klf7 promoted the expression of these same genes in the early postnatal period, and their knockdown led to reduced expression of these genes, particularly at P10 when their activating influence is strongest. Since promoters of shared KLF targets were enriched for KLF/Sp motifs but showed little change in chromatin accessibility, the authors propose a model in which distinct KLF family members function either as transcriptional repressors and activators that compete at constitutively accessible promoters and thereby act as a developmental transcriptional switch that coordinates the downregulation of axon growth programs and upregulation of synaptic maturation genes during cortical development.

      Strengths:

      The study addresses an interesting question and advances our understanding of the transcriptional regulation underlying postnatal cortical development. A major strength of the study lies in the innovative use of in vivo multiplexed CRISPR interference (CRISPRi), which allows for cell-type-specific, combinatorial knockdown of redundant TFs - this an elegant solution to a long-standing challenge in transcription factor research, and should be useful also for other neuroscience studies that require local and cell-type-specific gene loss-of-function. Also, the integration of RNA-seq and ATAC-seq across developmental time points provides a robust foundation for identifying direct targets of the KLF family, and the findings are reinforced by cross-species conservation and the identification of targets with clear neurodevelopmental relevance.

      Weaknesses:

      The major weakness of the study lies in its relatively narrow scope: the study focuses primarily on transcriptional mechanisms and largely lacks functional validation of the neuronal phenotypes that are predicted by the gene expression data (e.g. axonal morphology). For example, the authors analyzed the effects of KLF9/13 KD on the neurons' excitability and excitatory inputs, but did not assess the effects on inhibitory inputs and E/I-ratio or morphological parameters such as axonal length and axonal target fields - the manuscript would be strengthened considerably by such analyses (axonal projections could be analyzed e.g. via local injections of the gRNA AAVs and subsequent immunolabeling of brain sections). Similarly, the chromatin-based mechanisms underlying KLF activity remain relatively speculative, and the transcriptional mechanisms upstream of the KLFs remain unexplored (this could be addressed by analyzing existing datasets; see "Additional Point 1" below). Finally, the manuscript is too long (e.g., nearly five pages in the Discussion section are devoted to discussing various misregulated genes) and would benefit from presenting the Results and Discussion sections more concisely. However, despite these limitations, the paper offers an interesting model for a transcriptional switch during neuronal maturation in the cortex and establishes a powerful methodological framework for dissecting redundant gene networks in vivo.

      Shorten discussion (possibly results also).

      We appreciate the reviewer’s feedback and agree that linking transcriptional perturbations to cellular phenotypes of KLF activator and/or repressor KD would significantly strengthen this manuscript. However, we believe this is beyond the scope of this current manuscript and the effects of individual KLF activators and repressors on axon outgrowth/regrowth are known from prior in vitro and in vivo studies (Apara et al., 2017; Avila-Mendoza et al., 2020; Blackmore et al., 2012; Moore et al., 2009; Wang et al., 2018). Since there was no detectable effect of KLF repressor KD on excitatory synaptic transmission (Fig. S4.2), we did not believe it was likely that our excitatory neuron-specific knockdown would affect inhibitory synapses under our model where KLF targets have primarily axonal functions. In the absence of ATAC-seq from Klf9/13 KD neurons, the relationship between chromatin accessibility and KLF binding is limited to descriptions of chromatin accessibility around KLF targets in neonatal and mature neurons. However, we do offer two testable hypotheses in our Discussion of early developmental transitions driving the KLF switch (Thyroid Hormone and structured activity input). Finally, we recognize that the manuscript is too long and have edited or removed parts of the Discussion.

      Recommendations for the authors:

      Summary of Recommendations from Reviewing Editor: While the screen data is important and novel, there is consensus among the reviewers that the current study would greatly benefit by additional analyses as to the consequences of KLF knockdown on neuronal morphology and neurite outgrowth. It is recommended that this major open question be experimentally addressed. It is also recommended that additional comments/points below are addressed with changes to the figures/text as appropriate.

      Reviewer #1 (Recommendations for the authors):

      Minor points:

      (1) In Figure 2C2, it can be assumed that green represents KLF7 and red represents KLF9 but a legend indicating this would be helpful here. The Y axis is labeled as "Binned D-V Axis". Can the authors further clarify what this means? Is it referring to the different layers of the cortex? It looks like the expression pattern of KLF7 and 19 is not consistent throughout the entire dorsal-ventral axis at P7 and there is a region where KLF19 expression is higher than KLF7. Could this suggest that KLF family members may act on different types of neurons in distinct layers at different timepoints? Could the authors expand on this at all?

      We thank the reviewer for pointing this out and the appropriate legend has been added. A detailed description of our method for binning counts in the Dorsal-Ventral axis (y-axis) to normalize for differences in cortical thickness across ages is included in the methods section (see RNAScope Image Acquisition and Analysis), but text has been added to the results to clarify this. The non-uniform distribution of Klf7 across cortical layers at P2 and P7 likely indicates that that KLF switch, while conserved, occurs with variable kinetics across cortical cell types which could reflect their distinct rates of maturation in vivo (Gao et al., 2025). Minor edits have been made to the text associated with Figure 2 to call attention to this detail.

      (2) Figure 6C1 describes three clusters of DEGs. If I understand correctly, cluster 1 represents downregulated DEGs by KLF6/7 that are also developmentally downregulated between P10 and P20 and cluster 2 represents downregulated DEGs by KLF6/7 that do not change between P10 and P20. It could be interesting to see how these gene sets compare with one another. For example, are the genes in cluster 1 that decrease from P10 to P20 more related to axon regulation/cytoskeleton when compared to cluster 2 which consist of genes that do not change from P10 to P20 and perhaps may reflect other regulatory functions of KLF6/7.

      We appreciate this suggestion and we have updated Figure 6C1 to reflect this observation by highlighting genes that are part of significant enriched GO terms to allow the reader to see which cluster they belong to. Separate GO analyses performed for each downregulated gene cluster did not yield significant results and were therefore not included.

      Reviewer #3 (Recommendations for the authors):

      See above

      Additional points:

      (1) Upstream transcriptional regulation of KLFs:

      The idea that KLFs are regulated by Thyroid hormone (T3) is interesting, and the manuscript would be strengthened by exploring this idea further. This could be done, e.g. by analyzing existing snRNA-seq on genes regulated in cortical neurons by T3 (see PMID: 39178853). Similarly, KLFs were suggested to act together with AP1, and many KLFs were previously found to be activity-regulated - hence, the manuscript would be strengthened by analyzing the connection between KLFs and AP1. This could be done e.g. by re-analyzing transcriptomic and proteomic data generated e.g. by the Greenberg laboratory.

      We agree with the reviewer that the role of T3 in driving the KLF switch is intriguing and is an active area of investigation in our lab. However, these experiments are ongoing and are beyond the scope of this current manuscript. A preliminary analysis of the T3-treated snRNA-seq dataset present in Hochbaum et al., 2024 found that all major excitatory cortical cell types upregulate Klf9 upon T3 treatment while the expression of selected shared targets including Dpysl3 and Gap43 are downregulated in several cell types, supporting our hypothesis. These findings have been included in Figure S9.

      It was not our intention to suggest that there may be cooperation between AP1 and the KLF family, as AP1 motif enrichment was detected in shared upregulated genes while the KLF/Sp motif was found in shared downregulated genes. Furthermore, DEGs from both KLF knockdown experiments had exclusive enrichment for promoter KLF/Sp motifs, making it unlikely that there is widespread co-regulation of these genes by the AP1 family.

      (2) Figure 5D and Page 27, regarding Rac3:

      The authors state that " no change in accessibility in motif-positive peaks upstream of Plppr1 and Rac3 (Figures 5D and S5C)." This statement seems incorrect for Rac3 in L6 neurons where the ATAC signal is strongly reduced at P30 relative to P2. The authors might want to explain this or choose a better example.

      We acknowledge that Rac3 was a poor choice to represent a developmentally regulated Klf9/13 target with no change in promoter accessibility, and have updated Figure 5 with Atat1, a tubulin acetylase and shared KLF target, as our exemplar for this category of transcript.

      (3) Supplemental Figure 2:

      There seems to be a discrepancy between the expression levels in Panels A and C: the data in Panel A were generated from adult mouse cortex, and the levels of six KLFs are indicated as rather high - however, in Panel C, the adult levels of KLF6, 7 and 13 are rather low. How can this be explained?

      We thank the reviewer for detecting this apparent discrepancy. Most of this can be attributed to the log scale used in Figure S2A, which flattens expression differences in the moderate to high expression range. We elected to use a log scale in Figure S2A to highlight KLF family members with higher cortical expression relative to those with low or no expression. In addition, the counts displayed in this figure were batch-corrected to account for significant batch effects between libraries from additional excitatory cortical cell types included in Figure S2A, which were prepared and sequenced at different locations (Brandeis University v. Janelia Research Campus, as in Sugino et al., 2019). The additional processing step has been included in our methods section.

      (4) IGV plots in Figure 5D and Supplemental Figure 5B:

      It might be worth highlighting/shading the promoter region to see the position of the peaks relative to the promoters and TSSs.

      The TSS is indicated on all IGV plots by a dark arrow indicating the direction of transcription, so the promoter can be inferred to be immediately upstream.

      (5) Middle of page 20:

      Remove "saw"; this seems like a typo.

      (6) Page 27, reference to the plot for Plppr1:

      The plot for Plppr1 is in Supplemental Figure 5B2, not in S5C.

      (7) Page 28:

      There seems to be a typo/omission before "...general applicability of CRISPRi"

    1. Author response:

      Reviewer #1 (Recommendations for the authors):

      (1) Please report effect sizes and confidence intervals for the key comparisons in Figure 1. Where equality between conditions is central to the interpretation, use an equivalence test or soften the equal signs and associated mechanistic language. Please also clarify the permutation analysis: the Methods state that the null distributions were centered on the observed difference of deltas, whereas a permutation null would ordinarily test relative to zero.

      We thank the reviewer for these suggestions. We will report effect sizes and associated 95% confidence intervals for the key comparisons in Figure 1. We will also evaluate equivalence using defined margins based on control variability (e.g., standard deviation) and assess how conclusions depend on the margin definition. We will soften the associated mechanistic language and replace equal signs with approximate-equality symbols, clarifying that these indicate similarity without implying statistical equivalence. Finally, we will clarify the description of the permutation analysis in the Methods, including how the null distributions were constructed and centered.

      (2) Please report, for each experimental group, the numbers of cells or images, retinas, and animals, and indicate when multiple observations came from the same animal. Where observations are nested, the analysis should account for this structure using an animal-level or hierarchical bootstrap, a mixed-effects analysis, or animal-level summaries. It would be reassuring to confirm that the main pathway-specific conclusions remain robust when the animal determines the independent sample size.

      We thank the reviewer for highlighting the importance of accounting for nested observations. We will report the numbers of cells and images in the main figures. To supplementary table (Dataset S1), which already reports animal numbers, we will add the number of cells obtained from each animal. Finally, we will assess the robustness of our conclusions using animal-level summaries and mixed-effects analyses that account for observations nested within animals.

      (3) Please expand the discussion of how the changes in tOFFa temporal filtering, spatial organization, and input-output transformation are expected to affect visual coding, including responses to looming or approaching dark objects. If readily feasible, direct electrophysiological recordings of responses to an expanding dark stimulus would be informative. I do not consider these recordings necessary, however; predictions from the existing linear-nonlinear models or a more developed discussion of the expected coding consequences would be sufficient.

      We appreciate the suggestion and will use a linear-nonlinear model to predict how responses to a looming stimulus change following cone loss. We will use an expanding stimulus based on published looming protocols and apply the measured spatial and temporal filters and pass the resulting linear output through nonlinearities to predict the response. We will compare how the response changes over time in control and cone-DTR.

      Reviewer #2:

      Weaknesses:

      (1) Do the differences in compensation versus remodeling observed in ganglion cells reflect changes in the OPL? Different bipolar types may remodel their dendrites and form aberrant contacts with rods in the absence of cones. However, this would be challenging to test because there are currently no good markers for different bipolar types.

      We agree that bipolar cell remodeling could contribute to the differences in functional changes at ganglion cell level. Our functional measurements do not establish whether these differences originate in the OPL. We will address this possibility in the Discussion.

      (2) It would be interesting to determine whether these functional changes can be detected at the transcriptomic level or whether they are mediated primarily through post-translational modifications.

      We agree that determining the molecular mechanisms underlying these functional changes would be informative. As our current experiments do not distinguish between changes in gene expression and post-translational modifications, we will acknowledge the limitation and address these possibilities as directions for future work in the Discussion.

    1. Author response:

      Reviewer #1 (Public review):

      (1) The efficiency of the CRISPR/Cas9 knockout is illustrated qualitatively, but no sample size or penetrance value is reported, making it difficult for readers to judge how robust or reproducible this result is.

      We agree that quantitative documentation of the knockout experiments is needed. In the revised manuscript we will report the number of injected embryos, the number of independent experiments, and the penetrance of the albino phenotype obtained with the PKS guide RNAs, so that readers can judge the robustness and reproducibility of gene editing in M. globulus.

      (2) The gene annotation is reported to have complete PFAM domain coverage for only 75% of predicted genes, but no independent completeness metric (such as BUSCO scored against the annotated gene set rather than the assembly) is provided, leaving open whether the remaining genes are genuinely novel, partial models, or annotation artefacts.

      We have now scored BUSCO (metazoa_odb10, n = 954) directly against the annotated gene sets rather than the assemblies. The annotation of the blue male genome is 96.9% complete (S: 89.7%, D: 7.2%; F: 1.2%, M: 1.9%), and the annotation of the red female genome is 97.6% complete (S: 97.0%, D: 0.6%; F: 0.8%, M: 1.6%). These values indicate that the predicted gene sets show a high degree of completeness and that the genes lacking full PFAM domain coverage are not simply the product of fragmented or artefactual gene models. These metrics will be added to the revised manuscript.

      (3) Finally, at the time of review the NCBI BioProject accession cited for the genome and sequencing data (PRJNA1477966) could not be located, and it is not clear from the text whether this accession, once available, will include the gene annotation and RNA-seq datasets in addition to the raw genomic sequencing reads.

      We apologise for not commissioning the release of these records, the BioProject was still being processed by NCBI at the time of review. We confirm that the accession will include the gene annotations and the RNA-seq datasets in addition to the raw genomic sequencing reads, and we will verify that all records are publicly accessible before submitting our revision. We will also state explicitly in the Data Availability section which datasets are deposited under this accession.

      Reviewer #2 (Public review):

      (1) Genetic Background and Aquarium Trade Populations: A central argument of the manuscript is that M. globulus is attractive as a laboratory model because it is widely cultured in the aquarium trade and may exhibit reduced genetic variability due to captive propagation. [...] The manuscript does not provide sufficient information to evaluate these claims: how genetically representative are the sequenced individuals relative to natural populations; what is known about the provenance and breeding history of the aquarium trade stocks; are these animals derived from a small number of founder populations; is there evidence for substantial inbreeding or genetic bottlenecks within commercial brood stocks; and how similar are commercially available animals from different vendors and geographic sources?

      We thank the reviewer for raising this important point, which we have investigated directly. The two sequenced colour morphs have distinct provenances: the red individual derives from a line bred in captivity in North America, whereas the blue individual is wild-caught from the Indo-Pacific. We therefore expected the captive-bred red animal to show reduced polymorphism relative to the wild blue animal. We will thoroughly check differences of polymorphism between and across individuals from either population and discuss the results.

      (2) Presentation and Interpretation of HCR and Phalloidin Data: [...] The HCR images show detectable signal, but the expression domains are only minimally documented [...] Similarly, the phalloidin-labeled images provide limited anatomical information because the larvae are largely not labeled. I recommend: (1) adding labels identifying relevant embryonic regions and structures; (2) including arrows or overlays indicating key expression domains; (3) providing higher-magnification insets of relevant regions; (4) including selected optical sections rather than relying exclusively on 3D projections; (5) identifying known larval muscle groups in the phalloidin images; (6) improving image contrast and figure annotation where possible.

      We agree that the HCR and phalloidin panels should stand on their own for readers who are not sea urchin specialists. The revised manuscript will implement the reviewer's suggestions and provide more detailed labelling of embryonic territories and high-magnificant insets.

      Reviewer #3 (Public review):

      First, the authors should provide a thorough description of the methods used to cultivate and maintain M. globulus. This should include further details about the closed aquarium system; a schematic of the system would be insightful. Basic details about husbandry are needed, including (i) stocking densities of adults, embryos/larvae, postlarvae/juveniles; (ii) frequencies of level and water changes/top-ups; (iii) feeding regime at all phases of the life cycle.

      We agree that full documentation of the culture system is essential for uptake of the model, and we will expand the Methods accordingly. A construction schematic of the closed aquarium system, produced for us by Aquarium Connections (London, UK), is provided in Author response image 1 and will be included in the revised manuscript as a supplementary figure. The system is a three-tier rack (2000 x 1800 x 700 mm) comprising a brood-stock holding tank with egg-crate divisions, a row of settlement tanks, and a lower sump level with a UKS 200 protein skimmer, XHO algae lighting for live-feed culture, and a reverse-osmosis top-up reservoir, controlled via timer and switch boxes.

      The revised Methods will include the following husbandry parameters. Adults are stocked at 20 M. globulus per 300 L tank and fed kombu seaweed once per week. Embryos are stocked at 1 embryo per 4 mL, which corresponds to the density required at the larval stage; approximately 20% of larvae are lost by the time competency is reached. Once feeding begins, larval cultures are cleaned every three days, with water topped up at the same time. The day-by-day larval feeding and rearing schedule, from fertilisation through metamorphosis and the transition to the juvenile diet, is summarised in Author response table 1 and will be included in the revised Methods. We agree that this documentation will provide the foundation for future improvements, including shortening time to competence and maturity and standardising settlement.

      Author response table 1.

      Larval feeding and rearing schedule for M. globulus.

      Author response image 1.

      Construction plan of the closed M. globulus culture system (Aquarium Connections, drawing GC-001, rev. P1; scale 1:10 at A3). The three-tier rack (2000 x 1800 x 700 mm) houses the brood-stock holding tank with egg-crate divisions (top), settlement tanks (middle), and sump level with UKS 200 protein skimmer, XHO algae lighting, and RO top-up reservoir (bottom).

      Second, the description of the procurement and analysis of mRNA is brief, unreferenced and reads as protocols used for an established model species (e.g. what is PFA in this case - the concentration of paraformaldehyde and the buffer can vary markedly between organisms and life stages). Even the RNA extraction protocols can vary between and within species. [...] The HCR analysis, which is also scantily described, is restricted to embryonic and larval stages. Given the emphasis on the capacity of the M. globulus system to analyse all phases of the life cycle, it would be good to know if HCR can be performed on settled postlarvae, juveniles and adult tissues.

      We will substantially expand the Methods to give a complete, referenced account of fixation (including the paraformaldehyde concentration and buffer used at each stage), RNA extraction, and the HCR protocol as applied to M. globulus, rather than assuming familiarity with protocols from established models. We will also address the applicability of HCR beyond embryonic and larval stages in the revision. We agree that demonstrating in situ methods in post-settlement stages would reinforce the central premise of the model, and we will report our experience with postlarval, juvenile and adult material in the revised manuscript.

      Third, the authors should consider dividing Figure 2, which consists of confocal images of normal development, HCR results and CRISPR/Cas9 knockout results, into three separate figures that explore these studies separately. A figure on normal development could, for instance, include documentation of metamorphosis, with a suite of images of postlarval stages. A figure documenting HCR could be expanded to include more stages, higher magnification images and other genes. A figure on the Cas9 knockdown of a PKS gene can provide details on the normal expression of this gene.

      We agree that Figure 2 is currently overloaded. In the revision we plan to split it into separate figures: one devoted to confocal documentation of normal development. We intend to further include one (or more) juvenile stage presenting the HCR expression data together with the CRISPR/Cas9 knockout results, with the improved annotation described in our response to Reviewer #2.

      Fourth, there should be consideration of providing more characterisation about the protein-coding genes comprising the chromosomal region (Chr. 4) that has marked differences between sexes. This could go beyond Supplementary Table 4 and Supplementary Figure 5B, and include analysis of expression in adult tissues (this would be enhanced by having matching male and female tissues), and KEGG pathways and GO enrichments.

      We agree that the sex-differentiated region on chromosome 4 deserves fuller characterisation. In the revised manuscript we will extend the analysis of the protein-coding genes in this region beyond Supplementary Table 4 and Supplementary Figure 5B, including functional characterisation (GO and KEGG enrichment) and, where material permits, analysis of expression in adult tissues.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Hanako and colleagues demonstrated that glycolipid MPIase is essential for the TAT system, and they successfully reconstituted the TAT system in vitro for the first time. This will facilitate the understanding of the mechanism of the TAT system.

      My major points are listed below for the authors to consider:

      (1) The authors successfully reconstituted the TAT system using the purified TatA/B/C, but the translocation efficiency was much lower than that of native INV. The authors partly attributed this to the reason that "MPIase recovery would be too low to detect the TAT activity" in the Discussion part. So, what would happen to the translocation efficiency if you added more MPIase to the reconstituted system? How about the abundance of MPIase from the INV and reconstituted proteoliposomes?

      Fig. 3A shows that proteoliposomes reconstituted with active INV were inactive. However, TAT activity was observed when the proteoliposomes were fused with MPIase-containing liposomes. This indicates that MPIase was not sufficiently recovered during reconstitution. Typically, we added MPIase to proteoliposomes at 5% to phospholipids. However, increasing the amount of MPIase to exceed that of INV did not significantly increase the activity. The amounts of MPIase were stated in the text (P17, L3-4).

      (2) Why were only TatC levels measured in Figure 2C, whereas the expression levels of TatA were not detected? Also, from my observation, the amount of TatC in the third lane is lower than that in the previous two lanes.

      We also wanted to check the TatA levels, but the antibody was no longer available. Here, we confirmed that inactive INVs still contain Tat components.

      (3) The authors should explain why the TatA/B/C ratios in Figure 3C (1:1:1) and Figure 3D (10:1:1) are inconsistent.

      We examined various TatA/B/C ratios. We found that the ratio of 10:1:1 was better than 1:1:1. This was added to the text (P8, L11-16).

      (4) ~30% of the fluorescence was recovered in the membrane fraction (Figure 4A) both in the functional TAT signal sequence (RR) and in the inactivating mutant signal sequence (KK), which suggests that MPIase acts as a relatively broad recognition factor. Given that MPIase does not discriminate between RR and KK, why do un-translocated substrates remain in the cytoplasm rather than non-specifically adhering to the membrane when MPIase is depleted in vivo?

      Our results strongly suggest that MPIase acts as a signal sequence receptor. Therefore, in the absence of MPIase, the targeting of TAT precursors should be inhibited. Furthermore, because the TorA signal is less hydrophobic than the Sec signal, the level of nonspecific membrane binding would decrease.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigated the relationship between the Tat system and MPIase, a glycolipid that facilitates protein integration into the bacterial cell membrane. The TAT (twin-arginine translocation) system is a unique membrane transport machinery that exports fully folded proteins containing a twin-arginine signal peptide. Using both in vivo and in vitro approaches, the authors demonstrated that a sufficient amount of MPIase is required for Tat-dependent protein translocation. Furthermore, the authors successfully reconstituted the Tat transport system by combining recombinant TatA, TatB, TatC, MPIase, and FoF1-ATP synthase.

      Strengths:

      The reconstituted system clearly demonstrated the requirement for each component, as substrate translocation occurred only when all components were present. Based on these findings, the authors proposed a mechanistic role for MPIase in facilitating Tat-mediated membrane translocation. Previous studies have shown that MPIase is involved in Sec-dependent protein translocation and membrane protein integration, as well as YidC-dependent membrane insertion. The present study further demonstrated that MPIase also plays an essential role in the Tat translocation pathway. Overall, this work highlights the central importance of MPIase in bacterial membrane protein biogenesis and provides new insights into the molecular mechanism of Tat-dependent protein transport.

      Weaknesses:

      (1) To show the importance of the Tat system in bacterial cells, it would be good to describe in the introduction how many proteins are translocated via the Tat system.

      The E. coli K12 cells possess 27 TAT substrates. This was added to the text (P2, L14).

      (2) Figure 2B and D show that a sufficient amount of MPIase is important in SufI translocation. However, the reason why MPIase level was upregulated in the BL21 strain but not in the KS46 strain remains unexplained. The authors should address this point.

      CdsA is a rate-limiting enzyme in MPIase biosynthesis. When MPIase production needs to increase, the cdsA promoters are activated. In KS46, however, cdsA is under the control of the arabinose promoter on a plasmid, so CdsA cannot be induced even when MPIase is necessary. This was added to the text (P6, L26-27).

      (3) In Figures 4A and B, the authors explain that MPIase first works as a receptor of TorA-GFP without recognizing the RR motif. This conclusion is based on the results of the fractionation assays, where "sup" indicates the cytoplasmic and periplasmic fractions, and "ppt" indicates the membrane fraction. In Figure 4B, under the TatABC+++, (RR), +MPIase condition, the substrate is secreted most efficiently via the Tat pathway and should therefore be recovered in the periplasm fraction (sup). However, the authors point out that efficiently processed substrate was recovered in the ppt fraction rather than the sup fraction. The authors should explain why this occurred.

      We found a mistake in the processing of the results for the sample of the TatABC+++, (RR), +MPIase condition, in Fig. 4B. Therefore, we remeasured the sample and corrected the figure. The relevant text was also modified (P9, L16-29). It was found that a large part of fluorescence was recovered in the supernatant fraction. We also uploaded the raw data for the fluorescence values.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Overall, I agree with the reviewers that the conclusions are reasonably well supported by the data. I have additional scientific/editorial concerns and recommendations:

      (A) - Scientific:

      (1) P.4 lines 10-15 and Fig. 1 top-left (the effect of overexpressed TatABC).

      The authors state that the extent of Surf1 maturation increased upon overexpression of TatABC, but the band intensities for mature Surf1 in the TatABC-induced vs uninduced lanes are very similar.

      The expression was weakened: ' The level of the mature form increased....' was changed to ' The level of the mature form slightly increased....'

      (2) P.7 lines 23-24, Figure 3A. (the TAT system reconstituted in liposomes).

      The authors state that a small portion of Surf1 is successfully translocated and protected by ProK. But unlike the assay results on IMVs, the size of the mature form (i.e., translocated form) appears to be the same as the untranslocated form. Some explanation seems to be needed.

      In INV, the signal sequence is cleaved off by Lep, however, the cleavage is not coupled with translocation. In the reconstituted proteoliposomes, such cleavage is hardly observed since the membrane proteins were diluted by fusing with MPIase-containing liposomes. This was added to the text (P7, L26-29).

      (3) Based on the model (Figure 5), the interaction between the RR motifs and MPI appears non-electrostatic as the substrate binds to the sugar moieties of the MPIase, not to the pyrophosphate part. Then what is the molecular nature of the substrate-MPI interaction? Some explanation/speculation seems to be needed.

      MPIase has been identified as a factor that drives membrane protein insertion. Through the analysis of the mechanisms of insertion, we found that the glycan chain interacts directly with the transmembrane region of the membrane proteins through the numerous acetyl residues on the glycan. Moreover, we found that the positive charges of substrate membrane proteins interact with the pyrophosphate residue through the electrostatic interaction. Therefore, it is reasonable that the h region of the TAT signal binds to the MPIase glycan through the hydrophobic interactions, and the n region including the RR motif binds to the pyrophosphate residue. This was added to the text (P8, L25-27; P11, L14-17).

      (B) - Editorial:

      (1) P.2. lines 5-10. (Introduction)

      The description of the TAT-targeting signal needs to be more clearly described for a broader readership. For example, what are "h" and "c"?

      The signal sequence is composed of three regions, 'n', 'h', and 'c' from N-terminus. The n region contains positive charges including the RR motif. The next h region contains a hydrophobic stretch. The c region contains a cleavage site. This was added to the text (P2, L4-11).

      (2) P.7 lines 23-24.

      What is the rationale for using Pm-Fob-His as a control?

      This is a control for a membrane protein unrelated to the TAT system to reveal that MPIase is specifically interacts with TatABC.

      (3) P.8 lines 7-8.

      CCCP is not defined.

      CCCP (Carbonyl cyanide 3-chlorophenylhydrazone) is a protonophore. This was added to the text (P8, L18).

      Reviewer #1 (Recommendations for the authors):

      (1) The conclusions derived from Figures 1-3 heavily rely on representative immunoblotting images. Given that representative tracks can inherently introduce selection bias, how do the authors ensure the statistical robustness of these findings without quantitative bar graphs and rigorous significance analysis? To fully substantiate these interpretations, it is recommended that these immunoblotting assays be quantified across at least three independent biological replicates.

      We did not quantify the results in sections where we presented qualitative discussions. However, we performed all experiments at least three times.

      (2) Is it possible that the error values be added after the statistical values of translocation efficiency in Figures 2 and 3?

      We added the SD values to some results in Fig. 3.

      (3) "MPIase depletion of was then confirmed", an extra "of".

      Corrected.

      (4) The capitalization style of "TAT" in the whole text is suggested to be unified.

      Corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) The introduction is well articulated; however, including a paragraph on the known inhibitors might be helpful in understanding the current status. In addition, it might also help to introduce Dabs, FADDI variants, Gly-sar and MIPS.

      We thank the reviewer for the suggestion. We have included a paragraph on the substrates of hPepT2 (Lines 104-112).

      (2) The following details of modeling with AlphaFold2 should be included: how the final structure was selected, what the RMSD and structure alignment of the template are, and the final selected structure. A section on modeling with all the parameter details might be useful for reproducing the structure. In addition, specify how the alanine scanning was performed alongside the structure prediction of polymyxins.

      We appreciate the reviewer's careful assessment of our computational methodology. In the Materials and Methods section of the revised manuscript, we have added a dedicated subsection describing the detailed methods: "Structural modelling of hPepT2 and polymyxins" as shown below (Lines 435-447).

      (1) hPepT2 structure prediction: The inward-open conformation of hPepT2 was predicted using AlphaFold2 via the ColabFold implementation with default parameters, employing the MMseqs2 multiple sequence alignment pipeline. To obtain the physiologically relevant outward-open conformation of hPepT2 for substrate binding, homology modeling was performed using MODELLER with the cryo-EM structure of rabbit PepT2 in the outward-open state (PDB: 7NQK) as the template. The sequence alignment between hPepT2 and rabbit PepT2 was conducted using Clustal Omega. One hundred models were generated, and the final model was selected based on the lowest Discrete Optimized Protein Energy score and verified by Ramachandran plot analysis (with >95% of residues in favored regions).

      (2) Polymyxin structure and alanine scanning: The structure of polymyxin B1 was constructed and energy-minimized using the CHARMM36 force field. Computational alanine scanning of polymyxin B1 was performed by individually replacing each Dab (2,4-diaminobutyric acid) residue at positions 1, 3, 5, 8, and 9 with alanine using the mutagenesis wizard in PyMOL.

      (3) In the all-atom MD simulation method, detailing several parameters might help in reproducing the results: simulation time for each system, water model, system composition, protonation state, box type and dimensions, salt ions and concentration, membrane parameters and ligand parameterization methods. Also, the following details on energy minimization might be useful: minimization algorithm, number of steps for minimization and structure restraints in place.

      We thank the reviewer for the suggestion. In this study, all-atom molecular dynamics simulations were performed using the TIP3P water model for solvation. The system was placed in a rectangular simulation box with dimensions of 11 × 11 × 14 nm<sup>3</sup>. Neutralization was achieved by adding 0.15 M NaCl, resulting in a total of approximately 158,000 atoms. The membrane lipid composition consisted of 50% phosphatidylcholine, 25% phosphatidylethanolamine, and 25% phosphatidylserine, consistent with the coarse-grained molecular dynamics simulations. Force field parameters for polymyxin B1 were generated using CGenFF. Energy minimization was carried out using the steepest descent algorithm with default parameters. For production simulations, 500-ns simulations were run for sampling the outward‑open conformation, and 100-ns simulations were run for analyzing the polymyxin‑hPepT2 interaction. We have provided these methodological details in the revised manuscript (Lines 477-492).

      (4) On page 6, line 210, the MIC is used for the first time; although MIC is given in the abbreviation list, the first occurrence should have a complete name. A one-line explanation of MIC in the introduction or wherever suitable might be better but is not mandatory.

      We thank the reviewer for the suggestion. We have included the full term of MIC and its definition in the revised manuscript (Lines 236-238).

      (5) Similarly, Gly-sar is first mentioned on page 8, line 301, but its complete name is only mentioned later on page 10, line 368. This can be addressed if a short description is included in the introduction section.

      We apologize for the overlook. We have now referred to the full term of Gly-Sar and a short description on Page 3 (Lines 104-106), when it was first mentioned.

      (6) For coarse-grained MD simulation, why were 2 replicates performed? Most studies perform 3 replicates, which are also good in terms of statistics and error bar calculations. In addition, the authors should specify whether an independent minimization is done for each of the two replicates or whether the minimization step is common for both.

      In our study, coarse-grained MD simulations served as a preliminary exploration to identify the potential binding trajectory of polymyxin B to hPepT2, rather than the primary source of quantitative interaction data. The two independent CG-MD replicates yielded highly consistent final binding poses of polymyxin B at the lateral opening gate of hPepT2, serving this exploratory purpose. Furthermore, more detailed interaction energetics were subsequently characterized using all-atom MD simulations with four replicates (Fig. 4). We have now clarified in the revised Methods that independent energy minimization was performed for each of the two CG-MD replicates (Lines 465-469).

      (7) For MD simulation results, giving simulation movies in supplementary results might be a better way to show how the trajectories behaved.

      We thank the reviewer for the suggestion. We have prepared an animation of the coarse-grained MD trajectory showing the binding trajectory of polymyxin B to hPepT2, which has been submitted as a Supplementary Movie (.mp4 format). This movie visualizes how polymyxin B molecules spontaneously approach and bind to the lateral opening gate of hPepT2 over the 3-μs simulation. The Supplementary Movie is referenced in the revised Methods section (Lines 471-474).

      (8) The description of visualisation software such as VMD or PyMol is missing. The authors should specify if any visualization tool is used.

      We apologize for this omission. In the revised manuscript, we have specified in the Methods section that molecular visualizations, structural figures, and trajectory animations were prepared using PyMOL (Schrödinger) and ChimeraX (Lines 471-474).

      (9) For the mouse model study, the authors claim that FADDI-795 has no observable nephrotoxicity; however, the n=3 shows that a very small number of mouse models were used to make the assumption. In addition, the number of mice used in each experiment is not explicitly mentioned in the methods section.

      We have now added “n=3 each group” in the Methods section (Line 603). In this proof-of-concept study, we employed acute kidney damage analysis in a small number of mice to rapidly screen these analogues for nephrotoxicity. Three biological replicates were used because our mouse nephrotoxicity model is robust as demonstrated by a large number of animals in our drug discovery program (e.g. Nature Communications 2022, 13, 1625). Moreover, minimising animal use is a fundamental principle of the 3Rs (Replacement, Reduction, and Refinement).

      (10) In Table 2, the column 8 header is not visible.

      We apologize for the format error and have now fixed it in Table 2.

      Reviewer #1 (Recommendations for the authors):

      (1) The MD simulation trajectories or movies are not provided.

      As mentioned in our response to Comment 7 above, a Supplementary Movie (mp4 format) showing the CG-MD binding pathway has been provided with the revised manuscript and referenced in the Methods section.

      (2) Mouse model experiments are performed over n=3, which might not be statistically significant in case of an experiment with a large n.

      Please see our response to Comment 9 above.

      Reviewer #2 (Public review):

      (1) Several conclusions would benefit from a more cautious interpretation. A major limitation is that several transporter mutations substantially altered total or membrane protein expression, making it difficult to distinguish effects on substrate binding from indirect effects caused by impaired transporter stability or trafficking. The authors acknowledge this limitation in the Discussion, but some mechanistic conclusions remain stronger than the available evidence supports.

      We agree with the reviewer that it is difficult to draw definite conclusions when both K<sub>m</sub> and V<sub>max</sub> are altered in some mutants, especially when expression of the mutant transporter is impaired. However, we consider kinetic analysis to be more indicative of altered substrate binding, as changes in K<sub>m</sub> are generally more closely associated with substrate recognition and binding, whereas reduced transporter expression predominantly affects V<sub>max</sub>. Nonetheless, to avoid overinterpretation, we have moderated our conclusions by replacing “was/is” with “may be” (such as Lines 318, 319, 323, and 347).

      (2) Similarly, while the proposed binding model is biologically plausible and supported by mutagenesis, it remains an inferred model derived from molecular simulations rather than a direct structural determination. Statements describing the model as "validated" should therefore be moderated to indicate that the experimental data provide support rather than definitive structural confirmation.

      We thank the reviewer for the suggestion and have replaced “validated” with “explored” in the main text (Lines 298).

      (3) The translational implications are promising but remain preliminary. Although FADDI-795 demonstrated reduced nephrotoxicity in the mouse model while maintaining antibacterial activity, no pharmacokinetic studies were presented to demonstrate reduced renal accumulation or altered tissue distribution, and additional efficacy studies in infection models would further strengthen the therapeutic claims.

      We thank the reviewer for the suggestion. The current study is a proof-of-concept investigation to determine whether the structure-interaction relationship (SIR) model of hPepT2 and polymyxin can be used to design new lipopeptides with reduced nephrotoxicity. Following screening of the newly designed lipopeptides, we will select the most promising candidate for comprehensive pharmacological evaluations, including pharmacokinetics, toxicology, and efficacy.

      Reviewer #2 (Recommendations for the authors):

      (1) Strengthen the interpretation of the mutagenesis data. For transporter mutants that exhibit reduced total or cell surface expression, consider discussing more explicitly the extent to which reduced polymyxin uptake may result from altered protein stability or trafficking rather than direct effects on substrate recognition. Where possible, clarify this distinction throughout the Results and Discussion.

      We thank the reviewer for the recommendation. As suggested, we have expanded the discussion on the mechanisms underlying the altered protein expression of the hPepT2 mutants (Lines 325-333). 

      (2) Moderate statements regarding structural validation. The molecular dynamics simulations, mutagenesis, and functional studies provide strong support for the proposed structure-interaction model; however, wording such as "validated the first SIR model" could be softened to reflect that the binding model remains computationally inferred rather than directly resolved by structural methods.

      We thank the reviewer for the recommendation. We have revised the statement as described in our response to Comment 2 above.

      (3) Expand the discussion of study limitations. A more explicit discussion of the limitations associated with molecular dynamics predictions, the influence of altered transporter expression on functional interpretation, and the lack of direct pharmacokinetic measurements of renal accumulation would improve the balance of the manuscript.

      We thank the reviewer for the recommendation. We have expanded the discussion by elaborating the limitations of the molecular dynamics predictions (Lines 289-291), the effect of altered transporter expression on functional interpretation (Lines 324-333), and the lack of direct pharmacokinetic and renal accumulation measurements (Lines 393-401).

      (4) Provide additional methodological detail where appropriate. Please clarify the number of biological replicates used for each experimental approach, report exact P values where practical, and indicate whether assumptions for the statistical tests were assessed.

      We thank the reviewer for the recommendation. We have provided more information on experimental replicates in the figure legends. P values have been added (Lines 208-217), and the statistical tests used for every analysis are now specified in the corresponding table or figure legends.

      (5) Future validation of the lead compound. Although this may fall outside the scope of the present manuscript, it would be helpful to briefly discuss future studies aimed at evaluating the pharmacokinetics, renal exposure, and efficacy of FADDI-795 in relevant infection models, as these will be important for establishing its translational potential.

      We thank the reviewer for the recommendation. We have now incorporated this discussion into the revised manuscript (Lines 396-399).

      (6) Improve figure presentation. Several figure legends could include additional methodological information to make the figures more self-contained. In particular, indicating sample sizes, statistical tests used, and definitions of error bars would improve readability.

      We thank the reviewer for the recommendation. We have added this information to the figure legends as suggested.

      (7) Language and style. The manuscript would benefit from careful language editing to improve grammar, sentence structure, and readability. Several long sentences in the Discussion could be shortened, and minor typographical errors should be corrected throughout the manuscript.

      We thank the reviewer for the recommendation. We have revised the long sentences where possible and carefully proofread the manuscript to improve grammar, sentence structure, and readability. The revised manuscript has also been reviewed by a native English speaker.

      (8) Terminology. Please ensure that abbreviations such as "SIR" are defined at first use and are used consistently throughout the manuscript.

      We have confirmed that SIR is defined when first introduced in both the Abstract and main text (Lines 47 and 120).

      (9) Minor corrections.

      - Check for consistency in the naming of hPepT2/hPEPT2 throughout the manuscript.

      - Verify that all figures, supplementary figures, and tables are cited sequentially in the text.

      - Consider reporting confidence intervals alongside kinetic parameters (K<sub>m</sub> and V<sub>max</sub>) where appropriate to facilitate interpretation.

      We thank the reviewer for the recommendation. We have revised the manuscript accordingly. “hPepT2” is now used consistently throughout the text. We have revised the order of the supplementary figures to ensure that they are cited sequentially in the text; and included the 95% confidence interval of the K<sub>m</sub> and V<sub>max</sub> values of the D215A mutant compared to the wild type (Lines 208-212).

      Reviewer #3 (Public review):

      (1) Interactions of Polymyxin B with kidney proteins were not demonstrable in vivo, and with reliable technologies such as X-ray or NMR.

      We thank the reviewer for the comment. To the best of our knowledge, polymyxin-induced nephrotoxicity is primarily driven by renal tubular accumulation, followed by mitochondrial dysfunction, oxidative stress, inflammatory activation, and apoptosis, ultimately resulting in proximal tubular injury. Little is known about the molecular interactions of polymyxins with kidney proteins.

      Reviewer #3 (Recommendations for the authors):

      (1) Title: I suggest that you bring out the actual findings to convey the message.

      We thank the reviewer for the comment. We have changed the title to “Lateral opening site of human oligopeptide transporter 2 plays a key role in the interaction with polymyxins”.

      (2) Line 133: Provide data for these 10 molecules (possibly by supplementary file/table).

      We thank the reviewer for the suggestion. As described in the Methods, ten coarse-grained polymyxin molecules were randomly inserted into the upper water layer of the simulation system (Lines 452-456). Because these molecules are structurally identical and share the same force field parameters, they differ only in their initial random positions. Thus, we believe that presenting the coordinates or labelling the individual molecules would not provide additional information and can be misleading.

      (3) Line 135: Figure S3 mentioned prior Figure S2.

      We apologize for the error. We have now revised the sequence of the supplementary figures.

      (4) Lines 159 - 171: Cite Figure 5 a - b appropriately. Describe the obtained results fully.

      We thank the reviewer for the suggestion. We have expanded the relevant section to include a more detailed description of the results, with reference to Fig. 5a and 5b (Lines 187-193).

      (5) Line 208: Was cytotoxicity performed to validate this claim?

      We did not specifically evaluate the cytotoxicity of the lipopeptides in cultured cells because the cellular uptake data already indicated their potential toxicity. Importantly, histopathological assessment in animal models provides the most reliable evaluation of their in vivo toxicity (Nature Communications 2022, 13, 1625).

      (6) Line 509: Provide the approval number.

      The animal ethics approval number is AEC37419, which has been added to the “Ethics approval and consent to participate” section.

      (7) Table 2: Add the Standard deviation and the number of replicates performed.

      We thank the reviewer for the suggestion. We have added “n = 3 mice per group” to the table legend. As described in the Methods (Lines 590-597), the MIC of each lipopeptide was determined. MIC values are reported as absolute concentrations rather than continuous variables; therefore, reporting the mean ± standard deviation is not applicable. Likewise, the maximum kidney Semi-Quantitative Score (SQS) was determined by histological examination and is presented as the representative maximum score for each treatment group. Accordingly, the calculation of a mean and standard deviation is not appropriate for either MIC or SQS values presented in Table 2.

      (8) The authors need to indicate the number of replicates performed in the methods section.

      We thank the reviewer for the suggestion. As mentioned in our response to Recommendation 6 from Reviewer 2 above, we have added this information to the figure legends as suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the editors and reviewers for their thorough evaluation of our manuscript. We are grateful for their enthusiastic support of our comparative approach to understanding cardiac regeneration across fish lineages, and we appreciate the care with which they captured and summarized the key findings of our study. Their constructive suggestions have been invaluable in strengthening the work, and we are pleased to submit the revised manuscript enclosed herewith.

      The revised manuscript contains 7 main figures and 8 supplementary figures, representing the addition of two figures relative to the previous version. Specifically, we have included additional Podocalyxin immunostaining to clarify our characterization of ventricular vascularization, and Picrosirius Red staining to improve the characterization of fibrillar collagen in uninjured Xiphophorus hearts relative to zebrafish. We have also refined our RNA-seq analysis by manually curating well-defined markers of cardiomyocyte cell fate change and immune system activation, such that the manuscript no longer relies on a single-cell RNA-seq dataset. All remaining revisions have been made in direct response to the reviewers' comments, as detailed below.

      Reviewer #1 (Public review):

      Minor Weaknesses:

      Transcriptomic analysis was only done for one time point. Different time points could be included to validate whether some processes occur at different time points. But this can be done in the future for more detailed studies.

      We agree with this valuable suggestion. While RNA-seq at multiple time points would undoubtedly enrich our understanding of the temporal dynamics of Xiphophorus heart regeneration, such an expansion would constitute a substantial independent study extending well beyond the scope of this initial characterization. We agree that this represents a compelling avenue for future investigation, and we have acknowledged this limitation explicitly in the Discussion section of the revised manuscript (lines 357-363).

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 5B and quantification in Figure 5D seem inconsistent. The numbers of Mpx neutrophils at Day 7 seem the same as Day 14 and 30, but overall, there are very few Mpx+ cells. The authors should examine earlier time points, such as Day 1 and 3, to determine if neutrophil and macrophage infiltration is different in zebrafish.

      We thank the reviewer for this careful observation. We agree that the apparent similarity in Mpx<sup>+</sup> cell numbers across Day 7, 14, and 30 in the original figures warranted closer examination of earlier time points. We have therefore performed new Mpx and L-plastin immunostaining at 1 and 3 dpci in Xiphophorus hearts. These analyses confirm low immune cell infiltration at these early time points, with Mpx<sup>+</sup> cell numbers peaking at 14 dpci, which is a pattern strikingly different from the rapid and robust neutrophil infiltration observed in zebrafish within the first few days post-injury. These results reinforce our conclusion that the inflammatory response in Xiphophorus is both delayed and attenuated relative to zebrafish, and the updated data are now presented in Figure 6.

      (2) Their transcriptomic data at 7 days post injury suggest that TGFß signaling was not activated after injury, and tenascin C was not expressed in platyfish. The authors might check whether TGFB signaling is activated and tenascin is expressed at later time points, since platyfish show persistent scarring. This may help determine whether the molecular mechanisms of scarring in platyfish are the same as in zebrafish, and it just happens late.

      We think it is a good suggestion. Unfortunately, our anti-TnC antibodies do not produce a clear or specific signal in platyfish tissue, precluding immunofluorescence-based analysis of TnC expression at later regenerative stages. To partially address this limitation, we have included transcript abundance data for several genes associated with TGF-β signaling in the new Figure 5, which provides an initial view of this pathway's activity during Xiphophorus heart regeneration. Nevertheless, a thorough characterization of TGF-β activity and TnC expression, especially including protein-level validation and temporal profiling, would require dedicated methodological development and constitutes a new study. We therefore consider this to be an important avenue for future investigation rather than a component of the present initial characterization.

      (3) The PCNA staining BrdU labeling experiments suggest that proliferating cardiomyocytes are not maintained even though they re-enter the cell cycle. Since the platyfish hearts lack the proliferative compact cardiomyocytes that account for myocardial regeneration in zebrafish, do the authors suggest that their trabecular cardiomyocytes proliferate or that they only undergo DNA synthesis?

      This question touches on a central unresolved aspect of our findings. Our results indicate that a subset of cardiomyocytes re-enter the cell cycle and undergo DNA synthesis at 7 and 14 dpci, as evidenced by PCNA staining (new Figure 7D). However, BrdU incorporation combined with Tropomyosin immunostaining reveals an absence of BrdU<sup>+</sup>/Tropomyosin<sup>+</sup> cardiomyocytes within the border zone myocardium (new Figure 7F), indicating that newly synthesized DNA does not translate into efficient cardiomyocyte repopulation of the injured area. Whether the detected S-phase entry is followed by mitosis, or whether these cardiomyocytes undergo DNA synthesis without completing cell division, a phenomenon known as endoreplication, which has been described in other cardiac contexts, remains to be determined. Resolving this question would require live imaging or mitotic marker analyses beyond the scope of the present study, and we have highlighted this as an important open question in the revised Discussion.

      (4) Line 136-137. The authors might clarify what they meant by "N2.261 antibody recognizes different myosin types in zebrafish versus platyfish". Are these N2.261+ cardiomyocytes in the atria of platyfish also immature cardiomyocytes, but are there more immature cardiomyocytes in platyfish than in zebrafish?

      We thank the reviewer for this good question. In zebrafish, N2.261 has been established as a marker of immature cardiomyocytes in larvae and regenerating myocardium; however, we cannot currently conclude that N2.261-positive atrial cardiomyocytes in platyfish represent an analogous immature population. Rather, the broad atrial immunoreactivity in platyfish most likely reflects the presence of the N2.261 epitope within the dominant atrial myosin heavy chain isoform of this species. Thus, platyfish atrial myosin shares greater sequence similarity with the antibody's target epitope than the corresponding zebrafish atrial myosin does. In other words, the differential labeling pattern between species is more likely attributable to evolutionary divergence in myosin heavy chain sequences than to differences in cardiomyocyte maturation state. To directly identify which amino acid residues are essential for N2.261 immunoreactivity and to resolve these evolutionary differences, we have initiated a dedicated epitope-mapping study.

      To clarify this interpretation in the manuscript, we have replaced the previous concluding statement with the following: "Together, these findings indicate that N2.261 recognizes distinct myosin heavy chain isoforms in zebrafish and Xiphophorus: an embryonic ventricular isoform in the former and an atrial isoform in the latter. The differential labeling pattern between species is more likely attributable to evolutionary divergence in myosin heavy chain sequences than to differences in cardiomyocyte maturation state, reflecting the substantial lineage-specific reshaping of cardiac myosin repertoires that has occurred between cyprinids and poeciliids."

      (5) Line 212-"Interspecies comparison revealed that upon cryoinjury, 199 and 268 orthologous gene transcripts were more abundant in zebrafish than in platyfish, respectively". Does the author mean that "more abundant in zebrafish than in platyfish and vice versa"?

      We thank the reviewer for flagging this ambiguity. The original sentence was indeed unclear, and we have revised it to read: "Interspecies comparison revealed that upon cryoinjury, 199 orthologous gene transcripts showed a higher log₂FC in zebrafish than in platyfish, while 268 showed the opposite pattern, indicating that the two species mount distinct transcriptional responses to cardiac injury." We believe this phrasing now unambiguously conveys that the comparison is bidirectional.

      Reviewer #2 (Public review):

      Major comments

      (1) Title selection

      The title the authors chose suggests that platyfish and swordtails "partially regenerate," but I do wonder how much these animals truly regenerate. This may be a semantic discussion and a matter of personal preference. Still, based on other significant work on regenerative capacity (see, for example, the landmark cavefish regeneration paper PMID: 30462998 or work on medaka PMID: 24947076), the persistence of such a prominent fibrotic scar would be considered a minimal regenerative capacity. Measuring this "partial regeneration" more precisely by comparing zebrafish with platyfish and swordtails would also greatly strengthen the comparisons made here - see below.

      The same can be said about line 152-153 - do these hearts "regenerate" with deformation and partial scarring, or would it be more fair to say that they are "healed" or "repaired" with a process that involves fibrosis?

      We thank the reviewer for raising this conceptual point, and we appreciate the references to the cavefish and medaka literature. We acknowledge that the term "partially regenerate" requires careful justification given the persistence of a substantial collagenous scar at the injury site. We have retained this terminology for the following reasons:

      (1) The bulging wound undergoes resorption over time, suggesting that the initial structural deformation is transient rather than permanent.

      (2) Wound size is significantly reduced and the proportion of hearts retaining visible injury decreases over the course of the experiment.

      (3) Cardiomyocyte proliferation is detectably elevated in the myocardium at 7 and 14 dpci, indicating that some regenerative machinery is engaged.

      We would also note that partial heart regeneration accompanied by residual scarring has been described in newts, a classically regenerative vertebrate. This suggests that the boundary between regeneration and fibrotic repair is not always clear-cut. Rather than a binary distinction, the cardiac injury response may exist on a continuum, where hallmarks of regeneration, such as wound resorption, reduced scar size, and cardiomyocyte proliferation, can coexist with persistent fibrosis. We therefore believe that "partial regeneration with persistent scarring" accurately and honestly reflects our findings, and we have refined the relevant passages in the manuscript, including the indicated lines, to make this nuance explicit.

      (2) Cross-species comparisons

      Having two species of livebearers strengthens the findings of this paper, but the presentation of results from both species is inconsistent. For example, the reader should not be asked to assume that the architecture of the swordtail ventricle is similar to that of the platyfish (line 125). The same applies to the presence or absence of coronary vessels (Figure 1), the reduction in wound area over time (Figure 3), and the immune system's response (Figure 5). Most importantly, the authors miss an opportunity to move from qualitative observations to quantifying the "partial regeneration" phenotype they observe. Specifically, providing a side-by-side comparison between these new species and zebrafish would help define the extent of differences in regeneration potential. For instance, in Figure 6, while the authors provide excellent quantification of PCNA staining in platyfish, these data are less meaningful without a direct comparison with zebrafish results. The same applies to Figures 6E and 6F - although differences are noted, quantifying these results would enable a more rigorous assessment of the process.

      We thank the reviewer for this constructive critique. We agree that greater consistency in the cross-species presentation strengthens the comparative framework of the paper, and we have made several additions to address this.

      To document the cardiac architecture of swordtails explicitly, rather than asking the reader to assume similarity with platyfish, we have added the following data:

      (1) New supplementary figure (Figure S1) dedicated to the swordtail ventricle, incorporating 1) AFOG, 2) Picrosirius Red, 3) Alkaline Phosphatase Assay for the vasculature, 4) and immunostaining against Fibronectin, N2.261, and F-actin.

      (2) The dynamics of swordtail heart regeneration across 7, 14, 30, 60, and 90 dpci are presented in a separate supplementary figure (Figure S5), which also includes quantification of wound size over time. These data show that the wound, representing approximately 20% of ventricular area at 7 dpci, is reduced to approximately 2.5% by 60–90 dpci, providing quantitative support for our characterization of partial regeneration in this species.

      We appreciate the reviewer's point regarding side-by-side quantitative comparisons with zebrafish for markers such as PCNA. We respectfully note, however, that cardiomyocyte proliferation in zebrafish following cryoinjury is extensively documented across multiple independent studies, and we consider this body of evidence sufficient to contextualize our platyfish findings without requiring full parallel quantification. Nevertheless, in response to this comment, we have included side-by-side zebrafish and platyfish data for BrdU staining in the new Figure 7, with zebrafish quantification drawn from our previously published dataset (Sallin et al., 2015). We believe this addition meaningfully strengthens the cross-species comparison at this key figure while remaining within the scope of the present study.

      (3) Lack of coronary vasculature

      There is a growing body of evidence highlighting the importance of the coronary vessels during zebrafish heart regeneration (PMIDs: 27647901, 31743664). Surprisingly, this finding has not been integrated or discussed in the context of this literature.

      The results of the alkaline phosphatase assay and anti-podocalyxin-2 staining appear inconsistent. Specifically, in Supplementary Figure 1L-M, we can see some vessels covering the bulbus arteriosus and also what appears to be a signal in the ventricle. However, in Figures 1 K and 1L, we cannot see any vessels, even in the bulbus. The authors should also be more rigorous and add a description of how many animals were analyzed, their ages, and sizes. In zebrafish, the formation of the coronary arteries appears to depend on animal size and age. With the data provided, we cannot say whether this is a one-time observation or a consistent finding across many animals at different ages and across both species.

      We thank the reviewer for raising this point and for directing us to the relevant literature. We agree that the role of coronary vasculature in zebrafish heart regeneration is an important and growing area of research, and we have now integrated a discussion of this evidence into the manuscript, highlighting the contrast with the avascular Xiphophorus ventricle and its potential implications for regenerative capacity.

      Regarding the apparent inconsistency between the alkaline phosphatase assay and the anti-Podocalyxin (anti-Podxl) staining, we thank the reviewer for this careful observation. We have performed additional staining to resolve this discrepancy and conclude that the two approaches, rather than being inconsistent, reflect distinct and complementary aspects of ventricular organization.

      To better characterize the anti-Podxl signal, we performed immunofluorescence on thick (50 µm) sections of both platyfish and zebrafish hearts. In contrast to zebrafish, anti-Podxl immunoreactivity in platyfish is confined to a very thin outer layer of the myocardium. This subtle signal accounts for the weak staining previously observed in whole-ventricle preparations and was already visible, though not highlighted, in the original figures (see white arrow in new Supplementary Figure S3C, H, O–R). A direct comparison of anti-Podxl staining across species (Supplementary Figure S3A, D, N, R) clearly demonstrates the absence of coronary vascularization in the platyfish ventricle. Consistent with this, no vascular signal is detected in the platyfish bulbus arteriosus (Supplementary Figure S3I).

      To further clarify the nature of the outer myocardial layer in Xiphophorus, we performed Picrosirius Red (PSR) staining, which selectively labels fibrillar collagen, across all three species. Whereas the compact outer myocardium of zebrafish is PSR-negative, a thin but distinct PSR-positive layer is present at the ventricular surface of both platyfish and swordtails, closely mirroring the anti-Podx1 staining pattern.

      Taken together, these findings indicate that the outer ventricular layer of Xiphophorus fish is composed of a thin collagen- and Podxl-immunoreactive matrix, which is structurally distinct from the vascularized compact myocardium of zebrafish and is consistent with the absence of coronary vessels in these species.

      Finally, we acknowledge the reviewer's request for greater rigor regarding sample sizes, animal ages, and body sizes. These details have now been added to the Methods section. We note that, in line with the reviewer's observation regarding zebrafish coronary development, we have ensured that animals of comparable size and age ranges are described for each species to allow meaningful cross-species interpretation.

      (4) The link between livebearers' responses and pseudoaneurysms is overstated. This work is already extremely relevant without trying to make it medically oriented.

      We agree that the clinical parallel with pseudoaneurysms was overstated in the original manuscript, and that the comparative and evolutionary relevance of our findings stands on its own merits without requiring a medical framing. We have accordingly removed the term from the Results section and now invoke it only briefly at the close of the Discussion, where a concise note on broader medical relevance is appropriate without overshaping the narrative of the paper.

      Reviewer #2 (Recommendations for the authors):

      (1) The description of the N2.261 staining (Figure 1 and Supplementary Figure 1) is entirely irrelevant for the rest of the manuscript. One wonders why the authors have not used this antibody to characterize the presence or absence of "dedifferentiated" muscle after injury. Given that they make a point later about potential differences in dedifferentiation in livebearers, this is a missed opportunity to address it using tools the authors have characterized extensively in zebrafish.

      This is a valuable suggestion that we have now addressed. We performed N2.261 immunostaining on injured platyfish and swordtail hearts at 7 dpci, and included the results as an additional supplementary figure (Suppl. Fig. S4C). Importantly, N2.261 immunoreactivity was not detected in the peri-injury zone of the myocardium, suggesting that, unlike in regenerating zebrafish hearts, the embryonic cardiac myosin heavy chain isoform is not upregulated at the injury site in platyfish. This finding is particularly informative given the broad atrial N2.261 reactivity observed in intact platyfish hearts: the absence of enhanced staining in the injury zone argues against a dedifferentiation-associated upregulation of this isoform and instead suggests that platyfish cardiomyocytes may not undergo the same embryonic gene re-expression program that characterizes zebrafish cardiac regeneration. This result therefore directly informs our later discussion of potential differences in dedifferentiation between zebrafish and livebearers, and we have updated the relevant section of the manuscript accordingly.

      (2) Many of the markers highlighted here as part of the differential gene expression analysis are not the most canonical ones, and it is unclear how the authors selected them. For example, in Figure 5, anxa2a, hlx1, and nup153 are presented as macrophage markers. However, anxa2a appears to be expressed predominantly in the endocardium in response to injury (see PMID: 32341028), and other markers (mpeg, mfap4, etc) would have been more consistent as macrophage markers according to other literature. The same is true for cardiomyocyte proliferation markers and dedifferentiation markers in Figure 6. In this last case, N2.261 could have been used as reported by the authors before. L-plastin is a pan-leukocyte marker, not a macrophage marker. This should be corrected (Figure 5 and lines 258-279).

      We agree with this critique that several of the markers used in the original analysis were neither sufficiently canonical nor specific for the cell populations we intended to characterize. We apologize for this oversight. In response, we have moved away from scRNA-seq-derived marker sets and now rely entirely on manually curated markers drawn from the established literature, as detailed below.

      (1) Cardiomyocyte cell fate change (new Figure 7, Supplementary Figure S8): We now report the transcript abundance of genes encoding proteins with well-documented roles in the transcriptional reprogramming associated with cardiomyocyte dedifferentiation and redifferentiation. These include members of the Activator Protein-1 (AP-1) complex, the SWI/SNF chromatin remodeling complex, and the transcriptional coactivator cited4a, alongside established markers of cardiomyocyte proliferation (cx43) and sarcomere reassembly (the Rbfox family). We also include myh7 transcript abundance, which in zebrafish is upregulated in dedifferentiated cardiomyocytes.

      (2) Immune response (new Figure 6, Supplementary Figure S7): New Figure 6 focuses on pan-leukocyte markers, while Supplementary Figure S7 presents genes involved in innate myeloid activation and inflammation, including components of the NF-κB/TNF-α pathway, TLR signaling, and inflammatory regulation, as well as specific markers of neutrophils (mpx) and macrophages (mpeg1, mfap4), in line with the markers recommended by the reviewer.

      Regarding L-plastin, we thank the reviewer for this correction. The text has been amended throughout to describe L-plastin accurately as a pan-leukocyte marker rather than a macrophage/phagocyte marker.

      (3) In several instances, the authors reference papers without citing the original source. In many other instances, there are some oversights regarding citations. I would advise revising many of these:

      We thank the reviewer for drawing our attention to these citation oversights. We have carefully revised the reference list throughout the manuscript to ensure that original discovery papers are cited alongside, or in place of, review articles wherever appropriate.

      (4) Lines 49-51 - several very interesting reviews of cardiac regeneration are listed here, but referencing the source of the discovery is always most rigorous.

      For lines 49–51, we have replaced the previous review citations with three original research papers reporting the discovery of cardiac regeneration in zebrafish and three reporting it in axolotl. We have applied the same principle of prioritizing primary sources across all other instances flagged by the reviewer.

      (5) Line 76. When discussing the recovery of the muscle after cryoinjury, Poss et al. 2002 shouldn't be referenced. Sánchez-Iranzo et al 2018 should be replaced by González-Rosa et al. 2011, which is the contemporary manuscript to those of Chablais 2011 and Schnabel 2011.

      Done

      (6) Line 284 - when discussing cardiomyocyte dedifferentiation, it would be fair to reference also Kikuchi et al 2010.

      Done

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The Neuronal microtubule cytoskeleton is essential for long-range transport in axons and dendrites. The axon-specific plus-end out microtubule organization vs the dendritic-specific plus-end in organization allows for selective transport into each neurite, setting up neuronal polarity. In addition, the dendritic microtubule organization is thought to be important for dendritic pruning in Drosophila during metamorphosis. However, the precise mechanisms that organize microtubules in neurons are still incompletely understood.

      In the current manuscript, the authors describe the spectraplakin protein Shot as important in developmental dendritic pruning. They find that Shot has dendritic microtubule polarity defects, which, based on their rescues and previous work, is likely the reason for the pruning defect.

      Since Shot is a known actin-microtubule crosslinker, they also investigate the putative role of actin and find that actin is also important for dendritic pruning. Finally, they find that several factors that have been shown to function as a dendritic MTOC in C. elegans also show a defect in Drosophila upon depletion.

      Strengths:

      Overall, this work was technically well-performed, using advanced genetics and imaging. The author reports some interesting findings identifying new players for dendritic microtubule organization and pruning.

      We thank reviewer 1 for this assessment.

      Weaknesses:

      (1) The evidence for Shot interacting with actin for its functioning is contradictory. The Shot lacking the actin interaction domain did not rescue the mutant; however, it also has a strong toxic effect upon overexpression in wildtype (Figure S3), so a potential rescue may be masked. Moreover, the C-terminus-only construct, which carries the GAS2-like domain, was sufficient to rescue the pruning. This actually suggests that MT bundling/stabilization is the main function of Shot (and no actin binding is needed).

      Thank you for this comment. We agree that the rescue and overexpression experiments with UAS-Shot ΔCH1 (lacking only the first actin-binding domain) and UAS-Shot CTerm (containing only the MT binding domain) allow for other interpretations than ours that both features are necessary for Shot function in dendrites. To address this issue, we tested two more Shot domain deletion mutants, Shot ΔABD (lacking both actin-binding domains) and Shot ΔCTail, which lacks the MT binding domain (Figure 3). Neither of the latter two constructs can rescue the pruning and MT orientation phenotypes of shot mutants, confirming that actin binding is required for these functions in the context of full-length Shot. Importantly, in contrast to the previously used Shot ΔCH1, neither Shot ΔABD nor Shot ΔCTail inhibit pruning upon overexpression, ruling out other interpretations of overexpression toxicity (Fig. S2). The observation that Shot CTerm (containing only the MT binding domain) is sufficient to rescue indicates that actin binding is only required for Shot function in the context of full-length Shot. Interestingly, this is reminiscent of a proposed autoregulation mechanism where actin binding disinhibits Shot's MT binding domain (Applewhite 2013).

      Figure 3, Figure S2 (Shot domain analysis): We added two new constructs (Shot ΔABD, Shot ΔCtail) and a Shot BAC as additional control in Figure 3.

      (2) On the other hand, actin depolymerization leads to some microtubule defects and subtle changes in shot localization in young neurons (not old ones). More importantly, it did not enhance the microtubule or pruning defects of the Shot domain, suggesting these act in the same pathway. Interesting to note is that Mical expression led to microtubule defects but not to pruning defects. This argues that MT organization effects alone are not enough to cause pruning defects. This may be good to discuss.

      Thank you for this comment. We think that microtubule orientation defects are tightly linked to dendrite pruning defects, as all known conditions causing orientation defects either cause pruning defects, or they act as genetic enhancers of pruning mutants. We now clarify this in the introduction citing the example of EB1 RNAi (causes orientation defects (Matties et al., 2010), enhances pruning mutants (Herzmann et al., 2018)). This is also the case for Mical overexpression, which causes orientation defects and enhances the pruning defects caused by Shot RNAi. To make this point clearer, we now also show representative images of pruning defects caused by either Shot RNAi alone or Shot RNAi combined with Mical overexpression.

      Introduction - we explain better the genetic relationship between microtubule orientation and pruning. Figure 4 (phenotypic enhancement of shot knockdown by Mical overexpression): added representative images.

      (3) For the actin depolymerization, the authors used overexpression of the actin-oxidizing Mical protein. However, Mical may have another target. It would be good to validate key findings with better characterized actin targeting tools.

      Thank you for this comment. To our knowledge, there are no good tools to globally manipulate actin dynamics in vivo in a similar manner to, e. g., latrunculin in cultured cells. According to the literature, Micals are highly specific for actin. To enhance our analysis, we assessed the effect of Mical overexpression on actin distribution in c4da neurons (as assessed by LifeAct-GFP) and found significant effects at the first instar.

      Added new analysis of effect of Mical overexpression on Lifeact::GFP distribution in c4da neurons (new Fig. S3), added references for Mical targets.

      (4) In analogy to C. elegans, where RAB-11 functions as a ncMTOC to set up microtubules in dendrites, the authors investigated the role of these in Drosophila. Interestingly, they find that rab-11 also colocalizes to gamma tubulin and its depletion leads to some microtubule defects. Furthermore, they find a genetic interaction between these components and Shot; however, this does not prove that these components act together (if at all, it would be the opposite). This should be made more clear. What would be needed to connect these is to address RAB-11 localization + gamma-tubulin upon shot depletion.

      All components studied in this manuscript lead to a partial reversal of microtubules in the dendrite. However, it is not clear from how the data is represented if the microtubule defect is subtle in all animals or whether it is partially penetrant stronger effect (a few animals/neurons have a strong phenotype). This is relevant as this may suggest that other mechanisms are also required for this organization, and it would make it markedly different from C. elegans. This should be discussed and potentially represented differently.

      Thank you for this comment. We agree that the genetic interaction with pruning as readout does not prove that Rab11 and Shot act in the exact same pathway during establishment of microtubule organization. To address this question, we live-imaged Rab11::GFP vesicles in control and Shot knockdown neurons and found no significant difference in motile behavior (new Figure 7). As Rab11::GFP also is not visibly enriched at dendrite tips at the first instar stage, we cannot confidently state that a Rab11 tip MTOC exists in Drosophila neurons. Still, we only see EB1 comets originating from Rab11 puncta at the first instar, but not the third. We therefore rephrased and toned down our interpretation and state now that Rab11 may be a component of a developmentally transient MTOC but likely acts independently of Shot.

      Unfortunately, we do not see many comets per neuron with our fluorescently tagged EB1 constructs, indicating that only a fraction of them are labeled. This makes it very difficult to judge severity between single neurons. The impression is that usually only some comets per neuron have reversed orientation, which would be consistent with several independent mechanisms for orientation.

      Reviewer #2 (Public review):

      Summary:

      In their manuscript, the authors reveal that the spectraplakin Shot, which can bind both microtubules and actin, is essential for the proper pruning of dendrites in a developing Drosophila model. A molecular basis for the coordination of these two cytoskeletons during neuronal development has been elusive, and the authors' data point to the role of Shot in regulating microtubule polarity and growth through one of its actin-binding domains. The authors also propose an intriguing new activity for a spectraplakin: functioning as part of a microtubule-organizing center (MTOC).

      Strengths:

      (1) A strength of the manuscript is the authors' data supporting the idea that Shot regulates dendrite pruning via its actin-binding CH1 domain and that this domain is also implicated in Shot's ability to regulate microtubule polarity and growth (although see comments below); these data are consistent with the authors' model that Shot acts through both the actin and microtubule cytoskeletons to regulate neuronal development.

      (2) Another strength of the manuscript is the data in support of Rab11 functioning as an MTOC in young larvae but not older larvae; this is an important finding that may resolve some debates in the literature. The finding that Rab11 and Msps coimmunoprecipitate is nice evidence in support of the idea that Rab11(+) endosomes serve as MTOCs.

      Thank you for these comments.

      Weaknesses:

      (1) A significant, major concern is that most of the authors' main conclusions are not (well) supported, in particular, the model that Shot functions as part of an MTOC. The story has many interesting components, but lacks the experimental depth to support the authors' claims.

      Thank you for this comment. We agree that we cannot conclusively state that Shot is part of a classical MTOC where microtubules are nucleated and therefore tone down our conclusions regarding this.

      (2) One of the authors' central claims is that Shot functions as part of a non-centrosomal MTOC, presumably a MTOC anchored on Rab11(+) endosomes. For example, in the Introduction, last paragraph, the authors summarize their model: "Shot localizes to dendrite tips in an actin-dependent manner where it recruits factors cooperating with an early-acting, Rab11-dependent MTOC." This statement is not supported. The authors do not show any data that Shot localizes with Rab11 or that Rab11 localization or its MTOC activity is affected by the loss of Shot (or otherwise manipulating Shot). A genetic interaction between Shot and Rab11 is not sufficient to support this claim, which relies on the proteins functioning together at a certain place and time. On a related note, the claim that Shot localization to dendrite tips is actin-dependent is not well supported: the authors show that the CH1 domain is needed to enrich Shot at dendrite tips, but they do not directly manipulate actin (it would be helpful if the authors showed the overexpression of Mical disrupted actin, as they predict).

      Thank you for these comments. In response, we tested whether Shot knockdown affects Rab11 vesicle motility in first instar dendrites and find that this is likely not the case (new Fig. 7 K, L). We therefore remove the proposal that Shot directly regulates Rab11. Regarding the actin dependence of Shot tip localization, we based this conclusion on the observations that Shot ΔCH1 does not localize to tips, and that Mical overexpression significantly broadens the Shot tip signal (Fig. 5 E, F, J - previously Fig. 5 C, D, G). To further deepen this conclusion, we now add a localization analysis of a Shot mutant lacking both actin-binding CH domains (Shot ΔABD) and find that this also abrogates tip enrichment.

      (3) The authors show an image that Shot colocalizes with the EB1-mScarlet3 comet initiation sites and use this representative image to generate a model that Shot functions as part of an MTOC. However, this conclusion needs additional support: the authors should quantify the frequency of EB1 comets that originate from Shot-GFP aggregates, report the orientation of EB1 comets that originate from Shot-GFP aggregates (e.g., do the Shot-GFP aggregates correlate with anterogradely or retrogradely moving EB1 comets), and characterize the developmental timing of these events. The genetic interaction tests revealing ability of shot dsRNA to enhance the loss of microtubule-interacting proteins (Msps, Patronin, EB1) and Rab11 are consistent with the idea that Shot regulates microtubules, but it does not provide any spatial information on where Shot is interacting with these proteins, which is critical to the model that Shot is acting as part of a dendritic MTOC.

      Thank you for these comments. Due in part to low overall EB1 comet density in c4da neurons, we could rarely see comets originating from tip-localized Shot::GFP at the first instar stage - one example is shown in new Fig. S5C (previously Fig. 7D). At later stages, Shot is relatively evenly localized along the dendritic plasma membrane. As we cannot be certain about the relatively vague MTOC conclusions, we decided to rephrase our conclusions and state - based on genetic and functional data - that Shot's ability to stabilize microtubules is required, likely at dendritic tips.

      (4) It is unclear whether the authors are proposing that dendrite pruning defects are due to an early function of Shot in regulating microtubule polarity in young neurons (during 1st instar larval stages) or whether Shot is acting in another way to affect dendrite pruning. It would be helpful for the authors to present and discuss a specific model regarding Shot's regulation of dendrite pruning in the Discussion.

      Thank you for these comments. It is indeed our hypothesis that Shot's early role in setting up microtubule organization is also crucial for pruning, this is the most parsimonious explanation as pruning and orientation defects always go hand in hand. We added the following sentence in the Discussion: "As we do not have evidence for Shot regulation at the onset of the pupal stage (Fig. S4), the function of Shot crucial for pruning is most arguably its early role in setting up uniform microtubule orientation."

      (5) The authors argue that a change in microtubule polarity contributes to dendrite pruning defects. For example, in the Introduction, last paragraph, the authors state: "Loss of Shot causes pruning defects caused by mixed orientation of dendritic microtubules." The authors show a correlative relationship, not a causal one. In Figure 4, C and E, the authors show that overexpression of Mical disrupts microtubule polarity but not dendrite pruning, raising the question of whether disrupting microtubule polarity is sufficient to cause dendrite pruning defects. The lack of an association between a disruption in microtubule polarity and dendrite pruning in neurons overexpressing Mical is an important finding.

      Thank you for this comment. Microtubule orientation defects are tightly linked to dendrite pruning defects, as all known conditions causing orientation defects either cause pruning defects, or they act as genetic enhancers of pruning mutants. We now clarify this in the introduction citing the example of EB1 RNAi (which causes orientation defects (Matties et al., 2010), and does not cause pruning defects by itself, but enhances the effects of other pruning mutants (Herzmann et al., 2018)). This is also the case for Mical overexpression, which causes orientation defects and enhances the pruning defects caused by Shot RNAi. To make this point clearer, we now also show representative images of pruning defects caused by either Shot RNAi alone or Shot RNAi combined with Mical overexpression.

      Introduction - we explain better the genetic relationship between microtubule orientation and pruning. Figure 4 (phenotypic enhancement of shot knockdown by Mical overexpression): added representative images

      (6) The authors show that a truncated Shot construct with the microtubule-binding domain, but no actin-binding domain (Shot-C-term), can rescue dendrite pruning defects and Khc-lacZ localization, whereas the longer Shot construct that lacks just one actin-binding domain ("delta-CH1") cannot. Have the authors confirmed that both proteins are expressed at equivalent levels? Based on these results and their finding that over-expression of Shot-delta-CH1 disrupts dendrite pruning, it seems possible that Shot-delta-CH1 may function as a dominant-negative rather than a loss-of-function. Regardless, the authors should develop a model that takes into account their findings that Shot, without any actin-binding domains and only a microtubule-binding domain, shows robust rescue.

      Thank you for this constructive criticism. We agree that the rescue and overexpression experiments with UAS-Shot ΔCH1 (lacking only the first actin-binding domain) and UAS-Shot CTerm (containing only the MT binding domain) allow for other interpretations than ours that both features are necessary for Shot function in dendrites. To address this issue, we tested two more Shot domain deletion mutants, Shot ΔABD, lacking both actin-binding domains, and Shot ΔCTail, which lacks the MT binding domain (Figure 3). Neither of these two constructs can rescue the pruning and MT orientation phenotypes of shot mutants, confirming that both actin and microtubule binding are required for these functions in the context of full-length Shot. Importantly, in contrast to the previously tested Shot ΔCH1, neither Shot ΔABD nor Shot ΔCTail inhibit pruning upon overexpression, ruling out other interpretations of overexpression toxicity. The observation that Shot CTerm is sufficient to rescue indicates that actin binding is only required for Shot function in the context of full-length Shot. Interestingly, this is reminiscent of a proposed autoregulation mechanism where actin binding disinhibits Shot's MT binding domain (Applewhite 2013).

      Figure 3, Figure S2 (Shot domain analysis): We added two new constructs (Shot ΔABD, Shot ΔCTail) and a Shot BAC as additional control in Figure 3.

      (7) The authors state that: "The fact that Shot variants lacking the CH1 domain cannot rescue the pruning defects of shot[3] mutants suggested that dendrite tip localization of Shot was important for its function." (pages 10-11). This statement is not accurate: the Shot C-term construct, which lacks the CH1 domain (as well as other domains), is able to rescue dendrite pruning defects.

      Thank you for this comment. Shot ΔCH1 indeed does show a partial rescue of the pruning defects (but not of the polarity defects). To clarify whether actin binding is required for Shot function in dendrites, we tested the additional Shot construct Shot ΔABD and found that this construct could not rescue the defects at all (Fig. 3).

      Fig. 3, added Shot ΔABD.

      (8) The authors state that: "In further support of non-functionality, overexpression of Shot[deltaCH1] caused strong pruning defects (Fig. S3)." (page 8). Presumably, these results indicate that Shot-delta-CH1 is functioning as a dominant-negative since a loss-of-function protein would have no effect. The authors should revise how they interpret these results. This comment is related to another comment about the ability of Shot constructs to rescue the shot [3] mutant.

      Thank you for this comment. We agree that the rescue and overexpression experiments with UAS-Shot ΔCH1 (lacking only the first actin-binding domain) allow for other interpretations than ours - that actin binding is necessary for Shot function in dendrites. To address this issue, we tested another Shot domain deletion mutant, Shot ΔABD (lacking both actin-binding domains) (Figure 3). This construct did not rescue the pruning and MT orientation phenotypes of shot mutants, and also does not cause pruning defects upon overexpression, confirming that actin binding is required in the context of full-length Shot. As our domain analysis is reminiscent of a proposed autoregulation mechanism where actin binding disinhibits Shot's MT binding domain (Applewhite 2013), it is interesting to speculate that the overexpression toxicity of Shot ΔCH1 may stem from a residual ability of the second CH domain to relieve autoinhibition.

      Fig. 3, Fig. S2, added Shot ΔABD and Shot ΔCTail to domain analyses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Not enough info was given about how the statistics were performed in e.g., Figures 5E, 5H and 7C. Many data are highly significant, but this is not always directly clear to me from the images, e.g., 7C 6-10um. Or comparing 7I to 7L, how come the significance is so much different

      Thank you for this comment. For quantification of the localization data, we pooled the data points from individual intensity profiles for each dendrite segment and compared them statistically. Since this generates a relatively high number of data points, samples are quite likely to be significantly different. Acknowledging this, we now only compare the most distal segments (0-2 μm, 2-4 μm). We describe our approach now in more detail in the methods section.

      Better description of fluorescence quantifications in Figs. 5, 7 in Methods section.

      (2) Abstract "while plus-en out MT just grow from the soma" - I think this is probably only partially correct. There are various players important for organizing/nucleating these microtubules (e.g., Augmin for local MT nucleation).

      Thank you for this comment. This was oversimplified and we rephrased the abstract.

      (3) "We and others have previously shown that c4da neuron dendrite pruning depends on local microtubule disassembly in proximal dendrites (Herzmann et al, 2017)." Nice to add other references.

      Thank you for this comment. We added additional references.

      (4) Figure 2 EB imaging. Where was this performed along the dendrite? Close to the CB, this may reflect CB MT growth entering the dendrite.

      Thank you for this comment. We usually perform EB1 imaging in primary or secondary dendrites -- now mentioned in the text (p. 7). The idea that anterograde comets in this region could reflect microtubules entering from the soma is interesting, especially as the fast anterograde comets upon combined shot knockdown and Mical overexpression behave so differently. We take up this idea in the Discussion - thanks again!

      (5) Figure 3H-L I cannot see the magenta in the overlays. Separate channels seem to be needed. And where is the axon? Moreover, in the text it is stated that the I 3H fusion protein is exclusively in the axon. And control is not in 3M.

      Thank you for this comment. Our lacZ antibodies are relatively weak, so the axons are not always visible. We therefore changed the text to only mention the soma now. We also now include all shown genotypes in the quantification.

      (6) Figure 4A I'm not sure what is seen in the image. Is this a single dendrite? And if so, why is Shot only seen on one side?

      Thank you for this comment. In this panel (now moved to Fig. 5A) we had to take a fairly thin projection of a confocal stack because of strong surrounding tissue signal. Because of this, the Shot::GFP signal on one edge of the dendrite is not as well visible. We try to improve this by indicating the edges of the dendrite in the panel.

      Fig. 5A - D, indicated dendrite outline in confocal/SIM images.

      (7) Figure S2C description is not the same as in the graph (neurons / dendrites).

      Thank you for pointing this out. We corrected the legend (now Fig. 2C).

      (8) Figure S4 - is this a single dendrite or a bundle?

      The images (now Fig. S5) correspond to the microtubule signal from single primary dendrites, i. e., mostly microtubule bundles. We clarify now in the legend.

      Fig. S5, specified nature of samples in legend.

      Reviewer #2 (Recommendations for the authors):

      (1) In the Discussion, page 13, the authors state that: "On the one hand, we provide evidence that Shot anchors microtubules via actin in mature neurons." This is not supported by the authors' evidence; indeed, the authors themselves state on page 10 that "... the CH1 domain is unlikely to be required for Shot cortical localization in dendritic shafts." If the authors mean to say that the cortical localization of Shot is restricted to dendrite tips, then superresolution microscopy should be done on dendrite tips.

      Thank you for pointing this out. As the CH1 domain is clearly not required for Shot third instar localization, we agree with this notion and rephrase our interpretation accordingly: "Our data also do not rule out that Shot could anchor dendritic microtubules in mature third instar neurons, likely in cooperation with actin." (p. 17)

      (2) The authors characterize the localization of Shot during larval stages but not during pupal stages. The authors should characterize where Shot localizes during dendrite pruning, since it is possible that Shot may localize differently at this stage. For example, the authors mention that Shot localizes to dendrite tips during early larval stages but not late larval stages; it seems that the localization of Shot is dynamic.

      Thank you for this comment. In response, we analyzed Shot localization in a time-course experiment including the second instar larval stage and at 5 h APF in the early pupal stage (new Figure S4). This showed that Shot localizes almost exclusively to dendrite tips at the first instar, but is more broadly distributed in the soma and along dendrites shafts at all later developmental stages.

      New Fig. S4, time course of Shot::GFP localization including first to third instars and early pupal stage.

      (3) Figure 3: Khc-lacZ is not an ideal read-out of microtubule polarity per se. Better, read-out microtubule polarity using EB1-GFP.

      Thank you for this comment. Unfortunately, this experiment is not feasible in this situation. In our MARCM system, we use a red fluorescent tdTomato to label c4da neurons, and all Shot transgenes carry GFP tags, prohibiting the use of both fluorescent EB1 transgenes that we have, EB1::GFP and EB1::mScarlet3.

      (4) Figure 4 A-A' (also Figure S4): Superresolution microscopy data are presented but not analyzed. Analysis should be included before drawing a conclusion from these data.

      Thank you for this comment. In response, we quantified the width of the dendritic microtubule bundles in the STED experiments (now Fig. S5) and found that these bundles are significantly thinner upon shot knockdown.

      New Fig. S5 B, quantification of STED data.

      (5) Figure 5, A and B: These images are extremely difficult to interpret on their own. It is not clear what the authors are interpreting as a positive signal. Analysis and quantification of Shot-FL and Shot-deltaCH1 should be included, not just representative images. This seems like a missed opportunity.

      Thank you for this comment. In response, we show higher magnification images of Shot(endo)GFP, UAS-Shot::GFP, and the two actin-binding mutants (UAS-Shot ΔCH1::GFP and UAS-Shot ΔABD::GFP. For better interpretation, we indicate the boundaries of the dendrites as determined by tdTomato expression in the neurons (new Fig. 5 A - D).

      (6) Figure 6: If gamma-tubulin is knocked down, does this decrease the number of EB1-GFP comets that originate from Rab11(+)-endosomes?

      Thank you for this comment. This experiment is technically challenging because of the number and location of UAS transgenes involved. In addition, gamma-tubulin knockdown only has minor effects on pruning (no defects) or microtubule behavior (no orientation defects) in c4da neurons (Wang et al., 2019, eLife (Fengwei Yu lab).

      (7) Figure 6: Rab11(+) endosomes correlate with the initiation of EB1-GFP comets in first instar larvae, but the orientation of microtubules is analyzed in third instar larvae -is there an effect on microtubule polarity in younger larvae? What about second instar larvae? On a related note, what is the orientation of microtubules that originate from Rab11(+) endosomes? These should be quantified data.

      Thank you for this comment. We analyzed the effect of Rab11 knockdown at the first instar stage and found that this led to a non-significant increase in anterograde comets. These data are now included as new Fig. S6. We also checked the orientation and found that all comets from Rab11-cherry puncta were retrograde, this is now mentioned in the text.

      New Fig. S6, effect of Rab11 dsRNA on comet orientation in first instar neurons.

      (8) Figure 6: Panel G shows that Rab11-GFP and Msps-FLAG co-immunoprecipitate. What effect does Rab11[S25N] have on EB1-GFP comet frequency and site of origin?

      Thank you for this comment. Unfortunately, Rab11[S25N] does not cause c4da neuron dendrite pruning phenotypes in our hands, even though it has worked in another lab (Lin et al., 2020, PLoS Genetics), possibly due to subtle differences in the GAL4 transgenes used. We therefore did not try this transgene in our assays.

      (9) Figure 6: Does Patronin colocalize with Rab11 in neurons?

      Thank you for this comment. We had originally included an image of Patronin::GFP expressed in c4da neurons that appeared to show tip enrichment. When we tried to reproduce and quantify this result, we found that Patronin localization varied strongly with expression levels, and more focused or punctate signals seemed to reflect either aggregation or decoration of microtubules. We therefore removed the image and instead included a more detailed genetic analysis of Patronin, which shows similar phenotypes as Shot (MT orientation defects at the first instar, enhancement of Shot knockdown phenotypes (Fig. 6).

      Removed Patronin::GFP image (old Fig. 7H), added effect of Patronin dsRNA on comet orientation at first instar, enhancement of Shot dsRNA phenotype (new Fig. 6 D, E).

      (10) Figure 7D: These data showing EB1-mScarlet3 comets originating from Shot-GFP should be quantified. Also, which Shot domains are involved in this recruitment?

      Thank you for this comment. We observed only very few comets directly from tips in these experiments which precluded quantification. Acknowledging this, we moved the image to Supplementary Fig. 5.

      As EB1 likely binds to Shot via SxIP motifs in the C-terminal region, we also tested whether the ShotΔCTail construct lacking the GAS2 domain and adjacent regions can recruit EB1. This is not the case. We therefore conclude that Shot recruits EB1 via its C-terminal part (new Figure 6D, E).

      (11) Figure 2, A and B; Figure 4 B; Figure 6, B and D: Please include a scale bar for the time axis on these (and any other) kymographs.

      Thank you for this comment. We checked and added time scale bars if missing.

      (12) The authors should analyze the localization of GFP-tagged endogenous Shot in parallel to the exogenous expression of Shot under Gal4-UAS control.

      Thank you for this comment. We tried to image endogenously tagged Shot::GFP at the first instar timepoint, but this was unfortunately not feasible due to very high Shot::GFP expression in neighboring tissues at this stage (and relatively low expression in c4da neurons).

      (13) The authors should clarify at some point in the manuscript that EB1 comets can reflect either de novo microtubule growth (microtubule nucleation) or growth from the pre-existing ends of microtubules. EB1 comets alone do not indicate microtubule nucleation.

      Thank you for this comment. We added the following sentence to the description of the first EB1 experiment (p. 7): "EB1 binds to the plus ends of growing microtubules (both newly nucleated and re-elongating), and plus end-bound EB1 is visible as moving dots (also known as comets)."

      (14) It is unclear what the n in the graphs in the figures represent: neurons? Or dendrites/comets/etc? Please specify in the figure legend.

      Thank you for this comment. We now specify this for each graph.

      (15) Page 6, the authors write that "We and others have previously shown that c4da neuron dendrite pruning depends on local microtubule disassembly in proximal dendrites," but only cite their paper; please include the citations for the work by others.

      Thank you for this comment. We added additional references.

      (16) Page 9, the authors state: "Such phenomena are often seen when microtubules are not attached to specific anchoring sites such that they can be moved by microtubule motors." This is a fairly speculative interpretation of the data and should be saved for the Discussion.

      Thank you for this comment. We moved this speculation to the Discussion.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors tackle a long-standing question in developmental theory: given a gene-regulatory network that includes extracellular signaling, which topologies are even capable of transforming an initial spatial profile into a genuinely new pattern? Building on the classical reaction-diffusion framework in one dimension, but imposing biologically motivated constraints, they prove that every one-signal sub-network must be either Hierarchical (H), self-activating (L+), or selfinhibiting (L-). They further demonstrate that only three composite classes of full networks - pure H, a coupled L+ L- "Turing" pair, and an L- module fed by an intracellular positive loop ("noise-amplifying")-can create non-trivial spatial transformations. Analytical criteria and illustrative simulations are provided, together providing a closed taxonomy, which is supposed to be relevant for real systems.

      Strengths:

      (1) Useful classification framework. Reducing a vast number of possible gene circuits to three canonical patternforming motifs is a valuable organizing insight for both theorists and ---experimentalists.

      (2) Practical interpretability. Given a reaction network diagram, one can now decide (assuming the model applies to real systems) whether spatial patterning is even possible, saving experimental effort on in silico screens that could never succeed.

      Weaknesses:

      (1) After the resubmission, I still have concerns regarding the formal definition of "non-trivial transformations" (P1/P2) and its application to noisy or multi-dimensional systems. The criteria rely on counting "new" critical points (maxima/minima). In their response, the authors argue that the diffusion operator instantly smooths discontinuous white noise, allowing critical points to be properly defined. However, this very smoothing process passively generates a landscape of new, smooth local extrema from the initial noise. Consequently, trivial diffusive regularization could inadvertently fulfil the criteria for a "non-trivial" transformation, leaving the definition conceptually problematic.

      That is indeed the case: diffusion alone can generate new concentration maxima and minima; but these would be transient unless there are some self-activatory loops (as we detail over the article). If these concentration maxima and minima are transient in time, they do not count for pattern transformation. In the two version of the article we explicitly stated (in P1 in the introduction when defining pattern transformation) that we only consider resulting patterns that are stable in time. This implies that the concentration maxima and minima that may transiently arise from diffusion alone do not count for the definition of pattern transformation. In the current new version (the third) and after the reviewer’s suggestion, we insist (in red text) that the new concentration maxima and minima need to be stable in time (i.e. non-transient).

      Furthermore, when extending the framework to 2D/3D, the manuscript assumes that starting from a central "spike" will robustly preserve radial symmetry, yielding concentric rings or shells. This overlooks the fundamental nature of macroscopic mean-field models like reaction-diffusion equations. The realization of the final multidimensional pattern depends strictly on the stability of the solution against ubiquitous perturbations (including angular modes) rather than solely on the deterministic symmetry of the initial condition. It remains unclear how the current framework accounts for spontaneous symmetry breaking in cases where these angular modes become unstable, challenging the assumption that radial symmetry will strictly dictate the outcome. We note that the authors' use of noise as an initial condition does not resolve this fundamental issue. Reaction-diffusion equations inherently describe mean-field dynamics, meaning that microscopic fluctuations are continuously present in any real system, regardless of whether explicit stochastic terms are written into the equations. Ultimately, if a symmetric mean-field solution is structurally unstable to these inherent fluctuations, it simply cannot be realized in nature.

      If we understood right the reviewer is saying that there are always fluctuations everywhere and that, thus, the spike initial pattern should also have noise everywhere. In the discussion (and partially in the introduction) we have now added a discussion (in red) on how the resulting patterns possible from such spike-with-noise initial pattern can actually be understood, to a large extent, from those of the homogeneous-with-noise and spike-without noise. In that case, as the reviewer suggests, the resulting patterns do not necessarily have radial symmetry. We have kept the results on spike initial patterns without noise because they are very helpful to understand the pattern transformations possible from spike-with-noise initial patterns. Moreover, as we discuss now, the spatial fluctuations that occur everywhere can be really small compared to the concentration in the spike and, thus, the spike initial pattern without noise can be a good approximation in some cases, at least worth considering.

      (2) Theoretical limitations in the application of Linear Stability Analysis (LSA): I remain uncertain about the framework's reliance on LSA to categorize macroscopic transformations, especially those arising from large initial perturbations (spikes). In their rebuttal letter, the authors justify this by assuming the perturbation remains small over a short time interval. However, because the study aims to describe stationary, asymptotic states, applying a linear approximation that relies on transient t->0 conditions to predict long-term global stability is not fully resolved.

      We are not trying to predict long-term stability and we never intended to. We never claimed that to be the case. We are interested in stable in time spatial patterns but the long-term stability of resulting patterns is not something we intend to see from the LSA. The LSA just provides a necessary (but not sufficient) condition for non-trivial pattern transformations: any network unable to sustain the growth of small perturbations (i.e., any linearly stable network topology) cannot lead to a non-trivial pattern transformation, regardless of the nonlinear terms in the reaction term f(g). From the previous suggestions by the reviewer it could be the case that he/she thinks that since there is always noise everywhere and all the time we can never apply the LSA or that it cannot be applied along time. In fact, we do not try to applied over time (that would make no sense for our purposes). The LSA is only applicable and only informative when applied to the initial pattern (that is at time 0) to see gene network that cannot transform initial patterns into other patterns. As we discuss in the previous version and we further stress in the current one (in red in the LSA), large spikes do not invalidate this approach. Even if the spike would be large, it would only affect cells outside the spike through the diffusion of gene products from the spike. Thus, if a small enough time is considered, large spikes can be considered as small concentration perturbations outside the spike and, thus, in the the worse case scenario our whole approach may not be applicable inside the spike but outside of it (that is reasonable since spikes are by definition narrow).

      Here it is important to stress that many things that used the LSA in the original version of the article do not used it in the current version of the article. In fact, in this article we address two main questions: (i) which gene network topologies can produce non-trivial pattern transformations; and (ii) what can we say about the stationary patterns they produce. We acknowledge that a linear stability analysis alone is not enough to fully characterize the long-term behavior of the system, and this is why we only use LSA as an aid to answer question (i).

      This was not clear enough in the first version of the article but it was explicitly stated in the last version (in the Gene network classification and Linear stability analysis section). This allows us to discard many topologies, but further analysis is still required to assess whether linearly unstable network topologies can actually produce non-trivial pattern transformations.

      It is in this further study that question (ii) comes to play. Here we do not rely on LSA, but instead impose a series of requirements (R1-R5) on the reaction term f(g) that constrain the nonlinear dynamics in a biologically motivated way that prevents pathological behaviors (particularly, boundedness of solutions (R4) and monotonicity of the reaction (R5)). These requirements enable the qualitative analysis of the different unstable network topologies in order to say some things about the possible stationary patterns. This is complemented with numerical simulations and quantitative analysis of some prototypical examples in the supplementary information.

      (3) In the previous round of the review, I suggested that a biomolecular sink, such as A+B -> AB reaction, could break the approach. In their response letter, the authors defend their approach by arguing that such reactions can be accommodated by their abstract constraints (R1-R5) as long as the signs of the Jacobian elements remain invariant. However, the problem I see here is not the sign of the interactions, but the severe loss of spatial homogeneity.

      When a macroscopic initial perturbation (a "spike" of morphogen) is introduced into a domain with a strong bimolecular sink, it will inevitably cause massive local depletion of the consumed substrate near the source. Consequently, the background state of the system will rapidly evolve into a profile with macroscopic spatial gradients long before any spontaneous pattern-forming instability takes over. Mathematically, this dictates that the system no longer possesses a homogeneous steady state, and the Jacobian matrix becomes explicitly space-dependent, which should break the classical LSA approach.

      This criticism seems to be intimately related to the previous one. If we understand correctly, the argument is that in a very non-linear system, such as in the sink described, the spike will rapidly lead to a local change in the concentration of other gene products and that then the LSA is not applicable. This is true but what we care about is whether the initial pattern (that is the system at time zero) is actually stable or not. We care about it because as we explain, pattern transformation is only possible if the perturbation in the initial pattern (spike or noise) is unstable. This is simply a necessary condition (an initial pattern may be unstable to perturbation and still not produce nontrivially transformations). So whether the system will be suitable for a LSA some time after the initial pattern, as the reviewer suggests, is not something we need to know for our classification. Related to the other comments the reviewer may be concerned with whether other perturbations occurring everywhere (and all the time), that is noise, may actually affect the possible resulting patterns (that we described in the new section of the discussion).

      We want to thank the reviewer for this clarification, as we believe we had not fully understood their concern in the previous round of review. In our framework, the bimolecular sink the reviewer suggests corresponds to a three-gene-product network where A and B mutually inhibit each other and both activate AB.

      First it is important to consider that unless something else is specified an A+B→AB system with homogeneous initial pattern (with or without noise) is just not stable: the A and B gene products will decay to zero concentration and AB to a maximal concentration (that would be homogeneous over space if there is no noise). Besides the final stable state is totally stable since with no A or B left the concentration of AB cannot change. We explicitly state in the article (in the LSA section) that we apply the LSA to systems that, when unperturbed, are stable. For other systems we just wait for the system to stabilize and then ask whether pattern transformation is possible from that state if there is some perturbation, where we can now apply a LSA). So strictly speaking the system the reviewer is suggesting is outside the scope of the article and indeed unsuitable for LSA. Nevertheless, the system the reviewer proposes cannot lead to non-trivial pattern formation, as we detail below.

      The reviewer does not specify whether A, B or AB diffuse, we then consider all possibilities. We also assume that A is the gene product in the spike.

      Case in which no molecule diffuses. In this case a spike of A will simply lead to a valley of B (since A reacts with B to deplete B), and a spike of AB in the exact same location of the spike (i.e., no non-trivial pattern transformation occurs). If there is noise, each small fluctuation in the concentration of A or B would lead to a similar fluctuation in AB. Notice this case does not lead to non-trivial pattern transformations (the peaks in the initial pattern and resulting pattern are in the same places). This latter situation in fact we explain in the “Pattern formations from homogeneouswith-noise initial patterns in H networks section.”

      Case in which AB diffuses. Since nothing is promoting the production of A and B, their concentration will inevitably decay to zero and since AB diffuses its concentration on the long-term will inevitably become homogeneous (irrespectively of which initial pattern there may be).

      Case in which only A diffuses. In this case the spike of A will initially lead to a valley of B (since A consumes B) and a peak of AB (since this consumption leads to AB). However, since for each molecule of AB a molecule of both A and B are required and the concentration of B is homogeneous (since B does not diffuse), having more of A around the spike would not lead to more AB in the peak than elsewhere. The concentration of AB would thus become homogeneous over time. The same applies if B is the only molecule that diffuses.

      Case in which A and B diffuse and AB does not. In this case the spike of A will lead to a peak of AB. This would deplete B around the spike but since B can diffuse, new molecules of B would arrive and lead to a further growth in the peak of AB. As a result a stable peak of AB will form (even if A and B will ultimately decay to zero), just around the initial spike of A and both A and B will decay to zero (so no new peaks or valleys form and thus, no non-trivial pattern transformation).

    1. Author response:

      We thank the reviewers for their thoughtful and constructive feedback on our work.

      We have attempted to synthesize the main suggestions across both reviewers into four main areas, which we plan to address in a revised version of the manuscript.

      (1) The relationship between covert attention, overt attention, and choice

      While the current analyses address elements of relationship between covert attention, overt attention, and choice, the reviewers highlighted an opportunity to make full use of our rich gaze data to examine more explicitly how these processes are related to one another over the course of a decision. [Reviewer #1 Point 2, Point 3 i; Reviewer #2 Point 2 para 2]

      (2) The role of decision difficulty and decision time

      The reviewers point out that decision difficulty may i) account for the relationship between covert attention (probe response correctness) and choice; and ii) influence decision times, which are currently not analyzed; and that these considerations may affect how we interpret the relationship between covert attention and choice. [Reviewer #1 Points 2, 3 ii and iii; Reviewer #2 Point 1 paras 1 and 3]

      (3) Evidence for competition between covert and overt attention

      The reviewers requested more direct evidence for our claim that covert attention to higher-valued peripheral options occurs "at the expense of" overt attention [Reviewer #2 Point 2 paras 1 and 3, Point 3 para 3]

      (4) Figure 6: Presaccadic vs. non-presaccadic attention

      Reviewer 2 found the rationale behind the analysis distinguishing presaccadic from non-presaccadic covert attention difficult to follow and suggested and alternative to make this comparison more explicit [Reviewer #2 Point 3].

      In response to these suggestions, we will include additional analyses in the revised manuscript which more directly examine the relationship between covert attention, overt attention, and choice. We will also introduce new analyses which more simply and directly depict: why decision difficulty cannot wholly account for the link between covert attention and choice; a competition/trade-off between overt and covert attention; and that value modulation existed at non-presaccadic options. In addition to these four main areas, we will clarify several elements of the task design and analysis which may not have been sufficiently clear in the manuscript.

      Here, we offer a provisional response to each of the points above.

      (2) The relationship between covert attention, overt attention, and choice

      We agree that the relationship between covert attention, overt attention, and choice could be examined in more detail using the available eye-tracking data. In the revised manuscript, we plan to include further analysis looking at fixation and sampling measures before the decision is made.

      For example, to investigate the relationship between covert attention and choice (as reviewer 1 suggested), we plan to include an analysis which asks about how decision accuracy changes as a function of decision difficulty (relative option value) when the unfixated probe is correctly reported, versus when it is not. This will straightforwardly test for the influence of probe report on choice, controlling for the influence of decision difficulty.

      More generally, we will revise the manuscript to distinguish more carefully between associations demonstrated by the present data and stronger causal interpretations. To address this, we will either attempt more direct statistical tests of causality, or use more neutral language where the causal directionality between fixation, covert attention, and choice cannot be uniquely established by the present design.

      An important clarification on the dynamics of the task is relevant to several of the reviewers’ comments concerning the last-fixation analyses. Probe performance was generally measured earlier than the final fixation and choice – as it was always at the first to third option visit by the participant (or only first to second option visit in Experiment 1). Only 10.4% of probes occurred during the final option visit, and on average, participants made 2.8 further option visits after probe presentation. We can see that our original presentation of these results may have been difficult to follow. We will make this clarification in the revised manuscript, particularly where we refer to “last-fixated” or “last-unfixated” options. These terms refer retrospectively to which option was ultimately visited or not visited last before choice, rather than to the fixation occurring at the time the probe was presented.

      (2) The role of decision difficulty and decision time

      We agree with the reviewers that decision difficulty and in turn decision time should be considered more carefully.

      We will make explicit that relative option value was already controlled for in the existing analyses relating to covert attention and choice. Specifically, for the last-fixation analysis, relative value difference was included as a control variable in the regression. For the time-advantage analysis, we use a corrected dependent variable which is intended to remove possible influences of relative value (i.e., decision difficulty) on the fixation durations, following what has been done in previous work (e.g., Eum et al., 2023; Krajbich et al., 2010; Krajbich & Rangel, 2011).

      We additionally plan to include new analyses which more directly examine the role of decision difficulty and decision time. In particular, we will test whether the relationship between peripheral probe performance and choice remains after accounting for decision difficulty, and whether decision time differs as a function of probe performance.

      We would also like to take the opportunity to clarify that “time advantage” was operationalised as a relative measure: the difference in cumulative dwell time between the two relevant options, rather than the absolute cumulative dwell time on a single option. This definition will be stated more clearly in the main text and figure caption.

      Finally, to address the possibility that peripheral probe report errors reflect a general dual-task cost or increased response noise, we will examine decision accuracy as a function of decision difficulty separately for trials in which the peripheral probe was reported correctly versus incorrectly. This will help determine whether successful peripheral probe report is associated with poorer decision performance.

      However, we should note that a closer examination of Figure 4 offers clues to this question; it suggests that correctly reporting the probe at the last-unfixated item makes people more accurate, rather than happening at the expense of decision accuracy. To explain this:

      When objectively Option 1 is worse than Option 2 (i.e., to the left of x-axis center), choice for a last-fixated Option 1 (solid line) is unfairly boosted; but if the last-unfixated probe was correctly reported, then the choice proportion is correctly closer to 0%.

      When objectively Option 1 is better than Option 2 (i.e., to the right of x-axis center), choice for a last-unfixated Option 1 (dotted line) is unfairly discounted; but if the last-unfixated probe was correctly reported, then the choice proportion is correctly closer to 100%.

      That is, the bias-attenuation effect of correctly reporting the probe at the last-unfixated item seems to operate on 1) reducing choice boosting of the objectively-worse fixated option (left of x-axis center, solid line is further down on right panels than left panels); and 2) increasing choice boosting of the objectively-better unfixated option (right of x-axis center, dotted line is further up on right panels than left panels).

      (3) Evidence for competition between covert and overt attention

      To address more directly the question of whether covert attention occurs at the expense of overt attention, we will add an analysis testing whether successful report of a probe at an unfixated option negatively predicts report of the simultaneous probe at the fixated option. A negative relationship would provide a more direct test of the predicted trade-off between attentional allocation to the fixated and peripheral locations. We will revise the strength of the corresponding language in the manuscript according to what this analysis supports.

      (4) Figure 6: Presaccadic vs. non-presaccadic attention

      We agree that the rationale underlying the current Figure 6 analysis is too implicit in the current manuscript.

      To clarify the underlying rationale: The key logic of the analysis was to ask whether value modulation of probe performance is present at an unfixated option even when that option is not the target of the upcoming saccade. In the original analysis, we only compared attention at peripheral options that had previously been viewed: we compared attention at peripheral options that had previously been viewed and were subsequently fixated next, with those that had previously been viewed but were not subsequently fixated next. The requirement that the option had previously been viewed arose because, in the gaze-contingent “hidden” condition, the value information of an unvisited peripheral option would not yet have been available to the participant, and therefore there would be no possibility of any value modulation effects existing at that location.

      Although this analysis was intended to distinguish value modulation at likely presaccadic targets from value modulation at other peripheral locations, we now see that its rationale is not immediately transparent. In the revised manuscript, we therefore plan to replace or supplement this analysis with a simpler visualisation and analysis that addresses the same question more directly.

      We will also revise text associated with the time-to-saccade analysis. Our use of the term “monotonic” was imprecise. The intended point was not that a purely presaccadic account predicts a strictly monotonic change across the full time-to-saccade interval, but rather that the observed U-shaped relationship is difficult to reconcile with a simple presaccadic-only account, because probe performance remains elevated even when the upcoming saccade is relatively distant. We will clarify this point and more carefully distinguish evidence inconsistent with a simple presaccadic-only explanation from the stronger claim (which we did not intend to make) that presaccadic contributions have been completely excluded.

      (5) Additional clarifications

      We agree with Reviewer 1 that the task can reasonably be described as perceptual rather than strictly value-based, but we believe that our central concerns regarding how covert and overt attention interact with multi-alternative decision-making still stand.

      We will clarify why the secondary probe task was not incentivized. The probe task was intended to provide an incidental measure of covert attention while participants performed the primary decision task. Incentivizing probe performance could have encouraged participants to explicitly trade off performance on the two tasks or to alter their allocation of attention in response to the secondary task.

      In summary, the revisions will combine clarification of several existing analyses with the addition of new analyses that more directly test the reviewers’ central concerns. In particular, we will more explicitly address whether the relationship between covert attention and choice remains after accounting for option value/decision difficulty; whether peripheral and fixated probe performance show evidence of an attentional trade-off between overt and covert attention; how decision time and richer measures of overt sampling relate to covert attention; and whether value modulation is present at peripheral locations that are not the target of the upcoming saccade. We will also moderate causal and mechanistic language where the current data do not uniquely establish directionality.

    1. Author response:

      We thank the reviewers and the editor for providing a thorough and constructive review of the manuscript. Both reviewers indicated that the ordinal entropy production rate should be validated with a synthetic (surrogate) model of the calcium signal. We agree with this, and it will feature prominently in the revised version of the manuscript.

      As described (see Author response image 1), we have already implemented a synthetic GCaMP model. In the simulations, the strength of the calcium intensity varies monotonically with the intensity of tFUS. Applying our D<sub>sym</sub> measure to the synthetic data produces a monotonic dose-response for irreversibility. It does not reproduce the non-monotonic, moderate-dose peak that we found in our study. We report this preliminary finding here because it speaks directly to the reviewers’ central critique.

      Below we provide details of the planned revision.

      (1) A surrogate model of GCaMP’s response to tFUS does not reproduce the observed dose structure

      We built a generative model of the measured calcium signal according to the underlying physics (Author response image 1A):

      (1) A latent Poisson spike train whose rate during and after sonication is monotone in tFUS dose by construction, such that any non-monotonicity appearing downstream must follow from the measurement of irreversibility;

      (2) Convolution with an indicator kernel having fast rise and slow decay, matching the kinetic asymmetry mentioned by the reviewers;

      (3) A static saturating nonlinearity representing GCaMP6s saturation;

      (4) Additive measurement noise applied after the nonlinearity, with standard deviation calibrated to the variance of the real data’s baseline periods;

      (5) baseline z-scoring, windowing, and estimation of D<sub>sym</sub> using the same parameters as the manuscript (m = 3, τ = 8, f<sub>s</sub> = 32 Hz; 5 s baseline, 5 s stimulation, 10 s recovery).

      We ran four separate versions of the simulation (with and without saturation, noise before and after the nonlinearity). Simulated GCaMP traces for the variant with saturation and post-nonlinearity noise are drawn in Author response image 1B.

      Result. All four simulation variants produce ∆D<sub>sym</sub> that increases monotonically with dose and is largest at the highest dose (Author response image 1C). During the recovery window, the difference between moderate (2.3 to 3.7 W/cm<sup>2</sup>) and maximum (7.4 W/cm<sup>2</sup>) dose is −0.297±0.066 for the full version with the nonlinearity, −0.222±0.070 without the saturating nonlinearity, and −0.006±0.003 for the two versions in which noise precedes the nonlinearity (mean ± SD across seeds). In contrast, the D<sub>sym</sub> of the real data (reproduced for convenience in Author response image 1D) peaks at 3.7 W/cm<sup>2</sup>.

      Author response image 1.

      A synthetic model of the calcium measurement chain does not reproduce the observed dose structure. (A) Generative model. A Poisson spike train whose rate is proportional to acoustic dose is convolved with a fast-rise, slow-decay indicator kernel, passed through a static saturating nonlinearity, corrupted by additive noise, and then analysed with the same windowing and ordinal estimator used in the manuscript. (B) Simulated traces at low, moderate, and high dose for the variant with saturation and post-nonlinearity noise. (C) Simulated change in irreversibility during recovery, for all four variants of the simulation. Every variant increases monotonically with dose. Markers denote the mean across 20 seeds (error bars show standard deviation). (D) Real data as shown in the manuscript: observed change in irreversibility during recovery for on-target sonication. Markers denote individual animals; boxes show the mean ± 1 SEM. Irreversibility peaks at moderate dose.

      (1.1) Invariance of ordinal statistics

      Ordinal pattern methods are invariant under any strictly monotonic pointwise transformation of the signal, because such a transformation preserves rank order. Two consequences follow, and we will state them in the Methods of the revision:

      - The static component of the indicator nonlinearity, including saturation and the relationship between calcium concentration and fluorescence, cannot alter D<sub>sym</sub>.

      - The scaling of amplitude with dose does not affect D<sub>sym</sub>.

      Any potential confound of GCaMP kinetics on D<sub>sym</sub> enters through the dynamics, which is what the simulations above address.

      (1.2) Planned analyses

      In the revision, we will add:

      - A sweep over rise and decay constants, Hill coefficient, noise level, and spike statistics, reporting explicitly the region of parameter space, if any, in which a moderate-dose peak can be produced;

      - Burstiness and adaptation in the latent source;

      - Calibrating the amplitude of the simulated responses to the real data;

      - Finite-sample bias and variability of D<sub>sym</sub> as a function of window length, pattern count, and signal-to-noise ratio, addressing Reviewer 1’s questions about window duration and noise.

      (2) Aspects of the data that argue against a confound of kinetics

      Independent of simulation, certain results in the manuscript are not explainable from an effect of calcium indicator kinetics.

      (2.1) Non-monotonic irreversibility is reported for off-target sonication – without an accompanying amplitude response. Off-target sonication produces no appreciable fluorescence response, meaning that there is no evoked transient whose kinetics could generate ordinal asymmetry, yet irreversibility still varies non-monotonically with dose. We will elaborate on the off-target EPR effects in the revision.

      (2.2) The dose structure remains after adjustment for amplitude. Our nested model comparison shows that acoustic dose explains variance in D<sub>sym</sub> beyond baseline irreversibility and the windowmatched change in fluorescence. This is inconsistent with the view of D<sub>sym</sub> as a deterministic readout of the amplitude trajectory. In the revision, we will also include decay time constant as a covariate, as Reviewer 1 suggests.

      (2.3) Prediction from the prestimulation window. Baseline D<sub>sym</sub> is computed on a window that contains no stimulus and no evoked transient, and it predicts the subsequent calcium response after adjustment for baseline fluorescence and dose. The kinetics of a response that has not yet occurred cannot modulate the value of a quantity measured before it.

      (3) Reviewer 2, point 3: variation in baseline EPR across nominal intensity

      Reviewer 2 correctly observes that baseline irreversibility varies across nominal intensity by an amount comparable to the stimulation-induced change. This implies two candidate explanations, which we address separately.

      Systematic differences between intensity groups. If trials at different nominal intensities were not evenly distributed through a recording session, then drifts in bleaching, arousal, or habituation would produce spurious baseline differences between intensity groups. We tested this directly in the 408 analysed CMT 2.5 Hz trials. Session position, defined as the index of a trial within its recording, does not differ across the six intensity levels: Kruskal-Wallis H(5) = 3.51, p = 0.62 on-target and H(5) = 1.70, p = 0.89 off-target, with mean session position varying only between 18.3 and 23.3 across intensities. Within individual recordings, the rank correlation between session position and intensity has a median of +0.03 and does not exceed 0.34 in magnitude in any of the 10 recordings. Acoustic intensity is therefore not confounded with position in the session, and we will report this in the revision.

      Uncertainty larger than the plotted error bars. This is a fair concern and we will address it by reporting trial-resampling and pattern-count bootstrap confidence intervals for D<sub>sym</sub>. We also note that our primary inference already includes baseline D<sub>sym</sub> as a covariate in an ANCOVA specification.

      (4) Reviewer 1, remaining points

      Other conventional GCaMP metrics. The revision will add decay time constant and additional waveform descriptors to the amplitude-adjustment models described above in 2.2.

      Stimulation periods of differing duration. We will report the dependence of D<sub>sym</sub> on window length from the finite-sample bias analysis using synthetic data (see 1.2 above).

      Lower signal-to-noise recordings. We will measure the dependence of D<sub>sym</sub> on SNR with synthetic data.

      The auditory confound. We will add a Discussion paragraph citing Sato et al. (2018), Guo et al. (2023) and Kop et al. (2024), stating that the responses analysed here may reflect an indirect auditory mechanism.

    1. Author response:

      We particularly appreciate the reviewers’ detailed consideration of both the strengths and limitations of the study. The comments have helped us identify areas where the mechanistic interpretation can be further strengthened and where the conclusions should be more precisely defined.

      We have carefully considered all of the points raised and provide below our provisional responses and planned revisions. We intend to address these comments comprehensively in the revised manuscript.

      Reviewer #1 (Public review):

      An expanded description of the histological types of ovarian cancer used to generate survival curves and demonstration of the expression of total IRAK4 would strengthen the manuscript.

      If the information is available, it could be useful to indicate the percentages of different ovarian cancer histotypes analyzed in Figure 1.

      Providing panels demonstrating the expression of total IRAK4 in Figure 4 would be informative.

      We will expand the analysis/description of the ovarian cancer histo-types represented in the survival analyses and, where available, provide the percentages of the different histo-types analyzed in Figure 1. We will also provide additional panels demonstrating total IRAK4 expression in Figure 4. Indeed, these additions will provide greater clarity regarding the histological and molecular characteristics of the models used.

      Reviewer #2 (Public review):

      (1) UR241-2 decreases tumors at the injury site but apparently does not significantly decrease omental tumor burden in the syngeneic MiM model. It argues against simple nonspecific antitumor activity and supports a possible role for IRAK4 specifically within an inflammation/injury-dependent metastatic niche. The authors should make much more of this distinction-but also mechanistically prove it.

      We agree that the differential response to UR241-2 at the injury site versus the omentum is an important finding and will make this distinction substantially more prominent in the revised manuscript. We appreciate that reviewer considers this pattern arguing against interpreting UR241-2 simply as a nonspecific inhibitor of ovarian tumor growth. Rather, the preferential reduction of tumor formation at the injury site is consistent with a role for IRAK4 in the establishment of tumors within an injury-associated inflammatory niche.

      This interpretation is also consistent with our data that genetic IRAK4 depletion significantly prolonged survival than null control. This newly generated data will be added to the revised manuscript.

      We will also clarify an important limitation our study pertaining to the omental comparison. Because the omentum is thin and multilayered, direct needle implantation does not provide the same spatial control as implantation at an injury site in the peritoneum; needle penetration may result in tumor-cell deposition beyond the intended omental compartment and complicate quantitative assessment of omental tumor burden. Thus, while we agree that the differential response may be informative, we will avoid interpreting the absence of a significant omental response as definitive evidence that IRAK4 is not involved in omental disease. This is stated so because, importantly, clinical observations have been reported indicating that surgically injured or partially resected omental tissue can support tumor-cell adhesion (Dong X et al. https://doi.org/10.1186/s13048-024-01401-8). This provides an additional rationale for considering injury/inflammation, rather than anatomical site alone, as the relevant biological feature needed for inflammation-mediated seeding. We will incorporate this evidence and provide the published literature citations to show that residual injured omentum, when present, could constitute an inflammatory niche where seeding can occur. Nevertheless, omentum is generally resected in advanced EOC (omentectomy), so it is not available for seeding during the recurrence and hence is largely irrelevant to our study theme.

      (2) The manuscript does not yet establish that UR241-2's antitumor effects are mediated primarily through IRAK4. The compound has measurable activity against additional kinases, including MAP4K2 and LRRK2, and activity against other kinases is reported at higher concentrations. More importantly, there appears to be a substantial concentration disconnect across assays. IRAK4 phosphorylation is inhibited at nanomolar concentrations in some experiments, whereas colony formation/viability phenotypes occur largely in the micromolar range. For example, colony effects are reported at 5-20 µM and viability experiments at 20-60 µM. Thus, are the antiproliferative effects observed at 10-60 µM actually caused by IRAK4 inhibition? The manuscript needs a stronger pharmacological/genetic causality experiment. Ideally, the authors should test UR241-2 in IRAK4-knockdown/knockout cells. If UR241-2 retains essentially identical cytotoxic activity after IRAK4 loss, the mechanistic interpretation would need substantial revision. A rescue experiment with WT versus inhibitor-resistant IRAK4 would be even stronger.

      We agree that establishing pharmacological causality is central to interpreting UR241-2’ s mechanism of action and designating it as a drug-like candidate. In particular, the reviewer appropriately points out that biochemical inhibition of IRAK4 occurs at lower concentrations than some of the antiproliferative and colony-formation phenotypes and that UR241-2 has measurable activity against additional kinases. However, nanoBret assay (Figure-2M) identifies UR241-1 as a very selective IRAK4 kinase inhibitor. The homologue IRAK1 is inhibited at ~1092-fold higher doses than IRAK4. Similarly, FLT3, MAP4K1, MAP4K2, MAP4K3, MAP4K5 and LRRK2 are inhibited at over 18-800 folds. We agree that doses of UR241-2’s at which kinase inhibition and anti-proliferative actions occur do not match. We will attempt to carefully distinguish the concentrations required for direct IRAK4 pathway inhibition from those producing broader cellular phenotypes and will strengthen the pharmacological-genetic evidence linking UR241-2 activity to IRAK4. We will specifically address whether the cellular effects of UR241-2 are diminished when IRAK4 is genetically depleted or otherwise functionally absent. These analyses will help determine the extent to which the higher-concentration antiproliferative effects can be attributed to IRAK4 versus potential off-target activities.

      We will also revise the manuscript so that conclusions regarding UR241-2 mechanism are proportional to the available evidence and do not imply that all cellular effects at higher concentrations necessarily reflect IRAK4 inhibition.

      (3) The distinction between host IRAK4 and tumor-cell IRAK4 is insufficiently resolved. The Il1r1 experiments manipulate the host, whereas IRAK4 knockdown manipulates the tumor cell. UR241-2, meanwhile, presumably inhibits IRAK4 in both compartments. Consequently, the current experiments combine at least two mechanistically distinct possibilities: tumor-intrinsic IRAK4 versus host IRAK4. The manuscript would be considerably stronger if these compartments were experimentally separated. For example, IRAK4-deficient tumor cells implanted into WT versus pathway-deficient hosts, or pharmacological treatment of mice bearing IRAK4-deficient tumor cells, could determine how much of UR241-2 efficacy is tumor-intrinsic versus microenvironment-mediated.

      We agree that this is an important mechanistic distinction. The current experiments interrogate different compartments: IL1R1 manipulation primarily addresses the host inflammatory environment, whereas IRAK4 depletion in tumor cells addresses tumor-intrinsic signaling. UR241-2, in contrast, can potentially affect IRAK4 in both compartments.

      We will therefore more explicitly separate these mechanisms in the revised manuscript. In particular, we will examine the available genetic and pharmacological data in the context of a model in which host IL1R1/inflammatory signaling establishes a permissive injury-associated environment, while tumor-cell IRAK4 contributes to tumor-cell responses within that environment. Where feasible, we will incorporate additional experiments to better distinguish tumor-intrinsic from host-mediated contributions to UR241-2 activity.

      (4) The immune conclusions are presently associative. The increase in MHC-II-positive macrophages and neutrophils is interesting, but describing these populations as demonstrating an "antitumor immune response" is stronger than the evidence warrants. MHC-II expression does not itself demonstrate antitumor function. Likewise, neutrophils in ovarian cancer can be either tumor-promoting or tumor-suppressive depending on context. The authors show that UR241-2 changes immune composition/phenotype. They do not yet demonstrate that these cells mediate the therapeutic effect. This could be addressed by macrophage or neutrophil depletion, functional assays, cytokine profiling, T-cell activation measurements, or potentially single-cell profiling

      We agree that the changes in macrophage and neutrophil populations are currently associative, and do not by themselves establish that these cells mediate the therapeutic response of UR241-2. MHC-II expression, for example, indicates a change in macrophage phenotype but is not sufficient evidence of antitumor activity, and neutrophil function can be context-dependent.

      We will therefore revise the terminology to distinguish changes in immune-cell composition or phenotype from demonstrated immune-mediated tumor suppression. We will integrate these findings more appropriately into the inflammatory-niche model and will limit causal conclusions regarding macrophages and neutrophils unless supported by additional functional evidence.

      (5) The drug-development claims are somewhat premature

      The ADMET package is useful, but several features deserve more cautious interpretation. The compound shows: very high plasma protein binding, rapid mouse microsomal turnover, evidence of efflux, relatively rapid IV clearance, and measurable off-target kinase activity. The authors report mouse microsomal half-life of only ~8.7 min compared with ~209 min in human microsomes and an efflux ratio of ~4.14. These are not fatal problems for a proof-of-concept molecule, but they make language implying a near-clinical candidate premature. UR241-2 currently looks more convincing as a lead/tool compound demonstrating therapeutic tractability of IRAK4 than as an advanced drug candidate.

      We agree that, although the ADMET characterization of UR241-2 is comprehensive and addresses key parameters relevant to drug development, its potential as a drug candidate should be interpreted cautiously.

      Importantly, the available data supports appropriate cellular selectivity for IRAK4. Although the initial HotSpot kinase panel identified additional kinase interactions, most of these activities were substantially weaker than IRAK4 inhibition and were effectively dialed out in the cell-based NanoBRET assay, with target engagement generally requiring approximately 50–1000-fold higher concentrations than those required for IRAK4. Thus, the NanoBRET data provides a more physiologically relevant assessment of cellular target selectivity and supports preferential engagement of IRAK4 by UR241-2 at pharmacologically relevant concentrations.

      We also interpret the high plasma protein binding of UR241-2 in the context of its intended pharmacological profile. Rather than necessarily representing a liability, extensive plasma binding may be compatible with development of a formulation or dosing strategy designed to provide sustained, controlled exposure. This could be advantageous for IRAK4, an important component of innate immune signaling, because prolonged and complete suppression may increase susceptibility to infectious complications, particularly in patients with advanced or treatment-exposed EOC who may already have compromised host defenses. Accordingly, a sustained but incomplete degree of IRAK4 inhibition may be preferable to continuous, complete target blockade. This hypothesis will require direct validation during further drug-development studies.

      The relatively short microsomal half-life observed in mouse microsomes is recognized as a limitation for preclinical development and may complicate exposure and toxicity studies in rodents. However, UR241-2 demonstrated a substantially longer microsomal half-life in human microsomes (209 min) (the intended species), which supports further consideration of the scaffold for human drug development. Collectively, these findings support UR241-2 as a pharmacological proof-of-concept and a tractable lead scaffold, while acknowledging that additional medicinal chemistry and pharmacokinetic optimization will be necessary to establish a suitable development candidate

      However, we will revise the manuscript to position UR241-2 as an investigational lead/tool compound demonstrating the therapeutic tractability of IRAK4 while clearly acknowledging its current development limitations. We will also distinguish proof-of-concept efficacy from the properties required for progression toward a clinical candidate.

      (6) Exposure-response relationships need considerably more attention. Analysis of IRAK4 and/or NF-κB pathway in the treated tumors should be examined

      We agree that linking systemic exposure to tumor pharmacodynamics would strengthen interpretation of the in vivo findings. We will further examine the available tumor samples and pharmacodynamic data for evidence of IRAK4 and/or NF-κβ pathway modulation following UR241-2 treatment and relate these findings, where possible, to compound exposure and tumor response.

      This analysis should help distinguish target engagement and pathway modulation from nonspecific effects occurring at higher systemic or tissue concentrations.

      (7) Some mechanistic observations need deeper validation. The connections among IRAK4, adhesion, E-cadherin, WNT4 and ECM remodeling are intriguing but currently somewhat descriptive. At present, several pieces of this pathway appear adjacent rather than causally connected.

      We agree that these observations currently provide a mechanistic framework but that some relationships are not yet established as directly causal. We will therefore more clearly distinguish findings that demonstrate pathway activity from those that represent downstream associations.

      In particular, we will strengthen the discussion of how IRAK4 signaling may influence tumor-cell adhesion and establishment following tissue injury and will avoid presenting changes in E-cadherin, WNT4, or ECM-associated factors as independently proven causal steps unless directly supported by the data. Where appropriate, additional analyses will be used to clarify the relationship among these processes.

      We appreciate the reviewers' careful evaluation of the manuscript. The comments have helped us sharpen the central model: that tissue injury creates an inflammatory environment that promotes tumor establishment and that IRAK4 represents an important signaling node within this process. The revised manuscript will strengthen the mechanistic evidence for this model, clarify the distinction between injury-associated metastatic establishment and generalized tumor growth, and more appropriately define the translational potential and limitations of UR241-2.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigates low-affinity Ca2+ binding by WT calreticulin and mutant calreticulin associated with type I myeloproliferative neoplasms, as well as the impact on Ca2+ fluxes in suspension cultures of megakaryocyte-like cells in vitro in response to ER Ca2+ ATPase inhibitors that deplete endoplasmic reticulum (ER) Ca2+ store and open plasma membrane Ca2+ channels through STIM1-Orai interactions. The results are important in that they show that Ca2+ binding by calreticulin and store-operated Ca2+ entry are not fundamentally impacted by the type I deletion mutation in calreticulin, which rules out a direct effect of the calreticulin mutation on its own low-affinity Ca2+ binding and any broad impact on ER Ca2+ regulation. The strength of the data and methods used ranges from solid to convincing, although the use of suspension-based flow cytometric assays to investigate ER Ca2+ levels and Ca2+ entry can be challenged. High-affinity Ca2+ binding sites could be further considered, and possible confounding effects of Abl kinase activity in the megakaryocyte-like cell lines could be offset.

      The authors thank the editors and the reviewers for the summary, comments and many helpful suggestions. In the revised manuscript, we have used fluorimetry for more precise ER calcium measurements (new Figure 5), clarified the high-affinity calcium binding site concern based on our previous work, and addressed possible effects of BCR-ABL translocation kinase activity upon cellular calcium signaling using the drug Imatinib (new supplements to Figure 6 and 7).

      Public Reviews:

      Reviewer #1 (Public review):

      The researchers conducted their study using advanced techniques. They found almost no difference in calcium binding between the two proteins and observed no impact on calcium signaling, specifically store-operated calcium entry (SOCE). The study also noted an increase in ER luminal calcium-binding chaperone proteins. Surprisingly, the authors selected flow cytometry as a technique for measurements of ER luminal calcium. Considering the limitations of this approach, it would be better to use alternative approaches.

      Thank you for this suggestion. We have undertaken fluorimetry-based ER calcium measurements, which are shown in a new Figure 5. These also indicate similar ER calcium levels in CRT-KO HEK293T cells, compared to those reconstituted with wild-type CRT and CRT<sub>Del52</sub>.

      This is particularly important as previous reports, using cells from MPN patients, indicate reduced ER luminal calcium and effects on SOCE (Blood, 2020). This issue matters because earlier research with MPN patient cells reported reduced ER luminal calcium levels and altered SOCE (Blood, 2020). How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      We thank the reviewer for asking for these clarifications. We have revised the discussion to address some of these points and also clarify the findings of the referenced study (Di Buduo et al., 2020) which did not directly measure ER calcium levels. We also discuss findings from a related study with cultured megakaryocytes from patients that indicated different effects of type I vs type II mutations (Pietra et al., 2016). In the absence of engineered controls, sample-to-sample heterogeneities in primary cells make it difficult to attribute any measured differences as direct effects of CRT mutations. Different from these experiments, by using purified proteins and ITC, our studies show that the Del52 mutant has calcium-binding characteristics resembling that of the wild-type protein. Additionally, through genetic manipulations in cell lines, our studies directly address the effects of calreticulin KO and its Del52 mutation upon ER luminal and cytosolic calcium levels, and cellular SOCE signals. We did not measure significant differences in any of these parameters between the KO cells and those reconstituted with wild-type calreticulin or the Del52 mutant. As noted by the editors, these results show that Ca2+ binding by calreticulin and SOCE in a cell are not fundamentally impacted by the type I deletion mutation.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      These points are addressed in the discussion. Either protein misfolding in cells with wild-type calreticulin deficiency or the sensing of cellular calcium perturbations could induce the expression of ER calcium-binding proteins in calreticulin-deficient cells, although we favor the latter model for the reason specified in the discussion. Regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      Overall, it remains unclear how this work improves our understanding of MPN or clarifies calreticulin's role in MPN pathophysiology.

      Multiple studies referenced in the manuscript have suggested links between altered calcium signaling/binding by CRT mutants and MPN pathogenesis. Our studies indicate that ER and cytosolic calcium levels and SOCE are not directly impacted by the MPN type I CALR mutation, points noted in the abstract and discussion. Thus, calcium signaling may not play a specific role in MPN CALR mutant pathology via suggested mechanisms. We are confident that readers will find these results important for better understanding the role of calreticulin type I mutations in MPN.

      Reviewer #1 (Recommendations for the authors):

      This study aimed to express, purify, and evaluate low-affinity calcium binding by a calreticulin deletion mutant (CRTdel52) that is linked to myeloproliferative neoplasms (MPN). The researchers performed cell imaging, flow cytometry, and isothermal titration calorimetry to compare calcium binding between wild-type calreticulin and CRTDel52. They assessed cytosolic calcium levels and store-operated calcium entry (SOCE) in HET293T cells (CRT knocked-out background) and in megakaryoblastic MEG-1 cells. Additionally, they examined changes in the abundance of endoplasmic reticulum (ER) resident calcium-binding proteins in cells expressing either wild-type or mutant CRT.

      This study is well executed but lacks clear relevance to MPN, and it is not clear how this work advances our knowledge of calreticulin biology. No differences were found in calcium binding between wild-type calreticulin and CRTDel52, nor was SOCE impacted. They noticed, however, a compensatory increase in the abundance of some ER resident calcium-binding proteins. The lack of any significant changes in calcium behavior between wild type and CRTDel52 is not surprising based on the known amino acid sequence of calreticulin and calreticulin mutant and based on our knowledge about CRT calcium binding in general. Consequently, it is not clear how this work advances our understanding of the pathophysiology of MPN. Calcium may not play a critical role in the MPN pathology; instead, CRTDel52 secretion and receptor signaling appear more central. Further research should address how these findings relate specifically to MPN and cell biology, in general.

      Our current studies demonstrate increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G). These results, together with findings of maintained ER and cytosolic calcium levels and SOCE signals (Figures 4-7 and Figure 6, supplemental Figure 2 and Figure 7, supplemental Figure 1), indicate that altered calcium binding/signaling by Del52 does not directly contribute to MPN pathology.

      What is the biological or pathophysiological relevance of the CRTDel52-KDEL construct?

      The KDEL sequence is important for the ER retention of CRT (Sonnichsen et al., 1994), and its addition was expected to at least partially remedy the ER retention defect of CRT<sub>Del52</sub>. This point is clarified in the revised results section.

      The rise in ER calcium-binding proteins is noteworthy but anticipated, given likely genetic changes from UPR pathway activation in these cells. Is this relevant to MPN?

      We suggest that increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G) would contribute to the maintenance of the cell’s calcium signaling capacity.

      How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      The findings related to SOCE are addressed in the points discussed above and in the revised discussion. Related to ER luminal calcium, the study by Ibarra et al. (Ibarra et al., 2022) reported that CRT<sub>Del52</sub> overexpressed in U2OS cells (expressing endogenous CRT) had reduced ER calcium levels compared to the same cells expressing WT CRT or CRT<sub>Ins5</sub>. Those measurements did not use a ratiometric ER calcium probe, and additionally it is possible that the expression of compensatory calcium-binding proteins is more muted in cells expressing endogenous wild-type CRT.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      We agree that increased UPR could account for higher levels of some ER-resident calcium-binding proteins. As noted in the revised discussion, regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      It is not clear why flow cytometry was a choice of technique for measurements of ER calcium. Pacific Blue was detected at 405 nm excitation and 452-455 nm emission, while unbound probe signals appeared in the AmCyan channel (405 nm excitation/498 nm emission). It is unnecessary to mention fluorochrome labels (like Pacific Blue or AmCyan) for channels that are not in use. Only the channels actually utilized need to be specified: The calcium-bound GEM-CEPIA1er probe's signal was collected using the Pacific Blue channel (excitation at 405 nm, emission at 452-455 nm), while the unbound probe's signal was detected using the AmCyan channel (excitation at 405 nm, emission at 498 nm). Since it's not possible to monitor emission only at a specific wavelength with a Fortessa, could this be a different channel that is being recorded?

      Additionally, due to the similar spectra and potential for bleed-through between Pacific blue and Amcyan, compensation is likely necessary. Therefore, you should include single-stained control cells containing only one probe for proper reporting. Additionally, only the "GEM-CEPIA1er probe" is displayed, while the second probe is referred to solely as "unbound".

      A single genetically encoded GEM-CEPIA1er probe (Suzuki et al., 2014) was used for measuring both the bound and unbound signals. The GEM-CEPIA1er probe was excited with the 405 nm violet laser. In the methods section of the revised manuscript, the GEM-CEPIA1er probe wording is included for describing both the bound and unbound signal collections. Additionally, we have undertaken new spectrofluorimetric experiments (new Figure 5), which allow for the distinct emission peaks to be recorded corresponding to the Ca<sup>2+</sup>-bound and Ca<sup>2+</sup>-unbound signals. Similar results were obtained as reported for the flow cytometry-based experiments.

      Increased expression of wild-type calreticulin compared to parental cells should impact on ER calcium content and dynamics in back-transfected HEK293T-KO or MEG-1 cells. Direct ER calcium measurements in HEK293 cells with various calreticulin constructs would significantly strengthen this presentation.

      Our experiments were structured to compare calcium signaling in cells expressing only wild-type CRT or CRT<sub>Del52</sub> (resembling homozygous type I MPN CALR mutations) compared to CRT-KO cells. The over-expression of CRT in the reconstituted cells compared to endogenous expression level is a limitation of our study which we have acknowledged in the revised manuscript discussion. Understanding the effects of over-expression of wild-type CRT vs the CRT<sub>Del52</sub> mutant upon ER and cytosolic calcium signals and SOCE is interesting, but beyond the scope of the present study.

      The authors should examine the immunolocalization of CRTDel52 and wild-type protein in HEK293 cells.

      Previous published studies from another lab showed that CRT<sub>Del52</sub> is secreted from HEK cells and that the addition of a KDEL sequence to CRT<sub>Del52</sub> reduces secretion and induces its increased intracellular accumulation (Arshad and Cresswell, 2018). This point is noted in the revised results section and the reference is cited. This appears to be the general theme in primary cells and cell lines. Previous studies and our own prior published studies have shown that CRT<sub>Del52</sub> (but not wild-type CRT) is detectable in the media of cell lines and patient serum as well on the cell surface of primary cells and cell lines (Kaur et al., 2024, Venkatesan et al., 2021, Pecquet et al., 2023).

      SDS-PAGE of purified proteins is overloaded, and chromatograms show extra peaks or shoulders, making protein quality assessment uncertain.

      Representative chromatograms, peaks corresponding to protein monomers used for ITC analyses and the relevant gels are clarified in the revised manuscript. In new analyses since the original submission, intact protein mass spectrometry was undertaken for CRT<sub>Del52</sub>. The results indicate a 35-42 amino acid truncation in different preparations. The truncated proteins would still include acidic residues (between 340–351) previously implicated in low-affinity calcium binding by murine CRT that are shared between wild-type and CRT<sub>Del52</sub>. This new information is now included in the revised results section.

      Analysis of SOCE in calreticulin-deficient cells and cells reconstituted with calreticulin or overexpressing the protein has already been reported (PMID12324449).

      The indicated reference and additional related papers examining effects of CRT deficiency and overexpression on cellular calcium signaling (Arnaudeau et al., 2002, Bastianutto et al., 1995, Mery et al., 1996, Nakamura et al., 2001) are cited in the revised manuscript.

      Reviewer #2 (Public review):

      Tagoe and colleagues present a thorough analysis of the calcium (Ca2+) binding capacity of calreticulin (CRT), an endoplasmic reticulum (ER) Ca2+-buffer protein, using a mutant version (CRT del52) found in myeloproliferative neoplasms (MPNs). The authors use purified human CRT protein variants, CRT-KO cell lines, and an MPN cell line to elucidate the differing Ca2+ dynamics, both on the level of the protein and on cell-wide Ca2+-governed processes. In sum, the authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      First, the authors purify CRT protein and perform isothermal titration calorimetry to quantify the Ca2+ binding capacity of CRT. They use full-length human CRT, CRT del52, and two truncations of CRT (1-339 and 1-351, the former of which should lead to the entire loss of low-affinity Ca2+ binding). While CRT del52 has previously been shown to lead to a decrease in Ca2+ binding affinity in other models, the ITC data show that this is retained in CRT del52.

      Next, the authors utilize a CRT-KO cell line with subsequent addition of CRT protein variants to validate these findings with flow cytometric analysis. Cells were transfected with a ratiometric ER Ca2+ probe, and fluorescence indicates that CRT del52 is unable to restore basal ER Ca2+ levels to the same extent as CRT wild-type. To translate these findings to MPNs, the authors perform CRT-KO in a megakaryocytic cell line, where reconstitution with either CRT variant did not cause a difference in cytosolic calcium levels. The authors further test store-operated calcium entry (SOCE), an important process for maintaining ER Ca2+ levels, in these cells, and find that CRT-KO cells have lower SOCE activity, and that this can be slightly recovered with CRT addition.

      Finally, the authors ask whether other effects of CRT-KO/reconstitution can affect the cellular Ca2+ signaling pathway and levels. RNASeq analysis revealed that CRT-KO leads to an increase in various chaperone protein expressions, and that reconstitution with CRT del52 is unable to reduce expression to the same extent as reconstitution with CRT wildtype.

      Strengths:

      The authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      We thank the reviewer for the recognition that this study is important for our understanding of both normal and malignant cell biology.

      Weaknesses:

      (1) The authors should consider discussing the high-affinity Ca2+ binding site more in the introduction. Can they show a proof-of-concept experiment that validates that incubation of recombinant CRT reduces the function of that high-affinity Ca2+ binding site?

      In a previous study (Wijeyesakere et al., 2011), we showed that at a starting calcium concentration of 0 mM and with CaCl<sub>2</sub> injections to a final concentration of 70-80 mM the measured K<sub>D</sub> value was 16.6 mM for calcium binding to wild type murine calreticulin, (which has ~95% sequence identity with human calreticulin), corresponding to the high-affinity site. On the other hand, at a starting calcium concentration of 50-100 mM and CaCl<sub>2</sub> injections to a final concentration of 700-850 mM, the measured K<sub>D</sub> value for calcium binding to wild-type murine calreticulin was 590 mM (corresponding to the low-affinity sites). We did not observe the high-affinity sites when the starting calcium concentration was 50 mM and calcium injections were at 33 mM each; similar conditions are used in the present study. These points are clarified in the revised manuscript in the results section.

      (2) For Figure 2B, do you have an explanation for why the purified proteins run higher than predicted (48-52kDa) - are these proteins still tagged with pGB1?

      Yes, the purified proteins shown in Figure 2B retained a GB1 tag. This point is clarified in the revised methods.

      (3) The MEG-01 cell line has the BCR:ABL1 translocation, while CRT mutations are strictly found in BCR:ABL1 negative MPNs. Could these experiments be repeated in these cells treated with imatinib to decrease these effects, or see if basal MEG-01 Ca2+ levels/activity are changed with or without imatinib?

      Thank you for this important point. We have assessed cytosolic calcium levels in MEG-01 cells that were treated or not treated with imatinib in new Figure 6, supplemental Figures 1 and 2 and Figure 7, supplemental Figure 1) and show that the prior results hold in imatinib-treated cells.

      References

      ARNAUDEAU, S., FRIEDEN, M., NAKAMURA, K., CASTELBOU, C., MICHALAK, M. & DEMAUREX, N. 2002. Calreticulin differentially modulates calcium uptake and release in the endoplasmic reticulum and mitochondria. J Biol Chem, 277, 46696-705.

      ARSHAD, N. & CRESSWELL, P. 2018. Tumor-associated calreticulin variants functionally compromise the peptide loading complex and impair its recruitment of MHC-I. J Biol Chem, 293, 9555-9569.

      BASTIANUTTO, C., CLEMENTI, E., CODAZZI, F., PODINI, P., DE GIORGI, F., RIZZUTO, R., MELDOLESI, J. & POZZAN, T. 1995. Overexpression of calreticulin increases the Ca2+ capacity of rapidly exchanging Ca2+ stores and reveals aspects of their lumenal microenvironment and function. J Cell Biol, 130, 847-55.

      DI BUDUO, C. A., ABBONANTE, V., MARTY, C., MOCCIA, F., RUMI, E., PIETRA, D., SOPRANO, P. M., LIM, D., CATTANEO, D., IURLO, A., GIANELLI, U., BAROSI, G., ROSTI, V., PLO, I., CAZZOLA, M. & BALDUINI, A. 2020. Defective interaction of mutant calreticulin and SOCE in megakaryocytes from patients with myeloproliferative neoplasms. Blood, 135, 133-144.

      IBARRA, J., ELBANNA, Y. A., KURYLOWICZ, K., CIBODDO, M., GREENBAUM, H. S., ARELLANO, N. S., RODRIGUEZ, D., EVERS, M., BOCK-HUGHES, A., LIU, C., SMITH, Q., LUTZE, J., BAUMEISTER, J., KALMER, M., OLSCHOK, K., NICHOLSON, B., SILVA, D., MAXWELL, L., DOWGIELEWICZ, J., RUMI, E., PIETRA, D., CASETTI, I. C., CATRICALA, S., KOSCHMIEDER, S., GURBUXANI, S., SCHNEIDER, R. K., OAKES, S. A. & ELF, S. E. 2022. Type I but Not Type II Calreticulin Mutations Activate the IRE1alpha/XBP1 Pathway of the Unfolded Protein Response to Drive Myeloproliferative Neoplasms. Blood Cancer Discov, 3, 298-315.

      KAUR, A., VENKATESAN, A., KANDARPA, M., TALPAZ, M. & RAGHAVAN, M. 2024. Lysosomal degradation targets mutant calreticulin and the thrombopoietin receptor in myeloproliferative neoplasms. Blood Adv, 8, 3372-3387.

      MERY, L., MESAELI, N., MICHALAK, M., OPAS, M., LEW, D. P. & KRAUSE, K. H. 1996. Overexpression of calreticulin increases intracellular Ca2+ storage and decreases store-operated Ca2+ influx. J Biol Chem, 271, 9332-9.

      NAKAMURA, K., ZUPPINI, A., ARNAUDEAU, S., LYNCH, J., AHSAN, I., KRAUSE, R., PAPP, S., DE SMEDT, H., PARYS, J. B., MULLER-ESTERL, W., LEW, D. P., KRAUSE, K. H., DEMAUREX, N., OPAS, M. & MICHALAK, M. 2001. Functional specialization of calreticulin domains. J Cell Biol, 154, 961-72.

      PECQUET, C., PAPADOPOULOS, N., BALLIGAND, T., CHACHOUA, I., TISSERAND, A., VERTENOEIL, G., NEDELEC, A., VERTOMMEN, D., ROY, A., MARTY, C., NIVARTHI, H., DEFOUR, J. P., EL-KHOURY, M., HUG, E., MAJOROS, A., XU, E., ZAGRIJTSCHUK, O., FERTIG, T. E., MARTA, D. S., GISSLINGER, H., GISSLINGER, B., SCHALLING, M., CASETTI, I., RUMI, E., PIETRA, D., CAVALLONI, C., ARCAINI, L., CAZZOLA, M., KOMATSU, N., KIHARA, Y., SUNAMI, Y., EDAHIRO, Y., ARAKI, M., LESYK, R., BUXHOFER-AUSCH, V., HEIBL, S., PASQUIER, F., HAVELANGE, V., PLO, I., VAINCHENKER, W., KRALOVICS, R. & CONSTANTINESCU, S. N. 2023. Secreted mutant calreticulins as rogue cytokines in myeloproliferative neoplasms. Blood, 141, 917-929.

      PIETRA, D., RUMI, E., FERRETTI, V. V., DI BUDUO, C. A., MILANESI, C., CAVALLONI, C., SANT'ANTONIO, E., ABBONANTE, V., MOCCIA, F., CASETTI, I. C., BELLINI, M., RENNA, M. C., RONCORONI, E., FUGAZZA, E., ASTORI, C., BOVERI, E., ROSTI, V., BAROSI, G., BALDUINI, A. & CAZZOLA, M. 2016. Differential clinical effects of different mutation subtypes in CALR-mutant myeloproliferative neoplasms. Leukemia, 30, 431-8.

      SONNICHSEN, B., FULLEKRUG, J., NGUYEN VAN, P., DIEKMANN, W., ROBINSON, D. G. & MIESKES, G. 1994. Retention and retrieval: both mechanisms cooperate to maintain calreticulin in the endoplasmic reticulum. J Cell Sci, 107 (Pt 10), 2705-17.

      SUZUKI, J., KANEMARU, K., ISHII, K., OHKURA, M., OKUBO, Y. & IINO, M. 2014. Imaging intraorganellar Ca2+ at subcellular resolution using CEPIA. Nat Commun, 5, 4153.

      VENKATESAN, A., GENG, J., KANDARPA, M., WIJEYESAKERE, S. J., BHIDE, A., TALPAZ, M., POGOZHEVA, I. D. & RAGHAVAN, M. 2021. Mechanism of mutant calreticulin-mediated activation of the thrombopoietin receptor in cancers. J Cell Biol, 220, e202009179.

      WIJEYESAKERE, S. J., GAFNI, A. A. & RAGHAVAN, M. 2011. Calreticulin is a thermostable protein with distinct structural responses to different divalent cation environments. J Biol Chem, 286, 8771-85.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study revisits an important and controversial question in brain repair: whether NeuroD1 can convert brain immune cells into nerve cells in vivo. Using a virus-free genetic system, in vivo imaging, injury experiments, and single-cell profiling, the authors provide convincing evidence that NeuroD1-expressing cells do not become nerve cells under the tested conditions. Instead, these cells largely retain their original immune-cell identity, and some appear to undergo cellular stress or loss.

      Strengths:

      The main strength of the work is that it tests this question with a cleaner genetic strategy, avoiding some of the concerns associated with viral delivery and unintended cell labeling. Although the overall conclusion is consistent with the authors' previous work, the current study adds useful independent evidence, particularly through the virus-free fate-mapping system and live imaging in the brain.

      Weaknesses:

      There are some limitations. In the injury experiment, the labeled cells may include both resident brain immune cells and blood-derived immune cells recruited after injury, so the authors should be cautious when referring to all labeled cells as microglia. The level of NeuroD1 expression achieved by the genetic system is also not fully defined, which matters because the effects of such a cell-fate regulator may depend on expression level. Finally, the tested time window may not fully address very delayed or incomplete neuronal differentiation.

      Overall, this is a useful and careful study that supports the conclusion that NeuroD1 does not drive brain immune cells to become nerve cells in the tested settings. It should be valuable for researchers studying brain repair, cell fate conversion, and genetic fate mapping, and it provides a clear caution against overinterpreting reprogramming results based only on viral labeling.

      We’d like to thank the reviewer for providing these thoughtful suggestions and positive assessment of our study. We agree that CX3CR1 lineage can label both microglia and blood-derived immune cells after injury, To more specifically assess microglial lineage tracing and minimize the potential contribution of infiltrating peripheral myeloid cells, we performed additional experiments using TMEM119-CreER::LSL-NeuroD1-EGFP mice. These experiments showed no evidence of microglia-to-neuron conversion following TBI. To evaluate the expression level of NeuroD1, we performed qPCR to verify the NeuroD1 expression level in the CD11b<sup>+</sup> cell population of the NeuroD1-expressing mice, and found that NeuroD1 expression was increased approximately 6.21-fold compared with control mice, indicating robust NeuroD1 expression. For the question about the time window, we have revised our Discussion and Conclusion parts to avoid overgeneralizing our findings beyond the experimental conditions tested. We agree that our data cannot fully exclude the possibility of delayed neuronal differentiation or neuronal conversion under other pathological conditions. Besides, we have also revised the manuscript accordingly to address the concerns raised by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      In vivo glia-to-neuron conversion emerges as a potential regeneration-based therapeutic strategy for neural injuries and diseases. However, controversies exist in this exciting field, largely arising from the non-stringent methods employed for analyzing in vivo neuronal conversions. The study by Li et al. directly addressed this controversy regarding Neurod1-mediated microglia-to-neuron conversion. They took advantage of two transgenic mouse lines to specifically express Neurod1 in the microglia of adult mouse brains. Results from immunohistochemistry, in vivo live-cell imaging, and scRNAseq convincingly demonstrate that microglia cannot be converted in vivo to neurons by ectopic Neurod1 expression under both normal and injury conditions. Instead, it induces microglia death, consistent with their earlier findings. These solid results, though negative, are critical additions to the field and further support that stringent lineage tracing methods are essential for studying in vivo cell reprogramming. Overall, the studies are rigorously designed and executed. Only minor issues need to be dealt with.

      We’d like to thank the reviewer for providing these thoughtful comments and for the positive assessment of our study. We are glad that the reviewer agrees that our results from immunohistochemistry, in vivo live-cell imaging, and scRNA-seq support the conclusion that microglia cannot be converted into neurons by NeuroD1 expression under the tested conditions. We have carefully addressed the minor issues raised by the reviewer, including rechecking the grammar throughout the manuscript, revising the description of the TBI behavioral results, discussing the limitations of scRNA-seq for neuronal detection, adding the information of the promoters used in the study, and carefully revising the references. Besides, we have carefully revised the manuscript according to the reviewer’s suggestions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) TBI lineage tracing may label both microglia and infiltrating macrophages. Because tamoxifen was administered after TBI, CX3CR1-CreER may label not only resident microglia but also injuryrecruited CX3CR1+ monocytes/macrophages. The authors should avoid referring to all reporterpositive cells as microglia unless this is clearly supported. They should clarify the composition of reporter-positive cells in the TBI setting, ideally using existing scRNA-seq data or markers distinguishing microglia from infiltrating macrophages.

      The current data still support the conclusion that CX3CR1-lineage myeloid cells do not show obvious neuronal conversion in this TBI model, but the wording should be more precise.

      Thank you for pointing this out. We agree that, because TBI can recruit CX3CR1-expressing monocytes/macrophages, CX3CR1-CreER lineage tracing after injury may label both resident microglia and infiltrating myeloid cells. To more specifically assess the contribution of resident microglia, we therefore performed additional lineage-tracing experiments using TMEM119CreER::Ai14 (TMEM119-Ai14) and TMEM119-CreER::LSL-NeuroD1-EGFP(TMEM119-ND1) mice, with tamoxifen administration before or after TBI. In both experimental settings, we did not detect reporter-positive cells co-expressing the neuronal marker NeuN in either the lesion core or distal regions. These results provide no evidence of microglia-to-neuron conversion in the TBI model within the examined time window. We have also revised the text throughout the manuscript to distinguish resident microglia from broader CX3CR1-lineage myeloid cells where appropriate.

      Author response image 1.

      TMEM119-CreER::LSL-NeuroD1-EGFP animal also show no microglia-to-neuron conversion in TBI. (A)Tamoxifen induction after TBI shows there is no GFP<sup>+</sup>NeuN<sup>+</sup> cells in both lesion core and distal region. (B)Tamoxifen induction before TBI shows there is no GFP<sup>+</sup>NeuN<sup>+</sup> cells in both lesion core and distal region.

      (2) NeuroD1 expression level should be better characterized. The authors show that GFP-positive cells co-express NeuroD1, but the approximate NeuroD1 expression level is not clear. This is important because the effect of a fate-determining transcription factor may be dose-dependent. The authors should provide, if possible, a quantitative estimate of NeuroD1 expression using available IF, scRNA-seq, qPCR, or other data. This would help interpret the negative reprogramming result and would also be useful for future studies using this Rosa26-LSL-NeuroD1-IRES-GFP mouse line.

      Thank you for this valuable suggestion. To quantitatively assess NeuroD1 expression in the myeloid compartment, we isolated CD11b<sup>+</sup> cells from CX3CR1-CreER::Ai14 (CX3CR1-Ai14) and CX3CR1-CreER::LSL-NeuroD1-EGFP (CX3CR1-ND1) mice using MACS and performed qPCR analysis of Neurod1 expression. Neurod1 expression was increased approximately 6.21-fold in CD11b<sup>+</sup> cells from CX3CR1-ND1 mice compared with the corresponding control mice (Author response image 2). These results confirm robust NeuroD1 expression in the CD11b<sup>+</sup> cell population of the NeuroD1-expressing mice. This level of induction is broadly comparable to the approximately 4-fold increase in Neurod1 expression reported following NeuroD1 induction in the intestine, in which the same LSL-NeuroD1-EGFP transgenic mouse line was used [1].

      Author response image 2.

      Using qPCR verify the relative expression of Neurod1.

      (3) Figure 4 scRNA-seq should show Neurod1 expression. The scRNA-seq data show that reporterpositive cells retain microglial markers and lack neuronal markers, but Neurod1 expression itself is not shown. The authors should consider adding Neurod1 feature plots, violin plots, or average expression comparisons between control and NeuroD1-expressing groups. This would help validate the expression system and estimate the level achieved in sorted reporter-positive cells.

      Thank you for your suggestion. We examined the expression of Neurod1 in our scRNA-seq database, but the Neurod1 transcripts were rarely detected (not just in this database, also our previous dataset). This might because of Neurod1 is a transcription factor and single-cell RNA-seq method is not suitable for detect these low-expression transcript factor genes, so we utilized qPCR to evaluate the expression of Neurod1, as shown in Author response image 2.

      (4) The microglial loss phenotype may be related to NeuroD1/EGFP dosage. In Figure 3, the loss of GFP-positive microglia may reflect a specific effect of NeuroD1 in microglia, but it could also be related to excessive transgene expression or expression-system toxicity. The authors should discuss this possibility and, if possible, examine whether stress/death markers correlate with NeuroD1 or GFP expression intensity.

      Thanks for your suggestion. We agree with that, so we performed TUNEL staining on CX3CR1-CreER::Ai14 (CX3CR1-Ai14 for short) and CX3CR1-CreER::LSL-NeuroD1EGFP(CX3CR1-ND1 for short) mice at D4 and D18. The data shows that TUNEL<sup>+</sup> Reporter<sup>+</sup> cells were higher in NeuroD1-expressing microglia, so we’d like to say death cells (or apoptotic cells) are positively correlated with NeuroD1 expression. Besides, we also discussed more about this. See Figure 3 E, F, G and Fig. S3E

      (5) The time window and injury context should be more cautiously discussed. The study follows the genetic model up to 45 days. This is informative, but delayed or incomplete neuronal differentiation cannot be fully excluded, especially if conversion would require a longer repair phase after injury. In addition, except for TBI, the study does not include another neuronal-loss context that might provide a permissive regenerative niche. The authors do not necessarily need additional long-term experiments, but they should avoid overgeneralizing beyond the tested time window and injury model.

      Thank you for this thoughtful comment. We agree that our study evaluated the effects of NeuroD1 expression within a defined experimental time window (up to 45 days) and under the specific physiological and injury conditions examined. Therefore, our data cannot exclude the possibility that neuronal conversion might occur at later time points or under other pathological conditions that provide a more permissive regenerative environment. In response to the reviewer's suggestion, we have revised the Discussion and Conclusion to avoid overgeneralizing our findings beyond the experimental conditions tested. We now emphasize that NeuroD1 expression alone did not induce microglia-to-neuron conversion within the time frame and injury paradigms examined in this study, rather than concluding that such conversion can never occur under other conditions.

      Minor comments

      (1) In Figure 2, two panels are labeled "D"; the second should likely be "E."

      Thank you for pointing this out, we have corrected it in figure legend.

      (2) Several spelling and grammar errors should be corrected, such as "Represent images," "illutrating," and "GFP-postive."

      Thank you for pointing this out, we have corrected it in our manuscript.

      (3) More scRNA-seq details should be provided, including cell numbers per group, QC metrics, and cluster annotation information.

      Thanks for your suggestion, we have added more details about scRNA-seq data as shown in Figure S4.

      (4) The authors should maintain cautious wording throughout, especially when referring to "microglia" versus broader CX3CR1-lineage myeloid cells.

      Thanks for your suggestion, we added the data about we utilize TMEM119-ND1 mouse line to trace microglia-to-neuron conversion, which is specific to microglia not myeloid cell. Besides, we also carefully revise these minor comments in our manuscript.

      Reviewer #2 (Recommendations for the authors):

      The following are some required minor edits.

      (1) Please recheck the grammar throughout the manuscript.

      Thank you for pointing this out, we have carefully revised our manuscript.

      (2) Figure 4B-C: "indicating that NeuroD1 expression in microglia does not rescue TBI induced motor dysfunction". This should be rephrased, since dysfunction was not detected when comparing TBI and sham.

      Thank you for pointing this out, we changed our expression into “indicating that NeuroD1 expression in microglia does not improve the performance in TBI model”.

      (3) Interpretation of the scRNA-seq data needs to be cautious, since neurons normally do not survive well during this procedure.

      Thank you for your suggestion, we discussed about this in our discussion part.

      (4) It will be informative to list the sources or sequences of the promoters used.

      Thanks for your advice, all the virus information are provided in our method part.

      (5) Some of the references are repetitive.

      Thank you for your suggestion, we have carefully revised the reference.

      Reference

      (1) Li, H. J. et al. Intestinal Neurod1 expression impairs paneth cell differentiation and promotes enteroendocrine lineage specification. Sci Rep 9, 19489 (2019). https://doi.org/10.1038/s41598-019-55292-7

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editors and reviewers for their thoughtful evaluation and helpful recommendations. These comments helped us distinguish more clearly between two complementary advances in the study. First, the phenotypic, transcriptomic, and ontogenetic analyses refine the organization of the thymic macrophage compartment and identify CCR2 dependence of the TIMD4- VCAM1+ population. Second, the MaFIA fetal thymus organ culture (FTOC) experiments reveal that an intact Csf1r-expressing myeloid compartment is required for efficient progression across the DN3-to-DN4 checkpoint. The revised manuscript now presents both advances directly while matching the cellular specificity of each conclusion to the experimental system that supports it.

      The central functional result is supported by a coordinated set of observations: depletion of Csf1r expressing cells reduces CD4<sup>+</sup>CD8<sup>+</sup> double-positive (DP) ab-lineage thymocyte production, preserves γδ T cell output, causes reciprocal accumulation of DN3 and loss of DN4 cells, and reduces CD27 expression within the DN compartment. We now emphasize that convergence, which localizes the phenotype to the b-selection transition. At the same time, because the MaFIA system targets Csf1r-expressing myeloid cells rather than a single macrophage subset, the manuscript assigns the demonstrated requirement to the thymic myeloid compartment, this now reflected on the revised title as well. This framing preserves the biological importance of the result without attributing more cellular specificity than the experiment provides.

      The major revisions include:

      - A revised Title and an Abstract that leads with the principal conclusions: TM specialization, unequal developmental contributions, CCR2 dependence, and a requirement for Csf1r expressing myeloid cells during DN3-to-DN4 progression.

      - A clearer account of what Zhou et al. established and how the present work adds VCAM1-based prospective resolution, intravascular-labeling information, SpiC analysis, CCR2 dependence, and functional fetal thymus organ culture (FTOC) data.

      - A Results section that explains the convergence of DN3, DN4, CD27, DP cells, and γδ T cell measurements as part of developmental transition.

      - More precise interpretation of relative marker expression, pseudotime, SpiC deficiency, Tomato/GFP double-positive events, and thymocyte-associated transcripts in the scRNA-seq results.

      - A focused Discussion paragraph that acknowledges the cellular breadth of the MaFIA model while retaining the conclusion that the Csf1r-expressing myeloid niche supports early ab T-lineage development.

      - Expanded figure legends that make the gating, reporter, heatmap, MaFIA construct, and developmental interpretations easier to follow.

      The revision is based on fuller analysis and clearer presentation of the results and datasets.

      eLife Assessment

      The macrophage characterisation is interesting, although the evidence for the specific involvement of macrophages in beta-selection is incomplete, as alternative explanations have not been ruled out.

      We agree that the depletion experiment resolves a requirement for the Csf1r-expressing thymic myeloid compartment rather than for macrophages alone. We have revised the manuscript with this distinction in mind. Importantly, the developmental conclusion remains strong: DN3 accumulation, DN4 loss, reduced CD27 expression, and diminished CD4<sup>+</sup>CD8<sup>+</sup> double-positive (DP) output all point to impaired progression at the b-selection checkpoint, while preserved γδ T cell output and epithelial-cell numbers argue against nonspecific failure of the entire fetal thymus organ culture (FTOC). The revised text therefore states that Csf1r-expressing myeloid cells support this early ab T cell checkpoint, while discussing the relative contributions of TMs, monocytes, and DCs as the next level of cellular resolution.

      The Title, Abstract, final Results section, Discussion, Conclusion, and Figures 8-10 Legends were revised to make the positive compartment-level conclusion explicit and consistent.

      Public Reviews:

      Reviewer #1 (Public review):

      The thymic macrophage depletion experiments are not well controlled; DCs are also depleted, the fetal thymus has little or no medulla, and direct AP20187 toxicity to DN thymocytes in MaFIA mice has not been excluded.

      The reviewer identifies the key issue of cellular attribution. We have recast the experiment according to what the MaFIA system directly tests: the functional contribution of Csf1r-expressing myeloid cells within an intact FTOC. This interpretation is supported by efficient loss of both TM populations, preservation of EpCAM+ epithelial-cell numbers, and absence of TM depletion in AP20187treated non-transgenic C57BL/6 FTOCs. We no longer use adult cortical-versus-medullary anatomy to infer that macrophages must be the sole responsible population in the fetal thymus. Instead, we emphasize the experimentally secure result that perturbing the Csf1r-expressing myeloid niche produces a selective and internally consistent defect in ab T cell development at the DN3-to-DN4 transition. The possibility of contributions from DCs, monocytes, or direct transgene expression in a thymocyte fraction is addressed once, in a focused Discussion paragraph, as the rationale for assigning the conclusion at the compartment level.

      The MaFIA Results section now leads with the transgene-dependent depletion and the convergent developmental phenotype; the Discussion contains a balanced statement of cellular resolution; the Title and Legends consistently refer to Csf1r-expressing myeloid cells.

      Reviewer #2 (Public review):

      Zhou et al. previously reported similar TM heterogeneity, localization, and developmental characteristics. The manuscript should distinguish prior findings from new findings and address localization in relation to beta-selection.

      We have made this distinction explicit throughout. Zhou et al. established the two-population framework, their cortical versus medullary/cortico-medullary localization, their broad embryonic versus adult hematopoietic origins, and their age-associated remodeling. Building on that foundation, the present study contributes: (i) a prospective TIMD4/VCAM1 gating strategy linked to MafB and Csf1r reporters; (ii) intravascular-labeling evidence that both VCAM1+ populations are predominantly parenchymal; (iii) transcriptomic definition of efferocytic versus antigen-presentation/interferon programs; (iv) evidence that total TM abundance is maintained independently of SpiC; (v) selective CCR2 dependence of TIMD4- VCAM1+ macrophages and thymic monocytes; and (vi) functional evidence that the Csf1r-expressing thymic myeloid compartment supports DN3-to-DN4 progression. We use the anatomical localization established by Zhou et al. as the relevant biological context and reserve our new conclusions for the endpoints measured here. Accordingly, the FTOC phenotype is assigned to the Csf1r-expressing myeloid compartment rather than specifically to cortical TIMD4+ macrophages.

      The Introduction now clearly separates established knowledge from the questions addressed here, and the Results and Discussion explicitly identify the study-specific advances. Reference 27 has also been corrected to the final eLife publication.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Trajectory analysis

      The Abstract mentions trajectory analysis, but there is no corresponding Results section.

      We added a Results paragraph describing the Monocle3 analysis and integrated its interpretation with the fate-mapping and CCR2 experiments. Rooting the trajectory in Ly6c2+ Ccr2+ monocytes produces a transcriptional continuum toward macrophage states, supporting progressive acquisition of macrophage programs from a monocyte-like state. We also explain why pseudotime is complementary to, but not a substitute for, genetic lineage information: it models transcriptional relationships and therefore is not used to infer direct conversion between the two mature TM populations. This rationale allows the analysis to contribute meaningfully without asking it to resolve ontogeny on its own.

      The trajectory Methods were clarified and a new Results paragraph was added after the scRNA-seq specialization analysis; the Abstract now summarizes the result at the appropriate level.

      (2) Comparison with the Zhou et al. gating strategy

      Compare the TIMD4/VCAM1 gating strategy with the CD64/F4/80/TIMD4 definition used by Zhou et al., including the identity of TIMD4-VCAM1- cells.

      We now map the two VCAM1+ gates directly onto the established TM framework. TIMD4+ VCAM1+ cells correspond to the TIMD4+ cortical population, whereas TIMD4- VCAM1+ cells show the CX3CR1 enrichment expected of the medullary/cortico-medullary population. VCAM1 therefore adds a useful prospective discriminator within the CD64+ F4/80+ parent gate. The TIMD4- VCAM1- gate is MafB-low/negative, Csf1r-EGFP-high, and enriched for Ly6C, CCR2, and CX3CR1, supporting its designation as monocyte-enriched. We use “monocyte-enriched” because it accurately captures the dominant phenotype without implying that every event in the gate is developmentally identical.

      The first Results section and Figures 1-2 Legends now explain this correspondence and terminology.

      (3) Difference in IV-CD45 labeling

      Why are approximately 40% of Ly6C+CD11b+ cells labeled, but only approximately 10% of TIMD4-VCAM1- cells?

      The percentages arise from different denominators. Ly6C+ CD11b+ is a broad myeloid gate that contains both blood-exposed and parenchymal cells. TIMD4- VCAM1- is a narrower population defined within the CD64+ F4/80+ parent gate and therefore samples a different compartment. We now make that gating relationship explicit. The i.v.-labeling experiment consequently supports two positive conclusions: both VCAM1+ macrophage populations are predominantly parenchymal, and the broad Ly6C+ CD11b+ compartment contains a substantially larger blood-exposed component.

      The i.v.-labeling Results paragraph and Figure 2 Legend now describe the two gates and their distinct denominators.

      (4) SpiC requirement in individual TM populations

      Determine proportions and numbers of TM subpopulations in Spic-/- mice.

      The biological rationale for this suggestion is strong because Spic is enriched in TIMD4+ VCAM1+ macrophages. Figure 3C, however, quantifies the aggregate CD64+ F4/80+ TM compartment. We therefore revised the conclusion to the level directly supported by that experiment: total TM abundance is maintained in Spic-/- mice despite the expected loss of splenic red pulp macrophages. This is an informative distinction because it shows that the overall thymic macrophage compartment does not share the obligate SpiC dependence of red pulp macrophages, even though a subset-selective quantitative or functional effect remains a question for future work.

      The Figure 3 Results paragraph, Discussion, and legend now state preservation of total TM abundance rather than making a population-by-population claim.

      (5) Tomato/GFP double-positive cells

      Many cells appear to express both GFP and tdTomato. What does this mean?

      We expanded the explanation of the mTmG reporter. Flt3-Cre-mediated recombination initiates a switch from membrane Tomato to membrane GFP, but the pre-existing membrane Tomato protein need not disappear instantaneously. Tomato+ GFP+ events are therefore consistent with recent/incomplete reporter transition or persistence of stable Tomato protein after recombination. We do not treat them as a third ontogenetic lineage. The key comparative result is the distribution of reporter histories across populations: TIMD4+ VCAM1+ macrophages retain a large FLT3-independent fraction, whereas TIMD4- VCAM1+ macrophages and monocytes show substantially greater FLT3 history.

      The fate-mapping Results paragraph and Figure 6 Legend now define the reporter transition and the interpretation of double-positive events.

      (6) Direct AP20187 toxicity in DN thymocytes

      Test AP20187 toxicity in MaFIA and control DN thymocytes, for example by active caspase-3 staining.

      The proposed caspase-3 analysis is designed to determine whether a DN thymocyte fraction expresses sufficient MaFIA transgene to be directly affected. The existing controls establish two important features of the result: AP20187 does not reduce TM numbers in non-transgenic C57BL/6 FTOCs, demonstrating transgene dependence, and the developmental response is patterned rather than global, with preserved γδ T cell output, DN3 accumulation, DN4 loss, and reduced CD27 expression. These convergent observations support the conclusion that integrity of the Csf1r-expressing myeloid compartment is required for efficient DN3-to-DN4 progression. We now state that conclusion prominently and note the remaining question of thymocyte-intrinsic transgene activity once, in the focused Discussion paragraph.

      The MaFIA Results now emphasize the transgene-dependent control and convergent developmental measurements; the Discussion states the remaining cellular-resolution issue in one focused paragraph.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1 color annotation

      Clarify whether the color annotation in panel C is applied to panel F.

      Yes. Panel F evaluates MafB-mCherry and Csf1r-EGFP reporter expression within the same TIMD4/VCAM1-defined populations shown in panel C. We now state this explicitly so the reporter patterns can be interpreted as independent support for the identity of each prospectively defined gate. We also removed the previous reference to immunofluorescence localization because localization data are not displayed in Figure 1.

      Figure 1 Legend revised.

      (2) Figure 2 controls, thresholds, and population identities

      Provide staining controls and thresholds; reconsider categorical “lack” statements; clarify Ly6C+CD11b+ and CD64+F4/80+ populations and the apparent IV-CD45 discrepancies.

      The underlying interpretive point is whether marker expression is categorical or relative. We revised the text to describe Ly6C, CCR2, and CX3CR1 comparatively, which more faithfully represents continuous flow-cytometric measurements. This improves the biological conclusion: both VCAM1+ populations are Ly6C-low/negative and CCR2-low relative to the TIMD4- VCAM1- monocyte-enriched population, while CX3CR1 is enriched in TIMD4- VCAM1+ relative to TIMD4+ VCAM1+ macrophages. We also define the broad Ly6C+ CD11b+ comparison gate, the CD64+ F4/80+ macrophage parent gate, and the narrower TIMD4/VCAM1 subgates. With those denominators made explicit, the iv-CD45 measurements become internally consistent rather than apparently contradictory.

      The Figure 2 Results paragraph and legend now use relative marker language and explain the gate hierarchy.

      (3) Figure 3 colors, VCAM1, and SpiC populations

      Explain panel B colors and the apparent Vcam1 difference; provide individual TM-population data in Spic-/- mice.

      Panel B is a row-scaled expression heatmap: yellow and purple indicate relatively higher and lower scaled expression for each gene, respectively. These colors should not be read as absolute expression or compared directly with antibody fluorescence, because transcript abundance and cell-surface protein are regulated at different levels and have different dynamic ranges. The legend now makes this distinction explicit. Vcam1 transcript enrichment is nevertheless consistent with VCAM1 protein being a useful surface discriminator in our gating scheme. For SpiC, we now limit the conclusion to the measurement displayed in panel C, as preservation of total CD64+ F4/80+ TMs, while explaining the biological significance of their divergence from SpiC-dependent splenic red pulp macrophages.

      Figure 3 Results, Discussion, and Legend revised.

      (4) Figure 4 genes not shown

      H-2K, H-2D, and Runx3 are described but not shown in the figure.

      We clarified the division of information between Figure 4 and Table 1. Figure 4 displays pathway-level GO enrichment, which supports the higher-order conclusion that TIMD4- VCAM1+ macrophages are enriched for antigen-presentation and interferon-response programs. Individual genes contributing to the subset signatures, including H2-K1, H2-D1, and Runx3, are reported in Table 1. The revised wording no longer implies that those individual genes are plotted in Figure 4.

      Figure 4 Results paragraph and Legend revised.

      (5) Figure 6 double-positive reporter cells

      Explain Tomato+GFP+ cells and their relationship to single-positive cells.

      As described in our response to Reviewer 1, we now explain the kinetics of the mTmG reporter switch and interpret double-positive events as reporter-transition/persistence events rather than a separate lineage. This interpretation focuses the analysis on the biologically informative comparison— the different balance of FLT3-independent and FLT3-history labeling among TIMD4+ VCAM1+, TIMD4- VCAM1+, and monocyte-enriched populations.

      Fate-mapping Results and Figure 6 Legend revised.

      (6) CX3CR1 phenotype in Figures 6 and 7

      The text describes CX3CR1 phenotype without showing the data.

      We now describe Figures 6 and 7 using the markers actually displayed in those panels: TIMD4 and VCAM1. The relationship to CX3CR1 is established independently in Figure 2 and in the integrated scRNA-seq analysis. This separation makes the evidentiary chain clearer: Figure 2 links the surface-defined populations to CX3CR1 phenotype, Figure 6 compares their FLT3 reporter histories, and Figure 7 tests their CCR2 dependence.

      Corresponding Results language and Figures 6-7 Legends revised.

      (7) Figure 8 Csf1r-EGFP in thymocytes/DCs and “45.2”

      Show Csf1r-EGFP expression in thymocytes and DCs; determine whether their reduction is indirect or direct; clarify 45.2.

      The reviewer highlights why the MaFIA result should be interpreted at the Csf1r-expressingcompartment level. The reporter is demonstrably expressed by both TM populations and thymic monocytes, and AP20187 causes transgene-dependent TM depletion in FTOC. DC reduction may reflect direct transgene activity, dependence on the altered myeloid niche, or both; similarly, direct transgene expression was not measured in fetal DN thymocytes. We now make that cellular resolution explicit while emphasizing the developmental conclusion supported by the complete phenotype. We also clarify that CD45.2 denotes the congenic allele of the C57BL/6 control and not a distinct treatment or numerical value.

      MaFIA Results, Discussion, and Figure 8 Legend revised.

      (8) Figure 10 CD27 profiles

      Show CD27 flow-cytometric profiles.

      We revised the text and legend to make clear that Figure 10 reports quantified frequencies of CD27+ DN3 and DN4 cells. The CD27 measurement is interpreted in conjunction with, rather than in isolation from, the DN-stage data. Reduced CD27 among DN3 cells, reciprocal DN3 accumulation and DN4 loss, and reduced downstream DP output form a coherent sequence that localizes the developmental impairment to the b-selection-associated transition. We therefore use CD27 as a correlate of successful pre-TCR-associated progression rather than claiming direct biochemical measurement of pre-TCR signaling.

      Final Results paragraph, Discussion, and Figure 10 Legend revised.

      (9) Supplemental Figure 2 schemes

      Improve the schemes and explain dLNGFR and AP20187.

      The revised Figure Legend now explains each functional element. dLNGFR is the membrane-targeting low-affinity nerve growth factor receptor segment within the MaFIA fusion construct and is not used here as a lineage marker. AP20187 is a synthetic homodimerizer that binds the engineered FKBP domains, bringing the Fas intracellular domains together and initiating apoptosis in transgene-expressing cells. These definitions make the logic of the depletion system understandable without requiring familiarity with the original MaFIA construct.

      Supplemental Figure 2 Legend revised.

      Closing statement

      The revised manuscript now presents the study in clear, evidence-matched terms. It identifies the advances in TM phenotypic organization, transcriptional specialization, developmental contribution, SpiC independence of total TM abundance, and CCR2 dependence of the TIMD4- VCAM1+ population. It also emphasizes the convergent evidence that Csf1r-expressing myeloid cells support progression through the DN3-to-DN4 checkpoint, while accurately defining the cellular resolution of the MaFIA experiment. We thank the editors and reviewers for helping us sharpen both the significance and the precision of these conclusions.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      AIRE has been well known to contribute to immune self-tolerance in the thymus by expressing auto-antigens; in this manuscript, the authors describe unexpected findings about the interaction of AIRE with AID in B cells, and its function in the immune system, thereby contributing to a fundamental understanding of the broader functions of AIRE. The strength of this manuscript is that, by employing biochemical and genetic experiments, the authors convincingly show interaction between AIRE and AID and subsequent AIRE's function in the GC responses. However, two weak points exist: first, the connection between AIRE, auto-anti IL17 Abs, and IL17-positive effector T cells, and second, like the thymus, expression of auto-antigens by AIRE in the GC B cells has not been tested.

      We thank the journal editors for the thoughtful and constructive evaluation of our data and for recognizing the significance, strengths and weaknesses of our study. We have addressed these major weak points in the responses below and have also provided discussion on the important aspect of whether AIRE regulates autoantigen expression in GC B cells analogous to its role in mTECs, which certainly warrants future investigation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide in vivo and in vitro evidence for an interaction between AIRE and AID. This has implications for the dynamics of the germinal center response and autoimmunity related to the APSI disease.

      The manuscript describes an unexpected function of AIRE, which is more well known for its function to regulate negative selection of T cells in the thymus. Here, the gene has also been shown to be expressed by B cells (Immunity 2015: 26070482). They describe that AIRE interacts with AID, and in its absence, B cells acquire more hypermutations and also produce autoantibodies against IL-17. These autoantibodies have been described previously.

      Strengths:

      The study is interesting and provides some additional information about how AIRE regulates immune cell function. Several biochemical and in vivo experiments show the interaction and the function of AIREs in the regulation of AID activity in the GC response.

      Weaknesses:

      Some of the hypothetical consequences of this regulation are not investigated. This includes responses to model antigens and dynamics of the germinal center related to kinetics.

      We are grateful for the reviewer’s interest and careful and thoughtful review of the manuscript.

      Major Comments:

      (1) AID regulates both switch and somatic hypermutation. Switch is easier to achieve, so which of these processes does AIRE influence the most? Also, the switch is thought to occur before the B cell enters the GC. Looking at the histology, is AIRE also expressed at the early proliferative stage that has been described by Ann Haberman?

      We thank the reviewer for raising this provocative point. Yes, it has been recently shown that CSR occurs earlier after B cell activation largely prior to GC entry (Roco et al. 2019), and CSR is generally easier to induce in vitro than SHM. These data and observations further support well-established findings that CSR and SHM involve distinct mechanisms and pathways to resolve DNA lesions (Frossi et al., 2019; Masani et al., 2013; Roco et al., 2019; Schrader et al., 2023), even though both require AID. We observed a strong upregulation of AIRE expression upon CD40 signaling (Fig. 2C–G) and that AIRE in B cells negatively regulates both CSR and SHM (Fig. 3E– J and Fig. 4A–H), but it would be difficult to directly compare the magnitude of AIRE’s influence between them because they are mechanistically distinct and occurs at different stages and locations during antigen-specific B cell responses.

      We observed some AIRE-expressing B cells outside of the GC (Fig. 1A–C, J), some of which could represent newly activated cells about to enter the GC reaction, but we did not perform imaging or flow cytometry experiments or track their Ki67 or Bcl6 expression in the Aire<sup>Adig</sup> reporter mice. However, these cells appear to be largely IgD<sup>-</sup>, suggesting that they may have already undergone CSR and are perhaps not (entirely) the early proliferative B cells described by Dr Haberman.

      (2) In experiments determining anti-CD40-dependent upregulation of AIRE, naïve resting B cells were used from mice. A proportion of the B-cells got activated. Are these MZB or FOB cells as MZBs are more easily activated?

      We thank the reviewer for this careful interpretation of our data. Indeed, the methods used to purify naïve B cells do not exclude MZBs and, as they are more sensitive to activation, may also express AIRE upon stimulation with CD40L. However, although we cannot exclude MZBs in these cultures, they represent less than 10% of the total B cells in these cultures whereas we observed approximately 30% of the cells upregulating AIRE expression.

      Interestingly, Yamano et al. (Yamano et al., 2015) suggested that GC B cells may not express AIRE due to strong BCR signaling; however, it is known that GC B cells have attenuated BCR signaling to promote LZ to DZ transition, consistent with the timing of AIRE upregulation postCD40-CD40L engagement (Davidzohn et al., 2020; Khalil et al., 2012). In contrast, MZBs are known to exhibit greater baseline activation of pathways downstream of BCRs (Hampel et al., 2011), which may provide inhibitory signals for the upregulation of AIRE.

      (3) In the BM chimeric experiments in Figure 3. Do the AIRE+ and AIRE - populations distribute equally among B cell subpopulations?

      We thank the reviewer for this interesting comment. In our chimera experiment, although we did not analyze specific subsets such as B-1 or MZB, we observe that AIRE-deficient B cells outcompeted WT B cells in lymphoid organs as well as in the blood in spite of prior publications showing equivalent reconstitution of CD45.2 mice with CD45.1/CD45.2 B cells (Kalari Kandy et al., 2023).

      (4) Furthermore, in the NP-KLH experiments, one would expect that B cells with increased affinity would leave the GC earlier and become plasma cells. Thus, the kinetics of the AIRE+ vs AIRE- B cells within the GC would be different? Also, would they maybe take over at some point, as the increased affinity would favor help from Tfh cells that are known to be limited?

      We thank the reviewer for these insightful comments. Yes, we would expect to see that Aire<sup>-/-</sup> B cells will dominate GCs over time. Indeed, our chimera experiments (Fig. 3A–D) indicated that Aire<sup>-/-</sup> B cells were of higher frequency in the GCs compared to WT post-immunization, consistent with the expansion of higher affinity clones. However, although we did not quantify the development of PCs in these mice, our adoptive transfer experiments showed that mice receiving Aire<sup>-/-</sup> B cells developed higher affinity antibodies compared to those receiving WT B cells (Fig. 3G), indicating a higher affinity plasma cell pool compared to controls.

      (5) Given the previous studies on AIRE's function in regulating transcription (PMID: 34518235), how does this interaction fit into this picture?

      We thank the reviewer for raising this important point. Although we did not directly test the role of AIRE in the regulation of tissue-restricted antigens, we observed a clear interaction between AIRE and pSer5 Pol II (Fig. 6F), consistent with its function as a broad-spectrum transcriptional regulator (Fang et al., 2024; Giraud et al., 2012; Oven et al., 2007). This is particularly important for antibody diversification in B cells, as it is known that AID is targeted to sites of Pol II pausing (Chaudhuri et al., 2003; Pavri et al., 2010). Further, recent reports have shown that the CARD domain of AIRE promotes its polymerization and the formation of nucleation sites at which a positive feedback loop to create transcriptional hubs (Huoh et al., 2024). Interestingly, the CARD domain is also necessary for the interaction between AIRE and AID (Fig. 5F), indicating that either these condensates are also critical for preventing AID from being recruited to Pol II or that AIRE may perform alternative functions in B cells compared to mTECs. Nevertheless, in this current manuscript, we focus on the capacity of AIRE to utilize this interaction with pSer5 Pol II to prevent AID localization to its DNA substrates, however, future studies may focus on how this interaction may impact the expression of peripheral tissue antigens to promote T cell tolerance.

      (6) In the uracil experiments, the readout for AID to induce double-stranded breaks could be tested.

      We thank the reviewer for this suggestion for complementary data to our uracil analyses. In our manuscript, we tested the generation of uracil in Aire<sup>+/+</sup> and Aire<sup>-/-</sup> CH12 cells since this is a downstream function of AID’s activity. Therefore, our data focused on an immediate and direct impact of AIRE on AID’s activity. Further, although we observe an interaction between AIRE and AID as well as a downstream functional consequence, it is unclear whether AIRE may also impact other pathways in these cells after activation that may confound the results of downstream analyses, such as DSBs. We agree that DSBs would be an interesting readout further downstream, and future work may focus on the function of AIRE in these additional processes.

      (7) The candida experiments are a nice connection to the situation in patients. However, why is it mostly auto-antibodies against IL-17? How about other immune responses, as well as T cellindependent type I and II responses?

      We thank the reviewer for raising this important point. Indeed, we observed a significant increase in the generation of autoreactive, neutralizing antibodies against Th17-associated cytokines (Fig. 7D, E). However, although AIRE-deficient patients produce neutralizing antibodies against type 1 interferons as well, clearance of the fungal pathogen Candida albicans relies heavily on IL-17 to promote the upregulation of antimicrobial peptides and neutrophil infiltration (Conti et al., 2014), which may be a particularly pronounced response induced in our mouse models. Therefore, we targeted our assay towards these cytokines, but do not rule out possible production of autoantibodies against other factors. It would be interesting for additional studies to determine the production of autoreactive antibodies against cytokines in other infection models, such as viral infections, as APS-1 patients have been reported to display an increased susceptibility to viral infections as well (Bastard et al., 2021; Hetemaki et al., 2021; Oikonomou et al., 2021).

      These important discussions have been included in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this study, Zhou et al investigated the expression and function of AIRE in B cells in peripheral lymphoid tissues. First, they found the expression of AIRE protein in mature B cells in the follicles in human tonsils and spleens from healthy donors. Flow cytometry analyses using human samples as well as Aire-reporter mice demonstrated AIRE expression in germinal center B cells. The expression of Aire in B cells was induced by CD40 signals. Then, to investigate the impact of AIRE deficiency on B-cell function, the authors used a method of transplanting bone marrow cells from Aire-KO and WT mice into B-cell-deficient mice, comparing B-cell development and function reconstituted in the recipient mice. Their results showed that Aire-deficient B cells strongly responded to immunization with antigens, exhibiting enhanced class switching and somatic hypermutation of antibodies compared with WT B cells. The same phenomena were observed in CRISPRed B cell lines lacking Aire. The authors successfully utilized the Aire-deficient B cell line to demonstrate that Aire suppresses antibody class switching and somatic hypermutation via its interaction with AID. Finally, using B cell transfer into B cell-deficient mice demonstrated that mice harboring Aire-deficient B cells produced high levels of autoantibodies against Th17 cytokines and exhibited reduced resistance to Candida infection. This mirrors characteristic symptoms in AIRE-deficient patients. The findings of this study not only reveal an unexpected function of AIRE in B cells but also have the potential to contribute to understanding the pathogenesis of APECED and to offering a new direction for developing therapies.

      We are grateful for the reviewer’s careful and thoughtful review of the manuscript and appreciation of the significance of our findings.

      Strengths:

      The strength of this study lies in demonstrating the expression of the function of AIRE in B cells in both mice and humans. It also revealed the direct interaction between AIRE and AID, along with its binding mode (requiring CARD and NLS domains of AIRE), and showed that this interaction is crucial for AIRE function in B cells. It is also significant that the study demonstrated how B-cell-intrinsic dysfunction of AIRE leads to autoantibody production against cytokines.

      Weaknesses:

      As for loss-of-function analysis of Aire in B cells, in addition to the B cell transfer from Aire-KO mice performed in this study, generating B cell-specific Aire-deficient mice using Aire-flox mice (Dobes et al, Eur J Immunol 2018) would further reinforce the conclusions of this study. Furthermore, the relationship with Aire function in thymic B cells reported by previous studies remains unclear, posing an unresolved challenge. This study also failed to address whether Aire deficiency affects gene expression in GC B cells, in particular, whether it induces the expression of various self-antigens as reported in thymic B cells or mTECs.

      We thank the reviewer for these thoughtful critiques of our manuscript. As this study was largely carried out prior to the development of Aire-flox mice, we took advantage of the adoptive transfer model to generate mice with AIRE deletion specifically in B cells. Since Aire-flox mice were developed, these mice would be the new gold standard for analysis and would provide further strength to our existing data. Unfortunately, these mice are not readily available, and we do not currently have the resources to reconstitute this line and are unable to obtain them.

      In addition, although we do not directly test the role of AIRE in the regulation of tissue-restricted antigens, we observed a clear interaction between AIRE and pSer5 Pol II (Fig. 6F), consistent with its function as a transcriptional regulator (Fang et al., 2024; Giraud et al., 2012; Oven et al., 2007). This is particularly important for antibody diversification in B cells, as it is known that AID is targeted to sites of Pol II pausing (Chaudhuri et al., 2003; Pavri et al., 2010). Further, recent reports have shown that the CARD domain of AIRE promotes its polymerization and the formation of nucleation sites at which a positive feedback loop creates transcriptional hubs (Huoh et al., 2024). Interestingly, the CARD domain is also necessary for the interaction between AIRE and AID (Fig. 5F) indicating that either these condensates are also critical for preventing AID from being recruited to Pol II or that AIRE may perform alternative functions in this subset of B cells compared to mTECs. Nevertheless, in this study, we focused on the capacity of AIRE to utilize this interaction with pSer5 Pol II to reduce AID targeting to its DNA substrates. However, future studies may focus on how this interaction may impact the expression of peripheral tissue antigens to promote T cell tolerance.

      The discussion related to these important aspects raised by the reviewer are included in the original and revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) How do these findings connect to IgA responses and the ongoing GC response in the peyers patches?

      We thank the reviewer for raising this additional point. Indeed, we also observed AIRE expression in GC B cells from Peyer’s patches (Fig. S1J) and found that CSR to IgA in our in vitro models was increased by the deletion of AIRE in B cells (Fig. 4A–C). These data suggest that, in AIRE-deficient B cells, there may be an increase in IgA production, however, we did not directly test mucosal IgA levels in our adoptive transfer models. Future work may focus on the role of B cell intrinsic AIRE and its impact on the capacity of Peyer’s patch-associated B cells to produce IgA and its role on mucosal barrier immunity and microbiota coating.

      These important discussions have been included in the revised manuscript.

      (2) Some of the histology is rather dark, and it would be nice to see some more examples in the supplement.

      We thank the reviewer for bringing this to our attention. We have included an additional staining of a whole follicle with GC in the supplement as requested (Fig. S1C, bottom panel).

      Reviewer #2 (Recommendations for the authors):

      (1) In the Discussion section, could the authors explore Aire's functions and biological significance in greater depth? Although Aire was originally identified as a regulator of tissue-restricted antigen expression in mTECs, it has subsequently shown to play crucial roles in diverse processes, including germ cell development, antigen processing, and B cell function regulation (this study). In light of these recent findings, it is conceivable that Aire did not evolve specifically as an mTECassociated gene. Rather, the reverse may be possible: Aire may have originally emerged as a regulator of germ cell development and/or immune responses, with its expression and functions later co-opted by mTECs (analogous to how other transcription factors have been adopted for differentiation of mimetic mTECs). I would welcome the authors expanding on this evolutionary perspective in future work or revisions.

      We thank the reviewer for this insightful and fascinating perspective on the evolutionary and ancestral functions of AIRE as a transcription factor. We have incorporated additional discussion in the revised manuscript.

      (2) Personally, I knew that Aire is expressed in B cells, as indicated by several gene expression databases. But the ImmGen database shows that Aire mRNA expression is detected at higher levels in plasma cells than in GC B cells. What do the authors think?

      We sincerely thank the reviewer for this question and appreciate the personal comment in support of our results. Indeed, the ImmGen database shows little Aire expression in GC B cells, with most of the signal coming from plasma cells in the B cell compartment. Further, as the Reviewer suggested, reanalysis of previously published scRNA-seq (Duan et al., 2021) has shown Aire transcript in some, albeit a lower percentage of, GC B cells (Author response image 1), which was more abundant near the peak time of the immune response after immunization, in line with our and the Reviewer’s observations.

      Author response image 1.

      Analysis of Aire expression in the day 7 (upper row) and day 14 (lower row) post-immunization scRNA-Seq datasets (Duan et al., 2021), showing 200 Aire<sup>+</sup> GC B cells out of 3542 GC B cells (5.65%) on day 7 and 183 Aire<sup>+</sup> GC B cells out of 19757 GC B cells (0.93%) on day 14, and the distribution of these Aire<sup>+</sup> GC B cells in both dark zone (DZ) and light zone (LZ) B cells.

      A myriad of technical factors can impact the detection of lower-expressed genes by bulk RNA-seq and even scRNA-seq, and that looking at individual cells may give better resolution. This is the case for a considerable number of other genes we have studied or that have been reported in other investigators’ publications. Many biological factors can also have an important effect (e.g. protein levels may vary depending on cell states and do not necessarily correlate with transcript levels). Significant impact of the strain background, diet, microbiome and the living environment of the animals cannot be excluded either. Similarly, recent reports have shown that CGRP from spleen-innervating nociceptors signals to splenic B cells to promote germinal center reactions (Wu et al., 2024). However, Immgen showed no Ramp1 expression in splenic B cell compartments. In addition, Nur77/Nr4a1 has also been shown to be highly expressed by a small number of light zone cells and is critical for regulating GC reactions (Brooks et al., 2021; Mueller et al., 2015) even though ImmGen shows very little expression of this gene in its bulk sequencing. Therefore, we think that much of the data from ImmGen should be carefully validated with additional experimental methods. In this case, we were able to detect AIRE protein using methods including immunofluorescence, immunoprecipitation, western blot, and flow cytometry.

      (3) Line 343: "donor Aire-/-" may be a mistake for "donor Aire+/+"?

      We thank the reviewer for bringing this to our attention and have adjusted the manuscript accordingly.

      (4) Line 346: Is "data now shown" a spelling mistake? Including this, there are multiple occurrences of "data not shown" in the Results and Discussion sections. If the authors describe or discuss the data, it should be shown.

      We thank the reviewer for bringing this to our attention and have revised the manuscript accordingly.

      (5) Line 346-350: The authors should investigate more thoroughly whether Aire influences gene expression in GC B cells, particularly the expression of a set of self-antigen genes, as demonstrated in thymic B cells by Yamano et al (Immunity 2015). Regardless of the outcome, such data would be highly significant for this study and the broader research community.

      We thank the reviewer for raising this important point. Although we did not directly test the role of AIRE in the regulation of tissue-restricted antigens, we observed a clear interaction between AIRE and pSer5 Pol II (Fig. 6F), consistent with its function as a transcriptional regulator (Fang et al., 2024; Giraud et al., 2012; Oven et al., 2007). This is particularly important for antibody diversification in B cells, as it is known that AID is targeted to sites of Pol II pausing (Chaudhuri et al., 2003; Pavri et al., 2010). Further, recent reports have shown that the CARD domain of AIRE promotes its polymerization and the formation of nucleation sites at which a positive feedback loop to create transcriptional hubs (Huoh et al., 2024). Interestingly, the CARD domain is also necessary for the interaction between AIRE and AID (Fig. 5F) indicating that either these condensates are also critical for preventing AID from being recruited to Pol II or that AIRE may perform alternative functions in B cells compared to mTECs. Nevertheless, in this current study, we focused on the capacity of AIRE to utilize this interaction with pSer5 Pol II to prevent AID targeting to its DNA substrates, and future studies may focus on how this interaction may impact the expression of peripheral tissue antigens to promote T cell tolerance.

      (6) Line 1282-1293: The number of B cells used for Western blotting should be clearly indicated. Since the expression levels of Aire mRNA in B cells are lower than those in mTECs, detecting Aire protein in B cells is also likely difficult. It is important to specify the number of B cells required so that other researchers can reproduce the Western blotting results presented in this paper.

      We thank the reviewer for raising this important point and have revised the manuscript accordingly.

      References

      Bastard, P., E. Orlova, L. Sozaeva, R. Levy, A. James, M.M. Schmitt, S. Ochoa, M. Kareva, Y. Rodina, A. Gervais, T. Le Voyer, J. Rosain, Q. Philippot, A.L. Neehus, E. Shaw, M. Migaud, L. Bizien, O. Ekwall, S. Berg, G. Beccuti, L. Ghizzoni, G. Thiriez, A. Pavot, C. Goujard, M.L. Fremond, E. Carter, A. Rothenbuhler, A. Linglart, B. Mignot, A. Comte, N. Cheikh, O. Hermine, L. Breivik, E.S. Husebye, S. Humbert, P. Rohrlich, A. Coaquette, F. Vuoto, K. Faure, N. Mahlaoui, P. Kotnik, T. Battelino, K. Trebusak Podkrajsek, K. Kisand, E.M.N. Ferre, T. DiMaggio, L.B. Rosen, P.D. Burbelo, M. McIntyre, N.Y. Kann, A. Shcherbina, M. Pavlova, A. Kolodkina, S.M. Holland, S.Y. Zhang, Y.J. Crow, L.D. Notarangelo, H.C. Su, L. Abel, M.S. Anderson, E. Jouanguy, B. Neven, A. Puel, J.L. Casanova, and M.S. Lionakis. 2021. Preexisting autoantibodies to type I IFNs underlie critical COVID-19 pneumonia in patients with APS-1. The Journal of experimental medicine 218:

      Brooks, J.F., C. Tan, J.L. Mueller, K. Hibiya, R. Hiwa, V. Vykunta, and J. Zikherman. 2021. Negative feedback by NUR77/Nr4a1 restrains B cell clonal dominance during early Tdependent immune responses. Cell Rep 36:109645.

      Chaudhuri, J., M. Tian, C. Khuong, K. Chua, E. Pinaud, and F.W. Alt. 2003. Transcription-targeted DNA deamination by the AID antibody diversification enzyme. Nature 422:726-730.

      Conti, H.R., A.R. Huppler, N. Whibley, and S.L. Gaffen. 2014. Animal models for candidiasis. Curr Protoc Immunol 105:19.16.11-19.16.13.

      Davidzohn, N., A. Biram, L. Stoler-Barak, A. Grenov, B. Dassa, and Z. Shulman. 2020. Syk degradation restrains plasma cell formation and promotes zonal transitions in germinal centers. The Journal of experimental medicine 217:

      Duan, L., D. Liu, H. Chen, M.A. Mintz, M.Y. Chou, D.I. Kotov, Y. Xu, J. An, B.J. Laidlaw, and J.G. Cyster. 2021. Follicular dendritic cells restrict interleukin-4 availability in germinal centers and foster memory B cell generation. Immunity 54:2256-2272 e2256.

      Fang, Y., K. Bansal, S. Mostafavi, C. Benoist, and D. Mathis. 2024. AIRE relies on Z-DNA to flag gene targets for thymic T cell tolerization. Nature 628:400-407.

      Frossi, B., G. Antoniali, K. Yu, N. Akhtar, M.H. Kaplan, M.R. Kelley, G. Tell, and C.E.M. Pucillo. 2019. Endonuclease and redox activities of human apurinic/apyrimidinic endonuclease 1 have distinctive and essential functions in IgA class switch recombination. J Biol Chem 294:5198-5207.

      Giraud, M., H. Yoshida, J. Abramson, P.B. Rahl, R.A. Young, D. Mathis, and C. Benoist. 2012. Aire unleashes stalled RNA polymerase to induce ectopic gene expression in thymic epithelial cells. Proceedings of the National Academy of Sciences of the United States of America 109:535-540.

      Hampel, F., S. Ehrenberg, C. Hojer, A. Draeseke, G. Marschall-Schroter, R. Kuhn, B. Mack, O. Gires, C.J. Vahl, M. Schmidt-Supprian, L.J. Strobl, and U. Zimber-Strobl. 2011. CD19independent instruction of murine marginal zone B-cell development by constitutive Notch2 signaling. Blood 118:6321-6331.

      Hetemaki, I., S. Laakso, H. Valimaa, I. Kleino, E. Kekalainen, O. Makitie, and T.P. Arstila. 2021. Patients with autoimmune polyendocrine syndrome type 1 have an increased susceptibility to severe herpesvirus infections. Clin Immunol 231:108851.

      Huoh, Y.S., Q. Zhang, R. Torner, S.C. Baca, H. Arthanari, and S. Hur. 2024. Mechanism for controlled assembly of transcriptional condensates by Aire. Nature immunology 25:15801592.

      Kalari Kandy, R.R., X. Fan, and X. Cao. 2023. CD45.1/CD45.2 Congenic Markers Induce a Selective Bias for CD8+ T Cells during Adoptive Lymphocyte Reconstitution in Lymphocytopenia Mice. Immunohorizons 7:755-759.

      Khalil, A.M., J.C. Cambier, and M.J. Shlomchik. 2012. B cell receptor signal transduction in the GC is short-circuited by high phosphatase activity. Science 336:1178-1181.

      Masani, S., L. Han, and K. Yu. 2013. Apurinic/apyrimidinic endonuclease 1 is the essential nuclease during immunoglobulin class switch recombination. Mol Cell Biol 33:1468-1473.

      Mueller, J., M. Matloubian, and J. Zikherman. 2015. Cutting edge: An in vivo reporter reveals active B cell receptor signaling in the germinal center. Journal of immunology 194:29932997.

      Oikonomou, V., T.J. Break, S.L. Gaffen, N.M. Moutsopoulos, and M.S. Lionakis. 2021. Infections in the monogenic autoimmune syndrome APECED. Curr Opin Immunol 72:286-297.

      Oven, I., N. Brdickova, J. Kohoutek, T. Vaupotic, M. Narat, and B.M. Peterlin. 2007. AIRE recruits P-TEFb for transcriptional elongation of target genes in medullary thymic epithelial cells. Mol Cell Biol 27:8815-8823.

      Pavri, R., A. Gazumyan, M. Jankovic, M. Di Virgilio, I. Klein, C. Ansarah-Sobrinho, W. Resch, A. Yamane, B. Reina San-Martin, V. Barreto, T.J. Nieland, D.E. Root, R. Casellas, and M.C. Nussenzweig. 2010. Activation-induced cytidine deaminase targets DNA at sites of RNA polymerase II stalling by interaction with Spt5. Cell 143:122-133.

      Roco, J.A., L. Mesin, S.C. Binder, C. Nefzger, P. Gonzalez-Figueroa, P.F. Canete, J. Ellyard, Q. Shen, P.A. Robert, J. Cappello, H. Vohra, Y. Zhang, C.R. Nowosad, A. Schiepers, L.M. Corcoran, K.M. Toellner, J.M. Polo, M. Meyer-Hermann, G.D. Victora, and C.G. Vinuesa. 2019. Class-Switch Recombination Occurs Infrequently in Germinal Centers. Immunity 51:337-350 e337.

      Schrader, C.E., T. Williams, K. Pechhold, E.K. Linehan, D. Tsuchimoto, and Y. Nakabeppu. 2023. APE2 Promotes AID-Dependent Somatic Hypermutation in Primary B Cell Cultures That Is Suppressed by APE1. Journal of immunology 210:1804-1814.

      Wu, M., G. Song, J. Li, Z. Song, B. Zhao, L. Liang, W. Li, H. Hu, H. Tu, S. Li, P. Li, B. Zhang, W. Wang, Y. Zhang, W. Zhang, W. Zheng, J. Wang, Y. Wen, K. Wang, A. Li, T. Zhou, Y. Zhang, and H. Li. 2024. Innervation of nociceptor neurons in the spleen promotes germinal center responses and humoral immunity. Cell 187:2935-2951 e2919.

      Yamano, T., J. Nedjic, M. Hinterberger, M. Steinert, S. Koser, S. Pinto, N. Gerdes, E. Lutgens, N. Ishimaru, M. Busslinger, B. Brors, B. Kyewski, and L. Klein. 2015. Thymic B Cells Are Licensed to Present Self Antigens for Central T Cell Tolerance Induction. Immunity 42:1048-1061.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      eLife Assessment

      This study addresses an important question in liver biology: how zonal hepatocytes balance survival and proliferation following injury? The authors propose that a midzone Atf4-Chop axis to Btg2 program temporarily suppresses proliferation to promote survival after a variety of chemical and surgical liver injury models. The authors provide evidence that some zones mount tailored stress responses, which ultimately promote regeneration and liver healing; however, the "mid-zone" changes with different injury models, making it difficult to conclude that the ATF4-CHOP response is specific to this zone in all injury contexts. In addition, it is possible that Atf4 and Btg2 overexpression could lead to Cyp2e1 suppression, which could reduce the extent of injury after CCl4 or APAP. To some extent, these points make the strength of the evidence incomplete, but do not entirely detract from the significance of the study, which is underscored by the helpful observation that there are zone-specific stress responses that mediate liver regeneration and survival.

      We appreciate the editor’s thoughtful assessment. Regarding regional specificity, we agree that "mid-zone" patterns differ across injury models, particularly in HPx; to maintain consistency, we have removed the PHx data and now focus our conclusions on the APAP and CCl<sub>4</sub> models. Concerning the potential confound of Cyp2e1 suppression, our AAV overexpression resulted in only modest reductions in Cyp2e1 protein (Revised Figure 4-figure supplement 3A). Although enzymatic activity was not directly measured, basal CYP2E1 protein levels generally correlate with enzymatic activity. Furthermore, published studies indicate that protection from APAP toxicity requires >50% CYP2E1 suppression (PMID: 35145060; PMID: 30151903), whereas the reduction observed here was substantially smaller and therefore unlikely to account for the marked decrease in liver injury.

      Reviewer #2 (Public review):

      The manuscript reports protection of midlobular hepatocytes from APAP toxicity by activation of Atf4-CHOP (Ddit3)-mediated cell cycle arrest and stress response. The authors acknowledge that their finding is unexpected because CHOP typically induces cell death. Therefore, they functionally validate several aspects of the proposed Atf4-CHOP mechanism. Along these lines, the mitigation of APAP toxicity by AAV expression of Atf4 or Btg2, the latter identified as CHOP effector, is impressive. Whether Atf4 indeed acts through CHOP and whether midlobular hepatocytes are protected because of cell cycle arrest is less clear. These and other criticisms are described in the following.

      Major points:

      (1) The difference in baseline Cyp2e1 expression between F2A and F3B remains unexplained after revision.

      We apologize for the confusion. The apparent discrepancy reflects different visualization methods rather than a biological inconsistency. Figure 2A and Figure 3B are derived from the same spatial transcriptomic dataset but are visualized using different normalization strategies. Figure 2A uses row-wise Z-score normalization to emphasize the relative spatial distribution of each gene across liver zones. Consequently, even a moderate spatial gradient (e.g., ~1.5–2-fold) is visually enhanced to facilitate comparison of zonation patterns between genes. In contrast, Figure 3B presents the absolute log-normalized transcript abundance without rowwise scaling. Because Cyp2e1 is highly expressed in both the pericentral and adjacent mid-zone hepatocytes, the absolute difference between these regions is relatively modest, resulting in a less pronounced visual contrast. Thus, the different appearance of Cyp2e1 between Figures 2A and 3B reflects the visualization strategy rather than conflicting expression data. To avoid further confusion, we have added this clarification to the legends of both Figures 2A and 3B.

      (2) In contrast to the revised discussion, the abstract does not reflect that limited evidence for a cell cycle arrest in pericentral hepatocytes was found.

      We thank the reviewer for the continued effort in improving our manuscript. To reflect the limited support of ST data for cell cycle arrest, we have revised the Abstract (Revised manuscript, page 1, line 13-26)

      (3) Additional Btg2 knockout data support its proposed role in the revised manuscript.

      We appreciate the reviewer’s acknowledgement of the new functional data supporting the role of Btg2.

      (4) The BTG2 immunostaining remains weak, not only in in F6F but now also in F6D of the revised manuscript, which together with lack of high-resolution immunostaining of AAV-Ddit3-induced BTG2 in the absence of APAP results in limited support for the conclusion that APAP promotes nuclear localization of BTG2.

      We agree that the current immunostaining does not provide definitive evidence for nuclear localization. In response, we have deliberately avoided any claims regarding nuclear translocation in the revised manuscript, instead framing our findings strictly around increased BTG2 expression and accumulation in the peri-necrotic mid-zone region following APAP injury.

      (5) The extended list of transcription factors (from 30 to 50) includes Atf4 but direct evidence for an interaction with Ddit3 is missing from the revised manuscript.

      We thank the reviewer for this important point. We agree that our current dataset does not provide direct evidence for a physical interaction between Atf4 and Ddit3 proteins. Throughout the manuscript, we use the term "Atf4–Chop axis" to denote a functional regulatory pathway rather than a direct protein–protein interaction. To avoid ambiguity, we have revised the manuscript to characterize Atf4 and Ddit3 as co-activated transcriptional co-regulators and to restrict our conclusions to transcriptional co-regulation and pathway convergence, as supported by SCENIC, Cut&Run, and GO analyses. We now explicitly define "axis" as a functional pathway (revised manuscript, page 5, line 182–184; page 6, line 260–261; page 8, line 331333).

      (6) The ATF4 immunostaining after APAP challenge remains weak.

      We agree that the endogenous ATF4 immunostaining is relatively weak. This is consistent with the low basal abundance and transient induction of ATF4 protein during the integrated stress response. To improve visualization, we have added higher-magnification insets to Revised Figure 5A, which more clearly show nuclear ATF4 staining in hepatocytes adjacent to the necrotic region following APAP treatment. We hope these enlarged images better illustrate the spatial distribution of ATF4 while accurately reflecting its endogenous expression level.

      (7) S5A of the revised manuscript rules out loss of Cyp2e1 expression as a confounding factor.

      We thank the reviewer for acknowledging that Revised Figure S5A addresses this concern.

      (8) The revised manuscript continues to focus on rare spatial transcriptomics analyses of patients with APAP toxicity although more snRNA-seq analyses of such patients are available which should also allow for analysis of hepatocyte zonation.

      We thank the reviewer for this constructive suggestion and apologize for not describing our analyses more clearly in the previous revision. In response, we analyzed both the snRNA-seq and spatial transcriptomics (ST) datasets from GSE223561, which includes snRNA-seq profiles from 9 healthy and 10 APAP explants together with ST data from 3 healthy and 2 APAP patients (Revised Figure 4-figure supplement 3).

      The snRNA-seq data readily resolved hepatocyte zonation using established pericentral (Glul, Cyp2e1, Cyp2a5), periportal (Alb, Cyp2f2, Sds), and mid-zonal (Igfbp2, Hamp) markers (Figure 1A–D). Consistent with our mouse data, mid-zonal hepatocytes in APAP patients showed increased expression of stress-response genes (DDIT3, ATF4, and HMOX1). However, unlike the early mouse injury model, proliferation-associated genes (MKI67 and CCNB1) were also elevated (Revised Figure 4-figure supplement 3E), likely reflecting the end-stage nature of liver explants, in which regenerative responses are already well established rather than the early (3–12 h) injury phase examined in our study.

      The ST dataset provided complementary spatial information that cannot be obtained from dissociated nuclei alone. Although limited in sample number, one APAP specimen closely recapitulated our murine observations, showing mid-zonal stress enrichment accompanied by reduced proliferation, whereas the second exhibited a predominantly proliferative profile consistent with a later regenerative stage (Revised Figure 4-figure supplement 3F).

      Together, these analyses indicate that human mid-zonal hepatocytes exhibit a conserved stress-response program, while the relationship between stress signaling and proliferation is highly dependent on the stage of injury. We have incorporated both the snRNA-seq and ST analyses into the revised manuscript and explicitly discuss the temporal limitations of the available human datasets (page 5-6, lines 225239; page 9, lines 368-379).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) F2A includes Ddit3 in the revised manuscript and although it is not significant in S1G (former F2B), it is among the most highly expressed transcription factors in F4B and S3B.

      We thank the reviewer for this careful observation. As noted, Ddit3 is differentially expressed in the mid-zonal hepatocytes at both 3 and 6 h after APAP treatment (Revised Figure 2A). We also agree that Ddit3 does not reach statistical significance in the differential expression analysis shown in Revised Figure 2-figure supplement 1G under our predefined thresholds. However, our rationale for prioritizing Ddit3 was based primarily on transcription factor activity rather than differential expression alone. Differential expression analysis evaluates changes in the abundance of an individual transcript, whereas SCENIC regulon analysis infers the functional activity of a transcription factor from the coordinated expression of its downstream target genes. Consequently, transcription factors with relatively modest changes in mRNA abundance may nevertheless exhibit high regulon activity if their downstream regulatory network is broadly activated.

      Consistent with this, Ddit3 was identified as the highest-activity regulon in the mid zone, while Atf4 ranked seventh (Figures 4B and Figure 4-figure supplement 1B). Given the well-established cooperative roles of ATF4 and DDIT3 in the integrated stress response, these unbiased network analyses identified the ATF4–DDIT3 axis as a biologically relevant pathway for further investigation. To avoid overinterpretation, we have revised the manuscript to clarify that DDIT3 was prioritized based on its consistently high regulon activity and the activation of its downstream transcriptional network, rather than on differential expression alone (Revised manuscript, page 5, line 188-192).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      As discussed before, the authors employ a wide range of techniques (FOS IHC, FP for fine scale PVN OXT population dynamics, behavioural analysis, core and surface temperature tracking, physiological recordings to assess AAV specificity, optogenetic activation of PVN OXT neurons, and projection tracing) to address a clear question. The outcomes of these techniques seem to drive the same conclusion that PVN OXT neurons signal transitions from rest to arousal (behavioural and thermogenic) in a state-dependent manner:

      - FOS data identifies PVN OXT population activity following behavioural onset

      - Ca activity in these cells peaks at behavioural and thermogenic state transitions

      - Rump temperature and BAT activity increase at state transition points

      - Optogenetic stimulation of these cells recapitulates the thermogenic effects seen during physiological state transitions (in low body temperature animals) with a trending increase in physical activity

      Despite the inconclusive IHC results when validating the specificity of their AAV, the virgin female/ lactation experiment is convincing that they are specifically targeting PVN OXT neurons. The rationale for this experiment is clearer in the revised manuscript.

      Generally, in terms of the revised manuscript, the authors give strong responses to reviewer comments, either incorporating feedback, or giving clear explanations for the choices they made in the original manuscript. The revised manuscript is clearer about the question the authors aim to address, the reasons for their choice of experiments, and the limitations of the techniques used.

      We thank the reviewer for the close attention to the manuscript, the response to reviewers, and the revision, all of which have improved the manuscript.

      Criticisms:

      I appreciate and agree with the authors' point that this manuscript is more fundamental than simply social basis oxytocin neuron function. This is point is well made by their data, and in the revised text. However, I still believe more behavioural analysis would be welcome to any reader.

      They partly justify the lack of behavioural analysis in Figure 6 with the problem of "animal merging" on the SGBS images. However, in Figure 6C, they confirm that, in solo conditions, the SGBS readings are consistent with core body temperature readings. So why not stick to core body temperature, opto stimulate and analyse the social behaviour with DLC (with normal video recordings)?

      This is a good suggestion. Because we find that quiescent huddling (paired) bouts were associated with stronger body temperature regulation compared to solo quiescence and other behavioral states, and because PVNOT peak probability and frequency were higher in the paired compared to solo context, these experiments are warranted. We made the following edits to the discussion:

      “Future experiments should attempt to disentangle the effects of PVNOT light stimulation on social vs. non-social aspects of these behavioral state transitions; of particular interest would be to examine how light stimulation affects the duration and thermoregulatory control of social huddling.”

      The lactation validation still seems out of place in manuscript order. It is a very valuable validation, but it feels more like supplementary data for Figure 1. I feel the authors wanted it as a main figure because of how much work it must have been. In that case, it still makes more sense to include it in Figure 1.

      The purpose of the lactation experiment arose from the inadequacy of using histology to test whether AAV-transfected cells were oxytocin-immunoreactive. Because we observed intense oxytocin immunoreactivity in the fibres lining the ventricle, and less reactivity in the cell bodies than what we would have predicted from the Oxytocin-Cre-dependent AAV, we turned to the known physiological relationship between oxytocin-positive neurons and lactation. As such, this study is not associated with Figure 1, which demonstrates our initial, coarse-grained findings relating FOS activity in the PVN and in oxytocin-positive neurons during social thermoregulation.

      To your point, it typically does make sense to have the cellular validation “up front” as supporting or background information that enables the downstream experiments. However, what gives this data credibility as a standalone figure is the novel finding that PVNOT neurons display burst-like patterns of activity outside the context of lactation. Previous discussions with experts in the field, along with a review of the literature, unexpectedly led us to the observation that the burst-like patterns we observed during the transition from rest to wake and thermogenesis in virgin females represents a new aspect of oxytocin neuron physiology. Because we wanted to directly compare the new virgin female activity pattern (i.e., Figure 2) with the known lactation activity pattern, we decided it made the most sense to combine the validation aspect with the novel aspect into a standalone figure.

      Though their lactation experiment validates that they are targeting PVN OXT neurons, their optogenetic stimulation protocol may not be specifically inducing OXT release from these cells. PVN OXT neurons co-release glutamate but can also release glutamate independently of OXT following lower frequency tonic stimulation. OXT release from PVN neurons requires pulsatile stimulation at a higher frequency (Leithead et al., 2021; Piñol et al., 2014; Lincoln & Wakerley, 1975). In this paper, the authors use a low stimulation frequency (10Hz) and continuous pulse train (20s) to optogenetically manipulate the target PVN population which may bias the cells towards glutamate release over OXT. Therefore, though they find evidence that PVN OXT neurons are involved in driving the transition between states in their other experiments, their optogenetic stimulation may not necessarily involve OXT release/signalling. It may be valuable to separate this out to identify the signalling molecule underlying this behavioural/ thermogenic transition. This could be done by using an opto protocol that recapitulates physiological OXT release.

      The authors do however mention that isolating the specific contribution of OXT signalling compared to other co-transmitted molecules was not the aim of this study, so this is not an essential question for this manuscript.

      Thank you for this thoughtful point. We agree our optogenetic stimulation experiment should be interpreted as activation of PVNOT neurons rather than as selective evidence for oxytocin release or oxytocin signaling. PVNOT neurons can co-release glutamate (an idea we had also briefly touched upon in the Limitations and caveats section), and the stimulation pattern/frequency may influence the relative engagement of fast glutamatergic transmission versus peptide release. We agree the lactation literature, including Lincoln et al., highlights the importance of high-frequency pulsatile activity for oxytocin release, and that Piñol et al. provide evidence that PVNOT-linked glutamatergic transmission can interact with oxytocin-receptor-dependent modulation of downstream synapses–so thanks for pointing these out.

      We made revisions to support our protocol and now acknowledge this important aspect of the neuronal physiology. In Results, we now explain why we selected 10Hz: this frequency was grounded in the study by Fukushima et al. (2022), where 10Hz stimulation of PVNOT terminals in the rMR elicit thermogenic responses and 10Hz stimulation of PVNOT somata produce thermogenesis that’s dependent on oxytocin receptors in rMR.

      In the Limitations section, we now cite these three references to include broader context around stimulation frequency and differential release. We emphasize that our optogenetic data demonstrate sufficiency of PVNOT neuron activation, but do not establish whether the downstream thermogenic and behavioral effects are mediated by oxytocin, glutamate, or both. We note that resolving this issue will require future experiments using stimulation-pattern comparisons together with receptor-targeted pharmacology or genetic loss-of-function approaches.

      References

      Leithead, A. B., Tasker, J. G., & Harony-Nicolas, H. (2021). The interplay between glutamatergic circuits and oxytocin neurons in the hypothalamus and its relevance to neurodevelopmental disorders. Journal of neuroendocrinology, 33(12), e13061. https://doi.org/10.1111/jne.13061

      Lincoln, D. W., & Wakerley, J. B. (1975). Factors governing the periodic activation of supraoptic and paraventricular neurosecretory cells during suckling in the rat. The Journal of physiology, 250(2), 443-461. https://doi.org/10.1113/jphysiol.1975.sp011064

      Piñol, R. A., Jameson, H., Popratiloff, A., Lee, N. H., & Mendelowitz, D. (2014). Visualization of oxytocin release that mediates paired pulse facilitation in hypothalamic pathways to brainstem autonomic neurons. PloS one, 9(11), e112138. https://doi.org/10.1371/journal.pone.0112138

      A loss of function experiment to test for sufficiency would be a nice addition to further confirm their claims, but the authors mention that there were technical limitations to their attempts at inhibiting PVN OXT neurons. I appreciate the authors declaring that the DREADDs attempt suffered from unfortunate confounds. But for optogenetic attempts, I don't think they need a closed-loop system to get some useful results. They still can shine the light at "random" moments (that will correspond to random body temperatures) and then separate the data per body temperature.

      We thank the reviewer for this constructive suggestion. Such an experiment would strengthen our claims and complement the optogenetic activation (Fig. 6). Reviewer 3 brought up a similar concern.

      Building directly on the reviewer’s proposal, we now describe a loss-of-function experiment as an important next step. Optogenetic inhibition of PVNOT neurons can be delivered at pseudo-random times across light and rest phase. Because animals spend extended periods at rest during this phase, a substantial fraction will fall within established rest bouts, which can then be analyzed and stratified by body temperature, as the reviewer notes. The prediction is that silencing PVNOT neurons during rest should prolong the average duration of rest bouts and delay the onset of activity and thermogenesis, relative to matched unstimulated bouts.This provides a direct test of whether PVNOT activity is necessary for the transition from rest to activity. We have revised the Limitations and caveats section to describe this experiment.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Lastly, the mention of Raam et al. 2026 is insufficient. The authors just mention it regarding the potential differences with males, to be explored in future experiments. Even if not using males in the current study doesn't affect the stated conclusions, the fact that they chose females because "their thermo-behavioural states were readily discernible" is a considerable bias. Testing males in this very study might be out of scope, but more discussion is warranted.

      We thank the reviewer for this point. We agree that our decision to study females deserves fuller treatment, and we have expanded the Limitations and caveats section accordingly.

      We want to be clear about the rationale, because it was methodological rather than an assumption of sex specificity. Our previous study on behavioral thermoregulation in mice (Landen et al., 2024) showed that, during the light/rest phase, females–but not males–display clearly rhythmic episodes of rest and activity that align with transitions between thermoregulatory states, and are therefore well suited to the analyses that form the core of this study. This choice does constrain the generality of our findings to females, but it does not affect the validity of the conclusions we draw, all of which concern PVNOT neurons in females.

      At the same time, we agree that whether these mechanisms extend to males is a substantive open question and we now say so explicitly. A direct comparison in males, while beyond the scope of the present study, is an important next step, and the recently defined neural basis of collective thermoregulatory huddling (Raam et al. 2026) offers a useful framework for that work. We have modified the Discussion/Limitations and caveats as follows:

      “We focused on females for a practical reason: during the light and rest phase, females show clear, rhythmic bouts of rest and activity, which makes transitions between thermoregulatory states readily discernible and well suited to the analyses around each state transition used here (Landen et al., 2024). This choice constrains the generality of our conclusions, which pertain specifically to females. Because oxytocin signaling can differ between sexes (https://doi.org/10.1016/j.yfrne.2015.04.003), and because the neural control of thermoregulatory behavior may not be identical in males, whether the PVNOT dynamics we describe operate similarly in males remains an open question. Testing males directly was beyond the scope of the present study, but it is an important next step, particularly as the neural basis of collective thermoregulatory huddling has recently begun to be defined (Raam et al. 2026).”

      Reviewer #2 (Public review):

      Summary:

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and have implications for the thermoregulation field as well as the social/naturalistic behavior field. The findings are compelling and use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset.

      Comments on revised version.

      I appreciate the effort the authors have put into addressing all of my questions, and I have no remaining concerns.

      Thanks for the comments; they have greatly improved the manuscript.

      Reviewer #3 (Public review):

      Summary:

      This study investigates how the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons relates to physiological states in female mice, with a particular focus on behavioral states and thermogenic sympathetic activity. To address this question, the authors combined automated video-based behavioral classification with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes measured by infrared thermography, and the authors have made their custom analysis scripts available. The authors report that strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before transitions from resting to active states. This observation suggests that PVNOT neuronal activity may facilitate the transition from rest to activity. This phenomenon was observed in both pair-housed and individually housed animals. Taken together, these findings raise the possibility that the oxytocinergic system contributes to naturalistic behavior transitions even in the absence of social interactions. However, concerns regarding the selectivity of GCaMP expression in oxytocin-expressing neurons call into question the validity of the recorded PVNOT neuronal activity.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions. Elucidating these roles requires monitoring PVNOT neuronal activity under diverse behavioral contexts, as well as manipulating this activity to establish causal relationships. In this study, the authors present a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the monitoring and manipulation of PVNOT neuron activity. This setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      Thanks for the comments. We are encouraged to hear this framework will open new doors in understanding how the oxytocin system regulates behavior and energy homeostasis.

      Weaknesses:

      (1) Immunohistochemical validation of selective GCaMP expression in oxytocin-expressing neurons showed that only 24-51% of GCaMP-positive neurons expressed oxytocin. As an alternative approach, the authors demonstrate that GCaMP-expressing PVN neurons in virgin females exhibit calcium peaks during rest-wake transitions with kinetics similar to those observed in PVNOT neurons during early lactation. However, this comparison is based solely on population-level peak profiles and does not provide direct evidence for cell-type specificity of GCaMP expression in oxytocin neurons. This limitation substantially undermines the validity of the optical calcium imaging data. In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.

      We view our data as showing strong evidence that the recorded neurons include, but may not be limited to, PVNOT neurons for the following two reasons: (1) as the reviewer notes, our longitudinal experiment shows conservation in the physiological and biophysical profile of these neurons in females that went from virgins to parturition and lactation, and (2) as described in Discussion/PVNOT neurons in context of arousal and peptidergic PVN cell-types, non-OT cell-types in the PVN do not show this pulsatile busting profile.

      In the “Discussion/Thermal tracking and validation of PVNOT recording specificity” section we had stated “We note that the animals were perfused at ~ZT4–8, before we were aware that somatic OT immunoreactivity in PVN neurons reaches a daily low during the early light phase [56]”. We now add to this the idea, suggested by the reviewer, that “In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.”

      (2) Although the authors' interpretation is generally consistent with the data presented, their main conclusions rely heavily on observational findings. Moreover, optogenetic stimulation of PVNOT neurons failed to robustly recapitulate behavioral state transitions (Figs. 6D and S5B). Further interventional experiments will be necessary to more rigorously test the authors' interpretation and to establish mechanistic insight into the causal relationship between PVNOT activity and rest-to-active transitions. In particular, loss-of-function approaches targeting the PVNOT system, such as OXTR antagonism, inhibitory DREADDs, or cell-type-specific ablation, will be essential to determine whether perturbation of this system alters behavioral state transitions These points should be addressed in future studies.

      Reviewer 1 brought up a similar concern. We have added to the Discussion/Limitations and caveats to address this.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Note: as described in the previous response to reviewers, we have tried inhibitory DREADDs in this system and have concluded that it is of little value because delivering DREADD ligand requires handing the animals for an IP injection—a procedure that disrupts sleep/rest and induces stress hyperthermia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have answered our criticisms and can proceed as they chose. This is an important paper, and it is the author's choice whether to develop their research here or in a subsequent paper.

      Thank you.

      Reviewer #2 (Recommendations for the authors):

      I thank the authors for citing my pre-print, as suggested by Reviewer 1. The paper has now been published and the authors may like to cite the published version (doi.org/10.1038/s41593-026-02224-0).

      Thank you.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors now interpret their results as indicating that PVNOT activity biases the system toward state transition (from rest to active), rather than acting as a deterministic trigger. This interpretation is reasonable. However, the wording "PVNOT peaks (or neurons) predict transitions to behavioral arousal and thermogenesis" may be misleading. If arousal and thermogenesis occur in more than 80% of cases following PVNOT peaks, then such peaks could reasonably be described as "being predicted". Otherwise, the terminology should be revised for clarity.

      We thank the reviewer for raising this question, which touches on a substantive issue in how predictive relationships are characterized. We agree that "predicts" can misleadingly imply a high positive predictive value: i.e., that a large fraction of peaks are followed by transitions.

      This is not the claim we intend, nor is it the appropriate statistical criterion. A variable is predictive when it shifts the conditional probability (or, here, the conditional distribution) of the outcome relative to its base rate — the criterion underlying likelihood ratios, relative risk, and signal-detection measures — rather than when it exceeds an absolute occurrence threshold such as 80%. By this standard, a peak can be informative even if transitions do not follow the majority of peaks, provided transitions are substantially more likely (or thermogenically warmer) when a peak precedes them than when one does not.

      Our data support precisely this. The logistic regression shows peaks are much more probable immediately before rest offset than at other transitions or at baseline, and our new analysis shows that transitions preceded by peaks carry significantly larger post-offset Tb increases than those without. We are not claiming peaks act as a deterministic trigger, and we agree with the reviewer that they are not present before every transition.

      To keep our language aligned with these results, we have revised the wording to avoid "predict" where it could imply high hit-rate determinism, replacing it with comparative phrasing. Accordingly, we have revised the terminology throughout the manuscript: where a claim concerns timing, we now state that peaks “precede” transitions. We have removed “predict”/”predictive” from the section heading, figure legend, introduction and results as follows.

      “Then, we discovered that PVNOT calcium dynamics during huddling were associated with increased likelihood of transitions to body warming and arousal.”

      “PVNOT neuronal activity precedes transitions towards thermogenesis and behavioral arousal in social and non-social contexts.”

      Fig. 3 legend title: “PVNOT peaks are associated with increased likelihood of thermogenic rest-to-active transitions.”

      “Thus, PVNOT peaks are at least five-fold more likely to occur near the offset of quiescence/quiescent compared to onset, and signal an increase in physical activity—a correlate of behavioral arousal 53 and a means of increasing metabolic rate and Tb [26]”

      “Thus, for nesting and active huddling, PVNOT peaks are two- to three- fold more likely to occur at bout onset than offset.” Dropping flagged word here lol.

      “Together these results suggest that elevated PVNOT activity dynamics precede the offset of two rest states (quiescence and quiescent huddling) by approximately 100 seconds, and the onset of two post-quiescence active states (nesting and active huddling) by around 20 seconds, in solo and paired mice respectively.”

      “Moreover, PVNOT peaks aligned with the low point of a U-shaped body temperature profile: on average, Tb decreased before, and increased after, the time of the calcium peak in both solo and paired conditions (Fig. 3O,R). Together, these results suggest that PVN<sup>OT</sup> peaks occur during a low Tb trough and mark a subsequent rise in Tb.”

      (2) Regarding the 400-sec latency of BAT surface temperature increases following optogenetic stimulation, the authors now attribute this delay to slow peptidergic transmission. However, the authors should consider prior findings showing that BAT temperature increased immediately following optogenetic stimulation of PVN→rMR oxytocin neurons in anesthetized rats (Fukushima et al., 2022).

      My hunch is that doing this in anesthetized rats gives a stronger signal to noise… not sure if I can back that up though.

      At the least we can add a sentence that says “rMR oxytocin neurons immediately increases BAT temperature, while infusion of OXT or NMDA in the rMR results in BAT temperature increases after approximately one minute…” (see Fig. 3,4,5).

      We thank the reviewer for redirecting us to Fukushima et al. (2022). We note, however, that in that study the fast-responding variable was BAT sympathetic nerve activity, whereas the BAT temperature itself rose over several minutes following both optogenetic stimulation (their Fig. 4F, quantified at 5 and 10 minutes) and focal rMR infusion of oxytocin or NDMA (their Fig. 5, multiminute traces). This thermal timescale is comparable to the one we observe.

      The remaining difference could reflect methodological differences: we stimulated PVNOT somata rather than rMR terminals, measured intrascapular surface rather than BAT temperature directly, and recorded in awake, freely behaving animals (rather than anesthetized animals) in which competing thermoeffector and behavioral processes are active. Consistent with a methodological basis for the delay, focal infusion of oxytocin or NMDA into the rMR in that study increased BAT temperature over roughly a minute (Fukushima et al., 2022). Slow, diffuse peptidergic neuromodulation may further contribute, oxytocin is released from large dense-core vesicles and can act over extended time scales (Ludwig and Leng, 2006; Parmaksiz and Kim, 2025; Qian et al., 2023), although our data cannot isolate this mechanism from the factors above or from fast glutamatergic co-transmission that likely accompanies PVNOT activation (Hrabovszky and Liposits, 2008).

      (3) In the previous review, clarification was requested regarding the rationale and histological basis for intravenous FluoroGold injection. While the authors have now added methodological details, they should also incorporate the following explanatory text (previously provided in their rebuttal) into the manuscript for readers unfamiliar with PVN histological analyses:

      "Intravenous injection of FluoroGold (FG) was used to histologically differentiate between magnocellular and parvicellular oxytocin neurons in the PVN. Because the posterior pituitary is located outside the blood-brain barrier, i.v. FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVNOT neurons of interest."

      We thank the reviewer for this suggestion. We have added the explanatory text to the results subsection, “PVN<sup>OT</sup> cellular projections to the rMR”. The text now reads: “rMR cell types in mice, we used FluoroGold (FG to disambiguate magno- vs. parvocellular PVN<sup>OT</sup> projections [67] (Fig. S6A-C). Because the posterior pituitary is located outside the blood-brain barrier, intravenous FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVN<sup>OT</sup> neurons of interest.”

    1. Author response:

      Reviewer #1 (Public review):

      The manuscript from Zhu et al. identifies microbial riboflavin-derived MR1 ligands as potent pharmacological activators of human MAIT cells and provides evidence that MR1 ligand stimulation can enhance MAIT-mediated tumor killing across multiple solid tumor models. The study is conceptually interesting and supported by a broad combination of human primary samples, tumor cell lines, 3D models, SC transcriptomics, and xenograft experiments. Overall, the data largely support the central conclusion that MR1 ligand stimulation can strongly activate human MAIT cells and enhance anti-tumor cytotoxicity. However, the broader conclusions concerning endogenous MAIT mobilization, tumor specificity, and translational potential are not yet fully supported by the current data and should either be moderated or addressed with additional experiments.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Comments:

      (1) The authors use one-way ANOVA throughout the manuscript, but this may not be appropriate for some analyses, particularly when multiple experimental factors are present and their interaction effects need to be considered. For example, Figure 3f appears to involve multiple factors, for which a two-way ANOVA may be more appropriate. Similar issues may apply to other panels.

      We thank the reviewer for this valuable comment. We will carefully review the statistical analyses and revise the tests as appropriate, including the use of two-way ANOVA where multiple experimental factors are present.

      (2) In Figure 3f, the authors show data from patients #1 and #2 and state that the experiment is representative of three experiments. What does the reported "n=4" represent in this figure?

      We thank the reviewer for this valuable comment. We will revise the figure legend to clearly define what the reported n = 4 represents.

      (3) There appears to be a discrepancy between Figure 3f and Supplementary Figure 3b. The two panels appear to use the same treatment conditions and the same label, and both appear to use patient #1 samples, yet the reported values are different. Please clarify the experimental design and explain the reason for this discrepancy.

      In addition, the gating strategy used to define live tumor cells should be clearly described in the figure legend and/or Methods. The authors define "live tumor cells" as MR1/5-OP-RU tetramer-CD45- cells. However, in primary liver tumor samples, the CD45-/tetramer- population may contain other non-hematopoietic cells, such as fibroblasts, and therefore may not exclusively represent tumor cells. The authors should clarify whether additional tumor-specific markers or other criteria were used. The gating strategies for the relevant flow cytometry experiments should be provided in the Supplementary figures.

      We thank the reviewer for this valuable comment. Figure 3f (patient #2) and Supplementary Figure 3b (patient #1) were generated using samples from different patients. We will clarify this in the revised manuscript and provide the relevant gating strategies in the Supplementary Information.

      (4) I have some concerns regarding the claims of "selective activation of anti-tumor inflammatory pathways rather than generalized cytokine release" and "avoiding induction of tumor-supportive mediators." The authors show that MAIT cells stimulated with 5-OP-RU can substantially reduce tumor cell viability. Therefore, the cellular composition of the co-culture is likely to change considerably during the assay, which may affect the absolute levels of cytokines and other soluble mediators detected. For example, reduced tumor cell numbers could lead to lower production of tumor-derived factors such as VEGF, potentially confounding the interpretation that these mediators are not induced by MAIT activation. The authors should consider whether cytokine measurements have been normalized to viable cell numbers or otherwise account for differences in tumor cell abundance.

      We thank the reviewer for this important comment. We agree that differences in tumor cell abundance may affect cytokine measurements. We will moderate our claims accordingly and acknowledge this limitation in the revised manuscript.

      (5) The in vivo tumor models may show substantial variability between independent experiments. Rather than presenting a single representative experiment, the authors should consider showing pooled data from all independent experiments, with the total number of mice clearly indicated.

      We thank the reviewer for this valuable comment. We will provide pooled data from all independent in vivo experiments and clearly indicate the total number of mice.

      (6) Why did the authors use an MR1-overexpressing tumor cell line for the in vivo studies rather than the parental cells with endogenous MR1 expression, together with MR1-KO cells as a negative control? The authors demonstrate that MR1 is detectable across multiple tumor cell lines and that endogenous MR1 expression is sufficient to support MAIT-mediated killing in vitro. Moreover, MR1 overexpression substantially enhances tumor cell susceptibility to MAIT-mediated killing. Therefore, it is unclear whether the strong therapeutic efficacy observed in vivo reflects physiologically relevant MR1 expression or is driven by artificially elevated MR1 expression. An in vivo comparison using parental and MR1-KO tumor cells would substantially strengthen the translational relevance and establish whether the therapeutic effect can be achieved at endogenous levels of MR1.

      We thank the reviewer for this important comment. We agree that comparison with endogenous MR1 expression would strengthen the translational relevance of our findings. We will include new in vivo experiment comparing parental tumor cells.

      (7) How is tumor specificity of MAIT achieved? The authors propose that MAIT-cell activation by MR1 ligands provides an antigen-independent approach for tumor targeting. However, MR1 is broadly expressed and is not tumor specific. While the relative sparing of T and B cells in Figure 7B provides some evidence of cell-type selectivity, this does not establish tumor versus normal tissue specificity. It remains unclear whether activated MAIT cells can discriminate tumor cells from other normal MR1-expressing cells and tissues. This raises an important question regarding the potential systemic toxicity of MAIT cells activated by systemic administration of 5-OP-RU. In particular, could other MR1-expressing cells be targeted when a large number of MAIT cells are simultaneously activated? The authors should consider assessing systemic toxicity in vivo, for example by examining serum ALT/AST levels and tissue pathology, and/or by evaluating the effects of MAIT + 5-OP-RU in tumor-free animals. At least, the potential specificity and safety limitations of systemic MR1 agonism should be discussed.

      We thank the reviewer for this important comment. To further evaluate the potential safety concerns associated with systemic MR1 ligand stimulation, we will include a new experiment assessing the effects of MAIT cells plus 5-OP-RU in tumor-free animals. We will also discuss the potential specificity and safety limitations of systemic MR1 agonism in the revised manuscript.

      Reviewer #2 (Public review):

      The manuscript by Zhu et al. describes MAIT cell activation by riboflavin metabolites presented by MR1. The authors provide solid evidence for this activation and anti-cancer functional consequence using an array of selected cell lines, primary ex vivo and engineered xenograft models. Broadly, the results are thorough and well controlled, and provide a highly informative insight into the metabolite-MAIT-cancer cell interactions. However, the majority of this work is undertaken using models that preferentially express key targets, and whilst still useful, the (current) broader implications of this research are overstated. Additionally, the suggested MAIT modulation of the tumor microenvironment requires clarification.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Major Comments:

      (1) In Figures 2b-d, the authors suggest microbial metabolite stimulation of PBMC cultures increased MAIT cell frequency up to 60%. Whilst their flow data is compelling, the frequency of one population can be influenced by changes in other populations. A form of absolute or relative-to-total count should be used.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts and/or normalized data to more accurately assess changes in MAIT cell frequency.

      (2) The statements regarding cytokine induction in Figure 4e are too strong; many of those inflammatory cytokines are not automatically and consistently tumour-suppressive. The line 299 '...were not induced' may just reflect death of tumor cells. It would be useful to include tumour cell-only controls in Figure 4.

      We thank the reviewer for this valuable comment. We agree that the statements regarding cytokine induction should be interpreted more cautiously. We will revise the relevant claims.

      (3) Figure 7 is interesting, but the authors' conclusion that MAIT+5-OP-RU controls the tumor microenvironment is not robustly supported by their evidence.

      (a) It is not clear how CD14+ cells established a sustained suppressive environment.

      We thank the reviewer for this valuable comment. We will include additional experiments to further characterize the contribution of CD14+ cells to the observed suppressive environment.

      (b) It is not clear how the peritoneal addition of microbial metabolites 'significantly enhanced MAIT-mediated tumor control'. The authors show that the addition of 5-OP-RU reduced the number of GFP-expressing tumour cells present in peritoneal lavage fluid. There is limited evidence to suggest this occurs through MAIT cells or MR1 in this figure.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups to further determine the contribution of MAIT cells to the observed tumor control.

      (c) It is difficult to draw conclusions from peritoneal lavage flow when some experimental groups received cells IP, but then all groups were equally assessed for key populations, and all data are presented as frequencies. The authors should use absolute counts (or similar) to appropriately show changes in cell populations to account for varying total/live/cd45+ cell compartments.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts, in addition to frequencies, to account for differences in total and viable CD45+ cell numbers.

      (d) It would be necessary at a minimum to include 5-OP-RU-only controls, and ideally include MR1 blocking or the cancer line with MR1 removed. Alongside this, the authors should substantially reduce the strength of their statements on microbial metabolite-MAIT suppression of the tumor microenvironment.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups and will substantially moderate our statements regarding microbial metabolite-mediated modulation of the tumor microenvironment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This Review Article provides a thorough overview of whole-brain activity changes induced by brain stimulation and summarizes the current state of the field. However, it lacks integration across spatial and mechanistic scales, which limits the reader's ability to understand how the different findings relate to one another. In addition, several key concepts are not explained in sufficient depth for non-expert readers. The manuscript would benefit from the development of a cohesive conceptual framework to more clearly synthesize the existing literature.

      Thank you for the positive assessment. We fully agree, and as suggested we have added a new conclusion paragraph that outlines a synthesis of the paper and suggests a conceptual framework :

      “In this paper, we have reviewed aspects of neuronal responsiveness, from the microscale level of neurons and circuits, the mesoscale level of single brain areas, and the macroscale level of the whole brain. At the microscale, it is apparent that the circuit operating in an asynchronous mode displays the highest responsiveness, as seen in brain slices (D’Andola et al., 2018). The underlying mechanism is that the high levels of synaptic « noise » in asynchronous states set neurons in a high responsive mode, as seen in models of single neurons (Ho & Destexhe, 2000). This higher responsiveness is confirmed at mesoscale, and can be seen for example with Utah-array recordings comparing wake and anesthesia (Dwarakanath et al., 2025). Similarly, propagating waves occur in the asynchronous state in awake monkey (Muller et al., 2014), and sensory inputs evoke more propagating patterns (and higher PCI) in wakefulness with asynchronous states compared to slow-wave states of anesthesia in mice (Montagni et al., 2024). At the whole-brain scale, experiments also find that evoked responses are more complex and propagating compared to slow-wave states (Massimini et al., 2005), a situation which models can reproduce (Goldman et al., 2023; Sacha et al., 2025). Other measures, such as fluidity (Breyton et al., 2024) and reversibility (Camassa et al 2024) also point to the same conclusion. Collectively, these results show that asynchronous and irregular activity states set neurons in a high responsive mode, which in turn impacts network behavior and favors the propagation of activity as mesoscale propagating waves, or macroscale activity patterns that propagate across brain regions. It is therefore not surprising that the best correlate of conscious states is the asynchronous activity (Koch et al., 2016).”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper is a comprehensive review of perturbation studies and the state-dependence of the brain's response to perturbation at the circuit, mesoscale, and macroscale levels.

      Strengths:

      The strengths of the paper are the thorough description of many perturbation studies at different levels of organization, and the integration of both experimental and modeling studies. The review clearly communicates the need to consider (1) brain or local-population state, and (2) multiple levels of organization, in order to understand perturbation responses. Another major strength is the ability for the reader to reproduce figures using the EBRAINS platform.

      Weaknesses:

      Two major points of improvement should be resolved with the review, in order to make it useful for a broad audience.

      The first is that the review does not include a significant integration across scales, and as a result, reads like three separate (though comprehensive) reviews. Currently, the only integration across the scales is in the brief conclusion paragraph. I would recommend adding an additional section, in which the overarching picture is discussed. (i.e. a unifying view of state dependence, and what is learned by considering across scales). This need not be too long, but it should be longer than a single conclusion paragraph.

      Thank you for the positive assessment. We fully agree with the excellent suggestion of adding a concluding paragraph where the conceptual framework and overarching picture are presented. Please see the new conclusion paragraph that we added to the paper (copied above in the reply to Editors).

      The second major weakness is that there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      See our answer to the specific comments below for the list of unclear points.

      Reviewer #2 (Public review):

      Summary:

      In this review article, the authors discuss the whole-brain activity changes induced by brain stimulation. They review the literature on how these activity changes depend on the cognitive state of the brain and divide the results by the scale of the change being induced, from microscale changes across small groups of neurons, up to macroscale changes across the entire brain. Finally, they describe attempts to model these changes using computational models.

      Strengths:

      The review provides an overview of the results within this subfield of neuroscience, and the authors are able to discuss a lot of prior results. The framing of the changes in neuronal activity in terms of computational changes is also a helpful approach.

      Weaknesses:

      However, the authors are not able to contextualize these results within a single framework, i.e. explaining from first principles how different aspects of stimulus-induced changes interact to generate functional changes in the brain, and how different changes - at distinct spatiotemporal scales - combine to form larger effects. This is a significant weakness in generating a review of the literature, since the authors do not provide a cohesive conceptual framework on which to frame the results. Similarly, the authors do not explain how their different computational models fit together, and how one can get a singular computational understanding of the distinct mechanisms of brain activity changes due to stimulation under different brain states, by combining the results derived from each separate model.

      Thank you for the positive assessment. This is an excellent suggestion, actually also requested by Reviewer 1. We have now added a new conclusion paragraph where the conceptual framework is explained (copied above in the reply to Editors).

      Major Comments:

      (1) The authors have written this review as if it were intended for an audience who is already familiar with the topics. For example, they introduce concepts like complexity, spiral vs planar waves, without much explanation.

      Thank you for this helpful comment. We agree that the Introduction should be more accessible to readers who are less familiar with these concepts, and we have therefore revised the text to provide a clearer definition of complexity and a more explicit explanation of propagating wave patterns.

      Specifically, we now clarify that complexity can be understood as the richness of the set of accessible states of a system, which in our context can be related to the diversity of slowwave propagation modes and to the high-entropy, desynchronized activity of the awake brain. We also expanded the description of propagating slow waves to distinguish planar from spiral waves and to explain how their relative prevalence changes with anesthesia depth.

      In the Introduction we replaced the sentence “New methods … at various scales” with “New methods for characterizing the complexity of network dynamics and their response patterns have emerged, particularly recently (Krohn et al., 2023; Wolf et al., 2018), and are presented here at various scales. Here, complexity is associated with the set of accessible states of a system (Parisi, 2006). In the present context, this notion can be linked to the diversity of slow-wave propagation modes and to the richness (i.e., the entropy) of perturbation-evoked responses in brain activity.”

      While in Results (p. 13) the sentence “Spontaneous slow waves … administered (Huang et al., 2010).” has been expanded in “Spontaneous slow waves can also display propagating patterns, as shown in anesthetized mice (Huang et al., 2010; Mohajerani et al., 2010; Pazienti et al., 2022; Stroh et al., 2013). These patterns may take the form of planar waves, which travel across the cortex along a relatively regular front, or spiral waves, which rotate around a central core and therefore produce a more complex spatiotemporal organization. Under relatively deep anesthesia, spiral waves occur more frequently than planar waves, whereas the opposite imbalance is observed as anesthesia is lightened (Huang et al., 2010).”

      (2) Regarding complexity, the authors present a quantification termed PCI. However, in the associated box, they state that PCI could be implemented in a number of different ways, using analogous metrics (which are, nonetheless, not identical). Yet the authors simply claim that all these metrics are sufficiently similar to be grouped together as "PCI". The authors do not provide much intuition about this, and they also don't present any other potential quantifications. This makes any interpretation of their results strongly dependent on your understanding of the concept of PCI. It would be helpful to present some other, analogous metric to demonstrate that the results that the authors are focusing on are not somehow tied to the specific computational structure of the PCI metric.

      Thank you for pointing out to this inconsistency. We agree that the rationale for focusing on perturbational complexity was not sufficiently introduced in the original version of the manuscript.

      Broadly speaking, complexity measures used in consciousness research can be divided into two major classes. The first includes observational measures, which are computed from spontaneous ongoing activity and quantify statistical dependencies within neural time series. The second includes perturbational measures, which quantify the deterministic causal interactions revealed by a controlled perturbation of the system and their spatiotemporal propagation across the network (see Sarasso et al., 2021).

      The primary aim of the present Review was to discuss how complexity changes across spatial and temporal scales in response to perturbations. For this reason, we focused on perturbational complexity measures and, in particular, on the Perturbational Complexity Index (PCI), which remains one of the most widely adopted and validated approaches in this category.

      As described in Box 2, different implementations of PCI have been proposed. The two most established versions are PCI based on Lempel–Ziv complexity (PCI^LZ) and PCI based on state transitions in principal component space (PCI^ST). Although these implementations differ algorithmically, they were developed to operationalize the same theoretical construct and have been shown to correlate strongly when applied to the same datasets (Comolatti et al., 2019). For this reason, throughout the Review we use the term “PCI” as an umbrella label encompassing these related perturbational complexity measures.

      Importantly, all complexity measures discussed in the studies reviewed here belong to this broader class of perturbational approaches. While adaptations of the original algorithms are often required when dealing with different recording modalities and spatial scales, these modifications mainly concern preprocessing and signal representation rather than the underlying theoretical construct being quantified.

      To clarify this point, we have revised the Introduction to explicitly motivate our focus on perturbational complexity, to distinguish perturbational from observational complexity measures, and to explain why different PCI implementations can be discussed within a common conceptual framework. We believe that these additions make the rationale of the Review substantially clearer and reduce the impression that the conclusions depend on a specific implementation of PCI.

      (3) The authors divide the review into sections organized by the spatial extent of the effects that they are exploring (e.g. from microscale to macroscale). However, they don't bring together these insights into a cohesive structure - for example, by providing potential explanations of the macroscale effects by using the microscale changes.

      We agree – and this is now the focus of the newly-added conceptual-framework conclusion paragraph.

      (4) The authors completely ignore any aspect of cell-type specificity in their review, despite the known importance of specific cell types at the microcircuit scale. This makes it difficult to map their results onto the true biological system.

      We agree that cell-type specificity could be made more explicit. The revised manuscript now clarifies that several models already include cell-type specificity. For example, the AdEx-based models distinguish excitatory regular-spiking or pyramidal populations with adaptation from inhibitory fast-spiking populations without adaptation. This differentiation is not made with other models like leaky or quadratic integrate-and-fire models. At the mesoscale, mean-field approaches can be derived for different structures, such as cortex, thalamus, hippocampus, striatum, or cerebellum, and can incorporate the experimentally experimentally observed firing properties of relevant cell classes.

      (5) The authors introduce several different computational models, such as the Hopf model, the AdEx model, and the MPR model. However, they do not provide the reader with a conceptual understanding of the structure of each of these models (except through potentially more complex terminology, e.g. the Hopf model is a "phenomenological StuartLandau nonlinear oscillator"). Additionally, though they present the results of each simulation, they don't provide the reader with intuition about how these models compare against each other, and how best to interpret results derived from each model.

      Very good question, and the answer is not easy. If the goal is to capture large-scale phenomena with models as simple as possible, then Stuart-Landau, Hopf, or Jahnsen-Rit may be appropriate. This approach is rather top-down. But if the goal is to assess how microscopic changes (synaptic receptors for example) affect large-scale brain activity, then we need a bottom-up approach, where mean-field models are derived. We can better explain this.

      We agree that while the technical definitions of the whole-brain models (Hopf, AdEx, MPR) were provided, a clear conceptual framework comparing their underlying structures, specific trade-offs, and interpretation guidelines was missing. We have substantially revised the "Macroscale" section on Page 22 to provide immediate intuition regarding what each model represents structurally (e.g., macroscopic phenomenology vs. microscopic biological realism). We emphasized the structural assumptions of each of them, as well as the explicit utility in interpreting brain responsiveness. This ensures readers understand exactly why a researcher would choose one model over another depending on the mechanistic question at hand.

      (6) In several cases, the authors make statements that they appear to believe to be completely straightforward (and require no justification), but that do not appear so to the reader. For example, they mention: "In wakefulness and REM sleep, ..., the membrane potential is depolarized and close to the spike threshold, which explains why neurons respond more reliably and with less response variability compared with slow-wave sleep". However, this statement is not obvious to the reader and requires explanation (for example, in a system that is close to balance, bringing cells closer to the firing threshold can result in increased response jitter).

      We agree that the original statement was an over-simplification of a complex situation. We have revised it to avoid suggesting that depolarization alone monotonically increases reliability. The relevant mechanism is the combination of depolarization, desynchronized high-conductance synaptic input, balanced fluctuations, and reduced tendency to enter long silent Down states. In this regime, weak inputs are more likely to be converted into spikes and propagate through the network. However, too high conductance, excessive noise, can shunt inputs, enhance jitter, or saturate the network. We are now more explanatory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As stated in the public review, there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      Points needing clarification:

      (1) sPCI (slice Perturbational complexity index) - is this different from other PCIs in box 2? Regardless, the metric and its interpretation should be briefly explained in the main text.

      We use sPCI to refer to the PCI measure adapted for application to cortical brain slices (D’Andola et al., 2017; see Box 2). It relies on the same core algorithm as PCI, namely the Lempel–Ziv complexity of the spatiotemporal pattern of significant responses (Casali et al., 2013) but differs in the preprocessing steps required for slice recordings. We now added this clarification to the text.

      (2) Page 9 "by decreasing fast inhibition but also enhancing it" - What does that mean? More info about the model is needed.

      Thanks for raising this point, since this sentence was indeed confusing. We have revised it now.

      (3) Page 9 "balance between segregation and integration, a crucial ingredient on which sPCI relies" - How is this balance seen in the figure? All I see is sPCI and blockage of GABA.

      The comment is correct, and this mention of segregation and integration has now been eliminated.

      (4) Figure 3D, Page 11 "two different desynchronized (AI) states in a network of AdEx neurons"- What are the two different states? Why is the response different?

      The different AI states correspond to different synaptic strength parameters, we added this precision in the text.

      (5) Figure 3B - "Bifurcation diagram showing the different activity regimes displayed by spiking neuron network." Which model? Multiple are cited. This is a general issue throughout where multiple models are mentioned in the text, and it's unclear which is shown in the figure.

      We agree with the Reviewer's helpful remark. We have revised the manuscript to explicitly state the types of models depicted in the different panels of Figure 3. Corresponding details have also been incorporated into the relevant text in the Results section (previously pages 10–12)."

      A few editorial issues:

      (1) The text in many of the figure panels was too small to read. This is a significant issue that must be addressed.

      We will fix this at the next round, can you please let us know which figures are not visible?

      (2) I recommend reading through for writing flow. E.g. In the first paragraph of the introduction, there are two sentences that start with "importantly, ..." in a row.

      Thanks for noting this – this is now fixed.

      (3) Figure 1E - How does the color on the left relate to the right? What is the y-axis?

      The colour code corresponds to the latency of activation (light blue, 0 ms; red, 300 ms). The Y-axes is the global mean field power (voltage). It has now been included in the figure caption.

      (4) Figure 1C - What is the stimulus?

      The triangle corresponds to the electrical stimulation of the homotopic area 18 of the contralateral hemisphere. This information is now included in the figure legend.

    1. Author response:

      Reviewer 1 is concerned that our astrocyte enriched cultures have significant contamination of microglia or other myeloid cells, OPCs (and related cells) and neurons. Further, they assert that purified astrocytes do not express TNF.

      That astrocytes can’t make TNF directly contradicts our previous paper (Heir, et al., JNeurosci, 2024) showing that the TNF driving homeostatic plasticity is generated by astrocytes. The reviewer seems to want to dispute that paper, which is not really the topic of the current paper (which covers the regulation of TNF production, not whether particular cell types make TNF). The Nedergaard group also saw TNF release from human astrocytes (Wang, et al., 2006). The papers cited by the reviewer (all from the same group) rely on RNAseq data, which has limited depth and cannot distinguish if something is not expressed or simply below threshold. Further, as these datasets were generated from astrocytes isolated from brain (which has normal levels of activity), the astrocytic TNF expression would be very low. Plenty of data supports that astrocytes can express TNF when stimulated (by LPS or other activators), and our previous paper shows that this is also true when neuronal activity is blocked (or absent). But at baseline, astrocyte TNF is quite low and likely undetectable as assayed in those papers.

      Here we are using highly purified astrocyte cultures. The reviewer is perhaps unfamiliar with the type of cultures we are using. Given that we use mechanical disruption to remove neurons, followed after 1-2 weeks by shaking to remove microglia, and finally cell passaging, all before experiments, it is surprising that the reviewer thinks there could be neuronal contamination. Neurons cannot survive that procedure, and we do not observe them by morphology or immunostaining, nor see neuronal markers by qPCR.  The microglial contamination is also minimal, as noted in the manuscript, with qPCR for microglia markers is at noise levels (Iba1 Ct value of 34.6), while GFAP shows robust expression (Ct of 16.8; >100,00 fold more than Iba1). But it is possible, if unlikely, that some small number of microglia are making a lot of TNF. However, treating our cultures with the microglia toxin LME (used in Heir, et al., 2024) did not alter our results, further suggesting microglia are not contributing here. Other contaminating cell types (in the OPC lineage, for example) are also possible. However, the majority of cells in our astrocyte-enriched culture are positive for TNF by immunostaining (done while blocking protein export, to prevent any release of TNF). This makes it highly probably that astrocytes are producing TNF (and this production is regulated by g-protein signaling). To verify this, we will show TNF protein in cells co-labeled with astrocyte markers in our upcoming revision of the paper. This will definitively identify astrocytes as producing TNF in these rat cultures. With the human iPSC-derived astrocytes, microglial contamination is not possible (this requires a completely different differentiation protocol). We agree the ALDH1L1 labeling is not as expected, but it is unclear if this is an antibody issue or mis-localized protein. However, the cells also label with S100beta and GFAP, making the astrocyte identity the most likely option by far. We have additional qPCR data showing expression of ALDH1L1 by these cells, in addition to the other astrocyte markers (which will also be added to the revision). The in vivo situation is more complex, and we can’t exclude that astrocyte-DREADD signaling here indirectly alters TNF production in other cells. However, given the direct regulation of astrocyte TNF production in culture, the simplest explanation is that the same is occurring in vivo.

      Reviewer 2 was concerned that the Gs data was indirect and the use of pharmacological approaches with microglia. As for Gs signaling, it is a bit unclear what the reviewer is suggesting as an alternative hypothesis. We activate the Gs-coupled beta-adrenergic receptor to reduce TNF levels and get the same effect by activating adenylyl cyclase, the canonical downstream pathway from Gs-coupled receptors. While it is possible that beta-adrenergic receptors could have alternate coupling or that Gs activation acts on additional pathways, it seems odd to argue that Gs would not be working through adenylyl cyclase activation when activating the cyclase yields the same response. Certainly the most parsimonious explanation is that Gs-couple receptors act through adenylyl cyclase to reduce TNF production.

      As for the use of pharmacology with microglia, this was the more expedient solution to the difficulty of using AAV virus on microglia. Gathering the necessary Cre and conditional DREADD lines was an impractical solution in terms of time and resources. However, the pharmacology of these receptors is well characterized, as is the g-protein coupling. Given that the results are identical to the results from more specific manipulations in astrocytes, it seems reasonable to conclude that there is a common pattern of GPCR regulation of TNF production. The criticism that non-canonical pathways can be activated by these receptors seems equally true for the DREADDs, as these are just GPCRs with mutated binding sites. If anything, the forskolin experiment is the most specific, yet the reviewer dislikes this approach. The overall consistency of the responses, whether due to DREADD activation, native receptors or direct activation of adenylyl cyclase, is the strongest argument.

      This reviewer was also concerned about the limits of in vivo experiments. As noted above, we agree that the in vivo situation is less controlled and indirect effects are possible. However, since the direct action on astrocytes in a defined culture system is identical to what we observe in vivo, the most likely explanation is that the GPCR is having the same effect on TNF production, rather than leading to an unknown secondary signaling which then alters TNF production in microglia (or other cell types).

      The remaining concerns about sample size, statistics, etc will be fully addressed in an upcoming revision. All reported n’s are biological replicates. The iPSCs were generated from 3 distinct unrelated individuals.

    1. Author response:

      We would like to thank the editor and reviewers for their constructive and thoughtful feedback. We appreciate the reviewers' assessment that our work addresses a fundamentally important research question through a novel approach. We are also glad the reviewers found our data to be rigorously analysed, and that they valued our focus on the whole cortex rather than localised regions or electrodes. We are encouraged by the overall assessment of our work and welcome the suggestions for improving the manuscript. Below, we summarise how we plan to address the reviewers' comments in our revision:

      Analyses

      - We will include quantitative results for Fig. 1E that describe the distribution of the optimal alpha values across sensors and participants.

      - For Figures 2 and 3 we will include analyses of individual participants in the Appendix.

      - We will provide more detailed descriptions on how the timing of peak saccade curvature relates to saccade onset and to the identified optimal alpha value.

      Presentation of Methods and Results

      - We will phrase our claims and conclusions more carefully and nuanced throughout, ensuring direct coverage by the data and analyses.

      - We will revise the currently complex sections of the Methods and Results to improve clarity and readability.

      - We will be more explicit about how the data for scene onset were selected.

      Revision of the Discussion

      - We will extend the Discussion section to address possible mechanisms linking the timing of peak saccade curvature and ERF initiation. We will also provide a more thorough discussion of existing and more recent literature on the topic.

      - We will emphasise the main takeaway of the study: the observation that saccade-related processes are more important to the M100 than previously thought, and, reversely, that this component may be less directly related to fixation-locked responses. We will also present our observation of peak saccade curvature as a starting point for future research, as it was not intended as conclusive mechanistic insight into how and why this process relates to early cortical responses.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript seeks to make use of information about Ct values from PCR testing of mosquito pools for West Nile virus infection to make inferences about mosquito prevalence and West Nile risk. It does so through analysis of empirical data and simulated data with a realistic agent-based model.

      Strengths:

      This work is conceptually innovative for mosquito-borne viruses, building on ideas developed primarily during work on SARS-CoV-2. Exploring this topic is worthwhile regardless of the outcome. The use of data, testing in multiple labs, and the complementarity of modeling and empirical data analysis are all strengths of the approach.

      Weaknesses:

      Some of the primary weaknesses include a dependence of the results on relatively narrow model assumptions, and a lack of compelling improvement over existing methods. None of these weaknesses are fatal flaws; they are modest weaknesses that limit the potential of or excitement about the method.

      Thank you for the comment

      Reviewer #2 (Public review):

      Summary:

      The authors extend their previous population-based Ct-value framework for inferring community epidemic trajectories from human infections to vector infections, using mosquitoes as vectors for West Nile virus. They use agent-based modelling to distinguish virus-positive detections arising from non-active infection states from those reflecting active infections, and then apply this framework to mosquito surveillance data from Colorado and Texas.

      Overall, this is a well-designed and carefully evaluated study. The manuscript proposes a feasible and potentially valuable framework for vector infection surveillance. The findings are supported by both mechanistic agent-based simulations and applications to real-world mosquito surveillance data, which strengthens the biological plausibility and practical relevance of the proposed approach.

      Strengths:

      A major strength of the study is its clear methodological extension from human infection surveillance to vector infection surveillance. The agent-based modelling framework provides a useful basis for distinguishing active infections from virus-positive detections that may reflect non-active infection states. The application to surveillance data from two different geographic settings further supports the feasibility of the framework. Overall, the study is carefully designed, and the model schematic and main analyses are generally clear.

      Weaknesses:

      (1) It would be helpful if the authors could provide plots showing variation across locations and over time. This would further support the claim made in the paragraph at lines 101-107.

      Thank you for the comment. Our supplementary Material Figures S5 and S6 already included these visualisations. However, we note that these were not referenced in the manuscript. We have now referenced these within the lines:

      “First, the variation we observe is consistent across five trapping seasons and two states (Figures S5 and S6).”

      (2) Figure 2: The model schematic is clear in terms of workflow, but it would benefit from more information on model parameterization. In particular, it would be helpful to clarify which parameters or migration rates were estimated from the data and which were assumed based on prior literature.

      Thank you for the comment. All the parameters are used from the literature and recorded in the Supplementary Material. However, we have now added a note in the caption of Figure 2, referencing the Supplementary Material as below:

      “Overall structure of the agent-based model (parameters were derived from the literature; see Supplementary Material S2, S3 and S4)”

      (3) Figure 4: I wonder whether the authors examined how changes in the proportion of mosquitoes with static viral-kinetics trajectories would affect the observed bimodal distribution. Relatedly, it would be useful to know whether there is a threshold proportion at which the method becomes less able to distinguish active from static viral-kinetics patterns.

      Thank you for the comment. We have conducted this in analysis and have already included the relevant figures in the Supplementary Material, In particular, Figures S14 (in Section S7) and S22. We have referenced Figure S14 where we discuss the proportion of mosquitoes with static viral-kinetics trajectories that would affect the observed bimodal distribution. However, we had not included a reference to Figure S22, where we illustrate the proportions at which the method becomes less able to distinguish active from static viral-kinetics patterns. We have now included this reference in the same line.

      “We found that the simulated pooled Ct values aligned well with the observed data when the percentage viral load inherited from birds was 100% and the probability of a productive or non-productive infection in the mosquitoes was 0.5, capturing the bimodal distribution of low Ct values (from productively infected mosquitoes) and high Ct values (from non-productively infected mosquitoes) (Figure 4 (B), (C) & (D); Supplementary Material S7.1, Figure S14 for Ct distributions of viral inheritance probability vs. model change probability and Figure S22 for the accuracy and confidence-interval coverage across different productive infection proportions).”

      Conclusion:

      Overall, the evidence is reasonably strong for demonstrating the feasibility and biological plausibility of the proposed framework. Some conclusions would be further strengthened by additional sensitivity analyses on key assumptions, especially the proportion of static viral-kinetics trajectories and spatial-temporal heterogeneity across surveillance sites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for th e authors):

      (1) ll 70-73 - There are a lot of ideas in this sentence. It would be useful to support this with a schematic figure or something like that, which illustrates the conceptual predictions made here. Such a step is necessary given the novelty of what is being explored here.

      This portion of the introduction has been simplified to better introduce the core observations that formed our hypothesis:

      “The substantial variation in viral quantities observed during cross-sectional entomological surveillance suggests more complex vector/virus interactions, and the precedence set by population SARS-CoV-2 testing in humans suggests that Ct value data for WNV in mosquitoes could inform metrics of disease risk to humans. However, there are substantial differences in the epidemiology and biology of WNV infection in mosquitoes with SARS-CoV-2 in humans.”

      We have chosen not to include an additional schematic figure, as the core idea (population viral loads reflect the convolution of infection incidence and within-host viral kinetics) is illustrated in the referenced literature, and the main message of the current manuscript is the explore how this phenomenon is observed in arbovirus vector surveillance.

      (2) ll 134-137 - Mosquito species is another factor that could result in wide variation in Ct values due to differences in vector competence and infection kinetics. Given that 2-4 mosquito species are present in these pools with unknown frequencies, this seems like a potentially major source of unexplained variation.

      Importantly, Culex pipiens and Culex tarsalis mosquitoes are separated prior to testing for WNV. We have clarified in the legend for Figure 1 that Ct values are presented from pools of either Culex pipiens/restuans/salinarus or Culex tarsalis.

      Additionally, we have added the following text to the Materials and Methods section:

      “For identification purposes, Cx. Pipiens species mosquitoes are not separated from the Cx. Salinarius or Cx. Restuans, which are nearly identical morphological. However, Cx. Pipiens is far more abundant than either Cx. Salinarius or Cx. Restuans in Nebraska.”

      We also show in Figure S22, S23, and S24 that we observe similar variation in Ct values across mosquito species, location, and epi week, and thus we do not think that differences between species play a major impact in our findings.

      (3) ll 173-175 - I believe that this is a consequence of the trapping method. Could this please be spelt out a bit more?

      Indeed, all of the data in this manuscript were derived from mosquitoes collected in CDC Light Traps that attract host-seeking mosquitoes (i.e mosquitoes looking for a bloodmeal). We are not considering vertical transmission in our model as it has been reported to occur infrequently in laboratory studies. Therefore, WNV-positive mosquitoes collected in CDC Light Traps have been exposed to WNV through a previous blood meal from a bird. We have clarified the text to include this explanation:

      “The mosquito pool Ct value data in this study come from specimens collected using CDC Light Traps that are baited with CO2, specifically targeting host-seeking mosquitoes. Vertical transmission is not factored into our model, thus, for a WNV-positive mosquito to be captured in the pool, it must have already obtained one blood meal from an infected bird and be seeking its next blood meal, which introduces a delay between infection and being captured.”

      (4) ll 177-178 - Doesn't the temporal trend in Ct values primarily reflect temporal changes in mosquito infection prevalence?

      Thank you for the comment. We agree that the temporal changes in mosquito infection prevalence is the main factor influencing the distribution of the Ct values in pools, as the time-since-infection distribution of trapped mosquitoes does not vary sufficiently to lead to trapping mosquitoes at very different points in their viral kinetics trajectory. Our intended point from this sentence was, given the prevalence and pool size, the variation in the viral load of the infected mosquitoes does not vary in time as all infected mosquito are trapped after they have reached a constant high-viral load level. We have now revised this sentence to reflect this.

      “As the infected mosquitoes progress from increasing viral load to a high set-point viral load, temporal trends in pooled Ct values primarily reflect time-varying infection prevalence and the number of infected mosquitoes in each pool. The remaining non-temporal variation in pooled Ct values reflect individual-level variation in mosquito set-point viral loads.”

      (5) Section 2.2 - It would seem that the assumed viral kinetics in birds would be important to this line of reasoning, given that that determines initial viral load ingested by mosquitoes. I am unclear on what was assumed in the model regarding viral kinetics in birds.

      Thank you for the comment. We have discussed the viral kinetics of the birds in detail in Section 5.3 and Supplementary Material S3. However, we agree that we have not explicitly mentioned this in Section 2.2. Therefore, we have added a reference to these sections in the following paragraph:

      “The model assumes that the mosquito's initial viral load is proportional to the infector bird's viral load (see Section 5.3 and Supplementary Material S3 for further details on the bird viral kinetics model).”

      (6) ll 194-202 - Whilst you have shown that this hypothesis leads to predictions that are consistent with the data, this is a relatively narrow hypothesis, and others are neither discussed nor refuted.

      There are two features of the data which we discuss. First, the substantial variation in Ct values across pools. This is described in detail in Section 2.1. The second observation is the bimodal pattern, which L192-202 refers to. While we agree that we have not modelled alternative hypotheses, our point is that the distribution of pooled Ct values is bimodal, and capturing some mosquitoes with very low viral loads is the most plausible explanation for the mode at high Ct values. However, we contend that this is actually a fairly broad hypothesis, as there are many plausible mechanisms generating mosquito infections with low viral loads, which we already discuss (discussion section beginning “This could be explained by a variety of factors…”). No changes have been made to the manuscript.

      (7) Section 2.3, first paragraph - The problem with this approach is that these simulations depend on a number of assumptions and parameter settings that are not estimated as part of the model fitting process. Thus, the model is very narrow and contingent on these narrow and not compellingly justified assumptions.

      Thank you for the comment. While we agree with the reviewer that this is a potential limitation of our study, we have discussed this in detail in the discussion. As mentioned in the manuscript “the main objective of this study was not to formally fit the multi-scale agent-based model to the data, but rather to understand how individual-level viral kinetics in mosquitoes are reflected in pooled surveillance data, and to demonstrate the use of pooled Ct values in estimating WNV infection prevalence”, we believe the assumptions and model are sufficient to address the research objectives. Furthermore, the fact that simulated Ct value distributions from the ABM can be used directly with the prevalence estimation method to give similar estimates to the existing PooledInfRate package supports the validity of our assumptions, though we agree that this does not necessarily mean all of our assumptions are correct, nor that our model is generalisable to other settings. No changes have been made to the manuscript.

      (8) Section 2.3, second paragraph - So the newly proposed method using Ct values does no better than the existing method using binary data?

      Thank you for the comment. We agree with the reviewer that our method and the existing PooledInfRate package perform similarly at the estimated prevalence levels of WNV. However, the Ct-based method, as we have discussed and shown, is robust at all prevalence levels where the binary-only method fails, and our method can distinguish the productive and non-productive prevalence.. Thus, while the prevalence estimates are similar under both methods for the current dataset, the novelty lies in the ability to reconstruct prevalence using the data in an entirely different way, and the proof-of-concept for how Ct values may harbour more biological information than treating pools as positive/negative. We believe that these points are sufficiently discussed throughout the manuscript. We have not made any changes to the manuscript.

      (9) ll 252-254 - This may only be true because the simulation model and the inference model are identical. If the inference model were misspecified (due, for example, to incorrect assumptions about kinetics, etc), this result would likely weaken.

      Thank you for the comment. The difference in robustness between the binary-only method and Ct-based method is not a feature of the method, but rather of how the data is used. At higher prevalence, all pools are likely to have at least one positive mosquito in them, and thus all pools will be positive, removing all information to discriminate between different prevalence levels. In contrast, the Ct-based method is able to still discriminate between prevalence levels even when all of the pools are positive, as there is still information based on whether the positive pools have low or high Ct values. No changes have been made to the manuscript.

      (10) ll 272-273 - Again, this is highly dependent on built-in model assumptions.

      Thank you for the comment. We agree that the performance of a model can depend on the underlying assumptions and the structure of the model. This is true in general for any model-based inference technique (see White, 1982, for example). Therefore, our simulations, results and interpretations are intended to be evaluated under the model structures and underlying assumptions we have used throughout the manuscript. However, to be explicit, we have now added this line at the end of the paragraph that included the sentence.

      “These results are based on the model structure and the underlying assumptions we used and they may be affected by model misspecification, including incorrect assumptions (see White, 1982, for example).”

      (11) ll 273-281 - Can this be done with pooled data only, or does it require individual mosquito Ct values? The latter would seem to be less practical to obtain in real-world applications.

      Thank you for the comment. As we have cited the related work for SARS-CoV-2, in theory, these methods are applicable when individual Ct values are present. Both pooled data and individual-level data will work, but using pooled data requires the pooling and dilution process to be modelled explicitly. However, as the reviewer mentions, for mosquito surveillance, this is not a practical approach as mosquitoes are always pooled prior to testing to reduce effort and costs. No changes have been made to the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) The supplementary figures do not appear to be ordered according to their first mention in the manuscript, which makes them somewhat harder to follow.

      Thank you for the helpful comment. We have now made sufficient changes to the Supplementary Material and updated the references in the manuscript. Where possible, the supplementary figures and sections are now numbered and presented in the order of their first mention in the manuscript.

      (2) Lines 71-73: This sentence is somewhat vague, and I was not fully clear on the intended message. The authors may wish to revise it for clarity.

      Thank you for the comment. Similar comments have been made by Reviewer #1. We have revised this sentence for clarity.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a multi-modal study of associative learning and memory in humans, that combines scalp EEG, pupillometry and behavioral analysis to explore the construct of mnemonic prediction errors (MPEs), in terms of their relationship to attention and cognitive control. Across two pooled studies, participants performed associative memory tasks in which they learned the relationship between a cue word (action verb) and subsequent picture (animate or inanimate) with a strong vs. weak (4 or 1 repetitions) encoding manipulation. At test, participants were encouraged to generate a prediction following the cue word to determine whether the subsequently presented picture was a match or mismatch.

      The timecourse of pupillary responses during match decisions were decomposed using temporal principal components analysis, which identified 6 distinct and overlapping processes. Some of the components (PC3/PC4) exhibited sensitivity to both the strength and mismatch conditions, as well as behavior (both RT and accuracy) and retrieval success on the subsequent trial. Furthermore, relationships were also observed between pupillary responses (specifically for PC4) and both frontal theta and posterior alpha power measures obtained from scalp EEG in Experiment 2, as well as for frontal theta and subsequent learning from mismatch stimuli (assessed using subsequent memory findings from a surprise recognition test). The authors suggest the findings indicate that MPEs elicit changes in attention, arousal and cognitive control which impact subsequent learning.

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPEs.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also presents a clear limitation and challenge for interpretation. It is likely that readers, even those that are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain understanding of how they all fit together.

      Indeed, it seems like there many threads running together in the paper which make it challenging to find the through-line of the key findings. The authors do address some of the key questions motivating the paper in the Introduction, but the results are somewhat ambiguous with regard to the primary question of the study as to whether the detection of MPEs leads to interaction among cognitive control, attention, and arousal. To their credit, the authors tackle this question through both cross-correlation and formal mediation analyses, and summarize these in diagrammatic figures (Figure 3, Figure 6). Yet it is not resolved whether the results represent a clear answer pointing to independence, or rather a lack of statistical power, or ill-resolved formulation of the mediational relationship. In particular, the cross-correlation suggests that posterior alpha suppression in response to MPEs does precede frontal theta, yet this indirect relationship does not explain the variation in trial-by-trial RTs on mismatches. This suggests a potential model misspecification.

      In addition to the primary interaction issue mentioned above (between cognitive control, attention & arousal), the Introduction lays out a number of claims:

      (1) That pupil size will be more sensitive to strong than weak MPEs.

      (2) That MPE-linked increases in attention (indexed with posterior alpha suppression) and arousal (indexed with pupil size) will be linked to learning.

      (3) That MPE learning will vary as a function of prediction strength.

      Given the focus on learning, it is somewhat surprising that learning is not included in the mediation models. As the authors indicate in the Discussion, the use of trial-by-trial RT variation to drive the mediation model might be problematic, given that the RTs are sensitive to a range of factors beyond mnemonic prediction strength and also are under competing pressures (longer for mismatches than matches, due to surprise-linked slowing, but also faster following stronger rather than weaker mnemonic predictions). Thus, an alternative possibility might be to use trial-by-trial recognition of mismatches as the outcome variable in mediation models rather than trial-by-trial RT as the independent variable.

      A large component of the results (Sections 2 and 3) is devoted to analyses of cue-linked pupil and EEG processes that putatively reflect mnemonic predictions (i.e., occurring before picture probes are presented and match/mismatch detection, i.e., MPEs occur). Yet these Results and the subsequent pupillary PCA components (PC1 and PC5) that are elicited are not well-integrated with the primary themes of the paper or the causal hypotheses. One finding that does seem to figure prominently (in that it is mentioned in Abstract, Introduction & Discussion) relates to the amount of attention allocated to the mnemonic prediction generation. Yet this finding is not well emphasized in the Results themselves. Possibly it refers to the negative relationship between posterior alpha during memory retrieval and the magnitude of pupillary PC3 component, described in Section 3. But it was quite challenging to identify amongst the wealth of results described in this Section as well as the others. More generally, the large amount of findings described across all four lengthy Results sections makes it challenging for readers to discern what are the key ones that the authors would like to highlight.

      It is recommended that the authors do another pass through the paper to better highlight the most critical findings that they want to emphasize or which are most interpretable from a mechanistic and causal flow perspective and then de-emphasize or move other findings to the Supplemental Materials. Although the authors are to be commended for such a rigorous and comprehensive set of analyses, there are so many of them and findings, that the key points get buried and the reader needs to struggle potentially unnecessarily to identify the key take-away points.

      We thank Reviewer 1 for the helpful feedback on how the manuscript can be further strengthened. We recognize that the rich set of findings can overwhelm the reader, resulting in difficulty discerning the main take aways about the effects of mnemonic prediction errors. We particularly appreciate Reviewer 1’s encouragement to restructure the manuscript so as to focus on the findings reported in Sections 1 and 4 of the original revision; the current revision now focuses on these key observations.

      As part of this restructuring, Reviewer 1 also proposed moving the content from Sections 2 and 3 of the original revision to the Supplement. We agree with the Reviewer that the questions addressed in these sections on retrieval-related processes are not the main focus of the paper, but that they are informative in their own right. To avoid their getting lost in the Supplement, we decided that these results would be better served in a separate manuscript and thus we have removed them entirely. 

      We acknowledge in the revised manuscript that the mediation and cross-correlation analyses were exploratory and that these specific analyses may not be well powered in the current experiments. With respect to Reviewer 1’s concerns about the specification of the mediation models, we were motivated to test whether MPEs trigger an increase in cognitive control that in turn, triggers an increase in attention and/or arousal (Fig. 4a); as such, we designed the model to assess whether, on strong MPE trials, frontal theta mediates the relationship between prediction strength and attention/arousal. As noted in the manuscript and raised by Reviewer 1, mismatch RT here is an imperfect measure of trial-level prediction strength. Future experiments that selectively elicit strong MPEs and have a more controlled measure of trial-level prediction strength may be better equipped to address these questions about interactions between control, attention, and arousal. We agree with Reviewer 1 that models assessing subsequent memory as an outcome would be desirable. However, given that (a) we did not find strong evidence for interactions at the time of a strong MPE and (b) we only observed a relationship between frontal theta and subsequent memory (but not posterior alpha or pupil), subsequent memory mediation models do not appear to be well justified. Altogether, these findings illuminate open avenues for future research.

      Reviewer #2 (Public Review):

      Summary:

      The authors studied cognitive control and attention in response to mnemonic prediction errors (MPEs): situations in which the external reality violates internal memory-based predictions. The behavioral task first established strong versus weak predictions, and then either confirmed or violated these predictions. The authors examined markers of cognitive control (frontal theta) and attention (posterior alpha suppression, pupil response) while strong and weak predictions were confirmed or violated. They found increased cognitive control (frontal theta) for strong MPEs, which correlated with subsequent memory. Markers of attention (alpha suppression, pupil response) also accompanied strong MPEs but did not correlate with subsequent memory.

      Pupil response was investigated using an interesting approach that decomposes the response into different components, finding that different components respond earlier or later and show different correlations with MPEs and their strength. The authors also investigated how EEG, reaction time, and pupil responses correlated with one another, providing further insight into the mechanism underlying the response to MPEs. Together, the study points toward multiple control and attention mechanisms involved in MPE response and memory.

      Strengths:

      The study has a clear behavioral paradigm with multiple measures — behavioral, EEG, and pupillometry — that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      Weaknesses:

      The methods are rigorous, and the data support the claims. The weaknesses are minor and are offered here as avenues for future research.

      (4) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. Thus, the specificity of this relationship is untested.

      (5) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it’s not a fair comparison and could change participants’ overall criterion for old/new decisions. Future research could address this, e.g., by testing weak matches or potentially using a between-subject design.

      We appreciate Reviewer 2’s helpful feedback during the review process and encouraging comments on the strengths of the manuscript. We agree and note in the revision that it would be illuminating for future studies to contrast memory for events that violate and confirm mnemonic predictions.

      Comments on revised version

      The authors addressed all my concerns. I appreciate the authors’ thoughtful and detailed response.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It was challenging to read the paper with key findings happening at the time of the MPE presented first, and then to go “backwards in time” to examine process that occurred at preceding time periods (i.e., prior to probe presentation), and then again forward in time to examine learning related processes.

      (1) Restructuring the Results

      In this regard two distinct recommendations are made:

      - Move Sections 2 and 3 to Supplemental Materials, to maintain the focus on the key findings related to detection of MPEs and their effects on subsequent learning.

      - An alternative structure would be to present the cue-locked findings first, which relate to mnemonic predictions, then to those that occur when those predictions get violated, and finally to the learning process that occur following MPEs and which can be detected with subsequent recognition tests.

      We thank the Reviewer for this encouragement to restructure the manuscript. To address concerns regarding the density of the paper and the cohesiveness of the findings, we removed the content that was in Sections 2 and 3 of the original revision. We will publish those results in a separate paper with additional analyses to more directly address previously raised questions regarding the mechanisms indexed by frontal theta and posterior alpha at the time of retrieval.  

      (2) Integration of the summary figures

      At the minimum, a recommendation would be to better integrate Figure 6, and maybe various versions of Figure 3, earlier into the text, preferably even in the Introduction, and then repeatedly refer to them throughout the results. For example, in Figure 6, linking the leftmost panel of the figure to Sections 2 and 3 is critical, and the righthand panel to Sections 1 with the rightmost part related to learning explicitly linked to Section 4.

      We moved the summary figure up to now be Figure 1; we reference this figure in the Introduction and throughout the manuscript; and we additionally make reference in the figure to the association between frontal theta and PC3 at the time of a strong MPE as well as the cross-correlation outcome. We hope these modifications further aid the reader in identifying the main findings of the manuscript.

      (3) Specification of the mediation models

      Additionally, for the mediation models it is quite unclear why mismatch RT is treated as the index of MPE magnitude, as this seems to be where the problem may lie in model fitting. Why not think of this as an outcome variable (since would seem to be a causal outcome of the underlying functional processes elicited by mismatch detection)?

      We thank the reviewer for these thoughtful comments. Our goal in designing the mediation models in Fig. 4a-c was to test our hypothesis that strong MPEs trigger an increase in cognitive control that, in turn, triggers an increase in attention and/or arousal; thus, attention/arousal should be the outcome and cognitive control should be the mediator. Given that increases in attention and cognitive control were selectively observed for strong and not weak MPEs, the models were restricted to strong MPEs. We therefore needed a trial-level measure of MPE magnitude to determine whether stronger MPEs in the strong condition elicit greater attention/arousal by engaging more cognitive control. As such, we decided to use mismatch RT as a proxy measure of MPE magnitude; we acknowledge in the text that this measure is imperfect. While a model with mismatch RT as the outcome variable is possible, the relative timing of the attention/arousal effects (which largely occur following responses) would render interpretation to be more challenging. Had stronger evidence of indirect effects emerged in Figs. 4b-c, we could have conducted model comparison with RT as an outcome. We acknowledge in the text that these analyses were exploratory and characterization of potential indirect effects will require more data and a more precise measure of MPE magnitude.

      Alternatively, examining trial-by-trial recognition memory of mismatches as the relevant outcome variable would also seem to capture the functional process of interest. In this regard, have the authors examined whether trial-by-trial RT on mismatches predicts subsequent recognition of these items? If this direct relationship holds, it could be a target for mediation analyses in itself.

      We agree with the reviewer that in theory, a mediation model predicting subsequent memory would be a desirable test of an integrated model of the mechanisms underlying MPE-driven learning. However, the mediation analyses conducted to address the functional relationships at the time of a prediction error (Fig. 4) are not well-powered to begin with; this limitation is raised in the Results and the Discussion. A mediation model predicting subsequent memory would be similarly underpowered; given that there is not strong evidence for indirect effects at the time of a strong MPE, and that only frontal theta – and not posterior alpha nor pupil – predicts subsequent memory, we think that such a mediation model is not well justified to include. Such a model should be more directly tested by well-powered designs in future studies.

      (4) Hippocampal theta in the Introduction

      The Introduction discusses hippocampal theta as well as frontal theta yet also makes clear that the former is not really well-detected or analyzed using scalp EEG. Consequently, a recommendation would be to remove this paragraph from the Introduction, since it can be misleading and a “red herring” for the reader, and instead only bring up this point in the Discussion section, as a pointer to the need for future research using methods that may be more sensitive to hippocampal interactions with PFC regions.

      We appreciate this point and moved discussion of the hippocampus from the Introduction to the Discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction scheme for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      We thank the reviewer for all these positive comments.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, the description of the methods and results should be clarified as well. Accordingly, the authors should revise the manuscript to address the following items thoroughly within the revised manuscript (not merely in the response letter) with the additional references mentioned below included in the revised discussion:

      (1) In several places, e.g., on line 77 (p.2), the authors appear to suggest that "implicit-solvent representation" is the origin of the deficiency in commonly utilized coarse-grained potentials that this study is aiming to rectify. But desolvation barriers and solvent-separated minima are also features of implicit-solvent representations; they are just features that should be incorporated in more accurate implicit-solvent potentials. This point is stated quite clearly and accurately in the Abstract (p.1) but not consistently in the rest of the text. The authors should check the entire text carefully to ensure that a coherent, accurate perspective is presented.

      We thank the reviewer for pointing out this important issue. We agree that implicit-solvent representation itself is not the origin of the deficiency. Our intention is to incorporate desolvation-inspired effective terms within an implicit-solvent coarse-grained framework, because many commonly used implicit-solvent potentials do not directly account for the desolvation features, such as the desolvation barrier and the solvent-separated potential well.

      We have revised the Abstract, Introduction, Results, and Discussion to make this distinction consistent throughout the manuscript. The revised text now emphasizes that the model remains an implicit-solvent CG model, but contains additional effective terms inspired by desolvation features observed in all-atom PMFs.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model with added desolvation terms:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 2, lines 80–81) The Introduction retains the implicit-solvent description of existing residue-level CG models:

      "Despite these advances, most residue-level CG models rely on implicit solvent representations, in which individual water molecules are not explicitly represented."

      (3) (page 2, lines 87–89) The specific limitation is identified as the absence of a direct account of the multi-step desolvation process:

      "More importantly, conventional implicit-solvent CG models used for LLPS usually do not directly account for the multi-step desolvation process that accompanies the transition from a dilute solution to a dense condensate."

      (4) (page 4, lines 172–174) The added terms are described as part of a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (5) (page 15, lines 521–522) The Discussion restates the implicit-solvent model framework:

      "To address this challenge, we developed a desolvation-aware implicit-solvent CG framework that incorporates desolvation barrier and solvent-separated terms into the pairwise potential."

      (2) In the discussion of the importance of desolvation barriers and solvent-separated minima in the Introduction (pp.1-3), connections should be drawn to recent works that utilize these PMF features to rationalize hydrostatic pressure (P)-modulated effects on biomolecular LLPS, including the P-dependent reentrant phase separation of alpha elastin; see Cinar et al. (2019) Chem Eur J 25:13049 (https://chemistryeurope.onlinelibrary.wiley.com/doi/full/10.1002/chem.201902210) and references therein, especially discussions around Figures 10, 11 & 13 in this reference.

      We thank the reviewer for bringing this literature to our attention. We agree that pressure-modulated LLPS provides important context for the physical relevance of desolvation barriers and solvent-separated minima. We have therefore expanded the Introduction and Discussion to connect our model to prior work on hydrostaticpressure-dependent condensate behavior, including pressure-dependent reentrant phase separation of alpha-elastin.

      Corresponding changes:

      (1) (page 3, lines 93–95) The Introduction connects the PMF features to hydrostaticpressure-dependent LLPS:

      "Related studies on hydrostatic pressure effects have further suggested that desolvation barriers and solvent-separated minima can help rationalise pressure-modulated LLPS behaviors, including the pressure-dependent reentrant phase separation of α-elastin Cinar et al. (2019, 2018)."

      (2) (page 15, lines 533–537) The Discussion states the pressure-dependent implication conservatively:

      "These findings may also provide a useful physical basis for future studies of pressure-dependent condensate behavior, as pressure-induced changes in hydration, solvent-separated states, and desolvation barriers have been proposed to contribute to pressure-modulated and reentrant LLPS Dias and Chan, 2014); Cinar et al. (2019, 2018)."

      (3) In the lower panels of Figures 2D, E (p.5), what do the differently colored small circles in the double-minimum free energy profiles represent? Does the color shading have the same meaning as that in the upper panels? If so, what do the positions of the circles on the free energy profile represent? The authors should clarify this.

      We thank the reviewer for identifying this ambiguity. The small circles in the lower panels of Figures 2D and 2E are qualitative schematic representations of residue-pair configurations along the effective pair-potential profile. Their blue and green colors distinguish the barrier-variation and solvent-separated-well cases, respectively; they are not a quantitative scale and do not encode temperature or population magnitude. The positions of the circles indicate the direct-contact, barrierregion, or solvent-separated regions, while the density of circles schematically represents the population of configurations.

      We have clarified this interpretation in the Figure 2 caption and aligned the Results text with the redistribution among direct-contact, barrier-region, and solvent-separated states.

      Corresponding changes:

      (1) (page 6, Figure 2D and E lower panels) The schematics distinguish low and high ε_b or ε_ss and use the density and position of the circles to depict populations in the direct-contact, barrier-region, and solvent-separated regions; the blue and green colors distinguish the two parameter families and are not a quantitative scale.

      (2) (page 6, Figure 2 caption) The caption defines the population encoding used in the lower panels:

      "The lower panels schematically illustrate how changes in ε<sub>b</sub> and ε<sub>ss</sub> alter the distribution of residue-pair configurations. The small circles indicate schematic populations of residue-pair configurations along the potential profile, with denser circles representing a higher population."

      (3) (page 5, lines 210–213) The Results text connects the schematics to redistribution among the three residue-pair states:

      "These opposing effects suggest that the desolvation potential regulates macroscopic phase behavior by redistributing residue-pair configurations between direct-contact, barrier-region, and solvent-separated states (lower panels of Figure 2D and E)."

      (4) The discussion regarding entropy and enthalpy around Figure 2 is quite confusing as it stands. What do the authors mean exactly by the association of entropy or enthalpy with the desolvation barrier of the solvent-separated minimum? Are they referring to conformational entropy?

      We thank the reviewer for pointing out this ambiguity. We agree that our original wording around entropy and enthalpy could be misleading, because it might imply a rigorous thermodynamic decomposition of the PMF. In the revised manuscript, we have therefore clarified that the effect of the desolvation barrier refers to an entropyrelated configurational restriction of residue-pair configurations near the barrier region, rather than the overall conformational entropy of the entire chain. We also replaced the previous "enthalpic stabilization" wording with "effective free-energy stabilization" to avoid implying that the solvent-separated minimum is treated as a purely enthalpic contribution.

      Corresponding changes:

      (1) (page 5, lines 199–201) The barrier effect is described in terms of the sampled residue-pair population:

      "Analysis of residue-residue radial distribution functions showed that higher ε<sub>b</sub> suppresses the population of configurations near the barrier region (Figure 2—figure Supplement 1C)."

      (2) (page 5, lines 201–202) The entropy-related statement is restricted to configurational sampling near the barrier:

      "This reduction in the statistical weight of barrier-region configurations can be interpreted as an entropy-related configurational restriction and thus disfavors phase separation."

      (3) (page 5, lines 209–210) The solvent-separated minimum is described as an effective free-energy contribution:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (4) (page 6, Figure 2D and E lower panels) The schematic headings are "Barrier-mediated Restriction" and "Solvent-separated Stabilization", avoiding a strict entropy-enthalpy decomposition.

      (5) Do the authors assume that the PMF (effective implicit-solvent potential) is a purely enthalpic term? It appears to be the authors' assumption. If so, the assumption has to be stated clearly in their discussion of "entropy" vs "enthalpy" around Figure 2.

      We thank the reviewer for raising this important point. We do not assume that the PMF obtained from all-atom simulations is a purely enthalpic term. The PMF is a free-energy profile that contains enthalpic and entropic contributions. The current manuscript defines the PMF as −k<sub>B</sub> T lnP(r), uses the all-atom PMFs to motivate a nonbonded effective coarse-grained potential, and describes the solvent-separated minimum as providing effective free-energy stabilization. We do not perform a rigorous enthalpy-entropy decomposition, and the revised wording avoids implying such a decomposition.

      Corresponding changes:

      (1) (page 16, lines 594–595) The Methods define the PMF as a free-energy profile obtained from the radial probability density:

      "The potential of mean force (PMF) was computed as PMF(r) = −k<sub>B</sub>T ln P(r), where P(r) is the radial probability density obtained from the production trajectory."

      (2) (page 5, Figure 1 caption) The CG interaction is labeled as an effective potential rather than as an enthalpic PMF decomposition:

      "Pairwise effective potential incorporating desolvation-inspired terms. Different curves correspond to different desolvation parameters."

      (3) (page 4, lines 172–174) The parameters are described as shaping an effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 5, lines 209–210) The solvent-separated contribution is described using free-energy language:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (6) Closely related to points 3-5 above, it should be stated clearly that the "temperature" used in the authors' simulations does not represent experimental temperature if the authors are using purely enthalpic effective potentials because PMFs are in fact temperature-dependent. This clarification is necessary to avoid misunderstanding. In this regard, it should be noted that temperature-dependent effective interactions have been used for modeling biomolecular condensates in analytical theory (Lin, Song, Forman-Kay & Chan, J Mol Liq 2017, already in the citation list) as well as in coarse-grained molecular dynamics simulations [Dignon et al. (2019) ACS Cent Sci 5:821-830 (https://pubs.acs.org/doi/10.1021/acscentsci.9b00102); Chakravarti & Joseph (2025) Protein Sci 34:e70284 (https://onlinelibrary.wiley.com/doi/10.1002/pro.70284)]. The latter two studies, not cited currently, are particularly relevant and thus should be cited because the authors may wish to incorporate temperature-dependent features in their ongoing or future effort in constructing a more comprehensive coarse-grained interaction scheme for biomolecular LLPS simulation.

      We agree with the reviewer that the simulation temperature should be interpreted carefully. In the present simulations, the effective potential is temperature-independent within each chosen parameter set. Therefore, the reduced temperature primarily serves as a model temperature controlling the relative strength of thermal fluctuations, rather than as a direct experimental temperature. We have clarified this point in the revised manuscript and added relevant references on temperature-dependent effective interactions, which represent an important direction for future model development.

      Corresponding changes:

      (1) (page 4, lines 177–179) The manuscript states that the absolute simulation temperature is not an experimental temperature:

      "Because the effective interaction parameters used in the model are temperature-independent, the absolute simulation temperature should not be directly interpreted as an experimental temperature."

      (2) (page 4, lines 182–184) The reduced temperature is identified as a model temperature:

      "Accordingly, T<sup>*</sup> should be interpreted primarily as a model temperature that controls the relative strength of thermal fluctuations, rather than as having a direct quantitative correspondence with experimental temperature."

      (3) (page 14, lines 481–485) The FUS LC temperature comparison is framed cautiously:

      "It is worth noting that the residue-level coarse-grained models used here employ temperature-independent effective interaction parameters. As a result, the temperature values reported here cannot be interpreted as quantitatively equivalent to experimental temperatures, particularly when they deviate substantially from room-temperature conditions."

      (4) (page 15, lines 563–567; continues on page 16, lines 568–569) The Discussion identifies temperature-dependent effective interactions and a corresponding future extension:

      "In addition, effective interactions themselves can be temperature-dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025). Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (7) In tackling "entropy" vs "enthalpy", it should be noted that the temperature dependence of the effective interactions entails an entropic contribution (which is itself temperature dependent) in addition to conformational entropy. As for the effective potential with desolvation barrier and solvent-separated minimum, it should be noted that the decomposition into entropic and enthalpic contributions at the direct contact, desolvation barrier, and solvent-separated minimum can be dramatically different, see, e.g., MaCallum et al. (2007) PNAS 104:6206-6210 (https://www.pnas.org/doi/full/10.1073/pnas.0605859104) and references therein.

      We thank the reviewer for this important clarification. We agree that temperature-dependent effective interactions can contain entropic contributions beyond conformational entropy and that the balance of enthalpic and entropic contributions may differ among the direct-contact minimum, desolvation barrier, and solvent-separated minimum. The present model does not decompose the PMF into temperature-dependent enthalpic and entropic components; accordingly, we have avoided assigning those components to individual PMF features. The Discussion cites explicit-solvent PMF analyses when noting residue-pair and temperature dependence and identifies temperature-dependent desolvation parameters as an important future extension.

      Corresponding changes:

      (1) (page 15, lines 561–563) The Discussion cites explicit-solvent PMF work when noting residue-pair and temperature dependence:

      "Explicit-solvent PMF analyses have shown that desolvation barrier heights and solvent-separated minima can differ substantially among residue pairs and may also exhibit temperature dependence Cinar et al. (2019); Debiec et al. (2014); MacCallum et al. (2007)."

      (2) (page 15, lines 563–566) The manuscript states that effective interactions can themselves depend on temperature:

      "In addition, effective interactions themselves can be temperature dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025)."

      (3) (page 15, lines 566–567; continues on page 16, lines 568–569) Temperature-dependent desolvation parameters are identified as a future model extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (8) P.7, line 340: The proportionality relation follows directly from the standard FloryHuggins result T_c = T chi(T)/chi_c, thus the proportionality constant is exactly 1/chi_c. Is this the standard relation that the authors are invoking here? The authors should clarify this.

      We thank the reviewer for pointing out the missing intermediate steps. Yes, the relation we invoked is based on the standard Flory-Huggins critical condition. In the revised manuscript, we have expanded the derivation to explicitly show how the critical condition chi(T_c) = chi_c leads to the relation between chi(T_sim) - chi_c and the normalized thermal distance (T_c - T_sim)/T_sim.

      We also revised the wording to avoid presenting this as a universal law. The relation is now presented as a simulation-supported trend within the present model, rationalized by a simplified linear-response assumption between Delta R_g and the excess interaction strength.

      Corresponding changes:

      (1) (page 7, lines 263–264) The critical-condition substitution is now shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (2) (page 7, line 265) The resulting relation is written as Equation (2):

      "χ(T<sub>sim</sub>) − χ<sub>c</sub> = χ<sub>c</sub> (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (3) (page 7, lines 266–267) The fixed-chain-length assumption is stated explicitly:

      "For systems with the same chain length, χ<sub>c</sub> is a fixed constant. Thus, the deviation from the critical interaction parameter is directly related to the rescaled thermal distance (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (9) The study on dynamic consequences on pp.8-11 is interesting, but clarifications are necessary:

      (i) The vertical schematic in Figure 4A should be explained in detail in its entirety. As it stands, no explanation is provided either in the figure caption or in the text. In particular, what does "elasticity driven" refer to?

      (ii) The top snapshot in Figure 4A is labeled t_sim = 0 ns. Does it mean that the snapshot shown is the only chain configuration that the authors used to start the simulation, and that the snapshot does NOT represent the result of any time evolution, no matter how short the duration is? However, if that is the case, why is this snapshot identified with spinodal decomposition if it is not the product of a time evolution from a more homogeneous configuration?

      (iii) Related to (ii) - do the rectangular boxes shown represent the entire simulation box or just part of the box containing the polymer chains? One would imagine that if the top snapshot represents spinodal decomposition, the simulation would have been started at a more uniform distribution a short time prior? Why is this not the case?

      (iv) What precisely do the small yellow beads and black-colored springs in the zoomin image of Figure 4E represent?

      We thank the reviewer for all these inspiring comments and questions. We agree that the original Figure 4 schematic did not sufficiently explain the sequence of dynamical events and the meaning of several graphical elements. We have therefore revised both the Figure 4 caption and the Results text to make the schematic self-contained and to clarify how it relates to the quantitative analyses in Figure 4F and G.

      First, we replaced the phrase "elasticity driven" with a more precise description of "viscoelastic resistance". In the revised text, interfacial tension is described as favoring domain fusion thermodynamically, whereas transient inter-chain network connectivity within dense domains generates viscoelastic resistance to the deformation required for coalescence. This revision avoids implying that elasticity is the driving force and instead identifies it as a resistance that delays domain fusion kinetically during the plateau regime.

      Second, we clarified the meaning of t_sim = 0 ns and its relation to spinodal decomposition. The system was equilibrated at a supercritical temperature to obtain a homogeneous one-phase state and was then instantaneously quenched to the target temperature. The label t_sim = 0 ns denotes the first snapshot immediately after the quench. It is not the only initial configuration used in all simulations; the reported kinetic metrics were averaged over six independent slab simulation replicas.

      The t_sim = 0 ns snapshot is therefore described as a homogeneous but thermodynamically unstable post-quench state. Spinodal decomposition refers to the subsequent amplification of the initial density fluctuations after the quench, including the development of interconnected density fluctuations within 1-2 ns, rather than to a preceding evolution represented by the t_sim = 0 ns snapshot.

      Third, we clarified that the rectangular snapshots in Figure 4A show the entire simulation box viewed along the z-axis. The subsequent snapshots show how post-quench density fluctuations grow and reorganize into dense domains during spinodal decomposition.

      Finally, we clarified the symbols in the zoom-in schematic of Figure 4E. Yellow beads now denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts. These elements are meant to illustrate transient network connectivity and are not additional simulated particles or force-field terms.

      Corresponding changes:

      (1) (page 10, Figure 4A) The vertical schematic now labels the progression as "Thermodynamic instability", "Kinetic arrest (viscoelastic resistance)", "Domain coarsening (interfacial-tension dominated)", and "Dynamic equilibrium (chain self-diffusion)".

      (2) (page 10, Figure 4E) The plateau schematic labels the competing effects as "Interfacial Tension" and "Transient network resistance".

      (3) (page 10, Figure 4 caption) The caption defines the snapshots and the vertical schematic:

      "Upper snapshots show the simulation box along the z-axis at t<sub>sim</sub> = 0, 10, and 500 ns. The vertical schematic summarizes the dynamical progression described in the main text, from the post-quench spinodal instability to kinetic arrest, domain coarsening, and dynamic equilibrium."

      (4) (page 11, lines 368–370) The first recorded time point after the quench is defined explicitly:

      "Here, t<sub>sim</sub> = 0 ns denotes the first snapshot immediately after the temperature quench, corresponding to a homogeneous but thermodynamically unstable nonequilibrium state."

      (5) (page 10, Figure 4 caption) The yellow beads and black springs are defined:

      "In the zoom-in view, yellow beads denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts."

      (6) (page 12, lines 401–404) The Results explain the physical meaning of the transient network:

      "In the zoom-in schematic in Figure 4E, this transient network is represented by connections between residues involved in inter-chain contacts, illustrating how multivalent interactions can resist domain deformation during the plateau regime."

      (10) In discussing dynamic effects, it is useful to draw connections to related works on the effect of chain flexibility on "aging" of condensate [Biswas & Potoyan (2024) PRX 45:9222-9245 (https://journals.aps.org/prxlife/abstract/10.1103/PRXLife.2.023011)] and characterization of viscoelasticity in simulations of biomolecular condensates [Tejedor et al. (2023) J Phys Chem B 127:4441-4459 (https://pubs.acs.org/doi/10.1021/acs.jpcb.3c01292)], as the effects of desolvation can be explored further based on these prior works.

      We thank the reviewer for these important references. We have added connections to simulation studies of condensate viscoelasticity and aging. The revised manuscript now places our dynamic results in the context of transient network connectivity, chain flexibility, sticker lifetime, desolvation-associated rigidification, and viscoelastic or aging-like material behavior.

      We present these connections conservatively as relevant context and as future directions for extending the current model, rather than claiming a new universal dynamic mechanism.

      Corresponding changes:

      (1) (page 12, lines 397–399) The dynamics section cites simulation-based rheological analyses of condensate viscoelasticity:

      "Similar viscoelastic effects have recently been quantified in molecular simulations of biomolecular condensates using rheological analyses of time-dependent material properties Tejedor et al. (2023)."

      (2) (page 12, lines 399–401) The manuscript connects condensate aging to chain flexibility, sticker lifetime, and desolvation-associated rigidification:

      "Molecular simulations of condensate aging have further highlighted the roles of chain flexibility, sticker lifetime, and desolvation-associated rigidification in promoting more solid-like states Biswas and Potoyan (2024)."

      (3) (page 12, lines 423–427) The kinetic interpretation is connected to viscoelastic andaging-like behavior:

      "The sensitivity of kinetic arrest and coarsening dynamics to desolvation parameters underscores the importance of incorporating desolvation features into coarse-grained potentials for more physically plausible molecular simulations of LLPS, especially when connecting microscopic interaction lifetimes to emergent viscoelastic or ageing-like material behavior."

      (4) (page 16, lines 569–572) The Discussion identifies simulation-based rheological analysis as a future direction:

      "An additional direction would be to combine these potentials with simulation-based rheological analyses to quantify how desolvation reshapes condensate viscoelasticity, aging-like maturation, and long-time material relaxation Tejedor et al. (2023); Biswas and Potoyan (2024)."

      (11) Much of the present study is based on the original HPS formulation of Dignon et al. (2018). In this regard and also in anticipation of future development of improved interaction schemes, several issues should be stated and discussed, even if briefly:

      (i) The original HPS model has a basic shortcoming in accounting for the relative interaction strengths of, among others, arginine vs lysine residues [Das et al. (2020) PNAS 117:28795-28805 (https://www.pnas.org/doi/10.1073/pnas.2008122117)].

      (ii) Compared to 210-parameter pairwise interaction schemes, such as KH in Dignon et al. (2018) and Joseph et al. (2021), the 20-parameter interaction scheme is likely too restrictive to account for pairwise amino acid residue interactions [Wessén et al. (2022) J Phys Chem B 45:9222-9245 (https://pubs.acs.org/doi/10.1021/acs.jpcb.2c06181)].

      (iii) The height of the desolvation barrier may vary significantly for different amino acid residue pairs, see, e.g., Figure 11 of Cinar et al. (2019) mentioned above (and references therein). The authors should discuss these nuances in the revised version. They may also wish to take them into consideration in future investigations.

      We thank the reviewer for the suggestion to clarify these limitations. We have revised the Discussion to acknowledge explicitly the limitations of the 20-parameter hydropathy-scale representation relative to more flexible 210-parameter pairwise interaction schemes for describing amino-acid-pair interactions. We have also added discussion emphasizing that future desolvation-aware models should incorporate residue-pair-specific parameters for the desolvation barrier and solvent-separated potential well.

      Corresponding changes:

      (1) (page 13, lines 445–446) The scope of the averaged baseline parameterization is stated explicitly:

      "This uniform parameterization captures the generic desolvation features of the PMFs but does not resolve residue-pair-specific variations in desolvation energetics."

      (2) (page 15, lines 551–554) The Discussion identifies the limitation of the 20-parameter HPS representation, including Arg/Lys interactions:

      "In particular, HPS-type models use a 20-parameter hydropathy-scale representation, which is useful for capturing generic IDP phase behavior but is not flexible enough to resolve residue-pair-specific chemical effects, such as the distinct interaction patterns of arginine and lysine residues Das et al. (2020)."

      (3) (page 15, lines 554–557) The greater flexibility of 210-parameter pairwise schemes is described:

      "More general 210-parameter pairwise interaction schemes, such as KH-type and related residue-pair-specific models, provide greater flexibility for encoding amino acid-pair preferences and capturing sequence-specific interaction heterogeneity Dignon et al. (2018b); Joseph et al. (2021); Wessén et al. (2022)."

      (4) (page 15, lines 557–561) The limitation of using one averaged desolvation parameter set is stated:

      "Second, the present desolvation model employs a single set of averaged parameters (α<sub>b</sub>, α<sub>ss</sub>) for all residue pairs. While this simplification is effective for isolating the generic physical consequences of desolvation, it has limitations in describing the pair-specific variations in the desolvation barrier and the solvent-separated minimum."

      (5) (page 15, lines 566–567; continues on page 16, lines 568–569) Residue-specific and temperature-dependent parameters are identified as a future extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics. Strengths:

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce densephase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      Weaknesses:

      At the same time, there are several places where the manuscript would benefit from more careful framing. First, the desolvation terms are still effective coarse-grained parameters rather than a direct representation of water molecules. The language sometimes gives the impression that desolvation is being treated explicitly, whereas the model introduces desolvation-inspired effective interactions into an implicitsolvent framework.

      We thank the reviewer for the positive assessment and constructive suggestions. We agree that the desolvation terms should be described as effective coarse-grained parameters rather than explicit water molecules. We have revised the manuscript to describe the model as a desolvation-aware implicit-solvent coarse-grained framework with desolvation-inspired effective interaction terms.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 4, lines 151–153) The Results describe the added contributions as desolvation-related effective terms:

      "These observations underscore the importance of incorporating desolvation-related effective terms and exploring the effects of different desolvation strengths on the thermodynamics and kinetics of protein LLPS."

      (3) (page 4, lines 172–174) The pair interaction is described as a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 15, lines 541–543) The Discussion emphasizes that the framework retains water-mediated features within an implicit-solvent representation:

      "By retaining key water-mediated features while preserving the computational efficiency of implicit-solvent representations, this framework provides a mechanistic means to decouple overall phase-separation propensity from condensed-phase packing."

      Second, the conformational analysis is interesting, but the broader context of prior work on dilute-to-dense phase conformational reorganization of IDPs could be more clearly discussed. This would help clarify what is new in the present work, whether it is the conformational change itself, its dependence on desolvation terms, or the proposed scaling with distance from the critical point.

      We thank the reviewer for this suggestion. We agree that the conformational change itself should be placed in the context of prior work. The contribution of the present analysis is not simply the observation that IDP conformations can reorganize upon condensation. Rather, we examine how desolvation-inspired effective terms modulate dilute- and dense-phase conformations and how the conformational change correlates with thermal distance from the critical point within the present model.

      We have revised the Results section discussing Figure 3 to cite prior work and to state the interpretation of ΔR_g more clearly.

      Corresponding changes:

      (1) (page 6, lines 230–232) Prior work on conformational reorganization upon condensation is cited:

      "Previous studies have shown that IDP condensation can reorganize chain conformations by redistributing the balance between intra-chain and inter-chain interactions Wei et al. (2017); Hazra and Levy (2021); Tesei et al. (2021); von Bülow et al. (2025)."

      (2) (page 7, lines 241–242) The phase-dependent conformational response to the desolvation parameters is introduced:

      "In addition to the difference between dilute- and dense-phase conformations, varying the desolvation parameters further reveals a phase-dependent conformational response (Figure 3A–C)."

      (3) (page 7, lines 243–244) The stronger response in the dilute phase is stated directly:

      "Increasing ε<sub>b</sub> or decreasing ε<sub>ss</sub> shifts the dilute-phase R<sub>g</sub> distributions toward larger values, whereas the dense-phase R<sub>g</sub> remains comparatively insensitive to these parameter changes."

      (4) (page 7, lines 247–249) The source of the desolvation-dependent variation in ΔR_g is identified:

      "As a result, the desolvation-dependent variation in ΔR<sub>g</sub> = R<sub>g</sub><sup>dense</sup> − R<sub>g</sub><sup>dilute</sup> arises predominantly from the conformational changes of isolated chains in the dilute phase."

      (5) (page 7, lines 254–256) The observed relationship is presented as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      Third, the dynamic results are potentially useful, but the manuscript should more clearly articulate what is nontrivial beyond the expected slowing of local rearrangements by an added barrier in the potential.

      Overall, I think this is a useful and potentially important contribution.

      We thank the reviewer for this constructive comment and the positive overall assessment. We have revised the dynamics section to clarify that the nontrivial result lies in the competition between two effects: although the desolvation barrier directly slows local rearrangements, its reduction of dense-phase packing can reverse the net mobility trend at fixed temperature. At matched thermodynamic quench depth, the intrinsic slowing associated with energy-landscape roughness becomes evident. We also clarified that desolvation modulates transient kinetic arrest and domain-scale coarsening, not only local rearrangements.

      Corresponding changes:

      (1) (page 11, lines 353–357) The fixed-temperature and matched-quench-depth analyses are summarized as opposing contributions:

      "Together, the fixed-temperature and renormalized analyses in Figure 4C and D reveal two distinct and opposing contributions of desolvation to condensate dynamics. At fixed temperature, increasing ε<sub>b</sub> loosens dense-phase packing and thereby increases the measured diffusion coefficient, whereas at matched thermodynamic quench depth, the same parameter change suppresses chain mobility by roughening the microscopic energy landscape and slowing local rearrangements."

      (2) (page 11, lines 358–361) The multiscale interpretation is stated explicitly:

      "Condensate dynamics therefore emerge from a balance between density-regulated mobility and energy-landscape-regulated mobility, with macroscopic packing determining the dominant trend and microscopic barrier roughness imposing an additional kinetic modulation. This interplay highlights how desolvation reshapes condensate dynamics across multiple physical scales."

      (3) (page 12, lines 421–423) The dynamics section distinguishes the result from simple local slowing:

      "This picture shows that desolvation does more than slow down local chain rearrangements through an added barrier. It also regulates the balance between fluctuation growth, transient arrest, and domain coarsening, thereby shaping the evolution of phase-separated domains."

      Reviewer #2 (Recommendations for the authors):

      (1) The model is physically motivated and useful, but I would encourage the authors to be more precise in describing the added terms as desolvation-inspired effective interactions rather than explicit desolvation.

      We thank the reviewer for the comment and suggestion. We have revised the manuscript accordingly and describe the added terms as desolvation-inspired effective interactions within an implicit-solvent CG framework throughout the Abstract, Results, and Discussion. More detailed changes are provided in our response to the first point raised in Reviewer #2's Public Review.

      (2) The desolvation barrier is introduced as part of the equilibrium pair potential, and therefore it is expected to affect not only kinetics but also the phase boundary through changes in the configurational partition function. The manuscript would benefit from clarifying this point, since the term "barrier" may otherwise suggest a primarily kinetic role. In particular, the authors should explain whether the observed shift in T_c reflects a change in the effective pair attraction, for example, through the integrated Boltzmann weight or second virial coefficient, rather than only an entropic penalty associated with restricted configurations.

      We thank the reviewer for this important point. We agree that the desolvation barrier is part of the equilibrium pair potential and therefore affects the phase boundary through the Boltzmann-weighted sampling of residue-pair configurations, not only through kinetic slowing.

      Following this recommendation, we added a bead-level second virial coefficient analysis based on the effective pair potential. This analysis provides a pair-potentiallevel measure of the integrated effective attraction and clarifies why increasing the barrier lowers T_c, whereas stabilizing the solvent-separated minimum raises T_c.

      Corresponding changes:

      (1) (page 5, lines 203–205) The equilibrium role of the barrier is stated explicitly:

      "At the pair-potential level, the desolvation barrier modifies the equilibrium Boltzmann weight and thereby alters the integrated effective attraction, as quantified by the bead-level second virial coefficient B<sub>2</sub>."

      (2) (page 5, lines 205–207) The barrier-dependent second virial coefficient is connected to the shift in critical temperature:

      "Specifically, increasing ε<sub>b</sub> makes B<sub>2</sub>/σ<sup>3</sup> larger (Figure 2—figure Supplement 1G), indicating a weaker integrated effective attraction and providing a thermodynamic basis for the lower T<sub>c</sub><sup>*</sup>."

      (3) (page 5, lines 207–209) The solvent-separated-well trend is linked to a smaller second virial coefficient:

      "By contrast, deepening the solvent-separated well ε<sub>ss</sub> elevates T<sub>c</sub><sup>*</sup> (Figure 2E), which is associated with the enhanced population of solvent-separated configurations and a smaller B<sub>2</sub>/σ<sup>3</sup> (Figure 2—figure Supplement 1D, H)."

      (4) (page 18, lines 662–665) The Methods specify the integration range and the quantity used to compare integrated effective attraction:

      "The upper limit of integration r<sub>c</sub> is set as 3σ, which is sufficiently large to capture the full range of interactions while ensuring numerical convergence. The reduced value B<sub>2</sub>/σ<sup>3</sup> was used to compare the integrated effective attraction under different desolvation parameters."

      (5) (Figure 2—figure supplement 1G, H) The new panels report the integrated effective attraction:

      "(G, H) Bead-level second virial coefficient (B<sub>2</sub>/σ<sup>3</sup>) calculated from the effective pair potential under varying ε<sub>b</sub> at fixed ε<sub>ss</sub> = 0.02 kcal/mol (G) and varying ε<sub>ss</sub> at fixed ε<sub>b</sub> = 3.12 cal/mol (H)."

      (3) The conformational analysis in Figure 3 is interesting and potentially important. It would help to better place this result in the context of prior work showing dilute-todense phase conformational reorganization of IDPs, and to clarify what is new here beyond that broader observation.

      We thank the reviewer for the comment and suggestion. We have revised the Results section discussing Figure 3 to place dilute-to-dense conformational reorganization of IDPs in the context of previous studies and then to emphasize the specific contribution of the present work.

      The revised text clarifies that, within the present model, desolvation-inspired interactions mainly regulate chain conformations in the dilute phase, whereas dense-phase conformations remain comparatively insensitive. Detailed changes are provided in our response to the second point raised in Reviewer #2's Public Review.

      (4) The proposed scaling between ΔR_g and distance from the critical point is intriguing, but the argument relies on simplifying assumptions. I would present this more as an empirical scaling supported by a plausible theoretical argument rather than a general result.

      We thank the reviewer for this helpful suggestion. We agree that the correlation between Delta R_g and the distance from the critical point relies on simplifying assumptions and should not be presented as a general law. In the revised manuscript, we have softened the interpretation and now present this relationship as an empirical correlation supported by a simplified Flory-Huggins-based theoretical argument.

      Corresponding changes:

      (1) (page 7, lines 254–256) The relationship is described as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      (2) (page 7, lines 256–258) The interpretation is limited to an association with thermal distance from the critical point:

      "This result suggests that the conformational response to phase separation is closely associated with how far the system resides thermally from the critical point."

      (3) (page 7, lines 263–264) The critical-condition derivation is shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (4) (page 7, lines 268–272) The structural relation is explicitly introduced as a first-order linear-response approximation:

      "The thermodynamic driving force χ(T<sub>sim</sub>) – χ<sub>c</sub> can then be related to the structural observable ΔR<sub>g</sub>. Since ΔR<sub>g</sub> captures the structural transition from an intrachain-interaction-dominated state in the dilute phase to an interchain-interaction-dominated state in the dense phase, we assume, as a first-order approximation, that this conformational shift responds approximately linearly to the excess interaction strength, expressed as ΔR<sub>g</sub> ∝ [χ(T<sub>sim</sub>) − χ<sub>c</sub>]."

      (5) (page 8, lines 285–287) The unscaled relation is labeled as an empirical scaling approximation:

      "Although the complete relation in Equation (3) contains an additional T<sub>sim</sub> factor, the unscaled quantities remain strongly correlated over the simulated range. We therefore use T<sub>c</sub> − T<sub>sim</sub> ∝ ΔR<sub>g</sub> as an empirical scaling approximation."

      (5) The dynamics section would benefit from a statement of what is nontrivial, since a desolvation barrier is expected to slow local rearrangements.

      We thank the reviewer for the comment and suggestions. As described above, we have revised the dynamics section to clarify what is nontrivial beyond the expected slowing of local rearrangements by an added barrier. The revised text emphasizes that desolvation affects condensate dynamics through competing effects of macroscopic packing and microscopic energy-landscape roughness, and that it also regulates transient kinetic arrest and domain-scale coarsening. More detailed changes are provided in our response to the third point raised in Reviewer #2's Public Review.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editor and the reviewers for their time and efforts to evaluate our manuscript. We have taken into account all the comments and revised the manuscript accordingly which has considerably strengthened the message.

      Several parts of the manuscript, have been extensively rewritten to add explanations and clarify our hypothesis and claims, this also led us to add four new references.

      In addition, we have added the following new figures:

      - New part of figure 2. Figure 2E shows a platelet in a constrained clot after 4h of retraction with the fibrin cage around the platelet center still present. The actin staining of the platelet shows that radial actin fibers are present in each bulb extending to the platelet center. This observation supports our hypothesis that in each bulb an individual cytoskeletal swirling could take place resulting in the accumulation of fibrin fibers at the base of each bulb.

      - Modification of figure 3, to include the criteria used to define four categories of platelets and associated fibers in the 2D fiber retraction assay (new Fig. 3C).

      - New figure 13, illustrating the quantification of fibrin fibre compaction mediated by platelets in the 2D fiber retraction assay and the rotational movement of a fiber mass (video 9).

      - New supplementary figure 1, showing the result of a new model simulation in the absence of cytoskeleton swirling. Under this condition the fibrin fiber does not loop around the platelet bulb.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper reports a previously unrecognized mechanism by which platelets compact fibrin fibers during clot retraction. Rather than simply pulling on fibers, the authors propose that platelets generate swirling motions that wind and loop fibrin into dense structures.

      While the results are intriguing, the underlying physical mechanism remains unexplained. In particular, it is unclear how platelets generate swirling motion capable of inducing fibrin coiling, especially when suspended in 3d fibrin mesh. This raises concerns about the conclusions.

      The reviewer is right, it is difficult to imagine how platelets in a 3D fibrin mesh can accumulate fibers at the base of their extensions to form a cage-like fiber organisation around the center of the platelets. We therefore developed the 2D fibre-retraction assay, which we believe provides important insight for the coiled fiber accumulations above spread platelets in the 2D situation but also provides a framework for interpreting similar processes that may occur within a 3D clot. In response, we have placed greater emphasis on clarifying and strengthening the comparison between the potential mechanistic aspects in the 2D and 3D assays, in order to better support our proposed model (see Results, section: "Platelets, spread on a 2D surface, organize fibers above them", last paragraph). In addition, the Ideas and Speculations section of the discussion has been extensively rewritten to provide more detailed explanations about the potential mechanism leading to fibrin fiber accumulations around platelet bulbs in a 3D fibrin mesh.

      Also, does fibrin have inherent chirality or structural asymmetry that could promote coiling independently of platelet activity?

      Yes, double-stranded fibrin protofibrils have a helical twist [1]. Furthermore, a clot formed in the absence of platelets and other cellular components shows intrinsic tensile forces [2]. However, we show that inhibition of actomyosin actions prevents fibrin fiber accumulation in the 2D fibre-retraction assay providing evidence that platelet actions are necessary to observe the coiled fibers above spread platelets. This has been accentuated in the revised version and three references have been added.

      Furthermore, platelet retraction typically involves platelet aggregation rather than isolated cells, and it is unclear how fibrin coiling would proceed in clustered platelets.

      Under the in vitro fiber retraction conditions used in our study (constrained or unconstrained clots or even in the 2D assay) individual platelets are homogenously distributed within the forming clot or on the coverslip. Therefore, there are no big platelet aggregates or clusters of platelets under our experimental conditions and the results can only demonstrate how individual platelets act on fibrin fibers. This point has been emphasized in the revised version (Discussion, third paragraph).

      Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibers into compact coiled structures. Using a combination of 2D spread assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled-fibre morphologies, and platelet centred "rosette" structures present during fibre compaction. They further suggest that actomyosin-driven cytoskeletal dynamics, potentially involving rotational or swirling motion, underlie this proposed winding mechanism, analogous to DNA looping and compaction. The study addresses an important and longstanding question in thrombosis and hemostasis and offers a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength of this paper. In particular, the 2D fiber-retraction assay provides a useful model for understanding the spatio-temporal dynamics of platelet-mediated fibrin compaction, which can be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and offer valuable dynamic insight.

      Weaknesses:

      The primary weakness of this paper lies in its descriptive nature and its reliance on correlative rather than causal evidence. Several interpretations are not uniquely supported by the data presented. For example, the categorisation of fibrin accumulation in 2D assays as "fiber winding" and "fibre compaction" remains descriptive without establishing winding as a mechanism.

      When introducing the 2D fiber-retraction assay (figure 3) in the revised version, we now only mention the terms fiber accumulation and compaction to better align with the level of evidence, since wound-up fibers cannot be distinguished in this figure. The criteria to establish the four categories of platelets and associated fibers in the 2D fiber retraction assay have now been included in figure 3C.

      Nevertheless, coiled fibers above spread platelets are clearly visible in figure 4 and 8 and dynamic fiber rotations or winding-up are observed in figure 12 and video 9. These observations have been presented more cautiously, as indicative rather than definitive evidence of a winding mechanism.

      Alternative mechanisms, such as circular bundling, stacked fibers under tension, or fibrin crosslinking-induced aggregation, are neither excluded nor investigated.

      For fibrin fiber bundling, staggered or crosslinked protofilaments no platelet actions are necessary as described previously [2,3]. Since we observed a clear difference between +/- blebbistatin conditions in the 2D fiber-retraction assay, the fiber compaction we observe depends on platelet actions. Consequently, we consider these alternative mechanisms unlikely based on our data. This has been stated explicitly in the results section and discussion and three references have been added.

      Although the authors present compelling live imaging, establishing winding as a dynamic phenotype would require quantitative analyses, such as measuring angular velocities and coiling rates.

      We have incorporated quantitative measurements (new figure 13) about platelet mediated fibrin fiber compaction and angular rotation velocities to complement the observations obtained from live imaging. It is important to note, however, that angular velocities and coiling rates are likely influenced by the number of fiber–fiber contacts present at the time coiling occurs. Specifically, an increased number of contacts is expected to elevate tension within the network, thereby modulating the forces generated by platelets and, consequently, affecting both velocity and coiling dynamics.

      The use of a second fluorophore-labelled fibrin population could further strengthen evidence for rotational dynamics.

      These live videos are quite difficult to acquire because of the following reasons:

      - Small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time needed to adjust parameters for image acquisition, necessitates an arbitrary choice of the acquisition window and only one acquisition (90 min) per sample preparation is possible.

      - Furthermore, the laser-induced illumination can perturb the observed processes. We therefore use high-spatial-resolution 3D confocal time-lapse imaging, performed in photon-counting mode with very low laser excitation.

      For these reasons, the use of additional markers would be technically challenging and could perturb the delicate equilibrium and dynamics of the process under investigation.

      Similarly, the inference of rotational contractility or actomyosin "swirling", based on chiral actin organisation and blebbistatin treatment, is not sufficiently supported to conclude that platelets actively wind or loop fibrin fibers.

      Importantly, in the 2D fiber-retraction assay, we do not propose that the rotational actomyosin activity leads to a contractility of the platelets which would allow fiber retraction. Rather, we suggest that cytoskeletal actomyosin swirling (as demonstrated for nucleated cells by Bershadsky's team) can induce rotational dragging of extracellular bound fibrin fibers around the pseudonucleus of spread platelets thereby promoting accumulation of fibrin fibers (shown in figure 12C, video 9, third panel). Consistent with this interpretation, inhibition of myosin by blebbistatin prevents the accumulation of fibrin fibers above spread platelets in the 2D fibre retraction assay (Fig. 3).

      The mathematical model, while complementary and well-constructed, relies on multiple assumptions and lacks predictive validation.

      We thank the reviewer for this insightful comment and acknowledge that the proposed model relies on several important assumptions. In our view, the most significant assumption is that integrin molecules undergo rotational downstream motion as a consequence of their coupling to the swirling cytoskeleton. To assess the necessity and impact of this assumption, we provide an additional simulation performed in absence of the cytoskeletal swirling. Under this condition the fibrin fibers are not looped around the platelet bulb (this result has been added as supplementary figure 1). This analyses also provides further validation of the proposed model and underlying mechanism. At the same time, it is important to emphasize that the primary purpose of the model was to examine whether the hypothetical swirling dynamics of the cytoskeleton, together with the associated receptors, could in principle reproduce the experimentally observed fibrin organization.

      Appraisal:

      While the authors successfully document intriguing fibrin architectures and provide a compelling descriptive framework, they do not fully demonstrate a mechanistic model of active fibrin winding by platelets. The conclusions regarding platelet-driven winding and rotational dynamics are not sufficiently supported by direct or quantitative evidence. To substantiate these claims, the study would benefit from experiments that directly link platelet dynamics to fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement. As it stands, the results are suggestive but do not definitively support the proposed mechanism.

      Discussion and Impact:

      Despite these limitations, the study addresses an important question in thrombosis and hemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly for researchers interested in platelet mechanics and fibrin organisation. However, the overall impact will depend on whether the proposed mechanism can be more rigorously validated. In its current form, the study presents an interesting and thought-provoking model, but would benefit from either stronger experimental support for the proposed mechanisms or a more cautious interpretation of the findings.

      We agree that the proposed mechanism requires further validation. In the revised version we have added a new result (figure 2E) showing that radial actin filaments are present in each bulb of a platelet in a constrained clot, supporting the possibility that rotational cytoskeletal movements could take place in individual bulbs. In a new figure 13, we have also quantified fiber compaction and the angular velocity of a rotating fibrin mass observed in video 9. Furthermore, in the revised manuscript, we present a more cautious and explicitly hypothesis-driven interpretation of the mechanism. We hope that the publication of our observations will be of interest to researchers in the field of thrombosis and clot mechanics who possess the specialized tools and expertise necessary to rigorously evaluate and either substantiate or refute the proposed mechanistic model.

      Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing, and recent work has demonstrated that platelets play a critical role in generating the force required to drive the accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism, and further evidence is required to support their claims.

      We do not "propose that platelets actively 'wind up' these fibers by undergoing myosin-independent rotation" of the whole platelet, but rather of the cytoskeleton winding-up extracellular fibrin fibers attached to integrin receptors.

      Weaknesses:

      (1) The motivation to switch from the system used in Figures 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity', the differences between these assays appear to disrupt the 'cage-like' organization of fibrin around platelets shown in Figures 1 and 2 (compare images in Figure 2 with those in Figure 4). An indepth comparison of two methods is needed to support the conclusions from the 2D system.

      We agree that the cage-like fibrin organization around platelets is disrupted in the 2D fibre-retraction assay when platelets are completely spread on the coverslip before they have encountered fibrin fibers (Fig. 4). This has been explicitly stated in the revised version. However, some platelets in the 2D fiber-retraction assay form the same number of extensions as platelets in a 3D clot (Fig. 9 A, B) and are not completely spread on the glass surface. For these platelets a cage-like fibrin organisation is retained under the 2D conditions (Fig. 5 and 6). Nevertheless, the fiber density at the base of the bulbs is higher in the 2D assay than under the constrained 3D clot retraction conditions (Fig. 1C and Fig. 2), probably because in the 2D condition the fibers are less constrained and readily available for compaction.

      Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      We thank the reviewer for raising this point. We would like to clarify that Figure 2 and Figure 7 correspond to two distinct experimental setups: the constrained clot retraction assay (Figure 2) and the 2D fiber-retraction assay (Figure 7). As such, they are not directly comparable. We understand, however, that the reviewer is likely referring to the apparent differences between Figures 3–6 (lower plasma volume, higher fiber density) and Figures 7–8 (higher plasma volume, lower apparent fiber density).

      The reduced number of visible fibers in the latter condition is not solely a consequence of plasma volume per se, but rather results from the formation of a labile fibrin gel at higher plasma concentrations, which is lost during the fixation and aspiration steps. This effect was initially observed across samples from two donors with differing plasma fibrinogen levels. In one case, an unusually low fibrinogen concentration allowed the addition of higher plasma volumes without inducing gel formation. In contrast, in the other sample, a more typical fibrinogen level resulted in gel formation under the same conditions.

      Importantly, we performed all experiments using matched donor plasma and platelets. As a result, the precise fibrinogen concentration could not be determined prior to experimentation. Nonetheless, post hoc measurements confirmed that fibrinogen levels in most donor samples fell within the normal physiological range, which allowed us to always use the same plasma volumes for low and high plasma concentrations (4ul/ml PBS and 7 ul/ml PBS, respectively) except for one donor as mentioned above.

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      The reviewer probably refers to figure 3 and he is right; it is premature to mention fiber winding at this stage of the results section (see our response to reviewer #2). In the revised version, we have modified figure 3 to include the criteria used to classify the platelets into four different categories (Fig. 3C).

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between them with only qualitative descriptions of these phenotypes.

      The "gearwheel" is observed for completely spread platelets in the 2D fiber-retraction assay and a figure illustrating our hypothetical speculations to compare the 2D gearwheel with the 3D clot situation is presented in the discussion under the "Ideas and Speculations" paragraph (now Fig. 14). We have given a more comprehensive explanation of the proposed mechanism in the revised version.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      For the platelet shown in the middle panel of Figure 9B, the extensions cannot be clearly distinguished in the MIP (Maximum Intensity Projection) image because extension #8 is positioned above extension #7 and is therefore superimposed in the projection. However, the two extensions can be differentiated when examining the 3D image stack (Video 4, upper panel). As indicated in the figure legend, the number of extensions was determined manually by scrolling through the z-stack image sequence. In the revised version, we will also define the abbreviation “MIP” as Maximum Intensity Projection.

      (5) It is unclear what the modeling accomplishes, as there is no comparison between the results of these simulations and their experiments.

      We thank the reviewer for this valuable concern. We chose not to combine the experimental fibrin organization and the modeling results within the same figure panel, as the resulting image would be too complex and difficult to interpret. We have, however, added a supplementary figure 3 showing the results of a new simulation in the absence of cytoskeletal swirling. Under these conditions no winding of the fibrin fiber around the platelet bulb can be observed. It is also important to emphasize that the comparison between the model and the experimental data was intended to be primarily qualitative rather than quantitative.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      As mentioned already above, these live videos are quite tricky to acquire because of the following reasons: - small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time required to optimize imaging parameters, necessitate the selection of an arbitrary acquisition window. Consequently, only a single acquisition of approximately 90 min can be performed per sample preparation, with no guarantee that relevant platelet-fibrin interactions can be acquired in the acquisition window.

      - Furthermore, after blood donation, the first sample is usually ready to be acquired around 3 pm, acquisition time 90 min. At least 10 successful acquisitions per condition would be required to ensure statistical robustness, but maximal 4 can be acquired per donor, because platelet samples start to deteriorate within twelve hours after blood donation.

      Taken together, the intrinsic heterogeneity of the platelet population, the low likelihood of capturing informative events, and the limited availability of suitable imaging resources at our institute render a robust and quantitative comparison between conditions with and without blebbistatin extremely challenging, if not impractical, within a reasonable timeframe.

      In accordance with the reviewer's request, we have added a new figure 13 to the revised version, presenting quantitative data on the platelet-mediated fibre compactions and the speed of angular fibrin rotations observed in video 9.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Throughout the manuscript, it is difficult to map the data presented in the figures to the text in the results section. Often, many subpanels are referred to collectively (for example, 'Fig 4 AE and animation, Video 3' on line 150), and the reader is left to piece together how this data fits into the statements in the results section. More guidance from the authors would help to understand the connection between these data and their conclusions.

      In the revised version, we have provided clearer explanations to make it easier to understand the conclusions drawn from the data. Concerning the indication "Fig 4 A-E and animation, Video 3" just means that platelets shown in panels A-E of figure 4 can also be visualized in the animation video 3. We have also put an effort to clearly indicate which figure part is presented in the associated video.

      There are also many figures that contain redundant information. The authors should consider revising these figures and including some of these repeated images as supplemental figures.

      As noted by the reviewers, our study provides predominantly qualitative observations essentially because it is not obvious to choose parameters which would be pertinent and could be quantified accurately using expansion microscopy. A quantitative analysis would allow to show the quantification and a representative image to describe the phenotypes of platelet-mediated fibre organisations. Without a quantitative analysis, we consider it more appropriate to provide multiple examples, enabling the reader to assess the consistency as well as the variability across repeated observations.

      Additional References

      (1) Jansen KA, Zhmurov A, Vos BE, et al. Molecular packing structure of fibrin fibers resolved by X-ray scattering and molecular modeling. Soft Matter. 2020;16(35):8272-8283.

      (2) Spiewak R, Gosselin A, Merinov D, et al. Biomechanical origins of inherent tension in fibrin networks. J Mech Behav Biomed Mater. 2022;133:105328.

      (3) Ramanujam RK, Lavi Y, Poole LG, Bassani JL, Tutwiler V. Understanding blood clot mechanical stability: the role of factor XIIIa-mediated fibrin crosslinking in rupture resistance. Res Pract Thromb Haemost. 2025;9(4):102871.

      (4) Gaertner F, Ahmad Z, Rosenberger G, et al. Migrating Platelets Are Mechano-scavengers that Collect and Bundle Bacteria. Cell. 2017;171(6):1368-1382 e1323.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public review:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure. To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to neural changes across sessions. Repeated retrieval can induce memory updating (reconsolidation) or extinction, the latter involving the formation of a new association between the CS+ and safety. Although memory updating may have contributed to the ensemble reorganization observed here, several aspects of our findings argue against extinction as the primary explanation for the observed neural changes.

      First, we observed substantial neuronal ensemble turnover beginning with the first retrieval session. This early turnover is consistent with previous observations in the prefrontal cortex [1, 2] and with growing evidence that cortical memory representations remain dynamic throughout systems consolidation [3, 4]. Longitudinal studies have shown that neurons are continuously recruited into and removed from cortical memory ensembles while memory expression remains stable [1-4].

      Second, we calculated discrimination ratios to quantify discrimination of each tone relative to the CS+ across retrieval sessions (Figure S1). These analyses showed that discrimination increased, rather than decreased, over successive retrieval sessions, a pattern inconsistent with the behavioral profile expected if repeated testing had induced extinction.

      Finally, one of the most novel findings of our study is that ensemble turnover does not affect all neuronal populations equally. The graded neurons identified by our clustering analysis maintained their identity and functional organization across retrieval sessions, and their activity was better explained by tone threat value than by freezing behavior (Figure 8). This selective stability indicates that ensemble reorganization is not a uniform process but instead preferentially affects specific neuronal subpopulations while preserving a stable threat-value generalization gradient. Thus, although repeated retrieval may contribute to ongoing ensemble reorganization, our results demonstrate that this process is selective and largely spares the neuronal subpopulations that encode graded threat-value representations.

      Accordingly, we have revised the Discussion to explicitly acknowledge these points as follows:

      “The ensemble turnover observed here is consistent with previous studies demonstrating dynamic reorganization of cortical activity patterns over time [1-3, 5]. Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance [4, 6]. Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure [4, 6-8]. Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations [9], whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients [10-13]. Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles.” Pg. 19

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      This is an important point, which was also highlighted by the other reviewers. To directly address this concern, we implemented the generalized linear model (GLM) analysis suggested by Reviewer 3. We modeled the neuronal activity time series using both tone identity and freezing behavior as simultaneous predictors. Because tone identity was fixed across trials whereas freezing varied from trial to trial, the GLM allowed us to dissociate their independent contributions to neuronal activity.

      As described in the original submission, freezing was estimated from the miniscope's onboard inertial measurement unit (IMU), which measures body acceleration along three axes. Rather than classifying freezing using a fixed threshold, we estimated the continuous probability of freezing from the accelerometer signal using a Gaussian mixture model. This probabilistic estimate was incorporated directly into the GLM together with tone identity, providing a conservative test of whether neuronal activity was better explained by freezing behavior or by the auditory stimulus.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves (Figure 4) and to the graded and frequency-selective neuronal subpopulations identified by our clustering analysis (Figure 8). Across both experimental groups and all analyses, the median regression coefficients (β) associated with tone identity were consistently larger than those associated with freezing, indicating that tone identity contributed more strongly to neuronal activity. Moreover, tone coefficients exhibited graded monotonic profiles that closely tracked the learned threat value of each tone, with graded neurons showing the strongest gradients (Figures 4a, 8a, and 8e). Consistent with previous reports [14, 15] freezing accounted for a modest but significant component of PL activity. However, only 6–8% of graded neurons were classified as freezing-dominant, indicating that for the vast majority of these neurons, tone identity was the stronger predictor. Together, these findings demonstrate that the graded representation of learned threat value persists after accounting for freezing behavior, supporting our conclusion that PL activity reflects learned threat value rather than merely the behavioral expression of fear.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We now include measures of registration quality in the resubmission. Specifically, we calculated shifts in centroid distances, proportion of ROIs retained across all sessions, and representative examples of matched imaging fields over time (Fig, S3).

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We corrected correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      We clarified the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      For Figure 3, we maintained the previous one-way ANOVAs assessing changes in AUC per day to be able to note significance on the Figure panels. However, we added a three-way mixed-effects analysis, using group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30) as variables, with frequency and day of testing as repeated measures. The results were described as follows (statistical details Table S1):

      “To determine how AUC varied across groups over time, we performed a three-way mixed-effects ANOVA with group (CS+15, CS+3, and no shock), frequency (3, 7, 11, and 15 kHz), and time (test days 1, 15, and 30) as factors, with repeated measures on frequency and time. For positive responder neurons, the analysis revealed significant main effects of group (p < 0.001) and time (p < 0.05), as well as a significant group × frequency interaction (p < 0.001), whereas the time × frequency and group × time × frequency interactions were not significant (p > 0.05; Table S2a). Tukey-corrected post hoc comparisons showed that, in the CS+15 group, AUC differed between all frequency pairs except 11 and 15 kHz (p < 0.05). In the CS+3 group, the AUC at 3 kHz differed from those at 7, 11, and 15 kHz (p < 0.05), whereas no significant frequency differences were observed in the no-shock controls (p > 0.05). For negative responder neurons, the only significant effect was a time × frequency interaction (p < 0.01). Tukey-corrected simple-effects analyses revealed that, on day 30, the AUC at 15 kHz differed from those at 3, 7, and 11 kHz (p < 0.05; Table S2b). Because this pattern was observed across all experimental groups, including the no-shock controls, it is unlikely to reflect associative learning. These results indicate that although the AUC exhibited modest changes over time, these changes were not group-specific and therefore do not support learning-dependent alterations in neuronal responses. Together, these results show that despite substantial neuronal turnover, PL population responses encode generalization gradients, closely matching behavioral expression.” Pg. 9

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term inference may overstate the cognitive processes engaged by the current task. Accordingly, we revised the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli. The new GLM analyses further support this interpretation by demonstrating that, in the conditioned groups, neuronal activity at both the population and single-neuron levels is explained substantially better by tone identity than by freezing behavior (Figures 4 and 8). We therefore retained the term threat value, as our results indicate that PL activity primarily reflects learned threat value rather than simply the expression of freezing behavior, but removed inference.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We corrected the language and replaced valence for “threat value”

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for this thoughtful comment. We agree that our original wording may have implied that turnover was a spontaneous process. Repeated retrieval provides opportunities for updating the learned contingencies associated with both the conditioned and generalization stimuli, and therefore changes in ensemble composition across sessions need not arise independently of experience. We also agree that the stability of graded neuronal representations may be related to the animal's certainty about the learned contingencies. However, in our data the graded neuronal population remained remarkably stable across retrieval sessions, whereas changes occurred primarily within the dynamic, frequency-selective neuronal populations. This suggests that stable ensembles preserve representations of learned threat value while updating is concentrated in a distinct neuronal subpopulation. We have now incorporated these ideas into the Discussion.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific criticisms. We revised the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      (a) We hypothesized that the PL generates representations of learned threat value that support threat generalization and discrimination, and that these representations emerge from the coordinated activity of stable and dynamic neuronal subnetworks, preserving consistent relationships among stimuli despite ongoing cellular turnover.

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      (b) The summary statement was rewritten: " Together, these findings provide a neural framework for understanding how the PL supports adaptive threat generalization and discrimination.” pg. 4

      (c) 'In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS<sup>+</sup>15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We adopted the reviewer's suggested rewording: " In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting learned contingency value across testing days" pg. 9

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we followed Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using the time series activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis allowed us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we concluded that tone identity was a stronger predictor of neuronal activity than freezing. These results are summarized in Figures 4 for all cells contributing to population responses and Figure 8 for the main neuron types identified in the clustering analysis (frequency-selective and graded neurons). All details of this extensive new analysis are shown in red in the revised resubmission.

      In addition, we revised the text to remove the terminology of “learned and inferred valence” throughout the manuscript.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      We thank the reviewer for this helpful comment. We agree that our use of the term valence in describing the no-shock controls was imprecise. Because none of the tones was associated with reinforcement in this group, there was no learned valence that could modulate neuronal activity. Our intention was simply to convey that, although both positive and negative sound-responsive neurons were present, the population responses did not vary systematically across tone frequencies. We have revised this section accordingly.

      We also agree that our original wording overstated the interpretation of the graded population responses. Our data do not demonstrate that associative learning is required for sound responsiveness itself; rather, they show that associative learning is required for the emergence of graded population responses that distinguish tones according to their learned threat value. We have revised the text to make this distinction explicit.

      Finally, we agree that our previous references to "learning and inference" were not justified by the behavioral paradigm. We have removed this language throughout the manuscript and now describe the findings more directly as graded representations of learned threat value that closely parallel the observed behavioral generalization gradients.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We thank the reviewer for this thoughtful comment. We agree that repeated retrieval is an inherent limitation of longitudinal memory studies and that repeated non-reinforced presentations of the tones provide opportunities for memory updating. Accordingly, we have revised the Discussion to explicitly acknowledge that repeated retrieval may contribute to the ensemble reorganization observed across sessions through memory updating or reconsolidation processes (Discussion, pg. 19).

      We also agree that the progressive sharpening of the behavioral generalization gradients across retrieval sessions is consistent with memory updating. Both the behavioral data (increased discrimination ratios) and the neuronal data (progressively sharper population generalization gradients among neurons active after conditioning) indicate that the memory representation became more precise over time. We now discuss this possibility explicitly in the revised Discussion. We also agree that fear generalization often broadens with time; however, this is not universal. Under discriminative conditioning paradigms, repeated retrieval can instead produce progressively narrower generalization gradients [11]. We have revised the Discussion to clarify this distinction and added the appropriate references (pg. 19).

      While the reviewer's interpretation is therefore plausible, we do not believe it fully accounts for our observations. If repeated non-reinforced presentations were the sole driver of the observed neuronal changes, one might expect a more uniform reorganization across the neuronal populations engaged by the task. Instead, the reorganization was highly selective. Neurons encoding graded threat value remained remarkably stable across retrieval sessions, whereas neuronal turnover occurred primarily within the frequency-selective subpopulations. Thus, although repeated retrieval may update the memory representation, the neuronal substrate supporting graded threat-value coding is largely preserved while refinement occurs within a distinct neuronal subpopulation.

      Moreover, we observed substantial neuronal turnover beginning with the first retrieval session, consistent with previous longitudinal studies showing that cortical memory ensembles remain dynamic despite stable memory [1-4]. This early emergence of turnover suggests that repeated testing alone is unlikely to account for the continuous population dynamics observed throughout the experiment.

      Rather than viewing these findings as evidence exclusively for either memory updating or systems consolidation, we believe they are more consistent with current models proposing that these processes occur in parallel. Several influential frameworks argue that memories are continuously modified through retrieval while simultaneously undergoing systems-level reorganization [4, 6, 16, 17].We have therefore revised the Discussion to interpret the longitudinal changes more conservatively as reflecting the combined influence of retrieval-dependent memory updating and systems-level reorganization.

      In summary, we have revised the manuscript to better acknowledge the contribution of repeated retrieval while emphasizing what we believe is the principal finding of our study: despite substantial turnover within the overall ensemble, the neuronal population encoding graded threat value remained remarkably stable, whereas refinement occurred primarily within dynamic frequency-selective neuronal populations.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS<sup>+</sup>. In CS<sup>+</sup>15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      We agree that, because population similarity is highest for the 3/3, 15/15, and 15/11 tone pairs and freezing is also greatest for these same stimuli, the neural data could, in principle, reflect a correlate of behavioral expression rather than an independent representation of learned threat value.

      To directly address this possibility, we implemented a generalized linear model (GLM) to dissociate the contributions of tone identity and freezing behavior to neuronal activity. Across all analyses, tone identity consistently explained substantially more variance in neuronal activity than freezing behavior. Importantly, this finding held not only for the full population of sound-responsive neurons used to generate the population similarity analyses (Figure 4), but also for both the stable graded neurons and the dynamic tone-selective neuronal populations identified by our clustering analysis (Figure 8). Thus, although freezing behavior contributes modestly to PL activity, it cannot account for the enhanced similarity of population vectors across stimulus presentations or the graded population responses that form the basis of our conclusions.

      In addition, the temporal dynamics of the population vector similarity analysis are not entirely consistent with the interpretation that PL activity simply reflects the expression of freezing behavior. Population vector similarity peaked during the first 5 seconds following tone onset, whereas freezing occurred intermittently throughout the tone presentations. Although this temporal relationship does not establish causality, it is consistent with the interpretation that PL activity reflects the learned threat value associated with each tone rather than merely tracking the magnitude of freezing.

      Finally, we have revised the manuscript to more clearly acknowledge the correlational nature of these analyses. Specifically, we now state that population vector similarity is associated with, rather than determines, the degree of threat generalization.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We made this correction. (“These findings indicate that population-level similarity at stimulus onset scales with behavioral threat generalization”. pg. 13)

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We agree that the phrase "integrated learned valence" is unnecessarily opaque and we replaced it with more precise language “Our previous analyses demonstrated that threat-value generalization gradients are represented at the population level. However, these findings do not reveal how these representations arise. Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subnetworks that integrate information across tones to encode their learned threat-value.” (Pg. 13)

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      We thank the reviewer for pointing out that this section was unclear. We agree that our original wording was imprecise and could be interpreted as implying cognitive processes that were not directly tested in the present study. Accordingly, we have revised the terminology throughout the manuscript. We no longer refer to "inferred emotional valence" or "core memory content" and instead describe these neurons more specifically as exhibiting graded representations of learned threat value.

      This is not the interpretation we intended. To determine whether these neurons primarily reflected defensive behavior rather than learned stimulus value, we implemented a Generalized Linear Model (GLM) that dissociates the contributions of tone identity and freezing behavior to neuronal activity. Across the entire neuronal population, as well as within the stable graded and dynamic tone-selective neuronal subpopulations, tone identity consistently explained substantially more variance than freezing behavior (Figures 4 and 8). Furthermore, after accounting for freezing, the regression coefficients of the graded neurons continued to follow the learned threat value of the tones, exhibiting opposite monotonic gradients in the CS+15 and CS+3 groups. If these neurons simply tracked defensive behavior irrespective of the stimulus presented, this relationship would not be expected to persist after accounting for freezing. We therefore conclude that the activity of this stable neuronal subpopulation is better explained by graded representations of learned threat value than by defensive behavior alone, and we have revised the manuscript accordingly.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting that this section was unclear. We agree that the original phrasing was insufficiently precise. Our intention was to convey that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. This observation led us to hypothesize that neurons recruited into the active population after conditioning— likely dynamic, frequency-selective neurons—also contribute to these graded population responses through modulation of their firing rates.

      The reviewer correctly notes that the phrase "shaped by associative processes" was too vague. By this we meant that the firing properties of these neurons are modified by the animal's associative history, including both the original conditioning experience and any retrieval-dependent updating that may occur during subsequent test sessions. We have revised the manuscript to make this interpretation explicit rather than leaving it to the reader to infer.

      To test this hypothesis, the GLM analysis we implemented dissociated the contributions of tone identity and freezing behavior to neuronal activity. After accounting for freezing, tone identity (i.e., learned threat value) remained a significant predictor of neuronal responses. Importantly, this was also true for the dynamic, frequency-selective neurons (Fig. 8e–f), indicating that these neurons contribute to population-level representations of learned threat value through firing-rate modulation rather than simply reflecting defensive behavior.

      To clarify our interpretation, we have rewritten the relevant section as follows:

      "Graded clusters encode generalization gradients but constitute only a subset of the active neuronal population. Nevertheless, population-level representations, which incorporate all active neurons, remain robust and accurately preserve these gradients. This observation led us to hypothesize that neurons recruited over time (e.g., dynamic, frequency-selective cells) also contribute to threat-value representations. Consistent with findings in the hippocampus showing that neurons can encode task contingencies through firing-rate modulation despite responding selectively to a single location (Gagliardi et al., 2024; Huxter et al., 2003; Sanders et al., 2019), we tested whether dynamic, frequency-selective clusters exhibited firing-rate differences proportional to learned threat value." (page 15)

      Regarding the reviewer's suggestion that the characteristics of the newly recruited neurons may reflect learning during repeated non-reinforced test sessions, we agree that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population, and we now explicitly acknowledge this possibility in the Discussion. However, we do not believe that our findings are fully explained by repeated non-reinforced retrieval alone. First, no-shock control animals underwent the same repeated testing but failed to develop graded neuronal representations, indicating that repeated exposure in the absence of associative learning is insufficient to account for the observed changes. Second, both behavioral discrimination and the corresponding population-level neural gradients became progressively sharper over time, consistent with refinement of learned threat representations rather than an effect of repeated testing alone, which must lead to extinction.

      In summary, we thank the reviewer for highlighting both the ambiguity of our original wording and an important alternative interpretation. In response, we have clarified the text to explicitly define what we mean by associative processes, added a GLM analysis demonstrating that the newly recruited neurons encode learned threat value beyond freezing behavior, and revised the Discussion to acknowledge that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We modified the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We incorporated the new GLM analysis to address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We agree that the term "graded representational axis" was insufficiently defined and could be interpreted in multiple ways. Because this terminology was not essential to our conclusions, we have removed it and instead describe the observed phenomenon as a graded population representation of learned threat value across tone frequencies. We also removed the statement suggesting that recurrent connectivity provides a stable scaffold for these representations, as this mechanistic interpretation is not directly supported by our data.

      We also agree that some sections of the manuscript overstated the scope of our conclusions and have revised the wording accordingly. Our study uses neuronal activity recorded during memory retrieval after learning, an approach widely used in studies of systems consolidation to infer how learned information is represented within neural populations. Accordingly, we have revised the manuscript to explicitly state that our findings pertain to neural representations of learned threat value during memory retrieval rather than the broader content of emotional memories.

      Finally, we agree that our data are correlational and do not establish the causal role of the neuronal representations we identify. Throughout the manuscript, we now refer more precisely to population- and single-neuron correlates of learned threat value during memory retrieval following auditory fear conditioning.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this thoughtful comment. Our interpretation that these neurons reflect preexisting sensory-driven properties of PL cortex is based on two observations. First, tone-selective neuronal clusters were present in both conditioned and no-shock control animals, consistent with previous reports of sensory responsiveness in PL cortex [18, 19]. Second, these responses were already present during the first retrieval session, when the intermediate frequencies were presented for the first time. Thus, they cannot be explained by repeated exposure to those tones across subsequent test sessions.

      We therefore interpret the frequency-selective response properties as pre-existing features of PL circuitry that are present independently of conditioning. In contrast, associative learning modifies the firing activity of these neurons, allowing them to contribute to graded representations of learned threat value. This interpretation is supported by our GLM analysis, which showed that, after accounting for freezing, tone identity significantly predicted the activity of frequency-selective neurons in conditioned animals but not in no-shock controls. Thus, while the frequency-selective response properties are present independently of learning, associative learning modifies how these neurons encode learned threat value. We have revised the manuscript to clarify this distinction.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded neuronal responses may partially reflect freezing-related activity rather than representations of learned threat value. In the revised manuscript, we acknowledge that previous studies have identified PL neurons whose activity tracks freezing independently of stimulus identity or associative content. To directly address this possibility, we implemented the reviewer's suggestion by fitting a generalized linear model (GLM) to the neuronal activity time series derived from the Ca<sup>2+</sup> signals, using tone identity and freezing behavior as predictors. Because tone identity is fixed across trials, whereas freezing varies both during tone presentation and across trials (see below our answer to the Recommendations to Authors), this approach allowed us to dissociate their respective contributions to neuronal activity. We are grateful for this excellent suggestion, which has substantially strengthened both the manuscript and the conclusions that can be drawn from our data. The new analyses are summarized in Figures 4 and 8.

      In the points below we summarize the new findings.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such as combinations of optogenetic and two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      We examined inter-individual variability in freezing relative to the proportion of graded cells but the number of mice used in this study (CS+3= 5 and CS+15=7) did not give us enough power to reach significance.

      Therefore, we modified the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the correlational nature of our conclusions.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Many sections of the paper should be rewritten along the lines that I have suggested in my public review and below.

      Minor Comments:

      (1) INTRO - 'This broad accessibility reduces spatial specificity and increases learning variability...'.

      Broad accessibility of what, exactly? And how does the 'broad accessibility' reduce spatial specificity and increase learning variability? That is, I do not understand what the terms 'reduced spatial specificity' and 'increased learning variability' refer to at this point in the first paragraph...

      We rewrote the introduction and discussion to address the points raised by the reviewer.

      (1) INTRO - 'The prelimbic cortex (PL) contributes to the expression (Burgos-Robles et al., 2009; SierraMercado et al., 2011; Sotres-Bayon & Quirk, 2010) and the proper discrimination and generalization of threat memories (Rosas-Vidal et al., 2025; Stujenske et al., 2022).'

      What is achieved by calling it 'proper' discrimination and generalization? Can't one simply say that the PL contributes to the expression, discrimination, and generalization of threat memories?

      This was corrected.

      (3) INTRO - '... and that population-level firing rate similarity across stimulus presentations determines threat generalization'.

      Or, alternatively, that generalization of conditioned freezing responses from the tone CS to variants along the dimension of Hz values is reflected in systematic changes in the firing rate of PL neuronal ensembles; when the test stimulus is similar to the conditioned stimulus, the two elicit similar behavioural responses and evoke a similar population-level firing rate in the respective PL neuronal ensembles.

      We have revised the Introduction as stated above. However, as discussed in our detailed responses, freezing behavior cannot fully account for the observed patterns of PL activity.

      (4) METHODS - 'Memory retrieval was tested on days 1, 15, and 30 after conditioning to probe early, long-term, and remote memory (Bontempi et al., 1996). During retrieval, mice were tested in a novel context with the CS<sup>+</sup>, CS1<sup>-</sup>, and two intermediate frequencies (7 and 11 kHz), presented in semirandom order, with each tone repeated three times (Fig. 1a).'

      Why was testing conducted in a different context than that of conditioning? This is likely to result in an underestimation of generalization to the different tones...

      In tone fear conditioning, it is always customary to test in a different context to dissociate conditioning to the context vs conditioning to the tones, which usually take place simultaneously in the same context [20]. Therefore, testing generalization in a novel context gives the correct estimate of generalization to the tones in the absence of contextual conditioning confounds. Please note that while overall freezing levels may be lower in a novel context due to the absence of contextual conditioning, the relative generalization gradient across tones — which is what your study measures — is unlikely to be systematically distorted by context change.

      (5) RESULTS - 'No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that freezing reflected associative learning.'

      The inference doesn't follow from the result described. Was there more freezing among animals in the shocked groups compared to those in the no-shock group? I presume so - my point is that this comparison is the one that most directly speaks to the presence or absence of associative learning.

      Experimental animals exhibited not only higher overall freezing but also graded freezing responses across tone frequencies. It is important to note that no-shock controls did not display this pattern, ruling out the possibility that the different frequencies themselves elicited graded behavioral responses. To clarify this point, we revised the sentence as follows: "No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that the graded freezing patterns resulted from associative learning rather than the acoustic properties of the tones." (Pg. 6)

      (6) 'Across animals and sessions, we identified distinct neuronal populations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound (Fig. 2b)...'

      To be clear, do you mean to say that there were distinct neuronal populations that consistently [i.e., across all three sessions] increased their responses to the tones [positive modulation], decreased their responses to the tones [negative modulation], showed variable responses to the tones [mixed responses], and did not respond to tones [not modulated]?

      The sentence refers to neuronal populations identified within each recording session based on their responses to the tones, not to neurons that maintained the same response profile across all three sessions. We have revised the text to make this distinction explicit. The only stable patterns across sessions were observed in graded neurons that were stable across retrieval.

      “Across animals, we identified distinct neuronal subpopulations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound in each session (Fig. 2b)” Pg. 7

      (7) What does 'active' mean in relation to Figure 2? Does this refer to neurons that displayed either positive responses, negative responses, and/or mixed responses? In the text, it is stated that 'Sound responder neurons were classified using a test that detected modulation based on magnitude relative to baseline variability, allowing reliable identification of both transient and sustained responses while remaining robust to noise...'

      I can't work out if this is the same classification criteria used for the determination of positive modulation, negative modulation, and mixed responding.

      We thank the reviewer for pointing out this ambiguity. In Figure 2, the term "active" referred to neurons that exhibited significant sound-evoked modulation and were subsequently classified as showing positive, negative, or mixed responses. Thus, active and sound-responsive refer to the same population of neurons. We removed the word active to avoid confusion.

      The reference to transient and sustained responses describes the temporal profile of the calcium signals rather than separate response categories. Some neurons exhibited brief calcium transients that rose and decayed rapidly, whereas others displayed sustained activity throughout the tone presentation. The sound-response detection algorithm was designed to reliably identify both temporal response profiles. We have revised the manuscript to make these definitions explicit. We modified the sentence as follows: “Sound-responsive neurons were identified using a statistical test that detected activity modulation relative to baseline variability, allowing reliable identification of responses while remaining robust to noise. This approach was effective for neurons exhibiting either brief calcium transients that rose and decayed rapidly or sustained activity throughout the tone presentation.” Pg. 7-8

      (8) 'A moderate proportion of neurons was present across all retrieval sessions, with no differences between groups (p > 0.05).'

      Do you mean to say that 'A moderate proportion of neurons was ACTIVE across all retrieval sessions, with no differences between groups (p > 0.05)'?

      We replaced the word present and replaced it with “active”. Pg. 8

      (9) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'If neurons encoding graded responses carry core mnemonic information, they should exhibit enhanced stability over time. To test this hypothesis, we quantified the proportion of registered neurons that retained their cluster identity across at least two retrieval sessions and compared these values to a shuffled null distribution (10,000 iterations), with multiple comparisons controlled using the BenjaminiHochberg procedure.'

      What does the comparison to the shuffled null distribution tell us exactly? I accept that some neurons were stable positive responders across at least two sessions. The comparison to the shuffled null distribution creates a false impression about the robustness of this stability or the 'enhanced stability over time'.

      Our intention in comparing the observed stability to a shuffled null distribution was to evaluate whether the proportion of neurons retaining cluster identity exceeded chance levels expected from random assignment. The shuffled distribution therefore provides a statistical baseline against which the observed degree of stability can be evaluated. We agree, however, that the wording “enhanced stability over time” may be confusing regarding this finding. We rephrased this paragraph to clarify that a subset of neurons retained cluster identity across all retrieval sessions at levels greater than expected by chance as follows:

      “These data demonstrate that graded clusters remain consistently active at levels exceeding chance, preserving their cellular identity and providing a stable representation of learned contingencies and generalization gradients.” Pg. 15.

      (10) ABSTRACT. The abstract states that, 'Stimulus-evoked population similarity scaled precisely with behavioral generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing, indicating that population-level similarity predicts inferred threat.'

      I believe that the sentence could be rewritten as, 'Stimulus-evoked population similarity reflected the degree of generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing.'

      We revised the text according the reviewer’s suggestion; however, we had to shorten the sentence due to word limits. “Population similarity tracked behavioral generalization, whereas consistent population states emerged only for shock-associated or highly generalized tones.” Pg. 2

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The ensembles with graded activation in proportion to stimulus valence are described at various points in the manuscript as "maintaining the emotional 'gist'", "preserving core components of the memory trace", and "preserving core components of the memory trace". This conclusion is premature because there is an alternative interpretation. The graded response ensembles would also be consistent with coding for the freezing behavior itself, irrespective of the specific memory or stimulus association that drives it. An ensemble that encodes a behavior in this way would not be considered mnemonic, just as motor neurons in the spinal cord are not, even if they may fire during a conditioned response. Indeed, previous work has identified neurons in the prelimbic cortex that encode freezing independently from the stimuli that signal an aversive outcome (e.g., Kyriazi, Headley, and Pare 2020; Casanova, Pouget, ..., Vetere 2024).

      There are two ways the authors can address this point.

      (a) Fit a generalized linear model to the time series of inferred spiking activity from the Ca2+ signal and include stimuli and freezing as predictors. Since freezing behavior is inconsistent across trials, while stimulus presence is fixed, they can be disassociated. If, after accounting for freezing, responsiveness neurons still show a graded coding of stimuli that agrees with inferred aversiveness, this would strengthen their claim that they have identified an ensemble that corresponds with mnemonic or salience aspects of the stimuli.

      (b) Conduct no further analysis but cover the issue in the discussion as a limitation to their study and to dampen some of the language throughout the manuscript that implies that a memory trace has been identified.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an important alternative interpretation is that graded-response ensembles could reflect freezing-related activity rather than representations of learned threat value. To directly address this possibility, we implemented a Generalized Linear Model (GLM) analysis, as suggested by the reviewer. The GLM was fitted to the activity of every sound-responsive neuron included in the population analyses and simultaneously incorporated tone identity and continuous freezing probability (derived probabilistically from miniscope acceleration) as predictors, allowing us to quantify their independent contributions to neuronal activity.

      We want to note that freezing was quantified from the miniscope's inertial measurement unit (IMU) using a two-component Gaussian mixture model applied to the log-transformed body-acceleration signal. Rather than classifying freezing with a binary threshold, we used the posterior probability of the low-movement state as a continuous freezing regressor. This approach captures graded variations in immobility and provides a more conservative test of tone encoding, because it accounts for more behaviour-related variance than a binary classifier, making it more difficult to detect an independent contribution of tone identity.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves and separately to the identified frequency-selective and graded neuronal subpopulations. Across all analyses, tone identity consistently explained neuronal activity better than freezing. Furthermore, the freezing-corrected tone β coefficients scaled with learned threat value, with graded neurons exhibiting the strongest monotonic gradients, indicating that they provide the most robust representation of learned threat value. These findings demonstrate that the graded coding of learned threat value persists after accounting for freezing behavior and therefore cannot be explained simply by the behavioral expression of fear. The new analyses are presented in Figures 4 and 8. Notably, although freezing-dominant neurons were present in both the tone-selective and graded populations, they represented only a small fraction of each group and were least prevalent among graded neurons (6–8%), further supporting the conclusion that graded neurons primarily encode learned threat value.

      In addition, we revised the manuscript to avoid language implying that these neuronal populations constitute a mnemonic trace. Instead, we consistently describe them as encoding learned threat value, a more accurate interpretation that is directly supported by the new GLM analyses.

      (2) The title makes a seemingly causal claim by using the term 'arise', "Learned and inferred valence arise from interactions between stable and dynamic subnetworks". While it is true that the authors show that both stable and dynamic ensembles encode valence, they do not demonstrate that the behavioral expression of valence depends on these codes, nor their interaction. Experimentally testing this is beyond the scope of this study (holographic two-photon stimulation of transient and stable ensembles?), but they may be able to get closer to it by examining inter-individual variability. The authors could measure the proportion of neurons in each subject that participate in the stable (graded responding) and dynamic (stimulus-specific) ensembles, and see if they predict individual differences in the expression of freezing behavior or its generalization. Indeed, this correlation may change across testing days.

      We agree that the term “arise” in the title may imply a stronger causal relationship than is directly supported by the present data. We modified the title in the resubmission as follows: “Complementary stable and dynamic prelimbic ensembles encode learned threat value underlying generalization and discrimination”

      The new LGM analysis confirms that a large proportion of neural activity can be predicted by tone threat value; therefore, we think this title fully captures our findings.

      We also appreciate the reviewer's suggestion to examine inter-individual variability. In the revised manuscript, we tested whether the proportion of graded neurons correlated with freezing behavior. However, the limited number of experimental animals in each experimental group provided insufficient statistical power to reliably assess this relationship. Accordingly, we revised the manuscript to clarify that our conclusions are based on correlational observations rather than causal inferences.

      In summary, we revised the title and related language throughout the manuscript to avoid implying causal mechanisms beyond the scope of the current experiments.

      Minor points:

      (1) I was surprised by the absence of an ensemble in the No-shock group that responded uniformly to all stimuli. Can the authors confirm this?

      Yes, we confirm this finding. It was unexpected to us as well. We would like to clarify, however, that some control neurons may have responded to more than one frequency, but these responses were too infrequent or too weak to be classified as a distinct graded neuronal population by our clustering algorithm. Thus, while broadly responsive neurons may have been present in the control group, they did not form a robust, identifiable ensemble comparable to that observed after fear conditioning.

      (2) Several different approaches were used to analyze the same Ca2+ responses to stimuli across testing days. These were the "Sound responder classification", "Average stimulus-aligned trace procedure", "Population similarity over time across tone pairs", and the construction of "Stimulus response vectors". These feature differing alignment/binning/interpolation, normalization, and response quantification procedures, and it is unclear why they cannot all be in agreement, at least when it comes to alignment and normalization.

      We thank the reviewer for this careful reading of our Methods. All analyses were performed on the same underlying calcium imaging dataset, but they were designed to address different aspects of the data and therefore required different preprocessing steps. The analyses share a common initial pipeline leading to the calcium traces (all z-scored across the session). Differences in subsequent processing (e.g., use of ΔF/F versus CASCADE-deconvolved activity, normalization, baseline correction, temporal binning, and interpolation) were introduced only when required by the specific analysis method.

      To make this clearer, we have substantially revised the Methods. We added a new overview of preprocessing section that summarizes the common preprocessing pipeline and explicitly distinguishes the shared steps from those that are analysis-specific. We also included a summary table describing the input signal (ΔF/F or CASCADE-deconvolved activity), normalization procedure, and temporal processing used for each analysis. Finally, the individual Methods sections were revised to eliminate redundancies and more clearly describe the steps to avoid confusion. We hope these revisions make the rationale for the different preprocessing procedures and the overall analytical workflow more transparent (Pg. 22-23)

      (3) In the methods section "Window-wise response quantification" the Ca2+ signal was baselinesubtracted and divided by the standard deviation in the baseline across trials, but that data was already presumably z-normalized to the baseline of each trial ("Data alignment and normalization"). This second step of normalization seems excessive. Why is it not sufficient to just take the average peri-stimulus response across the z-normalized trials from the "Data alignment and normalization" section? This is simpler and would capture the effect size of the response relative to baseline.

      We thank the reviewer for this careful observation.

      The two operations are also not the same normalization applied twice; they standardize different sources of variability. During the alignment step, each trial is z-scored relative to its own baseline by dividing by the standard deviation of that trial's baseline across time. This places all trials on a common within-trial scale before averaging. In the window-wise step, the trial-averaged, baseline-subtracted response is expressed relative to the standard deviation of the per-trial baseline levels across trials—a distinct quantity that reflects trial-to-trial baseline stability rather than within-trial fluctuations. The purpose of this second term was to down-weight windows in cells with unstable baselines across trials, and it entered the analysis only as a significance criterion; the magnitude threshold defining a sound responder was applied to the trial-averaged baseline-relative response itself. We have revised the Methods to clarify the distinct roles of these two normalization steps.

      To further address this concern, we re-ran the sound-responder classification after removing the second (between-trial) normalization step, so that responder detection depended only on the per-trial baselinerelative response magnitude and its temporal persistence. Across all cells, tones, and sessions (n = 89,504 cell–tone–session classifications), the two procedures agreed on 95.3% of labels. The small fraction of cells whose labels changed were almost exclusively those lying immediately at the detection threshold: 86.6% of changes involved cells moving into or out of the "modulated" category, whereas direct reversals between excitatory and inhibitory classification occurred in only 8 of 89,504 cases (0.009%). Consistent with the between-trial standard deviation being a less stable quantity when few trials are available, label changes were approximately twice as frequent in the three-trial retrieval sessions (5.2%) as in the ten-trial conditioning sessions (2.1%). Overall responder proportions changed only minimally (positive responders +2.2%, negative responders +4.3%), and all population-level findings—including the graded threat-value gradient across tones, its absence in no-shock controls, and its persistence after controlling for freezing in the , as analysis—were unaffected. These analyses demonstrate that our conclusions are robust to this methodological choice.

      (4) It would increase confidence in the tracking of neurons across days if the authors showed some example images of neurons tracked across days.

      We added an example in the Supplement. Additionally, we now provide measures of registration quality (Fig. S3)

      (5) Table S2 is a bit confusing. I take it that Graded A/B were only for CS15, and Graded C/D/E were only for CS3. Also, Common B1-4 were the cells with positive responses to individual stimuli, and Common C1-4 were the cells with negative responses to stimuli. If this is the case, it should be explained in the figure legend (or even better, clusters should be named and numbered consistently in all figures.

      Thank you for pointing this out, we corrected the Table to indicate which test corresponds to which figure and cluster, specifying which ones were positive or negative modulated. Please note old Table 2 is now Table 5

      (6) The term network and subnetworks implies some connectivity between neurons, but in this study, it is used to refer to the ensembles of cells activated in a similar manner. Since connectivity is never assessed, it would be better if the authors stuck to the terms ensembles or populations.

      We changed the wording and now use ensembles or populations

      (7) The legend for Figure S3 has the text 'eded', which seems to be a typo.

      We corrected this typo.

      References

      (1) Kitamura, T., et al., Engrams and circuits crucial for systems consolidation of a memory. Science, 2017. 356(6333): p. 73–78.

      (2) DeNardo, L.A., et al., Temporal evolution of cortical ensembles promoting remote memory retrieval. Nat Neurosci, 2019. 22(3): p. 460–469.

      (3) Tome, D.F., et al., Dynamic and selective engrams emerge with memory consolidation. Nat Neurosci, 2024. 27(3): p. 561–572.

      (4) Mau, W., M.E. Hasselmo, and D.J. Cai, The brain in motion: How ensemble fluidity drives memory-updating and flexibility. Elife, 2020. 9.

      (5) Gallego, J.A., et al., Long-term stability of cortical population dynamics underlying consistent behavior. Nat Neurosci, 2020. 23(2): p. 260–270.

      (6) Zaki, Y. and D.J. Cai, Memory engram stability and flexibility. Neuropsychopharmacology, 2024. 50(1): p. 285–293.

      (7) Lacagnina, A.F., et al., Distinct hippocampal engrams control extinction and relapse of fear memory. Nat Neurosci, 2019. 22(5): p. 753–761.

      (8) Sangha, S., Plasticity of Fear and Safety Neurons of the Amygdala in Response to Fear Extinction. Front Behav Neurosci, 2015. 9: p. 354.

      (9) Bouton, M.E., S. Maren, and G.P. McNally, Behavioral and Neurobiological Mechanisms of Pavlovian and Instrumental Extinction Learning. Physiol Rev, 2021. 101(2): p. 611–681.

      (10) Jenkins, H.M. and R.H. Harrison, Effect of discrimination training on auditory generalization. J Exp Psychol, 1960. 59: p. 246–53.

      (11) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (12) Herzog, K., et al., Reducing Generalization of Conditioned Fear: Beneficial Impact of Fear Relevance and Feedback in Discrimination Training. Front Psychol, 2021. 12: p. 665711.

      (13) Lommen, M.J.J., et al., Training discrimination diminishes maladaptive avoidance of innocuous stimuli in a fear conditioning paradigm. PLoS One, 2017. 12(10): p. e0184485.

      (14) Casanova, J.P., et al., Threat-dependent scaling of prelimbic dynamics to enhance fear representation. Neuron, 2024. 112(14): p. 2304–2314 e6.

      (15) Kyriazi, P., D.B. Headley, and D. Pare, Different Multidimensional Representations across the Amygdalo-Prefrontal Network during an Approach-Avoidance Task. Neuron, 2020. 107(4): p. 717–730 e5.

      (16) McKenzie, S. and H. Eichenbaum, Consolidation and reconsolidation: two lives of memories? Neuron, 2011. 71(2): p. 224–33.

      (17) Winocur, G. and M. Moscovitch, Memory transformation and systems consolidation. J Int Neuropsychol Soc, 2011. 17(5): p. 766–80.

      (18) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (19) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

      (20) Phillips, R.G. and J.E. LeDoux, Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning. Behav Neurosci, 1992. 106(2): p. 274–85.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      We have addressed the outstanding points made by the reviewers and provide a detailed description of the additional analyses performed & key results below. We have also updated the manuscript to reflect these additional results, and to contain a more detailed consideration of alternative plausible models.

      We also note that we have corrected one figure panel (Fig.2 panel B, Exp. 2 only), where we identified a small bug in the visualisation code whereby the data of either one or two participants was not correctly plotted in some conditions. This makes no difference to the reported effects.

      Please note that the reviewers acknowledged that your introduction now more broadly refers to the previous work from various groups on motor beta lateralisation (MBL).

      (1) Evaluating the correlation between CPP and MBL, which is key for supporting the claim that CPP is feeding MBL. If, as you are alluding to in your rebuttal, single-trial estimates of CPP are too noisy, trials could be binned based on CPP.

      As requested, we now provide additional analyses binning the data by CPP amplitudes, for the high-low coherence conditions at P1. Full details are provided below. In both experiments we find that, for a given coherence, greater CPP amplitudes at P1 correlate with stronger motor beta lateralisation.

      (2) Examining the possibility of down-weighting (in line with previous studies) compared to your current bounded integration description. Specifically, does the your model predict a bi-modal CPP-P2 distribution that is not evident in the data?

      We have now fit 3 additional models investigating alternative mechanisms that might account for the behavioural results. In particular, we have explored 4 different ways in which flexible weighting of the second pulse might account for both the behavioural and neural data. A full account of the results is provided below, and has been included in the manuscript. In sum, we find that a model which directly and uniformly downweighs evidence from the second pulse (as opposed to indirectly through little or no distance remaining to bound, as in our model) can account for the behavioural data well, but it cannot recapitulate CPP-P2 results unless an accumulation-terminating bound is also included in the model, and the additional complexity of a model with these two free parameters is not supported by model comparison. An alternative model where P2 is downweighed as an inverse function of P1 strength (i.e., stronger downweighing for P1-high coherence pulses) could recapitulate both the behavioural and neural data, but again, this was not favoured by model comparison metrics that account for complexity in the current dataset. We have added a piece on this in the discussion, noting how previous studies mentioned by the reviewer such as Cheadle et al., 2014 and Glickman et al., 2022 find a consistency bias where later evidence is boosted when it agrees with the earlier evidence, opposite to the dampening suggested by the model here, but that a key distinction in our task is that P2 always agreed with P1, so that a dampening might be plausible if subjects tend to withdraw some of their attention from the confirmatory P2 based on the strength of P1.

      Regarding the CPP-P2 distribution, our original bounded model does indeed predict a bimodal CPP-P2 distribution with a peak at 0 arising from the early termination trials, which does not appear in our data (see Fig. S19 and related reply below). However, that EEG noise precludes the detection of any such bimodality in single trial amplitude distributions is demonstrated by the fact a bimodal distribution is strongly predicted for CPP-P1 amplitudes due to the two coherences, most strongly in fact for the unbounded model since there would be nothing to cap the higher-coherence, yet no trace of such bimodality is evident there either, due to EEG noise.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.

      - Innovative task design, including parametric manipulation of gap duration.

      - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - A direct characterization of how the centroparietal positivity and motor beta lateralization interact is missing, which limits the novelty. In their reply to reviewers, the authors argue that the signal-to-noise ratio of EEG signals is insufficient for such analyses at the single-trial level. If so, a binned or trial-averaged approach could still be attempted.

      As requested, we have now performed an additional analysis binning trials according to single-trial CPP-P1 amplitudes. To this aim, we sorted trials according to P1 coherence, and median-split them within condition according to the CPP-P1 amplitudes integrated in a time window around the grand-averaged peak [0.4 to 0.6s] after pulse onset, on the same subset of electrodes as in the manuscript. We then plotted motor beta lateralisation (MBL) as the difference in [Contra - Ipsi] hemispheres. Stronger negativities thus indicate stronger lateralisation towards the correct response. In all 4 cases, (both experiments and both coherence levels), higher CPP amplitudes were associated with stronger lateralisation from 0.5s post-pulse onwards (Author response image 1).

      Author response image 1.

      MBL (bottom) traces aligned to P1 onset (time = 0), median split by CPP amplitude [0.4-0.6s] post pulse onset, within P1 coherence condition. Trials with stronger CPPP1 potentials were linked to stronger MBL lateralisation toward the correct response.

      - An exhaustive characterisation of sensors and frequency bands is also missing. In their reply to reviewers, the authors suggest that this would detract from their hypothesis-driven focus. I disagree: the main hypothesis and figures could remain centred on the centroparietal positivity and motor beta lateralization, with a more comprehensive mapping of sensors and frequencies placed in supplementary material. Since the purpose of the paper is to examine EEG-based decision signals in a novel behavioural context, a broader characterisation of the underlying EEG landscape would seem appropriate.

      To broaden our characterisation, we have now included an additional supplementary figure that describes another distinct, relevant EEG signal. Fig. S12 shows the lateralised readiness potential (LRP), a lateralised motor preparation signal that has long been used as an index of relative motor preparation with high temporal resolution (Eimer, 1998; Kelly & O’Connell, 2013; Vidal et al., 2015).The LRP is typically computed as the difference in voltage between [IpsiContra] lateral motor electrodes with respect to eventual response, and it captures the fact that the contralateral motor cortex exhibits more pronounced negative ramps than the ipsilateral one immediately preceding action execution. The EEG landscape characterised in our paper thus comprises four distinct signals that are all functionally relevant to the task, including occipital alpha power, relevant for attention & temporal expectation encoding, which was included both in the main manuscript (Fig. 2) and the supplement (Figs. S9, S11). Given our already extensive supplementary material (18 figures) focused on our main research questions, we feel that a full, hypothesis-free exploration across the dimensions of frequency, space (sensors) and time, considering that there are 60 experimental conditions among which differences may be tested for (2 directions x 3 gaps x 4 coherence pairings in exp 1, plus 2 directions x 4 gaps x 4 coherence pairings in exp 2, plus single-pulse trials), would render the supplemental materials excessive in volume. Again, the data will be shared publicly for future exploration of these many dimensions.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Some readers may be left questioning why behaviour in this experiment is so different from previous experiments which use almost exactly the same design (Kiani et al., 2013; TohidiMoghaddam et al., 2019; Azizi et al., 2021; 2023). Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while the authors present a model which describes the data based on pre-mature decision-commitment, they do not examine explanations from the existing literature, that evidence is flexibly weighted, and do not provide any analyses which could be used to compare these descriptions. While their model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

      The revised version of the manuscript includes a detailed discussion about possible reasons why our stimulus characteristics, task design & experimental protocol may have led to the observed behavioural results (lines 605 onwards). Furthermore, we have now included an extended model comparison as a supplementary note which examines alternative models that could account for the observed data and an additional discussion section that links it to the existing literature.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have responded to each of the comments in the previous review. The manuscript introduction and discussion have been substantially improved, and now more adequately address the previous literature. Limited improvements were made to the analysis, although the authors acknowledged why the suggested improvements from the reviewers were unlikely to be successful, but did not attempt to address the comments using other methods.

      One common theme to both reviews was that, although the model broadly describes the data, it is not fully tested, and alternative descriptions are not fully considered.

      We thank the reviewer for their careful consideration of our data & their reply. We agree it is important to formally test alternative descriptions, most particularly those involving downweighting of the processing of P2, and we have now done so. If participants were simply downweighting P2 by implementing generally smaller drift rates regardless of P1, we would expect the CPP-P2 to exhibit the same classical pattern as in P1, with higher amplitudes following high-coherence P2. The interaction pattern we observe, whereby CPP-P2 amplitudes are systematically lower following high-coherence P1, within each P2 coherence, can only be explained by some dependency between the processes occurring at P1, and those following at P2. What if, as the reviewer suggested originally, “additional evidence from the second pulse was down-weighted according to certainty following the first pulse?” We thus consider this also.

      To formally test this, we have now fit four further models and compare both the fit quality to accuracy data and the predicted EEG results to our original implementation. All models were fit to the grand-averaged accuracy data for all conditions, using 10K simulations, and for both experiments separately (as was done for the bounded model presented in the manuscript). For EEG simulations in unbounded models, we assumed that the CPP signal fell back down to zero upon the dots turning blue, as for the simulations in the main manuscript. Fig. S15 illustrates the fit quality as measured by means of Bayesian Information Criterion (BIC), and we go into more detail on each model in turn below.

      “(1) Unbounded model + P1-independent P2 drift rate reweighting (DR<sub>P2</sub>)

      First, we fit a model with no bound in which the drift rate for P2 was estimated by reweighting (positive or negative) the P1 drift rate via an additional free scaling parameter w. This model thus had the same complexity as our original one (k = 3 free parameters): two drift rate parameters for high and low coherence of P1 (d<sub>high</sub>, d<sub>low</sub>), plus a scaling parameter, w, dictating the strength of P2 relative to same-coherence P1. The drift rate for P2 (d<sub>P2</sub>) was computed as the d*w, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, multiplied by the scaling parameter w, so that if w < 1, P2 drift rates would decrease compared to the same coherence in P1. This model implements the reviewers’ suggestion that systematic downweighting of the second pulse might also account for the results (while also allowing for upweighting (w > 1) for the sake of flexibility).

      This first model (DR<sub>P21</sub>) yielded a similar fit quality to the behavioural data compared to our original bounded model (Bnd), as measured by BIC (Fig. S15). The new model broadly recapitulated the key behavioural results, including order effects, (Fig. S16,A) and could also recapitulate the generally lower CPP-P2 amplitudes (Fig. S16,B). However, this model failed to capture the key coherence-based pattern observed in the CPP-P2 data. Namely, while our EEG results showed that CPP-P2 in trials following P1-low coherence pulses reached overall higher amplitudes than that in trials following P1-high coherence pulses (see manuscript Fig. 3), this model’s simulations predicted that CPP-P2 should scale only with P2 coherence, showing higher amplitudes and steeper build-up rates for P2-high trials, regardless of P1 coherence (Fig. S16C). This is at odds with our empirical results.”

      “(2) Unbounded model + P1-dependent P2 Drift rate reweighting (invDR<sub>P2</sub>)

      Next, we tested a model in which P2 downweighting could depend on P1 strength. That is, we made the P2 drift rate scaling parameter w inversely proportional to P1 coherence so that P2 drift rate d<sub>P2</sub> = d*w/d<sub>P1</sub>, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, and d<sub>P1</sub> indicates the preceding P1 coherence. This implements a kind of certainty weighting, whereby evidence following a strong P1 is more strongly dampened than evidence following a weak P1. This model could recapitulate the key behavioural findings (Fig. S17A), and also qualitatively captured the CPP-P2 effects (i.e. P1-low trials reaching overall higher amplitudes than P1-high trials, Fig. S17C), although the magnitude of this effect was substantially smaller than predicted by the original simple bounded model with no drift rate modulations. However, BICs indicated that this model provided an overall worse fit to the accuracy data compared to the original bounded model (Fig. S15).”

      (3) Bounded models + P2 drift rate reweighting (Bnd + DR<sub>P2</sub>, Bnd+ invDR<sub>P2</sub>)

      Finally, we investigated how well models with both a bound and either of the two P2 drift rate scaling methods we investigated above (uniform downweighting, P1-dependent downweighting) could capture the data, thus effectively testing two extensions of our original implementation that allowed for flexible reweighting of P2.

      The systematic reweighting model with a bound (Bnd + DR<sub>P2</sub>) could recapitulate all key behavioural and EEG findings (Fig. S18), but the additional complexity of the model was not supported by BIC (Fig. S15). Crucially, the reason that this model could recapitulate the CPPP2 results was still the presence of a bound, although we note that the proportion of trials that were predicted to terminate early was reduced in this model compared to the original bounded model presented in the manuscript (c.f. Fig. 4). Yet, this relatively small fraction of trials where accumulation ended early meant that 1) in some trials no accumulation was allowed to occur at all during P2, and 2) where it occurred, the DV was closer to the bound following P1-high coherence trials, thus needing to accumulate less further evidence before reaching a bound and yielding smaller CPP-P2 amplitudes overall in those trials. The inclusion of a bound was thus key to allow a model with systematic P2 downweighting to account for both behavioural and EEG data.

      The model with P1-based scaling of P2 drift rates with a bound could also recapitulate all key behavioural and EEG findings (Fig. S18), but again, the additional complexity of the model was not supported by BICs (Fig. S15).

      “Conclusion

      This extended modelling exercise suggests that 1) a simple bounded model is favoured by model comparison, 2) an alternative model of equal complexity which includes a P1dependent systematic downweighting of P2 rather than a bound can produce qualitatively similar results, the common feature of both viable models being the push-pull relationship between P1 and P2, and 3) more complex models including both a bound and P2 modulations can also account for both the behavioural and EEG results, but the additional complexity is not supported by the current data. While it is possible that some flexible weight modulations occur, these are not sufficiently influential to justify its inclusion in the model. Future work would nevertheless be warranted to explore this possibility in more detail using tailored task paradigms. For the scope of this paper, we have maintained the bounded account in the main manuscript as it is the one supported by the model comparison in the current dataset, but we have also included the alternative P1-based downweighting account as a supplementary figure, along with some additional discussion.”

      In response to my previous comment 3, the authors show their model predicts that there should be no CPP-P2 if the bound is reached before P2, otherwise CPP-P2 is similar to CPPP1 (Figure R2). The argument in the manuscript is that the lower CPP-P2 is because of this bound. The distribution of CPP-P2 amplitudes should therefore have higher variance than CPP-P1 amplitudes, and one might even predict a second mode in the distribution, around 0 amplitude (those trials that terminated before P2). The authors do not show this.

      In Figure R3, it looks like the data have been normalised independently for CPP-P1 and CPPP2 (since the means are approximately the same); normalisation also prevents a comparison of the variance. However, it is apparent that there is no bimodality in the CPP-P2 distribution - were there substantially more trials with 0 CPP-P2 amplitude than CPP-P1? Is the EEG data actually more consistent with a model that systematically downweights P2?

      In the previous Figure R3, data were normalised across CPP-P1 and CPP-P2, not separately. We plot the non-normalised values here (Author response image 2), for comparison, along with median, variance and skewness values for CPP-P1 and CPP-P2. Additionally, we attach the single-participant plots at the bottom of this document (Author response image 3)

      We reanalysed the non-normalised data, excluding outliers (defined as values exceeding the mean +/- 3 times the standard deviation, computed for each pulse & for each participant separately). We found, in both experiments, lower median amplitudes (Exp. 1: t(21) = 2.68, p = 0.013; Exp. 2: t(20) = 4.37, p < 0.001), higher variances (Exp. 1: t(21) = 0.59, p = 0.55; Exp. 2: t(20) = 2.33, p = 0.03) and more positive skewness (Exp. 1: t(21) = 1.79, p = 0.08; skewness: 0.027 vs. 0.69; Exp. 2: t(20) = 3.85, p <0.001; skewness: 0.027 vs. 0.214) in CPP-P2 compared to CPP-P1, although variance and skewness effects were only significant in Exp. 2.

      The reviewer argued in the previous review as well as here that increased variance would be predicted by the model – this is correct, but we believe that the mere presence of higher CPPP2 variances in our empirical data does not, on its own, necessarily support the model. That is because this increased variance could be explained by other factors, such as the increased EEG signal complexity of data at P2 compared to P1. This increased complexity naturally arises from overlapping potentials from CPP-P1 (which can be corrected for, but will increase data noise and thus variability nonetheless), as well as the various gap durations across conditions, which would also affect pre-P2 dynamics. Thus, while our data (partially - in Exp. 2 only) support the reviewer’s interpretation, we would be cautious in using the observation of increased variability as evidence for or against our model given the considerations above.

      Author response image 2.

      A. Empirical CPP–P1 and CPP-P2 amplitude [450-550ms post-pulse] distributions, pooled across coherences. Data were not normalised within-participant. Data were baselined 100ms before pulse onset prior to CPP-P2 amplitude extraction. In both experiments, CPPP2 amplitudes had a lower median (vertical line) amplitude and higher variance than CPP-P1. B. CPP amplitude simulations based on the original bounded model. A high number of trials where CPP-P2 amplitude should equal 0 due to early terminations. C. CPP amplitude simulations based on the inverse weighting model (invDR<sub>P2</sub>). The model did not predict any trials with zero CPP-P2 amplitude because early terminations were not allowed. Rather, the mean distribution shifted towards lower predicted amplitudes, because P2 was downweighted proportionally to P1 coherence.

      The reviewer asks: “Were there substantially more trials with 0 CPP-P2 amplitude than CPPP1?”. Given the noisiness of single-trial EEG data due to high-frequency artifacts, and/or spurious signal drifts (as illustrated by the raw CPP value distributions in Author response image 2A above), it is not be possible to directly detect trials on which CPP = 0. Instead, this must be inferred by other means. If the CPP is in reality at 0 on a larger proportion of trials, then on average, this should manifest as a higher fraction of trials with lower amplitudes, resulting in a more positively skewed distribution of CPP-P2 amplitudes compared to CPP-P1. As reported above, in both experiments we find that skewness is higher in CPP-P2 than in CPP-P1, in line with this hypothesis.

      Regarding the question “Is the EEG data actually more consistent with a model that systematically downweights P2?”, we point to the additional modelling we conducted in response to the comment above. To recapitulate, we find that a model that systematically downweights P2 could account for behavioural findings, but could only recapitulate the key EEG CPP-P2 patterns if an accumulation-ending bound was also included in the model. Instead, a model where P2 is downweighted as a function of P1 strength could qualitatively capture both behavioural and EEG findings, but was not strongly supported by goodness of fit measures in this dataset. We note, however, that the latter model does not predict a bimodal CPP-P2 distribution, but rather a shifted mean and more positive skewness for CPP-P2 trials (Author response image 2C; skewness: CPP-P1 = 1.08; CPP-P2 = 1.24). Thus, in that respect, it does appear to provide a better qualitative recapitulation of the single-trial CPP-P2 data. However, given the extent of EEG noise, the bimodal underlying distribution of our bounded model would also translate to a unimodal, skewed distribution as observed, so this does not provide a strong basis for adjudication. Underscoring this, it is noteworthy that an unbounded model in fact predicts a more separated bimodal distribution for P1 than a bounded model, yet, again, with EEG noise, we are not able to identify any such bimodality in the empirical P1 amplitude distribution.

      Minor:

      In the discussion, the authors write "We also used a narrower range of coherences than the previous studies, which possibly lends itself to calibrating a bound to achieve acceptable accuracy while saving cognitive effort." (Page 22). Perhaps this should be reworded. The range of coherence in this study was ~26-44% in Exp 1 (a difference of 18%) in previous experiments the coherence was 3.2-12.8% (a difference of ~10%). The ranges of performance were similar.

      We agree that the phrasing could be improved. We meant to say that we used only two coherences (high-low) with less than a twofold difference between them, instead of multiple levels of evidence strength in previous studies (e.g. 0,3.2,6.4,12.8) – we have clarified this.

      The authors mention in their rebuttal "However, in contrast to previous studies, we did not include any feedback on a trial-by- trial basis, instead only providing feedback at the end of each block indicating the average accuracy." Actually, Kiani et al., 2013 also only gave feedback at the end of each block. This is also implied in the discussion. I suggest this be removed as the common feedback in Kiani et al., 2013 suggests this cannot explain the difference.

      The methods in Kiani et al. 2013 state “At the end of motion stimulus, a 400–1000 ms delay period (truncated exponential) was imposed before the Go signal, disappearance of the fixation point, was presented. The subject was required to report the net direction of motion within 1 s after the Go signal by pressing a left or right key. Distinctive auditory feedback was delivered for correct and error responses. On trials with 0% coherence, the type of feedback was chosen randomly.“ We understand this means feedback was provided after every trial. If this interpretation is wrong, we would like to kindly ask the reviewer to point us to the relevant methods section so that we can correct the manuscript.

      Author response image 3.

      Individual CPP-P1 (blue) and CPP-P2 (orange), for both experiments (non-z-scored). Vertical lines indicate median CPP amplitudes for each pulse, respectively.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos, which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to:

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species has not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend repeating some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

      (R1. A1) We thank the reviewer for their time and effort to thoroughly review our paper and for the positive comments on our manuscript. In the revised manuscript we have addressed the concerns that you have raised in the Joint recommendations above (Pages 1-4).

      Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard, and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch, but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult-like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent, with a low-to-high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive, but, as the authors point out, they are particularly relevant and timely. Since 2016, there have been a series of studies published in high-profile journals that have presumably shown the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for prenatal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post-hatch are functionally deaf. Since it is highly improbable that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that the authors do not discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your setup and using the same calibration. The ~20dB difference might be real, or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch. But adding a (novel) behavioral assay in young birds could further resolve the issue.

      (R2. A0) We thank the reviewer for their time and effort to thoroughly review our paper, and for the positive comments on our manuscript.

      In our revision, we have added a new figure (Fig 5) and three new paragraphs (Lines 387-422) in the discussion to compare all published ABR and behavioral audiograms and the differences between and among these datasets. Our adult data is consistent with the reported findings from four other labs despite differences in stimulus design, setups and genetic background of the animals, providing strong support for our findings.

      Furthermore we have more clearly presented our argument why we think juvenile and embryos cannot detect heat whistles. For clarification, we have added a new figure (Fig 4) and four new paragraphs (Lines 271-355) in the discussion.

      We agree with the reviewer that more data is needed on the development of hearing in songbirds and zebra finches especially; both anatomical data and functional, such as innervation. We emphasize this need in our discussion (Lines 352-355 and Lines 405-411).

      More Minor Points:

      (1) As mentioned in the main text, the duration of pips (from pips to bursts) affects the effective bandwidth of the stimulus. I believe that the authors could give an estimate of this effective bandwidth, given what is known from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call, which can now also be estimated.

      (R2. A1) Please see answer A3.4 under Joint Recommendations.

      (2) Figure 5b. Label the green and pink areas as song and heat-call spectrum. Also note that in the legend the authors say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what they meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches, but they probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case, the figure and the legend need clarification.

      (R2. A2) We thank the reviewer for pointing out these issues. We have changed the figure and legend accordingly. In the meantime, we published a paper measuring the in vivo source levels of the heat whistles (Anttonen et al., Current Biology 2025), and we have carefully gone through this manuscript to incorporate those findings and adjust the text accordingly. We have therefore adjusted the analysis in Fig 3B to correct for the narrower frequency distribution of the heat whistles.

      (3) Figure 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". The authors would not want readers to think that they used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (R2. A3) We thank the reviewer for pointing out these issues We mention the frequency sweep in the legend for panel A, but decided against including this in the figure to prevent too much clutter. We have changed the figure = legend to (new text underlined):

      Legend Fig 3. “A Setup used to measure sound-induced vibrations of eggs. A 94 dB, 0.25-to-10 kHz frequency sweep was played at the eggs to determine the vibration transfer function.”

      (4) Methods. In the description of the stimulus, the authors describe "5ms long tone bursts", but these are the tone pips in the main part of the manuscript. Use the same terms.

      (R2. A4) Thank you for catching this, we have changed this into “5ms long tone pips".

      Reviewer #3 (Public review):

      Summary

      Following recent findings that exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird, this study attempts to estimate the hearing capacities of zebra finch nestlings and the perception of high frequencies in that species. It also tries to estimate whether airborne sound can make zebra finch eggs vibrate, although this is not relevant to the question.

      Strength

      That prenatal sounds can affect the development of altricial birds clearly challenges the long-held assumption that altricial avian embryos cannot hear. However, there is currently no data to support that expectation. Investigating the development of hearing in songbirds is therefore important, even though technically challenging. More broadly, there is accumulating evidence that some bird species use sounds beyond their known hearing range (especially towards high frequencies), which also calls for a reassessment of avian auditory perception.

      Weaknesses

      Rather than following validated protocols, the study presents many experimental flaws and two major methodological mistakes (see below), which invalidate all results on responses to frequencyspecific tones in nestlings and those on vibration transmission to eggs, as well as largely underestimating hearing sensitivity. Accordingly, the study fails to detect a response in the majority of individuals tested with tones, including adults, and the results are overall inconsistent with previous studies in songbirds. The text throughout the preprint is also highly inaccurate, often presenting only part of the evidence or misrepresenting previous findings (both qualitatively and quantitatively; some examples are given below), which alters the conclusions.

      Conclusion and impact

      The conclusion from this study is not supported by the evidence. Even if the experiment had been performed correctly, there are well-recognised limitations and challenges of the method that likely explain the lack of response. The preprint fails to acknowledge that the method is well-known for largely underestimating hearing threshold (by 20-40dB in animals) and that it may not be suitable for a 1-gram hatchling. Unlike what is claimed throughout, including in the title, the failure to detect hearing sensitivity in this study does not invalidate all previous findings documenting the impacts of prenatal sound and noise on songbird development. The limitations of the approach and of this study are a much more parsimonious explanation. The incorrect results and interpretations, and the flawed representation of current knowledge, mean that this preprint regrettably creates more confusion than it advances the field.

      (R3. A0) We thank the reviewer for their detailed and critical assessment. We agree that establishing auditory sensitivity in very young altricial birds is technically challenging and that careful interpretation of ABR data is essential. We also appreciate the reviewer’s recognition of the importance of obtaining direct physiological data on auditory development.

      However, we respectfully disagree with the reviewer’s central claim that our methodology is flawed or that our conclusions are unsupported. Many of the concerns raised reflect misunderstandings of ABR methodology, selective interpretation of the literature, or assumptions that are not supported by empirical evidence. Below, we address the main points in turn.

      As also stated in our Provisional response, the reviewer’s critique can be distilled into four main arguments:

      (1) ABR cannot be reliably measured in very small animals.

      (2) Our stimulus design (especially 25 ms tone bursts) invalidates frequency-specific results.

      (3) ABR thresholds should be corrected to behavioral thresholds, which would alter conclusions.

      (4) Our findings are inconsistent with prior studies in songbirds.

      We address each of these below before responding point-by-point.

      (1) Suitability of ABR in small animals.

      Reviewer claim: ABR may not be suitable for very small hatchlings.

      This claim is not supported by existing evidence. ABR measures summed neural activity, and signal amplitude depends in part on the distance between neural tissue and recording electrodes. In smaller animals, this distance is reduced, which can increase signal amplitude and improve signal-to-noise ratio.

      Consistent with this, ABR has been successfully recorded in animals substantially smaller than zebra finch hatchlings, including zebrafish (Jørgensen et al., 2012), 10 mm froglets (Goutte et al., 2017) and 5 mm salamanders (Capshaw et al., 2020). It is in fact much more surprising the technique still provides robust signals even in extremely large animals such as Minke whales, where the distance between electrodes and brain is on the decimeter scale (Houser et al., 2024). We have extensive experience of recording ABRs in such small systems.

      Thus, there is no principled reason why ABR would be an invalid method to study auditory sensitivity in zebra finch hatchlings.

      (2) Stimulus design and tone duration

      Reviewer claim: Use of 25 ms tone bursts invalidates frequency-specific results.

      We agree that stimulus duration affects frequency specificity and ABR detectability. However, the reviewer’s assertion that there is a single “correct protocol” (≤5 ms) is inaccurate. In avian ABR studies, stimulus duration varies depending on experimental goals.

      Our choice of 25 ms tone bursts was intentional and necessary to accurately represent low frequencies (down to 250 Hz), ensuring sufficient cycles per stimulus in the plateau segment of 15 ms and minimizing spectral splatter (see auditory brainstem response design considerations discussed in Lauridsen et al., 2021) and our responses below.

      Key clarifications:

      a) Click-evoked ABRs form the basis of our conclusions about onset of hearing, not tone bursts.

      b) Tone bursts were used primarily to assess frequency-dependent maturation, not detect earliest sensitivity.

      c) We explicitly demonstrate that:

      - 25 ms bursts yield higher thresholds (lower sensitivity)

      - 5 ms pips yield lower thresholds and align with published ABR audiograms

      We have now:

      - Further clarified the rationale of stimulus design in the methods (Line 520-530) and added a section in the discussion (Lines 397-411).

      - Included additional comparison between burst and pip datasets (Lines 387-395).

      - Clarified that conclusions about early hearing do not depend on tone-burst data (Fig 4 and Lines 271-294).

      (3) ABR vs behavioral thresholds

      Reviewer claim: Failure to correct ABR thresholds (20–40 dB) invalidates conclusions.

      We agree that ABR thresholds typically overestimate behavioral thresholds. However, we disagree that this invalidates our conclusions.

      Importantly:

      a) We do not replace measured ABR data with corrected values, as this would be methodologically inappropriate.

      b) Instead, we:

      - Present measured ABR thresholds transparently (Fig 1-3)

      - Compare them directly to published behavioral audiograms (Fig 5)

      - Explicitly discuss the expected offset (Lines 413-422)

      In the revised manuscript we:

      a) Add a new figure (Fig. 5) compiling all published ABR and behavioral audiograms

      b) Show that:

      - ABR and behavioral audiograms have similar shapes (Fig 5, new discussion Lines 387-422)

      - Offsets are typically ~20 dB (Line 413-422)

      Crucially, even under conservative corrections:

      - Early hatchlings remain far less sensitive than adults (>54 dB SPL) to clicks.

      - Heat whistle levels remain at or below detection limits even in adults (new figure Fig 4)

      - The developmental gap (>50 dB between adults and 2 DPH hatchlings) remains decisive.

      Thus, incorporating ABR–behavioral differences does not change the central conclusion.

      (4) Consistency with prior literature in developing songbirds.

      Reviewer claim: Results contradict previous studies in developing songbirds.

      We respectfully disagree. The cited studies fall into three categories:

      (1) Behavioral studies (e.g., alarm-call responses)

      (2) Gene expression studies (e.g., ZENK activation)

      (3) Different species with different developmental trajectories

      None of these directly measure auditory sensitivity thresholds in zebra finch embryos or hatchlings.

      We emphasize:

      - Behavioral responses do not provide threshold measurements

      - Neural activation (e.g., ZENK) does not demonstrate functional perception thresholds

      - Cross-species comparisons must consider differences in developmental timing.

      We have now expanded the Discussion to explicitly address these studies and clarify how they relate to our findings (Lines 357-367). Our data are consistent with what is known about the physiology of auditory development in all birds studied so far.

      (5) Final statement

      We have revised the manuscript extensively to:

      - Clarify methodology and experimental design

      - Expand discussion of ABR limitations

      - Incorporate additional literature and comparisons

      - Correct inconsistencies in reporting

      We maintain that our central conclusions—that early zebra finch hatchlings lack detectable auditory brainstem responses and are unlikely to perceive parental heat calls at natural levels—is robust and supported by the data. We go into more detail in our point-by-point rebuttal below.

      Detailed assessment

      For brevity, only some references are included below as examples, using, when possible, those cited in the preprint (DOI is provided otherwise). A full review of all the studies supporting the points below is beyond the scope of this assessment.

      (A) Hearing experiment

      The study uses the Auditory Brainstem Response (ABR), which measures minute electrical signals transmitted to the surface of the skull from the auditory nerve and nuclei in the brainstem. ABR is widely used, especially in humans, because it is non-invasive. However, ABR is also a lot less sensitive than other methods, and requires very specific experimental precautions to reliably detect a response, especially in extremely small animals and with high-frequency sounds, as here.

      (1) Results on nestling frequency sensitivity are invalid, for failing to follow correct protocols:

      (R3. A1.1) We disagree that our protocol is invalid. There is no universal ABR protocol standard in birds. Our approach is consistent with established principles of stimulus design and is validated by:

      - Robust click-evoked responses

      - Consistent developmental trajectories

      - Agreement between pip-based ABR and behavioral audiograms

      We now clarify this explicitly. See response R3. A0 (Stimulus design and tone duration).

      The results on frequency testing in nestlings are invalid, since what might serve as a positive control did not work: in adults, no response was detected in a majority of individuals, at the core of their hearing range, with loud 95dB sounds (Figure S1), when testing frequency sensitivity with "tone burst".

      This is mostly because the study used a stimulation duration 5 times larger than the norm. It used 25ms tone bursts, when all published avian studies (in altricial or precocial birds) used stimulation of 5ms or less (when using subdermal electrodes as here; e.g., cited: Brittan-Powell et al 2004; not cited: Brittan-Powell et al 2002 (doi: 10.1121/1.1494807), Henry & Lucas 2008 (doi: 10.1016/j.anbehav.2008.08.003)). Long stimulations do not make sense and are indeed known to interfere with the detection of an ABR response, especially at high frequencies, as, for example, explicitly tested and stated in Lauridsen et al 2021 (cited).

      (R3. A1.2) ABR with long-duration stimuli were shown previously to work perfectly well in birds and do not interfere with the detection of an ABR response. Longer stimuli have been used, e.g. in the following bird papers:

      - Amin et al., J Neurophysiol 2007 (cited) on zebra finch ABR: 20 ms tone bursts

      - Korneeva et al. (2006), evoked responses from field L in flycatcher (cited by the reviewer): 20 ms

      - Saunders et al. (1973) (cited): 60 ms tone bursts.

      - Larsen ON, Wahlberg M, Christensen-Dalsgaard (2020) Amphibious hearing in a diving bird, the great cormorant (Phalacrocorax carbo sinensis), J Exp Biol, doi:10.1242/jeb.217265: 25 ms tone bursts

      Human ABR has been measured even with long-duration speech signals (duration 40 ms and longer). See for example Binkhamis et al., Ear and Hearing 40: 659-670, 2019.

      Furthermore, the reviewer unfortunately misunderstood some aspects in Lauridsen et al., which tested a specifical method exploiting neural phase locking, and showed that this method has a low-frequency bias because neural phase locking decreases at high frequencies.

      Also, Lauridsen et al clearly show the reason for using 25 ms bursts: because we aimed to measure frequencies down to 250 Hz, we need to have a sufficient number of cycles (3) in the plateau segment to represent the frequency adequately (Lauridsen et al. fig 1). A 15 ms plateau contains 3 cycles, plus 5 ms rise/fall time (1 cycle) equals 25 ms. We have clarified this in our methods (L520-525) and added a section in the discussion (L397-411).

      Thus, long-duration stimulations make sense and have been used before successfully. See also response R3.A0 (Stimulus design and tone duration).

      Adult response was then re-tested with a correct 5ms tone duration ("tone-pip"), which showed that, for the few individuals that responded to 25ms tones, thresholds were abnormally high (c.a. by 30dB; Figure 2C).

      Yet, no nestlings were retested with a correct protocol. There is therefore no valid data to support any conclusion on nestling frequency hearing. Under these circumstances, the fact that some nestlings showed a response to 25ms tones from day 8 would argue against them having very low sensitivity to sound.

      (R3. A1.3) Please see answer A3.1 under Joint Recommendations and R3.A0 (Stimulus design and tone duration).

      (2) Responses to clicks underestimate hearing onset by several days:

      Without any valid nestling responses to tones (see # 1), establishing the onset of hearing is not possible based on responses to clicks only, since responses to clicks occur at least 4 days after responses to tones during development (Saunders et al, 1973). Here, 60% of 4-day-old individuals responding to clicks means most would have responded to tones at and before 2 days post-hatch, had the experiment been done correctly.

      (R3. A2.1) We disagree that clicks necessarily underestimate onset.

      Clicks are broadband stimuli that:

      - Recruit large neural populations

      - Are commonly used to detect early auditory responses

      The cited delay between tone and click responses reflects stimulus energy differences, not an inherent limitation of clicks. We have clarified this in the revision and softened language to refer to “no detectable ABR response” rather than absolute deafness.

      The report that Saunders could only see responses to clicks later than to tones only reflects that he used click amplitudes that were insufficiently high. The ABR responses reported were to extremely intense tones (110 dB SPL) of long duration (60 ms).

      If Saunders had used clicks (duration 60 µs) with comparable sound energy, they would have had been very difficult to produce. He should have used clicks with an amplitude 1000 times (60 dB) higher than the tones to produce the same sound energy. This would have been clicks at 170 dB SPL, equivalent to the sound at the mouth of a medium-sized military cannon. Applying this pressure would not be a recommendable method for hearing assessment, but instead lead to irreversible hearing damage.

      In budgerigars, hearing onset occurs before 5 days post hatch, since responses to both clicks and tones were detectable at the first age tested at 5dph (Brittan-Powell et al, 2004).

      (R3. A2.2) This is not how we interpret the cited paper. They state that ‘Responses were first obtained from 1-week-old at high stimulation, and their click responses (Fig. 1) show no wave 1 peak at 6 days post-hatch. Also, their conclusion (p 3101) states that ‘budgerigars probably cannot hear at hatching’.

      (3) Experimental parameters chosen lower ABR detectability, specifically in younger birds: Very fast stimulus repetition rate inhibits the ABR response, especially in young:

      (a) The stimulus presentation rate (25 stim/ sec) is 6 times faster than zebra finch heat-calls, and 5 to 25 times faster than most previous studies in young birds (e.g., cited: Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec).

      Faster rates saturate the neurons and accordingly are known to decrease ABR amplitude and increase ABR latency, especially in younger animals with an immature nervous system.

      In birds, this occurs especially in the range from 5 to 30 stim/sec (e.g., cited: Saunder et al 1973, Brittan-Powell et al 2004). Values here with 25 rather than 1-4 stim/min are therefore underestimating true sensitivity.

      (R3. A3a) Please see answer A3.3 under Joint Recommendations.

      (b) Averaging over only 400 measures is insufficient to reliably detect weak ABR signals: The study uses 2 to 3 times fewer measures per stimulation type than the recommended value of 1,000 (e.g., Brittan-Powell et al 2002, 2024; Henry & Lucas 2008). This specifically affects the detection of weak signals, as in small hatchlings with tiny brains (adult zebra finches are 12-14g).

      (R3. A3b) Please see answer A3.2 under Joint Recommendations.

      (c) Body temperature is not specified and strongly affects the ABR:

      Controlling the body temperature of hatchlings of 1-4 grams (with a temperature probe under a 5mm-wide wing) would be very challenging. Low body temperature entirely eliminates the ABR, and even slight deviance from optimal temperature strongly increases wave latency and decreases wave amplitude (e.g., cited: Katayama 1985).

      (R3. A3c) Please see answer A2 under Joint Recommendations.

      (d) Other essential information is missing on parameters known to affect the ABR: This includes i) the weight of the animals,

      (R3. A3d-i) These important details have now been added in Table S6.

      (ii) whether and how the response signal was amplified and filtered,

      (R3. A3d-ii) Signal was amplified 500 times (74 dB). We have included these important details to the methods (Line 508).

      (iii) how the automatised S/N>2 criteria compared to visual assessment for wave detection,

      (R3. A3d-iii) There is no universally accepted/fitting protocol for performing ABR recordings in various animals, we decided to perform both visual and automated criteria detection of thresholds. The automated criterion in our experience is more strict approach than visual detection of thresholds, because using an automated criteria for threshold detection will remove potential experimenter bias from the results. We added this in our methods (Lines 583-585).

      (iv) what measures were taken to allow the correct placement of electrodes on hatchlings less than 5 grams.

      (R3. A3d-iv) We have placed electrodes in much smaller animals than 5 grams, and the common landmarks (ear opening, midline of skull) could easily be identified in the hatchlings.

      (4) Results in adults largely underestimate sensitivity at high frequencies, and are not the correct reference point:

      (a) Thresholds measured here at high frequencies for adults (using the correct stimulus duration, only done on adults) are 10-30dB higher than in all 3 other published ABR studies in adult zebra finches (cited: Zevin et al 2004; Amin et al 2007; not cited: Noirot et al 2011 (10.1121/1.3578452)), for both 4 and 6 kHz tone pips.

      (b) The underlying assumption used throughout the preprint that hearing must be adult-like to be functional in nestlings does not make sense. Slower and smaller neural responses are characteristic of immature systems, but it does not mean signals are not being perceived.

      (R3. A4) We acknowledge variation across studies and now include a comprehensive comparison of all zebra finch audiograms (new Fig. 5 and discussion Lines 387-422).

      Importantly:

      - Our pip-based audiogram aligns with previous ABR studies

      - Differences in high-frequency sensitivity likely reflect methodological variation or population differences.

      Our conclusions rely on relative developmental changes, not absolute thresholds.

      (5) Failure to account for ABR underestimation leads to false conclusions:

      (a) Whether the ABR method is suitable to assess hearing in very small hatchlings is unknown. No previous avian study has used ABR before 5 days post-hatch, and all have used larger bird species than the zebra finch.

      (R3. A5a) As stated above (R3.A0), small animals should give better signals, and we have been able to measure ABR in much smaller animals previously.

      (b) Even when performed correctly on large enough animals, the ABR systematically underestimates actual auditory sensitivity by 20-40 dB, especially at high frequencies, compared to behavioural responses (e.g., none cited: Brittan-Powell et al 2002, Henry & Lucas 2008, Noirot et al 2011). Against common practice, the preprint fails to account for this, leading to wrong interpretations.

      (R3. A5a) See our answer to R3.A0 above.

      For example, in Figure 1G (comparing to heat call levels), actual hearing thresholds would be 3040dB below those displayed. In addition, the "heat whistle" level displayed here (from the same authors) is 15dB lower than their second measure that they do not mention, and than measures obtained by others (unpublished data). When these two corrections are made - or even just the first one - the conclusion that heat-call sound levels are below the zebra finch hearing threshold does not hold.

      (R3. A5a) Our conclusion that heat whistles are unlikely to be perceived does not rely on a single dataset or method, but on the convergence of three independent constraints: (i) signal amplitude, (ii) adult auditory sensitivity, and (iii) developmental immaturity of the auditory system.

      First, heat whistles are low-amplitude signals. Our in vivo measurements show levels of ~33 dB re 20 µPa at 10 cm and ~14 dB at 1 m (Anttonen et al., 2025, Curr Biol). Even allowing for uncertainty in near-field estimation, a conservative upper bound at very close range (<5 cm) is ~40 dB SPL.

      Second, the most sensitive available measure—behavioral audiograms—places adult zebra finch thresholds at ~40 dB SPL at ~6 kHz (Okanoya and Dooling, 1987, J Comp Psychol), increasing steeply toward higher frequencies. Thus, even under optimal conditions, heat whistles fall at or below the detection threshold of adults, and only potentially at very close range.

      Third, auditory sensitivity in early development is substantially reduced. Our ABR data show a ≥40–60 dB decrease in sensitivity in hatchlings relative to adults for click stimuli, which provide the most favorable conditions for eliciting responses. Because frequency-specific sensitivity develops later, thresholds at 6–8 kHz are expected to be even higher in hatchlings and embryos.

      Taken together, these constraints define a narrow and unfavorable detection window: a low-amplitude, high-frequency signal positioned at the edge of adult sensitivity, combined with a large developmental decrease in auditory sensitivity. Under these conditions, it is unlikely that heat whistles are detectable by hatchlings or embryos.

      Importantly, this conclusion does not depend on precise correction factors between ABR and behavioral thresholds. Even when considering the most sensitive behavioral data and conservative estimates of sound level, the signal remains at or below the limits of detection in adults, and far below expected sensitivity in early developmental stages.

      We have included this argument more clearly in our discussion (Line 271-367), illustrated by new Fig. 5.

      (c) Rather than making appropriate corrections, the preprint uses a reference in humans (L180), where ABR is measured using a much more powerful method (multi-array EEG) than in animals, and from a larger brain. The shift of "10-20dB" obtained in humans is not applicable to animals.

      (R3. A5c) Again our conclusions rely on relative developmental changes, not absolute thresholds. The clinical practice in humans to measure ABR is with 4 electrodes, not a multi-electrode EEG array.

      Animal studies where ABR audiograms have been compared directly to psychophysical audiograms show differences of around 20 dB. For example in Brittain powell et al 2002, audiogram comparisons between behavioral and ABR in budgerigars were made within the same lab, same animal population and by the same people, leading to 20 dB difference. Our discussion includes a new paragraph on this topic (Lines 413-422).

      (6) Results are inconsistent with previous findings in developing songbirds:

      (R3. A6) We now explicitly discuss all cited studies. Key points:

      - Early behavioral responses do not imply high-frequency sensitivity

      - Studies in other species do not directly translate to zebra finches

      - None of the cited work provides direct measures of auditory thresholds in embryos

      As expected from all of the above, results and conclusions in the preprint are inconsistent with findings in other songbirds, which, using other methods, show for example, auditory sensitivity in: a) zebra finch embryos, in response to song vs silence (not cited: Rivera et al 2018, doi: 10.1097/WNR.0000000000001187)

      (R3. A6a) We thank the reviewer for pointing out this study. We agree that the question of auditory responsiveness in embryos is important, and that a range of approaches have been used to address it. However, the study cited (Rivera et al., 2018) does not directly measure auditory sensitivity, but instead infers auditory processing from differences between treatment groups exposed to different acoustic conditions. As such, it is not directly comparable to physiological measures of hearing sensitivity, such as ABR or behavioral thresholds.

      In addition, interpretation of these results is complicated by limited characterization of the acoustic environment and differences in experimental handling between groups, which may introduce confounding factors unrelated to auditory perception. Given these considerations, and because our study focuses specifically on quantifying auditory sensitivity using established physiological methods, we have chosen not to include a detailed discussion of this work.

      (b) flycatcher hatchlings at 2-3d post hatch (first age tested), across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB) (cited: Aleksandrov and Dmitrieva 1992, not cited: Korneeva et al 2006 (10.1134/S0022093006060056)).

      (R3. A6b) Korneeva et al. (2006) and Aleksandrov and Dmitrieva (1992) report evoked responses in very young flycatcher hatchlings across a broad frequency range. Notably, these measurements were obtained using more invasive recording approaches (e.g., implanted electrodes in Field L) in unanesthetized birds, which are known to yield lower thresholds compared to far-field ABR recordings under anesthesia. These methodological differences likely account for part of the higher sensitivity reported.

      Importantly, even in flycatchers, auditory sensitivity shows substantial postnatal improvement: thresholds decrease by up to ~40 dB over the first days after hatching, and the upper frequency limit expands from ~4 to ~7 kHz. Thus, while absolute sensitivity may differ across species and methods, the overall developmental trajectory—gradual improvement in sensitivity and progressive extension toward higher frequencies—is consistent with our findings and with broader patterns reported in songbirds.

      We have added this paper in our discussion (Lines 362-365).

      (c) songbird nestlings at 2-6d post hatch, which discriminate and behaviourally respond to relevant parental calls or even complex songs. This level of discrimination requires good hearing across frequencies (e.g., not cited: Korneeva et al 2006; Schroeder & Podos 2023 (doi: 10.1016/j.anbehav.2023.06.015)).

      (R3. A6c) The species mentioned are different species from our study species. In the Pied flycatchers (Korneeva et al. 2006) experimental conditions were different: recordings were made from unanesthetized nestlings with implanted electrodes directly in the brain (field L), so likely with better SNR. The audiograms show a 10 dB SPL threshold after day 11, so the species may be considerably more sensitive than the zebra finch. The swamp sparrows in the Schroeder and Podos (2023) behavioral study were exposed for 4 days starting at 4-7 days post-hatch, so the study does not address embryonal hearing.

      (d) zebra finch nestlings at 13d post-hatch, which show adult-like processing of songs in the auditory cortex (CNM) (Schroeder & Remage-Healey 2021, doi: 10.1002/dneu.22802).

      (R3. A6d) This study does not conflict with our data. Even though sensitivity is lower at 10 days than in adults, cortical processing could still be ‘adult-like’.

      (e) zebra finch juveniles, which are able to perceive and learn song syllables at 5-7kHz (fundamental frequency) with very similar acoustic properties to heat calls, and also produced during inspiration (Goller & Daley 2001, doi: 10.1098/rspb.2001.1805).

      (R3. A6e) This result is not in conflict with our data. First, the onset of song learning occurs earliest at 20 DPH as discussed in the paper and our work demonstrates that click-evoked ABR thresholds are adult-like at 20 DPH. In the cited paper, the tutoring experiments were initiated at 35 DPH so the auditory system of the studied juveniles is mature.

      Second, even though Goller & Daley 2001 do not report the source level of the specific syllable or the playback sound pressure levels, the source level of the inspiratory notes is comparable to other syllables, and thus around ~67 dB and ~34 dB louder than heat whistles.

      NONE of these results - which contradict results and claims in the preprint - are mentioned.

      Instead, the preprint focuses on very slow-developing species (parrots and owls), which take 2-4 times longer than songbirds to fledge (cited: Brittan-Powell et al 2004; Köppl & Nickel 2007; Kraemer et al 2017).

      (R3. A6f) We have included papers in our discussion that reflect the known data (to our best knowledge) on the developmental neurophysiology and neuroanatomy of the auditory system and not proxies thereof.

      (7) Results in figures are misreported in the text, and conclusions in the abstract and headers are not supported by the data:

      For example:

      (a) The data on Figure 1E shows that at 4 days old, 8 out of 13 nestlings (60%) responded to clicks, but the text says only 5/13 responded (L89).

      (R3. A7a1) We apologize for this typo. Corrected.

      When 60% (4dph) and 90% (6dph) of individuals responded, the correct term would be that "most animals", rather than "some animals" responded (L89).

      (R3. A7a2) We have rephrased this sentence into “observable in most animals during” as suggested.

      Saying that ABR to loud sound appeared "in the majority only after one week" (L93) is also incorrect, given the data.

      (R3. A7a3) We have rephrased this sentence into: “Thus, sounds at loud, yet physiologically relevant SPLs do not evoke ABRs in the first days after hatching, but do so in all animals at 8 DPH.” (Lines 97-99).

      It follows that the title of the paragraph is also erroneous.

      (R3. A7a4) The paragraph title supports our conclusions and we will keep it.

      (b) The hearing threshold is underestimated by 40dB at 6 and 8Kz on Fig 2C, not by "10-20dB" as reported in the text (L178).

      (R3. A7b) We have changed the title of this section and moved the last sentence to the discussion to remove the focus on heat whistles. We added a paragraph in the discussion to specifically address the difference between ABR and behaviorally measured audiograms (Lines 413-422).

      (B) Egg vibration experiment

      (8) Using airborne sound to vibrate eggs is biologically irrelevant:

      (R3. A8.1) We agree that parental contact could influence vibration transmission.

      However, (1) prior studies assume airborne sound transmission, and (2) our experiment tests this assumption directly. We now clarify this scope (Lines 328-342 and Figure 4) and discuss contact-based transmission as a potential future direction.

      The measurement of airborne sound levels to vibrate eggs misunderstands bone conduction hearing and is not biologically meaningful: zebra finch parents are in direct contact with the eggs when producing heat calls during incubation, not hovering in front of the nest. This misunderstanding affects all extrapolations from this study to findings in studies on prenatal communication.

      (R3. A8.2) The definition of bone conduction is the response to sound that is not mediated by a functional middle ear, but through the skull. In the earlier study, the eggs were stimulated by sound from a headphone, so that is the reason for using the same stimulation here. See also joint response A3.5 above.

      (C) Misrepresentation of current knowledge

      (9) Values from published papers are misreported, which reverses the conclusions:

      (R3. A9) We thank the reviewer for identifying inconsistencies and have:

      - Corrected heat whistle frequency ranges consistently through our paper

      - Added a comprehensive comparison figure gathering all available audiograms (Fig 5)

      - Expanded discussion of high-frequency hearing.

      These revisions do not alter our conclusions.

      Most critical examples:

      (a) Preprint: "Zebra finch most sensitive hearing range of 1-to-4 kHz (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)" (L173).

      Actual values in the studies cited are:

      1-to-7kHz, in Amin et al 2007 (threshold [=50dB with ABR] is the same at 7kHz and 1KHz).

      1-to-6 kHz, in Okanoya and Dooling (the threshold [=30dB with behaviour] is actually lower at 6kHz than at 1KHz).

      1-to-7kHz, in Yeh et al (threshold [=35-38dB with behaviour] is the same at 7kHz and 1KHz).

      (R3. A9a.1) In this sentence presenting our results (“sensitive hearing range of 1-to-4 kHz”) we originally wrote that “these are consistent with the following papers (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)". This latter part was left out during the writing process. This explains the different numbers. We apologize for this mistake.

      To avoid confusion, in our revision we have placed all ABR curves together into new Fig 5 and have included a new paragraph to discuss the differences (Lines 387-422).

      Note that zebra finch nestlings' begging calls peaking at 6kHz (Elie & Theunissen 2015, doi: 10.1007/s10071-015-0933-6), would fall 2kHz above the parents' best hearing range if it were only up to 4kHz.

      (R3. A9a.2) Of course that is possible. However this representation is incorrect because begging calls are harmonic sounds with a fundamental frequency around 500 Hz and formant at 6 kHz. Begging calls thus contain lots of energy at frequencies below 6 kHz, while the heat whistles do not. The peak frequency of heat whistles is also their lowest frequency component.

      (b) The preprint incorrectly states throughout (e.g., L139, L163, L248) that heat-calls are 7-10kHz, when the actual value is 6-10kHz in the paper cited (Katsis et al, 2018).

      (R3. A9b) The authors in Katsis et al. 2018 provided a range of 6-10 kHz estimated from the spectrogram without any further specification of methods. In another manuscript, we have quantified the heat whistle frequency (Anttonen et al Curr Biol https://doi.org/10.1016/j.cub.2025.08.054) to be 6.8 ± 0.6 kHz. We have changed this accordingly throughout our manuscript.

      (c) Using the correct values from these studies, and heat-calls at 45 dB SPL (as measured by others (unpublished data), or as measured by the authors themselves, but which is not reported here (Anttonen et al 2025), the correct conclusion is that heat calls fall within the known zebra finch hearing range.

      (R3. A9c) Please see our answer R3.A5a. We have included this argument more clearly in our discussion (Lines 271-355), illustrated by new Fig. 4.

      (10) Published evidence towards high-frequency hearing, including in early development, is systematically omitted:

      (a) Other studies showing birds use high frequencies above the known avian hearing range are ignored. This includes oilbirds (7-23kHz; Brinklov et al 2017; by 1 of the preprint authors, doi: 10.1098/rsos.170255) and hummingbirds (10-20kHz; Duque et al 2020, doi: 10.1126/sciadv.abb9393), and in a lesser extreme, zebra finches' inspiratory song syllables at 57kHz (Goller & Dalley, 2001).

      (R3. A10a) We agree that some bird species produce or use acoustic signals extending into high frequencies. However, signal production is not evidence of perceptual sensitivity. Many animals, including birds and mammals, produce signals that contain harmonic or broadband components extending beyond their most sensitive hearing range without implying functional detection at those frequencies.

      The cited examples (oilbirds, hummingbirds, inspiratory song syllables in zebra finches) concern signal production or ecological specializations in different species, not measured auditory sensitivity in zebra finches, and particularly not during early development. As such, they do not provide evidence that zebra finches—adults or embryos—can detect low-amplitude, narrowband signals in the 6–7 kHz range.

      Our study explicitly addresses auditory sensitivity using physiological measurements, which is the relevant metric for evaluating detectability.

      (b) The discussion of anatomical development (L228-241) completely omits the well-known fact that the avian basilar papilla develops from high to low frequencies (i.e., base to apex), which - as many have pointed out - is opposite to the low-to-high development of sensitivity (e.g., cited: Cohen & Fermin 1978; Caus Capdevila et al 2021).

      (R3. A10b) We agree that the avian basilar papilla develops from base to apex (high to low frequency). We have now added a sentence in the Discussion to acknowledge this (Lines 406411).

      Importantly, morphological development does not directly translate to functional sensitivity. Functional hearing depends critically on factors such as hair cell innervation, synaptic maturation, and central auditory processing, which are known to develop over time.

      Our data show a low-to-high frequency progression in functional sensitivity, consistent with previous physiological studies. This apparent mismatch between anatomical gradients and functional onset has been noted in other systems and likely reflects the later maturation of neural encoding rather than hair cell differentiation per se. We now clarify this distinction in the revised manuscript (Lines 406-411).

      (c) High frequency hearing in songbirds at hatching is several orders of magnitude better than in chickens and ducks at the same age, even though songbirds are altricial (e.g., at 4kHz, flycatcher: 47dB, chicken-duck: 90dB; at 5kHz, flycatcher: 65dB, chicken-duck: 115dB; Korneeva et al 2006, Saunders et al 1974). That is because Galliformes are low-frequency specialists, according to both anatomical and ecological evidence, with calls peaking at 0.8 to 1.2kHz rather than 2-6kHz in songbirds. It is incorrect to conclude that altricial embryos cannot perceive high frequencies because low-frequency specialist precocial birds do not (L250;261).

      (R3. A10c) We agree that species differ in their auditory ecology and frequency specialization, and we do not claim that all altricial birds share identical developmental trajectories.

      However, the cited comparisons involve different species, methodologies, and developmental timelines, which limits their direct comparability. In particular:

      Developmental staging is not directly comparable across species using days post-hatch alone.

      - Different methods (e.g., invasive recordings vs. ABR vs behavioural assays) yield systematically different thresholds.

      - Ecological specialization (e.g., low-frequency vs. broadband species) influences adult audiograms and likely developmental trajectories.

      We have revised the Discussion to explicitly acknowledge these limitations and to avoid overgeneralization across species. Importantly, our conclusions are based on within-species comparisons (adult vs. hatchling zebra finches) combined with measured signal levels of heat whistles. These constraints are sufficient to evaluate detectability without relying on cross-species extrapolation.

      (11) Incorrect statements do not reflect findings from the references cited For example:

      (a) "in altricial bird species hearing typically starts after hatching" (L12, in abstract), "with little to no functional hearing during embryonic stages (Woolley, 2017)." (L33).

      There is no evidence, in any species, to support these statements. This is only a - commonly repeated - assumption, not actually based on any data. On the contrary, the extremely limited evidence to date shows the opposite, with zebra finch embryos showing ZENK activation in the auditory cortex in response to song playback (Rivera et al, 2018, not cited).

      The book chapter cited (Woolley 2017) acknowledges this lack of evidence, and, in the context of song learning, provides as only references (prior to 2018), 2 studies showing that songbirds do not develop a normal song if the song tutor is removed before 10d post-hatch. That nestlings cannot memorise (to later reproduce) complex signals heard before d10 does not mean that they are deaf to any sound before day 10.

      Studies showing hearing in young songbird nestlings (see point 6 above) also contradict these statements.

      (R3. A11a) We agree that the precise onset of hearing in altricial embryos is not well established. We have therefore revised the wording in the Abstract and Introduction to avoid categorical statements and instead reflect the limited available evidence (Lines 13-16 and 33-37).

      Our data provide direct physiological measurements showing extremely low sensitivity immediately after hatching, which constrains the likelihood of functional hearing in earlier embryonic stages.

      Regarding the cited ZENK study, we note that immediate early gene expression indicates neural activation but does not provide a measure of auditory sensitivity or detection thresholds. As such, it cannot be directly compared to physiological or behavioral measures of hearing.

      (b) "Zebra finch embryos supposedly are epigenetically guided to adapt to high temperatures by their parents high-frequency "heat calls" " (L36 and L135).

      This is an extremely vague and meaningless description of these results, which cannot be assessed by readers, even though these results are presented as a major justification for the present study. Rather than giving an interpretation of what "supposedly" may occur, it would be appropriate to simply synthesize the empirical evidence provided in these papers. They showed that embryonic exposure to heat-calls, as opposed to control contact calls, alters a suite of physiological and behavioural traits in nestlings, including how growth and cellular physiology respond to high temperatures. This also leads to carry-over effects on song learning and reproductive fitness in adulthood.

      (R3. A11b) We thank the reviewer for raising this point. In the revised manuscript, we have replaced the previous phrasing with a more precise and neutral summary of what these studies report, namely that embryonic exposure to heat-call playbacks has been associated with differences in physiological and behavioral traits.

      Our study, however, addresses a distinct question—whether such acoustic signals are detectable by embryos given known constraints on signal amplitude and auditory sensitivity. The cited studies do not directly quantify auditory perception or the physical sound environment experienced by embryos. As a result, they do not provide a direct test of the sensory mechanism required for acoustic communication. A detailed evaluation of experimental design and interpretation in those studies is beyond the scope of the present manuscript, and we therefore limit our discussion to assessing the biophysical and physiological plausibility of the proposed mechanism.

      (c) "The acoustic communication in precocial mallard ducks depends specifically on the lowfrequency auditory sensitivity of the embryo (Gottlieb, 1975)" (L253)

      The study cited (Gottlieb, 1975) demonstrates exactly the opposite of this statement: it shows that duckling embryos, not only perceive high frequency sounds (relative to the species frequency range), but also NEED this exposure to display normal audition and behaviour post-hatch. Specifically, it shows that duckling embryos deprived of exposure to their own high-frequency calls (at 2 kHz), failed to identify maternal calls post-hatch because of their abnormal insensitivity to higher frequencies, which was later confirmed by directly testing their auditory perception of tones (Dimitrieva & Gottlieb, 1994).

      (R3. A11c) We thank the reviewer for this clarification and have revised the relevant text. Our intention was to highlight that embryonic auditory experience can shape postnatal behavior, not to imply strict low-frequency limitation. Therefore we already included the actual frequency in the original sentence. We have removed the non-descriptive term “low-frequency” (Lines 330-332).

      (12) Considering all of the mistakes and distortions highlighted above, it would be very premature to conclude, based on these results and statements, that altricial avian embryos are not sensitive to sound. This study provides no actual scientific ground to support this conclusion.

      (R3. A12) We respectfully disagree with the reviewer’s conclusion.

      Our study does not make a general claim that altricial embryos are incapable of perceiving sound. Rather, we evaluate a specific hypothesis: whether zebra finch embryos and hatchlings can detect sound and parental heat whistles.

      Our conclusions are based on the convergence of:

      (1) Measured low sound pressure levels of heat whistles,

      (2) Established adult auditory thresholds (behavioral data),

      (3) A large developmental decrease in auditory sensitivity demonstrated by our ABR measurements.

      Even under conservative assumptions, these constraints place heat whistles at or below adult detection thresholds and far below expected sensitivity in hatchlings and embryos.

      Thus, our conclusion is not based on absence of evidence, but on quantitative constraints that make detection unlikely under biologically realistic conditions.

      Recommendations for the authors:

      Joint recommendations:

      In response to the joint recommendations, we have:

      - Expanded methodological transparency (temperature, electrode setup, stimulus parameters),

      - Added new data (Fig S3) and figures (Fig 4 and 5),

      - Clarified ABR limitations and interpretation,

      - Strengthened the separation between measured results and interpretation,

      - Reframed conclusions to avoid overstatement.

      These revisions leave the central two conclusions unchanged: 1) zebra finch hatchlings and embryos are functionally deaf, and 2) under biologically realistic conditions, heat whistles are unlikely to be detectable by zebra finch hatchlings or embryos.

      (A) Reviewers 1 and 2:

      Much of the reviewer discourse revolved around providing clarifications of methodology for measuring the ABR and caveats for interpretation. There was near consensus with reviewers 1 and 2 on issues related to the ABR, which should be addressed.

      We appreciate the reviewers’ consensus that the main conclusions are supported, while requesting clarification of methodological details and interpretation of ABR measurements.

      (1) Please address all of the issues raised by reviewers 1 and 2 above.

      (A1) All points raised by Reviewers 1 and 2 have been addressed in detail in our point-by-point rebuttal below. In addition, we have revised the manuscript to improve clarity, added new figures (Fig. 4, 5), and substantially expanded the Discussion with eight new paragraphs to better contextualize our findings.

      (2) Please also

      - clarify all aspects of experimental details of the ABR that were missing, including temperature control (estimate body and ambient temperatures during ABR recordings,

      - please address the possibility of hypothermia of hatchlings that could have reduced ABR responses,

      - and potential local head cooling due to surgical exposure and its likely effect on highfrequency response depression).

      (A2) In our revision, we have expanded the Methods section (L479-485) and added the following new data:

      Body and ambient temperature/hypothermia

      We have now included the body temperatures during ABR recordings in new table S6. These data show that:

      (1) Body temperature was stable throughout recordings,

      (2) Temperatures were within the physiological range,

      (3) Conditions were consistent across all age groups.

      Importantly, even the youngest hatchlings maintained stable temperatures and showed no indication of hypothermia. Therefore, differences in ABR responses cannot be attributed to temperature effects.

      Potential cooling due to surgical exposure

      This concern does not apply to our experiments. We used subdermal needle electrodes, which do not require surgical exposure. Therefore, no local cooling of the head occurred, and no tissue exposure could affect high-frequency sensitivity. We have added a clarifying sentence in the Methods section to explicitly state this (Line 501-503).

      (B) Reviewer 3 also had additional requests for clarification that should also be addressed:

      (3.1) Stimulus duration too long: The study used 25 ms tone bursts instead of the standard {less than or equal to} 5 ms "pips." Could this prevent reliable ABR detection, especially at high frequencies?

      (A3.1) We agree that stimulus duration affects ABR characteristics and now clarify our rationale in the manuscript.

      - The 25 ms tone bursts were deliberately chosen to ensure sufficient cycle representation at low frequencies (down to 250 Hz) and to avoid frequency splatter.

      - Using a constant duration across frequencies ensures comparable stimulus energy.

      Importantly:

      - The 25 ms data yield audiogram shapes consistent with both click responses (Fig 2C) and published behavioral data (new Fig 5).

      - To address potential high-frequency limitations, we included a dataset using 5 ms tone pips, which produced thresholds consistent with published ABR studies (new Fig 5).

      Thus, both stimulus types support the same conclusion: a gradual maturation of hearing sensitivity from low to high frequencies. We have expanded the Discussion with three paragraphs to clarify these methodological trade-offs (Lines 387-422).

      (3.2) Were 400 sweeps enough averaging? Might a signal appear at 1000 or more?

      (A3.2) Signal-to-noise ratio improves with the square root of the number of averages. Increasing from 400 to 1000 sweeps would therefore reduce thresholds by at most ~4 dB. This magnitude is small relative to the >54 dB developmental differences observed, and the large gap between signal levels and detection thresholds. Thus, increasing sweep number would not alter the conclusions. We now clarify this explicitly in the Methods (L530).

      (3.3) Was the repetition rate too high? How does the stimulus presentation affect the ABR? Might a signal have emerged with 1-4 per second?

      (A3.3) We have clarified stimulus presentation rates in the revised manuscript:

      - Clicks were presented at 25 Hz. Control measurements (now included as Supplementary Fig. S3) show no effect of this rate on ABR amplitude or threshold.

      - Tone bursts and pips were presented at ~3 Hz, consistent with commonly used rates that avoid neural adaptation. We apologize for leaving this out in our original submission.

      We now explicitly describe these parameters and their rationale in the Methods (Lines 543-552).

      (3.4) If possible, provide an estimate of the effective bandwidth of the tone pips and compare it with the bandwidth of the parental heat-whistles.

      (A3.4) We agree that stimulus bandwidth differs between tone pips and heat whistles, and that broader signals may stimulate multiple auditory filters. Shorter stimuli (e.g., 5 ms pips) have broader bandwidth and may stimulate multiple filters—particularly at low frequencies—potentially lowering thresholds, whereas longer stimuli (25 ms bursts) are more frequency-specific and may yield higher thresholds. At higher frequencies (including the heat whistle range), this effect is expected to be smaller.

      However, quantitative correction is currently not possible due to a lack of species-specific data on auditory tuning curves in zebra finches. The only available avian data (budgerigar; Saunders et al., 1979) suggest auditory filter bandwidths (Q10 dB, i.e., the bandwidth 10 dB below the peak divided by peak frequency) of ~1.4 at low frequencies and ~1000 Hz at higher frequencies, but how multifilter stimulation affects thresholds is unknown and likely species- and frequency-dependent.

      Given these uncertainties, direct comparison between tone stimuli and heat whistles requires strong assumptions. We therefore suggest that future studies should measure responses to natural heat whistles directly.

      (3.5) Egg Vibration Experiment. Address the possibility that if a parent were physically lying on top of an egg and generated a heat call, parental body vibration could significantly communicate some perceptual vibrotactile signal to the egg. Reviewer 3 raised the possibility that the experiments in this paper tested the extent to which an auditory input can vibrate the egg - what if a vocalizing bird was on the egg?

      (A3.5) We agree that embryos may receive multiple types of sensory input from parents, including direct mechanical cues.

      However, our experiment specifically tests the hypothesis proposed in prior work: that airborne sound (heat whistles) induces egg vibrations sufficient for perception. Our findings show that airborne sound-induced vibrations are orders of magnitude below known vibrotactile sensitivity thresholds.

      Regarding parental contact:

      - Heat whistles are produced by an aerodynamic whistle mechanism, not tissue vibration (Anttonen et al., Curr Biol 2025), meaning most respiratory energy is radiated as sound rather than dissipated as heat/vibration in the body.

      - A parent sitting on the egg would attenuate airborne sound transmission, not amplify it.

      We now clarify in the Discussion that other cues (e.g., respiration, direct contact, temperature) may exist and need to be included in new experiments (Lines 349-352). Even so, these are distinct mechanisms and were not the hypothesis tested in prior playback studies.

      In our paper, we will not add a detailed discussion of these prior papers as this is outside the scope of this paper. Instead, we added a paragraph what would be a constructive way forward (Lines 349-355). Hopefully somebody in the community will have the good fortune to secure research funding to continue this benchmarking work.

      (4) Finally, all reviewers agreed that some more context on the ABR and its relationship to functional hearing could be provided, with less direct focus on the heat-call experiments.

      (A4) We agree and in the revision discussion have compiled all published ABR and behavioral audiograms (Fig. 5) and added new paragraphs on functional hearing (Lines 261-269), and ABR vs behavioural audiograms (Lines 413-422).

      Furthermore, to remove focus on the heat whistles, we have moved all heat-whistle-specific interpretation out of the Results into a single, focused Discussion section (Line 271-355).

      The behavioral studies mentioned below show auditory responses (e.g., begging suppression). However, these behaviors are tested between ~5–10 days post-hatch (consistent with our findings), in different species, and do not provide quantitative sensitivity thresholds, nor do they address detectability of low-amplitude, high-frequency signals like heat whistles.

      In our revision we added a new discussion paragraph including these behavioral studies (Lines 357-367).

      For example, there are ample cases in the literature of altricial birds exhibiting behavioral evidence of auditory sensitivity by reducing begging calls in response to parental alarm calls:

      Platzen & Magrath (2004) - Playback of parental alarm calls nearly abolished nestling non-begging calls and reduced begging in scrubwrens. Proc. R. Soc. B 271:1271-1276.

      Different species: scrubwrens. Playback age: 5-, 8- and 11-DPH nestlings.

      Magrath, Haff, Horn & Leonard (2006) - Review and experiments on the developmental shift to silence/freeze after aerial alarm calls as chicks become fledglings; documents nestling quieting to alarms. Proc. R. Soc. B 273:2335-2341.

      Different species: scrubwrens. Playback age: 7-9 DPH nestlings, and 2- 4 days after fledging.

      Magrath, Pitcher & Dalziell (2007) - Nestlings respond to the sound of a predator's footsteps and parental food/alarm calls; includes begging suppression following predator sounds. Anim. Behav. 74:1117-1129.

      Different species: scrubwrens. Playback age: 8 DPH nestlings.

      Haff & Magrath (2012) - Nestlings suppress calling after heterospecific alarm calls (when acoustically similar to conspecific alarms), indicating generalized auditory danger recognition. Anim. Behav. 84:e.g., 495-505 (article).

      Different species: scrubwrens. Playback age: 5-6 and 10-11DPH nestlings. They show that 10-11 days old suppress calling while 5-6 days old do not.

      Barati & McDonald (2017) - Noisy miner nestlings suppress begging after conspecific alarm calls and some heterospecific cues; stronger/longer suppression for terrestrial-predator alarms. Sci. Rep. 7:9563.

      Different species: Noisy miner (Manorina melanocephala ). Playback age: 14 DPH. Nestlings started to vocalise at 5 DPH.

      Suzuki (2011) - In Paridae, parental alarm calls encode predator type; prior work (cited within) shows young of altricial species suppress vocalizations to alarms. Curr. Biol. 21:15-20.

      Different species: great tits. Playback age: 17 DPH.

      Can you please contextualize the present results about the timing of auditory development with the above body of work with respect to the timing of alarm call-induced begging call suppression?

      In our revision we have added a new paragraph in the discussion on these papers (Lines 357-367), and highlight the need for comparative work on hearing development in different species (Lines 352-355 and Lines 405-411).

      (5) Strictly speaking, a flat ABR does not equal deafness - at the extreme, an average of 10,000 trials may pull out a minuscule signal. Thus, the more rigorous path would be, in the results section, to ensure that statements summarize the data as they are, representing an absence of a neural signal.

      Save the interpretation of what this may mean for the discussion, and provide alongside this interpretation the necessary caveats related to temperature, rendition rate, averaging, etc.

      Clarify the conditions where a flat ABR demonstrates or fails to demonstrate immature deafness.

      Expand clarification for how the known 20-40 dB difference between ABR and behavioral thresholds can exist if a flat ABR can be interpreted as deafness.

      Consider refraining from concluding deafness from a flat ABR. Discuss that behavioral, single-unit, or alternative physiological assays might detect responses below the ABR threshold. If such cases exist, cite.

      (A5) We thought about this considerably before starting our measurements. What constitutes the absence of a signal? Even with intracellular recordings of all but one of the auditory neurons, the last one could still contain a signal and theoretically transmit information to the nervous system. We agree with the reviewers that absence of an ABR response should not be equated with absolute deafness. We have revised the manuscript accordingly and removed all statements implying “deafness” from the Results. The Results now strictly report presence or absence of detectable ABR responses.

      However, in both clinical and comparative contexts, absence of ABR responses at high SPLs (e.g., 90–95 dB) is widely interpreted as functionally non-responsive hearing. The developmental shift we observe (>54 dB) is far larger than typical ABR–behavioral offsets (20–40 dB). In the Discussion, we have added a new paragraph arguing that we think that the term functional deafness is reasonable here (Lines 261-269).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary.

      In this manuscript, the authors investigate the mechanisms underlying macrostructure formation in a freshwater filamentous cyanobacterium strain, F. draycotensis, focusing on how its ability to aggregate and form these structures depends on the physical properties of the filaments. Using experimental observations, they demonstrate that the cyanobacterium actively captures and surrounds particles, a process driven primarily by gliding motility.

      To explain these physical dynamics, the authors present a 3D model indicating that particle collection relies on filament length, as well as a specific mechanical response, namely, filament buckling and the subsequent formation of loops of bundles of filaments. While the authors have previously documented the buckling and looping characteristics of this strain, this study provides new insight by demonstrating that these physical phenomena are essential for particle capture and collection.

      Strengths:

      This manuscript benefits from a rigorous and detailed quantitative analysis of video recordings, which clearly documents the motility, buckling behaviour, and particle collection dynamics of the filaments.

      The authors effectively validate their hypothesis by using a naturally shorter filamentous strain, which fails to collect particles, suggesting that filament length is indeed a critical parameter.

      To further confirm the length dependency within the same species, the authors experimentally generated shorter filaments of F. draycotensis. The fact that these shortened filaments also lose the capacity to collect particles provides strong evidence supporting their proposed mechanism.

      We thank the reviewer for the accurate summary of our work and for their identified strengths of the study.

      Weaknesses:

      There is a conceptual concern. The authors linked the specific physical properties of this strain to evolutionary data, highlighting that the studied lineages diverged approximately two billion years ago. This creates a misleading impression that particle collection via flexible looping filaments is a recent evolutionary adaptation. However, particle collection has been observed in other cyanobacteria, such as Trichodesmium, which features short, rigid filaments. Therefore, the term ”emerging” does not seem appropriate for the title and text. The capacity to collect particles in the studied strain F. draycotensis appears to be primarily a function of physical characteristics (filament length and flexibility) rather than evolutionary age. Any cyanobacterial strain possessing similar physical properties is likely to exhibit comparable behaviour, rendering the evolutionary timeframe largely irrelevant to the core mechanism.

      We would like to first clarify our use of the term “emergent”. It seems that the referee took this in an evolutionary context, whereas we are using this term in the context of its use in systems dynamics, and referring to: “a complex entity displaying behaviors that its components do not have on their own, and emerge only when they interact in a wider whole”. Here, particle collection and dynamic aggregate formation “emerges” from the buckling and interaction of many filaments.

      With regards to the evolution of particle collection behavior, our comment on the evolutionary distance between F. draycotensis and Pseudoanabena sp. was meant to highlight the point that particle collection seems to be a function of physical characteristics and motility: Despite a large evolutionary distance, and possibly many biological differences, a physics-based argument is capturing the difference between the particle collection ability of these two organisms. Thus, we are in agreement here with the reviewer. We did not intend to make any arguments about “evolutionary age” of the particle collection behavior.

      We see that the short, evolutionary comment in the Introduction has confused the reviewer and potentially is confusing to other readers too. We will therefore remove this evolutionary comment from the Introduction section of the revised manuscript and make the point in more detail in the Discussion section.

      In addition, the phylogenetic tree presented in Figure S5 does not reflect the current consensus on cyanobacterial evolution and systematics and does not align with modern phylogenomic frameworks (see, for example, Strunecky et al., 2023 https://doi.org/10.1111/jpy.13304). There is also no such order Cyanobacteriales, which has been mentioned in a few older publications but is clearly outdated.

      We thank the reviewer for this comment, as it has made us realise that we never explained our choice of taxonomic framework in the manuscript, and perhaps this is the source of the confusion.

      The order Cyanobacteriales does exist: it is the order-level name applied in the Genome Taxonomy Database (GTDB) [5, 6], currently the most comprehensive and actively curated genome-based taxonomy of prokaryotes. GTDB classifies taxonomic groups algorithmically, as monophyletic groups in a concatenated marker-protein phylogeny with ranks normalised by relative evolutionary divergence. This has produced a number of re-groupings and new names relative to the older, morphology-derived classifications; many of these have since been formally proposed under the International Code of Nomenclature of Prokaryotes and the SeqCode [2], and are progressively being adopted by the NCBI. The placement of Cyanobacteriales, and of the other orders shown in Figure S5, can be inspected directly on the GTDB “Taxonomy Tree” (see here for the orders within the class Cyanobacteriia).

      We would also like to note that we do not see our tree and the framework of Strunecky et al. as being in conflict. Strunecky et al. constructed their phylogenomic backbone using GTDB-Tk and the same 120-marker concatenated alignment that the GTDB itself uses. What differs between the two schemes is therefore not the underlying phylogeny but the nomenclature applied to the resulting clades: Strunecky et al. work within the botanical tradition and combine the phylogenomic tree with phenotypic characteristics, thereby proposing ten new orders and fifteen new families, whereas GTDB assigns rank boundaries purely by evolutionary divergence and so draws broader order limits. In practice, the GTDB order Cyanobacteriales spans several of the families (e.g. Oscillatoriales and Coleofasciculales) and orders (e.g. Chroococcales and Nostocales), that are proposed within the Strunecky et al. work. Our reason for adopting the GTDB nomenclature is for practical reasons specific to this study. F. draycotensis is a recently described organism [3] that is not included in Strunecky et al. and has no placement in their tree. In GTDB it falls within a family-level lineage (placeholder name JAAUUE01) inside the Cyanobacteriales, with the sequenced members of the Coleofasciculaceae as its closest relatives. We could not have assigned it to one of the Strunecky orders without inventing a placement. The same applies to some of the other, recent metagenomically described cyanobacteria [10], which similarly have no assigned names in the literature. GTDB, by contrast, provides a reproducible, algorithmic assignment for all of these genomes, and is now widely used for this reason in genome- and metagenome-based studies of cyanobacteria (e.g. [1]). We therefore used it consistently throughout.

      Finally, with regards to the reviewer’s point about the tree itself, we would like to note that Figure S5 was intended only to convey the evolutionary distance between F. draycotensis and Pseudanabaena sp., and it was built from a modest set of six concatenated ribosomal protein markers using an approximate maximum-likelihood method with SH-like local support values. This is considerably less rigorous than the 120-marker RAxML and Bayesian analysis of Strunecky et al., and we agree that a stronger tree may be preferable. For the revised manuscript we are recomputing the tree from a substantially larger set of concatenated single-copy marker genes, using IQTREE with model selection and non-parametric bootstrap support. We would note, however, that the specific conclusion drawn from this figure — that the two strains we use for our experimental work, namely F. draycotensis and Pseudoanabena sp. belong to deeply divergent cyanobacterial lineages — is supported by the deep backbone of the cyanobacterial tree, which is stable across marker sets and inference methods, and is equally supported by the tree of Strunecky et al.

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript, as well as the Figure S5 legend.

      Another concern is that the authors nearly completely ignore the role of type IV pili in the gliding motility of cyanobacteria, including filamentous strains. For a long time, there was a misconception that the gliding motility of cyanobacteria was due to slime protrusion. Slime plays a role in this process. However, several studies have shown that filamentous strains also use type IV pili to glide on surfaces. The authors should discuss this and include it in their model.

      The reviewer is correct that we did not include molecular details of gliding motility in our biophysical model. They are also correct to point out that pili and slime biosynthesis genes are shown to be involved in gliding motility [8]. It is, however, still unclear how these factors interact to produce mechanical gliding forces that can result in filament rotation (observed only in some filamentous cyanobacteria), filament reversal, as well as decoordination during such reversals, which we have previously shown in F. draycotensis [9]. Therefore, we have chosen to keep the biophysical model at a coarse-grained, phenomenological level. Instead of explicitly modelling the detailed molecular mechanisms behind force generation, we model only the minimum necessary physical forces and torques needed to reproduce the observed rotation and translation of the filament during gliding under de-coordinated conditions. This model is able to reproduce the experimentally observed buckling and twisting of filaments, and is therefore sufficient and useful to achieve a coarse-grained understanding of mechanical forces and their relationship to buckling, twisting and entanglement, which are the main processes we focus on here. As molecular details behind force generation in rotating, filamentous cyanobacteria become available, more detailed physical models can be constructed. We also note, in this context, that the two filamentous cyanobacteria we compare both encode the type IV pilus machinery, so the presence of a pilus motor does not by itself distinguish a particle-collecting from a non-collecting strain (see our response to the reviewer’s next point).

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript.

      In addition, the authors concluded that gliding motility is responsible for particle collection by Fluctiforma draycotensis. Although I believe that their conclusion is correct, there might be several limitations to the experiments which allow for other reasons to be considered. Their conclusions were based on the use of a non-motile strain and an unspecified community without the motile Fluctiforma draycotensis strain. The problem I see here is that it is not clear why this strain is not motile; it could be because of the lack of type IV pili, mutations which alter their functionality, defects in slime secretion, any other mutation (e.g. in chemoreceptors), cellular structure, metabolism, or combinations of these. Furthermore, it is possible that the community changes its composition and behaviour when it lives without the cyanobacterium with a rich carbon source (glucose) or with a non-motile cyanobacterium which may not secrete slime or, for example, a signalling component which controls behaviour of the bacteria in the community. For that reason, the authors should be more cautious with their conclusion that solely motility behaviour of Fluctiforma draycotensis is responsible for particle collection. Additional factors might be responsible for these effects.

      Our conclusion that gliding motility is the main factor underpinning particle collection is based on several observations.

      Firstly, on the macroscopic scale we present several control experiments where we did not observe particle collection: (i) in the community featuring a non-motile F. draycotensis, and with mostly the same other bacterial species as the community featuring the motile F. draycotensis, (ii) in a bacterial community derived from the original F. draycotensis community but lacking any cyanobacteria, (iii) in the original community with physically shortened F. draycotensis, and (iv) in another cyanobacterial community featuring different bacteria and a naturally shorter, filamentous gliding cyanobacteria Pseudanabaena sp. A straightforward, parsimonious explanation that satisfies all these observations is that particle collection is underpinned by physical characteristics of gliding filamentous cyanobacteria.

      Secondly and more directly, in time-lapse microscopy imaging we repeatedly observe clusters of beads being moved by gliding filaments, and thereby being collected into larger clusters. Thus, whilst factors such as slime secretion also contribute, the primary mechanism driving the observed particle motion seems to be that particles stick to filaments and are carried around with them as they glide. We cannot rule out a contribution of pili to bead attachment and transport. We note, however, that both cyanobacteria compared here encode the type IV pilus machinery. In a homology survey of the two genomes, Pseudanabaena sp. and F. draycotensis both carry orthologues of the core T4P components — the assembly ATPase PilB, the retraction ATPase PilT, the inner-membrane platform protein PilC, the prepilin peptidase PilD, and the alignment-complex proteins PilM and PilF — together with the hormogonium-associated hmpD, hmpF and hmpG. Pseudanabaena sp. is therefore not pilus-deficient, and it does glide, yet it does not collect particles. The difference between the two organisms consequently cannot be attributed to the presence or absence of the pilus motor, which we would argue supports the physical argument we make here. Consistent with this, we have not identified mutations in pilus-related genes in the mutant, non-motile F. draycotensis.

      We are currently in the process of preparing another manuscript describing the mutations that led to motility loss in the non-motile F. draycotensis, as well as the proteins that are differentially expressed in the motile and non-motile F. draycotensis. These analyses will shed more light on the molecular mechanisms abolishing motility and how they might be influencing particle collection.

      In the revised manuscript, we will make these points clearer in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      The authors studied aggregation, buckling, and particle collection by the filamentous cyanobacterium Fluctiforma draycotensis, as well as by the filamentous Pseudanabaena sp. (order Pseudoanabenales). They performed a range of experiments, from imaging individual gliding filaments to multiple-day experiments showing the formation of large aggregates around a particle formed from a precipitate. They also developed a model of buckling filaments to argue that the ability of elastic filaments to collect particles and form macrostructures is confined to a part of the filament phase space in terms of length and flexibility, meaning that gliding combined with certain filament length and flexibility naturally reproduces the observations.

      Strengths:

      This is an impressive study that uses multiple tools to connect macrostructure formation with filaments’ gliding motility and buckling. It adds an important perspective on the biological and physical factors at play in the emergence of aggregates.

      We thank the reviewer for the accurate summary of our work and highlighting the strengths of the study.

      Weaknesses:

      The authors ignore the possibility that filament behavior plays an important role in the emergence of the observed patterns. Cyanobacteria have been shown to control their gliding motility (Pfreundt et al Science 2023; Kurjahn et al Nature Comm 2024), and their molecular motors are known to be regulated by chemotaxislike signaling pathways (Risser ARM 2025). As far as I know, how the coordination between the pulling agents along an individual filament works is actively debated, but there seems to be little doubt that it exists.

      To illustrate this point better, note that the aggregation observed by the authors is consistent with the length-dependent ability of filaments to coordinate gliding (I’m not saying this is how it works in Fluctiforma draycotensis; I’m saying it’s consistent). Suppose the coordination requires sufficiently long filaments, which could be the case when signaling molecules travel along the filament, propagating information about when individual pulling agents should reverse. In such a model, short filaments act randomly because they fail to coordinate gliding by the time they glide off nascent aggregates, whereas longer filaments can perform informed reversals because they have more time for coordination. Such behavior then explains the lack of aggregation in Pseudanabaena sp. (via behavior, not lack of stiffness). Note that Trichodesmium is stiff; its filaments do not buckle, yet Trichodesmium forms organized aggregates via tightly controlled motility. Note also that, as the authors report, since Pseudanabaena sp. is both shorter and faster, its filaments have relatively (to the time needed to glide the filaments’ length) little time to coordinate reversals. In my opinion, whether the observed patterns passively emerge from gliding and buckling or result from active behavior remains an open question.

      We appreciate the comment by the reviewer. We certainly agree that behavioral responses exist in filamentous cyanobacteria and will interplay with the physical aspects to produce exciting, complex dynamics. Besides the exemplar ideas that the reviewer provides, there can be many other scenarios involving behavioral responses, such as responses to light and to quorum sensing molecules or photosynthesis-generated radicals. For example, in F. draycotensis we have observed photo-responses at the aggregate level, which we are are currently studying. Photoresponses are also observed in Trichodesmium aggregates [7]. In general, a full understanding of the interaction of the biological (i.e. behavioral) and the physical aspects will require several future studies.

      In the current study, however, we focus on characterising the physical aspects of gliding motility alone, combined with experimental observations. We believe that this approach is important to establish a form of “null expectation” from the physics of gliding, elastic filaments alone. Currently, the molecular mechanisms responsible for coordinating the reversal behaviour of multiple filaments are still unclear, so it is difficult to experimentally demonstrate behavioural contributions to aggregate formation, e.g. via experiments where such behaviour is switched off. In the meantime, simulations such as those presented here allow us to test more precisely the potential role of activity, coordinated reversals and the elastic properties of the filament. In future it will be interesting to scale up the presented model to include multiple interacting filaments, and to systematically test the respective roles of active coordination behaviour for one individual filament (reversals) and for multiple interacting filaments (where contacts modulate activity), as well as the physical properties (length and flexibility). Such modelling studies can then identify if a ‘purely physical’ model can or cannot generate realistic aggregates, and pinpoint whether additional coordination mechanisms are needed to regulate aggregation. By testing the combination of different physical and biological coordination mechanisms, it would then help to indicate how much of a role is played by various potential active coordination behaviours.

      We will bring out this point more clearly in the Discussion section of the revised manuscript.

      I also have a small suggestion regarding this statement on model novelty:

      The essential novelty of this model is that the filament itself is active and out of equilibrium, and additionally, the forces and torques are applied locally along its centreline, and not at its extremities as in previous steady-state mechanical studies of elastic, twistable filaments such as DNA [31-33] (see Methods and SI).

      This statement needs to be revised as it ignores a substantial body of work on self-organization of active filaments: (R. E. Isele-Holder, J. Elgeti, G. Gompper, Soft Matter 2015; Pfreudnt et al, Science 2023; Faluweki et al PRL 2023; Kurjahn et al Nature Comm 2024).

      We agree with the reviewer that there is a significant literature on active filaments, some of which we have already cited and will now discuss in more details, as well as adding and discussing the suggested additional references. Our statement on “model novelty” refers to the analysis of buckling instabilities of biological filaments, and in particular DNA, due to a combination of forces and torques. To our knowledge, this has only be studied explicitely by [4], and only in the local (resistive force theory) limit. The elastohydrodynamic simulations coupled to local active forces and torques, as we implemented here, are therefore novel and will expand the analysis of both microbial filaments and other biological polymers. We will clarify these points in the Methods and Discussion sections of the revised manuscript.

      Last point: the authors often say that their observations are reproducible (’...reproducibly forms macroscopic granules...’). What is meant? Different experiments on different days, different aliquots?

      The “replicability” statement was in reference to different experiments started on different days using cultures obtained from serial transfer experiments, as well as cultures re-initiated from cyrostocks. This point will be made clear in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors report and characterize the formation of aggregate microstructures by the motile filamentous cyanobacterium Fluctiforma draycotensis, which exhibits gliding motility accompanied by rotation along the long axis while excreting EPS. In experiments with motile F. draycotensis cultures, they observed the formation of granular structures composed of cyanobacteria and other material (iron, polystyrene beads, etc.), with macrostructures on the scale of 1mm within 24 hours. The structures were motile at speeds comparable to that of the cyanobacteria filaments, resulting in their growth through coalescence over time. Notably, such macrostructures were absent in nonmotile F. draycotensis, pointing to the role of filament motility in their formation. Through experiments examining the micro-scale dynamics, inert material such as small polystyrene beads was found to be transported by the gliding, buckling, and plectoneme dynamics of the filaments, pointing to the underlying mechanism by which particles are collected into larger-scale microgranule structures.

      To interrogate the properties that drive the cyanobacteria filament buckling, plectoneme formation, and entanglement, the authors develop a mechanical model for filaments as nearly inextensible, slender bodies with resistance to twisting and bending under active gliding forces and torques and responding to fluid flows and surface adhesion. They derive expressions for the thresholds for buckling and twisting instabilities, which are additionally demonstrated and interrogated through simulation via the Immersed Boundary Method. Most importantly, bending and plectoneme formation only occur with sufficiently long filaments, and the threshold is shorter for bending than for plectoneme formation. Experimental observations with wild-type filaments agree with the model-predicted thresholds. The authors perform additional experiments with shorter filaments below both thresholds, including the filamentous bacterium Pseudanabaena, which fail to collect particles (though can in principle form macrostructures).

      Strengths:

      This work appears to be novel (notably, the discovery and characterization of the particle collection behavior of a filamentous cyanobacterium) and has interesting implications for both naturally observed cyanobacterial macrostructures as well as the controllable parameters in engineering them. The experimental and modeling work is well motivated, contributing to the broader understanding of macrostructure formation and material aggregation through active filament dynamics (not exclusive to cyanobacteria), as well as the underlying physical properties governing important filamentous cyanobacterium dynamics. As such, I would expect the results of this paper to be of broad interest to both biophysicists and microbiologists. Generally, the manuscript is well written with clear, compelling figures that illustrate the important conclusions of this study.

      We thank the reviewer for the accurate summary of our work and recognising the broad relevance of the study.

      Weaknesses:

      In the section on “Shorter gliding filaments cannot collect particles nor form granule macrostructures”, the filamentous cyanobacteria considered “all” fall below the predicted thresholds for bending and twisting. The “long” F. draycotensis are 60 microns in length, notably less than the 120 and 320 micron thresholds derived in the previous section as well as the lengths of filaments considered in Figure 3D, yet these “long” 60 micron filaments form macrostructures. How can this be understood in the context of the model predictions? Is the nature of the macrostructures in Figure 4B, the microscale parameters, or the collection of particles somehow different than those with filaments an order of magnitude longer in earlier parts of the paper? The paper would be stronger if these sorts of questions were addressed in the text and/or with supplementary figures.

      We thank the reviewer for this point. Indeed as we mention in the text, the ‘long’ population has a mean length of 60 micron. However, as shown in the length distribution plot in Fig 4A, the maximum filament lengths observed in these populations (within the samples used for microscopy) are 560 microns for the long filaments, versus 240 microns for the short filaments. Thus, we expect the long population to contain multiple filaments that can buckle and a few that can form plectonemes, whilst the short population might have some buckling filaments and none that form plectonemes. We stress that Fig 4A only shows the length distribution for what we believe to be a representative sample taken from the long and short populations, not the full data from the entire population.

      We will revise the main text to include the maximum filament lengths of the two populations as well as the mean values. We will also add lines to Fig 4A to indicate the buckling and plectoneme threshold lengths from the analytical estimate for the F. draycotensis filaments (same values as in Fig 3), to make it clear that the long population contains more buckling/plectoneming filaments than the short population.

      References:

      (1) E. S. Cameron et al. “Diversity and specificity of molecular functions in cyanobacterial symbionts”. In: Sci Rep 14.1 (2024), p. 18658. issn: 2045-2322 (Electronic) 2045-2322 (Linking). doi: 10.1038/s41598-024-69215-8. url: https: //www.ncbi.nlm.nih.gov/pubmed/39134591.

      (2) M. Chuvochina et al. “Proposal of names for 329 higher rank taxa defined in the Genome Taxonomy Database under two prokaryotic codes”. In: FEMS Microbiol Lett 370 (2023). issn: 1574-6968 (Electronic) 0378-1097 (Print) 0378-1097 (Linking). doi: 10.1093/femsle/fnad071. url: https://www.ncbi.nlm.nih.gov/ pubmed/37480240.

      (3) S. J. N. Duxbury et al. “Niche formation and metabolic interactions contribute to stable diversity in a spatially structured cyanobacterial community”. In: ISME J (2025). issn: 1751-7370 (Electronic) 1751-7362 (Linking). doi: 10.1093/ismejo/ wraf126. url: https://www.ncbi.nlm.nih.gov/pubmed/40577531.

      (4) Raymond E. Goldstein, Thomas R. Powers, and Chris H. Wiggins. “Viscous Nonlinear Dynamics of Twist and Writhe”. In: Physical Review Letters 80.23 (June 1998), pp. 5232–5235. issn: 1079-7114. doi: 10.1103/physrevlett.80.5232.

      (5) D. H. Parks et al. “A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life”. In: Nat Biotechnol 36.10 (2018), pp. 996–1004. issn: 1546-1696 (Electronic) 1087-0156 (Linking). doi: 10.1038/ nbt.4229. url: https://www.ncbi.nlm.nih.gov/pubmed/30148503.

      (6) D. H. Parks et al. “GTDB release 10: a complete and systematic taxonomy for 715 230 bacterial and 17 245 archaeal genomes”. In: Nucleic Acids Res 54.D1 (2026), pp. D743–D754. issn: 1362-4962 (Electronic) 0305-1048 (Print) 03051048 (Linking). doi: 10.1093/nar/gkaf1040. url: https://www.ncbi.nlm.nih. gov/pubmed/41123020.

      (7) U. Pfreundt et al. “Controlled motility in the cyanobacterium Trichodesmium regulates aggregate architecture”. In: Science 380.6647 (2023), pp. 830–835. issn: 1095-9203 (Electronic) 0036-8075 (Linking). doi: 10.1126/science.adf2753.

      (8) Douglas D Risser. “Motility in Filamentous Cyanobacteria”. In: Annual Review of Microbiology 79 (2025).

      (9) Jerko Rosko et al. “Cellular coordination underpins rapid reversals in gliding filamentous cyanobacteria and its loss results in plectonemes”. In: eLife 13 (2025), RP100768.

      (10) A. Scarampi et al. “Enrichment of convergent metabolic functions in microbial communities through imposed and emergent environmental niches”. In: bioRxiv (2026). doi: 10.64898/2026.02.11.705344.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian-optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year-round. The authors further claimed that such an association could be able to provide an early warning in the community.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We would like to express our sincere and heartfelt gratitude to all of you for your exceptionally thorough, constructive, and intellectually rigorous evaluation of our manuscript. The breadth and depth of the feedback we have received reflect a high standard of scientific scrutiny that we deeply respect and appreciate.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 Mpro enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the Mpro monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how Mpro inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      None. The requested mutagenesis data have been provided in the revised manuscript, and all of my previous concerns have been satisfactorily addressed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      None. The overall quality of the manuscript has been substantially improved. The authors have added supportive mutagenesis data in the revised manuscript to validate the proposed role of C300. All of my concerns have been adequately addressed.

      We appreciate the reviewer’s recognition of the improvements made in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on revised version:

      The authors have very nicely addressed most of the previous comments raised. But one comment remains to be clarified relating to original point 5 and the authors' response:

      "We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296-306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation."

      Two questions remain for the statements in line 376-380. First, while it is stated "residues 296-304 in the C-terminal region of Mpro were more flexible upon ebselen binding", the segment of 296-306 is shown Figure 4c. Second, the HDX change for this segment upon ebselen binding is very subtle in the figure (in contrast to the significant HDX change of the same segment in the protein upon PF-07321332 binding), thus making the strong conclusion that "This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity" not convincing. The reviewer would suggest the authors either delete this conclusion or largely tone it down.

      We thank the reviewer for the recognition of our efforts and agree with the reviewer’s suggestion. We have corrected “residues 296–304” to “residues 296–306” in Line 377 and removed the statement “This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity.”, as suggested.

    1. Author response:

      We thank the reviewers for their careful and constructive evaluation of our study. We are encouraged that the reviewers recognized the significance of DSB-induced genomic amplification (DIGA) and the evidence implicating DNA-end processing and recombination-associated DNA synthesis in this response.

      To our knowledge, this study provides the first description of DIGA as a large-scale increase in genomic DNA content following the induction of DSBs in cancer cells and represents an initial effort to define factors that regulate this phenomenon. The present work shows that DIGA can be induced by several sources of DSBs, involves de novo DNA synthesis, is genetically distinguishable from canonical CDT1-dependent origin re-licensing, is regulated by pathways controlling DNA-end protection and resection, and requires RAD51, RAD52, POLD3, and POLD4.

      At the same time, we agree that many important questions remain regarding the physical organization and genomic distribution of the additional DNA, the sites from which synthesis originates, the length and number of synthesis tracts, and the full determinants that render some cancer cells more susceptible to DIGA than others. We view these as important questions that arise from the initial characterization of this previously unrecognized phenotype and that will require substantial additional investigation.

      Reviewer #1:

      We agree that p53 status alone does not explain the considerable variation in DIGA observed among the cancer cell lines examined. Our experiments using isogenic HCT116 cells identify p53 as one factor capable of limiting DIGA, while the genetic studies implicate DNA-end protection and resection pathways as additional determinants. The present data, however, do not establish which of these or other pathways account for the differences among individual cancer cell lines. Defining the molecular basis for this variability will require systematic comparison of DIGA-prone and DIGA-resistant cells.

      The reviewer also raises an important question regarding the temporal relationship between DIGA and normal S-phase DNA replication. Our conclusion that the increase in DNA content involves de novo synthesis within the same cell-cycle interval is supported by BrdU incorporation in cells with >4N DNA content, the persistence of DIGA when progression through mitosis is blocked by nocodazole, and the detection of newly synthesized DNA in synchronized irradiated cells. These experiments do not, however, define precisely when DIGA-associated synthesis begins relative to normal S-phase replication. More detailed time-resolved analysis will be required to establish this relationship.

      We also agree that the present findings support a BIR-like mechanism rather than providing a complete physical reconstruction of classical BIR. The dependence of DIGA on DNA-end resection, RAD51, RAD52, POLD3, and POLD4 provides genetic evidence for recombination-associated DNA synthesis with features of BIR. Direct determination of synthesis-tract architecture, template usage, and genomic distribution will be required to define the underlying synthesis mechanism more completely.

      Reviewer #2:

      We agree that direct characterization of the additional DNA represents an important next step in understanding DIGA. The current study demonstrates a substantial increase in cellular DNA content, de novo DNA synthesis within the >4N population, and dependence on factors involved in DNA-end resection, strand invasion, and BIR-associated synthesis. These experiments do not determine which genomic regions are amplified or the length of individual synthesis tracts. Genomic analysis of cells undergoing DIGA should help determine whether the observed increase in DNA content reflects numerous amplification events, extensive synthesis from a subset of sites, or a different organization of the additional DNA.

      We appreciate the reviewer highlighting the study by Costantino et al. (2014), which provided important evidence that BIR-associated repair of damaged replication forks can generate segmental genomic duplications in human cells. The genomic alterations characterized in that study, however, arose under a substantially different experimental setting. Costantino et al. induced replication stress through cyclin E overexpression and analyzed copy-number alterations accumulated over a three-week period in clonally derived cells. Among these alterations, amplifications smaller than 200 kb were reduced following depletion of POLD3 or POLD4, leading the authors to propose that this subset of segmental duplications may represent BIR events, whereas larger amplifications and deletions could involve other repair mechanisms.

      DIGA, however, differs from these previously described alterations in several readily observable respects. DIGA develops over approximately one to three days following acute induction of DSBs by IR or AsiSI and produces increases in total cellular DNA content sufficiently large to be detected directly by flow cytometry. Thus, the two phenomena differ in their mode of induction, kinetics, and scale. At the same time, the involvement of POLD3 and other recombination-associated factors in both settings raises the possibility that they share aspects of the underlying DNA-synthesis machinery. The present data do not establish whether the additional DNA in DIGA consists of numerous segmental duplications, substantially longer synthesis products, or another genomic configuration.

      We also agree that our experiments do not directly demonstrate that DIGA-associated DNA synthesis initiates precisely at individual DSB sites. The ability of AsiSI-generated DSBs to induce DIGA, together with its dependence on DNA-end resection, RAD51, RAD52, POLD3, and POLD4, links the phenomenon closely to DSB processing. Direct mapping of newly synthesized DNA relative to defined DSBs will ultimately be required to determine where DIGA-associated synthesis originates.

      The reviewer asks whether MLN4924-induced re-replication and DIGA have been examined simultaneously. We have not examined this combination. Our distinction between these processes instead rests on their different genetic requirements. In particular, depletion of CDT1 strongly suppresses MLN4924-induced re-replication but does not suppress IR-induced DIGA, and DIGA is stimulated while rereplication is inhibited by the depletion of SET8. These observations argue against canonical CDT1-dependent origin re-licensing as the mechanism underlying DIGA, although they do not exclude more complex interactions between replication and DSB-associated DNA synthesis.

      We agree that the effects of XRCC4, XLF, and LIG4 are mechanistically intriguing and not yet fully understood. The present experiments establish that loss of XRCC4 or XLF, and to a lesser extent LIG4, suppresses DIGA, whereas loss or inhibition of DNA-PKcs enhances it. Stabilization of broken DNA ends by XRCC4/XLF is one possible interpretation, but effects on end resection, DSB persistence, repair-pathway choice, or other functions of these proteins could also contribute. The opposing effects of different components of the NHEJ machinery therefore identify an important mechanistic question that remains to be resolved.

      Finally, we agree that the correlation between DIGA and radiation sensitivity across the melanoma cell-line panel does not by itself demonstrate causality. The data establish an association between the propensity to undergo DIGA and sensitivity to IR. Because these cell lines differ in multiple additional properties that may influence the radiation response, matched models in which DIGA can be selectively altered will be important for determining the extent to which DIGA itself contributes to radiation-induced loss of proliferative capacity.

      Reviewer #3:

      We agree that the magnitude of the increase in DNA content is one of the most interesting unresolved features of DIGA. Previous analyses of BIR-associated synthesis at defined mammalian lesions have generally described synthesis events considerably smaller than the total increase in DNA content observed here. Our experiments do not establish the length or number of individual synthesis events responsible for DIGA. We therefore use the term BIR-like to describe the genetic requirements of the process rather than to imply that each DSB gives rise to a single exceptionally long BIR tract. Determining how many genomic sites participate and how much DNA is synthesized at individual sites will be important for understanding how the large increase in total DNA content is generated.

      Regarding the AsiSI BrdU experiments, it is important to note that BrdU was provided as a one-hour pulse immediately before harvesting. BrdU signal at 48, 72, or 96 hours therefore reports DNA synthesis occurring during that particular one-hour interval and does not measure the cumulative DNA synthesis that preceded the measurement. Consequently, relatively modest BrdU incorporation in cells that have already accumulated high DNA content does not indicate that the preceding increase occurred independently of DNA synthesis. Conversely, these experiments alone do not define the physical mechanism by which the additional DNA accumulated.

      We agree that the requirement for XRCC4, XLF, and LIG4 is unexpected under a simple model of BIR. As noted above, the present experiments establish this genetic relationship but do not define its molecular basis. The differential effects of DNA-PKcs and downstream NHEJ factors suggest that individual NHEJ components may influence DIGA through functions that are not adequately represented by viewing the pathway simply as a linear ligation reaction.

      Finally, we agree that the similarity between the flow-cytometric profiles produced by IR and MLN4924, as well as their common sensitivity to aphidicolin, does not by itself distinguish the underlying mechanisms. The distinction in the current study instead derives from their different genetic requirements, particularly the dependence of MLN4924-induced re-replication on CDT1 compared with the lack of such a requirement for DIGA, together with the differential effects of factors involved in DNA-end protection, resection, and DSB repair. These observations argue that DIGA is not simply the consequence of canonical origin re-licensing (i.e., rereplication or endoreduplication), while leaving open the possibility that additional replication-associated mechanisms contribute to the phenotype.

      In summary, we appreciate the reviewers highlighting several important mechanistic questions raised by our findings. We view these questions as natural extensions of the initial discovery and characterization of DIGA. The present study identifies a large-scale DSB-associated increase in genomic DNA content and establishes important roles for DNA-end protection, resection, strand invasion, and BIR-associated factors in regulating this response. Determining the genomic architecture of the additional DNA, the sites and molecular intermediates from which synthesis originates, and the cellular determinants of DIGA susceptibility will be important goals for future studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dong et al. present an in-depth analysis of mutant phenotypes of the Rab GTPases Rab5, Rab7, and Rab11 in Drosophila second-order olfactory neuron development. These three Rab GTPases are amongst the bestcharacterized Rab GTPases in eukaryotes and have been associated with major roles in early endosomes, late endosomes, and recycling endosomes, respectively. All three have been investigated in Drosophila neurons before; however, this study provides the most detailed characterization and comparison of mutant phenotypes for axonal and dendritic development of fly projection neurons to date. In addition, the authors provide excellent high-resolution data on the distribution of each of the three Rabs in developmental analyses.

      Strengths:

      The strength of the work lies in the detailed characterization and comparison of the different Rab mutants on projection neuron development, with clear differences for the three Rabs and by inference for the early, late, and recycling endosomal functions executed by each.

      Weaknesses:

      Some weakness derives from the fact that Rab5, Rab7, and Rab11 are, as acknowledged by the authors, somewhat pleiotropic, and their actual roles in projection neuron development are not addressed beyond the characterization of (mostly adult) mutant phenotypes and developmental expression.

      We would like to thank Reviewer #1 for their appreciation of our characterization of distinct Rab mutants.

      Reviewer #2 (Public review):

      Summary:

      This study by Dong et al. characterizes the roles of highly-expressed Rab GTPases Rab5, Rab7, and Rab11 in the development and wiring of olfactory projection neurons in Drosophila. This convincing descriptive study provides complementary approaches to Rab expression and localization profiling, conventional dominantnegative mutants, and clonal loss-of-function mutants to address the roles of different endosomal trafficking pathways across circuit development. They show distinct distributions and phenotypes for different Rabs. Overall, the study sets the stage for future mechanistic studies in this well-defined central neuron.

      Strengths:

      Beautiful imaging in central neurons demonstrates differential roles of 3 key Rab proteins in neuronal morphogenesis, as well as interesting patterns of subcellular endosome distribution. These descriptions will be critical for future mechanistic studies. The cell biology is well-written and explanatory, very accessible to a wide audience without sacrificing technical accuracy.

      Weaknesses:

      The Drosophila manipulations require more explanation in the main text to reach a wide audience.

      We appreciate Reviewer #2’s analysis of our work and thank them for their suggestions to improve the clarity of our manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors aimed at a comprehensive phenotypic characterization of the roles of all Rab proteins expressed in PN neurons in the developing Drosophila olfactory system. Important data are shown for a number of these Rabs with small/no phenotypes (in the Supplements) as well as the main endosomal Rabs, Rab5, 7, and 11 in the main figures.

      Strengths:

      The mosaic analysis is a great strength, allowing visualization of small clones or single neuron morphologies. This also allows some assessment of the cell autonomy of the observed phenotypes. The impact of the work lies in the comprehensiveness of the experiments. The rescue experiments are a strength.

      Weaknesses:

      The main weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos. The insights thus do not extend far beyond what is already known from other work in many systems.

      We thank this reviewer for their feedback and appreciation of our genetic manipulations.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Consensus suggestions after discussion of all three reviewers:

      All three reviewers agree that the morphological and phenotypic analysis of the fly olfactory neurons is a strength of the manuscript. The shared perceived weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos; the findings are in line with a large body of literature.

      The three reviewers feel that the manuscript could be strengthened greatly by adding data on an actual cargo (cell surface proteins?) and a more detailed analysis of the actual developmental origin (what happens when during axon and dendrite development) with respect to sorting of such cargo in the neurons they analyzed.

      We appreciate the time and effort of all three of our reviewers and share their interest in both identifying Rab-regulated cargos as well as determining the developmental origins of the Rab phenotypes. We have added three additional main figures (new Figure 4, Figure 8, and Figure 9), two supplemental figures (Figure 1—figure supplement 1 and 2), two additional supplemental tables (Table S2 and S3), and five additional panels (in Figure 3 H–L) of mutant developmental phenotype analysis.

      Regarding cargos, we also share the reviewers’ desire to identify cargos regulated by each Rab and made attempts to do so but were ultimately unable to achieve this goal. The main obstacles to this were: (1) it is not known which cell-surface proteins are most robustly endocytosed in PNs; without this knowledge it is difficult to identify candidates whose localization would predominantly reflect endosomal rather than plasma membrane distribution, making it challenging to detect changes in compartment-specific localization upon Rab perturbation; (2) reagents to evaluate cell-surface proteins in PNs are not cell-type-specific making it difficult to evaluate changes in their distribution in PNs; (3) tagged overexpressed proteins are either unavailable or expressed at levels too high to sensitively detect changes in their distribution. We have elaborated on each of these points below and feel that cargo identification, while an important future direction, is beyond the scope of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      There are a number of experiments and ideas that the authors might consider to further improve on this work.

      (1) The idea, introduced by the authors in the introduction, that Rab-mediated recycling of cell surface proteins back to the membrane versus degradation is, of course, excellent and interesting. It is less clear how this applies to the present study. The functions of Rab5, Rab7, and Rab11 are so widespread, potentially affecting many signaling roles resulting in primary or secondary effects on membrane and even cytoskeletal regulation, that it remains unclear whether the mutant phenotypes are related to the recycling or degradation of cell surface receptors. To link the idea to experimentation, the authors have an excellent opportunity in their system to look at endogenously tagged cell surface proteins (or at least one example), many of which the Luo lab has characterized in these neurons, to minimally correlate cell surface protein defects to the observed developmental defects.

      We understand this critique and share this reviewer’s interest in identifying the specific cargos regulated by each Rab during development. We attempted to use antibodies to evaluate changes in cell-surface protein localization in response to disrupting individual Rabs but were unable to reliably distinguish shifts in association with specific endosomal compartment as many available antibodies label cell-surface proteins expressed in antennal lobe cells beyond projection neurons (such as olfactory receptor neurons, glia, or local interneurons) which complicates analyses.

      Additionally, although we have, in other work, generated multiple 'flp-on' tags for PN cell-surface proteins, these cannot be used in combination with the MARCM system, as it relies on a heat-shock-inducible flp to label singlePN clones. Heat shock would simultaneously induce tag expression in other cells expressing the tagged gene, preventing PN-type-specific detection. This incompatibility thus prevents us from simultaneously perturbing individual Rabs and tracking corresponding changes in surface-protein localization with single-cell resolution.

      Moreover, for proteins that are not highly endocytosed, it is difficult to separate plasma-membrane from endosomal localization, and we currently do not know which cell-surface proteins are most robustly endocytosed in PNs. Thus, while we share the reviewer’s interest in identifying candidate cargos, technological limitations make it difficult to achieve this goal within the scope of the current study.

      (2) The mutant phenotypes are mostly characterized based on adult outcomes. Maybe a little more can be learned about when and how Rab5, Rab7, or Rab11 function is locally required by characterizing the developmental processes that lead to, e.g., aberrant dendritic development in Rab5 and Rab11.

      We also feel that charting the developmental origins of Rab mutant phenotypes is important. Prior to mid-pupal stage (around 48 hours after puparium formation), glomeruli in the antennal lobe have not yet assumed their stereotyped positions, which complicates analyses and interpretation; thus, many of our analyses are conducted at the adult stage. For Rab11 mutants we did perform many developmental analyses to evaluate the origins of the axonal development (Figure 6—figure supplement 1) and dendrite elaboration phenotypes (Figure 5 J–L) we observed at the adult stage. We realize that the developing axonal analyses were in supplemental material where they could be missed. We have moved these data to the main figures (Figures 8 and 9) and emphasized these analyses. Further, we extended our Rab5 mutant analyses to evaluate developmental phenotypes (Figure 3H– L and Figure 4). We believe that these new analyses have strengthened the manuscript.

      (3) Regarding the subcellular localization analyses: a collection of endogenously tagged Rabs in Drosophila has been generated by Dunst et al. (2015), which is surprisingly not cited. Maybe the authors could consider looking at the endogenous localization of Rabs using the tagged version in parallel to their overexpressed tagged versions.

      We have now cited and discussed this paper (line 82) and thank the reviewer for pointing out this omission. We previously attempted to evaluate these endogenously tagged Rab proteins in PNs; however, since PN dendrites project into the antennal lobe, a dense neuropil region containing PN dendrites, ORN axons, glial processes, and neurites of local interneurons, we are unable to resolve individual Rab puncta from cytosolic (non-vesicle associated Rabs) or evaluate Rab localization in a cell-type-specific manner. For this reason, we focused on evaluating the localization of tagged Rab proteins from UAS-transgenes using a MARCM rescue strategy. We directly addressed this in the text (starting on line 83).

      (4) It is maybe not entirely surprising that Rab5 and Rab11 have the strongest phenotypes, as these have been implicated in early and recycling endosomal processes in basically all eukaryotic cells with major implications for signaling throughout development and function, often causing cell death (and in the case of Rab5, tumorigenic phenotypes in flies). By contrast, the Drosophila brain can develop in the absence of Rab7 (Cherry et al., 2013; also the reference for the Rab7 null mutant, not Chan et al., 2011). A key concern in any developing fly cell rendered mutant using clonal analysis is the perdurance of RNA or protein (ultimately even maternal contribution), which could be addressed by discussion or experimentally.

      We thank the reviewer for pointing out our citation error, which we have now corrected.

      However, we note that Cherry et al. (2013) found that loss of Rab7 causes pupal lethality at stages prior to completion of 50–80% of development and can also cause embryonic lethality when maternal Rab7 contribution is blocked. This indicates that the whole organism cannot fully develop in the absence of Rab7. And while Cherry and colleagues did evaluate overall brain morphology in Rab7 mutant pupae, they did not look at the development of individual cell types. So, it is still unclear how loss of this GTPase affects the development of individual central nervous system neurons.

      Thus, to understand whether Rab7 has functions in PN development, we used the QMARCM system to perform Rab7 LOF analyses in PN clones. While we did not observe any phenotypes in single-cell MARCM clones (Figure 6), we did see mild defects in neuroblast clones (Figure 6—figure supplement 1A–C). Since single-cell MARCM clones are more susceptible to RNA/protein perdurance, we further evaluated Rab7 function by expressing a Rab7 dominant-negative transgene in DL1-PNs using a DL1-specific GAL4 driver (Figure 6—figure supplement 2), which circumvents potential perdurance issues mentioned by this reviewer. Importantly, this same transgene produces dendrite targeting defects when expressed in all PNs (Figure 1I), confirming its efficacy. However, no phenotypes were observed when expression was restricted to DL1-PNs, suggesting that Rab7 may not be required in DL1-PNs for their dendrite targeting. Given that both Rab7 mutant neuroblast clones and pan-PN expression of Rab7 dominant negative causes PN dendrite targeting defects we conclude that Rab7 is nonautonomously required for dendrite targeting of DL1-PNs.

      We have softened our language with regards to the Rab7 analysis and have emphasized, and strengthened, our previous discussion of these points beginning on line 268 in the results section and on line 427 of the discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1B, it would be useful to show the circuit over multiple developmental stages, rather than just in its final form.

      We have added this to Figure 1; it is now panel C. Thank you for this suggestion.

      (2) Expression analysis of Rabs in Figure 1C - how do these levels and ratios compare to the whole brain? Whole body?

      Unfortunately, we are unable to evaluate how Rab expression in PNs compares to all other cells in the brain as there is no sequencing data available for this organ at this time point. We did compare the expression of endosomal Rabs between PNs and their presynaptic targets, ORNs. We found that many Rabs displayed similar expression patterns between these two cell types during development. We have added a new paragraph on this, beginning on line 100 and we added two additional supplemental figures (Figure 1—figure supplement 1 and 2).

      (3) The authors should include at least a few sentences comparing the current approach and results to previous comprehensive Rab protein expression analysis, for example, in PMID 17409086, 22000105, 22844416, and 33666175.

      Thank you for pointing out this omission, we have amended it beginning on line 82.

      (4) For the non-Drosophila reader (for example, a cell biologist working on endosomal traffic in cultured neurons), the paper is less accessible. Some examples:

      We thank this reviewer for their suggestions for ways to clarify our work for the non-Drosophila reader. We have addressed each of their points.

      (a) The severity difference between Rab5, Rab11, Rab7 and Rab4, Rab 21, Rab35 isn't immediately obvious from the images to someone who doesn't work with this system - does the brightness of the ectopic growths indicate the number of ectopically grown processes? It might help to have half a sentence to make this difference more accessible for the readers who aren't familiar with this system.

      We have clarified this beginning on line 112.

      (b) Figures 1D-E require more extensive description of the experimental setup with orthogonal expression systems than is provided briefly in the cartoon, figure legend, and supplement. For example, it should be noted what white vs blue represents in the marked glomeruli.

      We have clarified this point beginning on line 114.

      (c) There should be at least one sentence introducing what is marked and what it means when MARCM clones are first shown in Figure 2B, in addition to the supplemental figure.

      We have added a detailed explanation of MARCM on line 155.

      (d) It's not clear to a non-expert what the meaning is of no innervation of non-adPN glomeruli in wild-type in Figure 2E. This requires a sentence of explanation.

      We have added additional details about this on line 159 and 169.

      (e) Can a control image be shown for the experiment in 3B?

      We have added an additional set of control images in Figure 3B on the left.

      (5) The experiment measuring axonal projection to the lateral horn in Rab5 clones in Figure 3 J-L is underpowered (n=3 for mutant). While this may be due to the frequency of an overall projection defect as shown in Figure 3B, it makes it difficult to assess the robustness of the terminal phenotype. Further, for clarity, similar measurements (e.g., bouton width) should be aligned vertically between E-G and J-L.

      We have performed additional analyses on Rab5 axons in the lateral horn and added them to a new Figure 4. The n’s are now n=10 for controls and n=7 for Rab5 mutants. Additionally, we have aligned similar measurements in the figure panels and standardized the axes of each graph so that it is easier to compare between developmental stages.

      (6) The argument that cell-type-specific phenotypes are due to distinct cargoes is weak. The same cargo could have different functions or signaling properties in different cell types (e.g., "Taken together, the distinct branching phenotypes observed in the mushroom body versus lateral horn suggest that Rab5 may regulate the trafficking of a distinct set of cargos in each axonal compartment"). Similarly, this argument is just one of many possibilities, as the effect could be quite indirect (eg via mis-regulated signal transduction): "Yet, the terminal boutons in Rab5 mutants were nearly 2-fold larger than those of controls (Figure 3G), suggesting Rab5 regulates the trafficking of cell-surface proteins that normally restrain bouton growth " and "Page 12 "Rab7mediated degradation does not have a major role in regulating axon or dendrite development" - should be softened since Rab7 may easily play an important but redundant role.

      We have made all of these changes and removed references to trafficking of specific CSPs.

      (7) Statistics need to be added to: Figure 3B, Figure 5D-E, Figure 2 - figure supplement 1C, Figure 4 - figure supplement 1E, F, Figure 5 - figure supplement 1A, E, Figure 6 - figure supplement 1K.

      We thank this reviewer for pointing out this omission, we have added statistical measurements to our graphs.

      As to not visually overwhelm readers with statistical measurements on already dense graphs (such as Figure 7B), we have added two supplemental tables (Table S2 and S3) that display all the results of all of the comparisons performed in the statistical tests.

      We have cited this table in the figure legends and in-text figure references.

      In addition to the methods, we have also added the exact statistical test and post-hoc tests (when applicable) to the figure legends.

      (8) The BSDC identifier for the UAS-Rab11-mCherry stock may be incorrect.

      It appears that some of the values in the ‘Identifiers’ column of the Key Resources table shifted downward. We have fixed this and appreciate that the reviewer pointed this out.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Our revision includes:

      (1) The generated data from long-read whole-genome sequencing of 1000 Genomes Project samples, including FASTQ files, SV calls, and the imputation panel, are now openly available via ENA and OpnMe. The imputed structural variant data have been submitted to UK Biobank for release through the UK Biobank Research Analysis Platform, subject to UK Biobank release procedures. SV-WAS summary statistics have been made available via OpnMe.

      (2) Clarification of analyses and methods, addition of two new Supplementary Figures, and correction of minor issues throughout the manuscript.

      (3) A significantly expanded Discussion to address the reviewers’ comments and better contextualise our methods and results.

      eLife Assessment

      This fundamental work significantly enhances our understanding of how structural variants influence human phenotypes. The conclusion is convincingly supported by rigorous analyses of long-read sequencing data. If the raw data are made publicly available, these high-quality datasets and findings will further advance our knowledge of genetic variation in the human population.

      We thank the editors for this positive assessment of our work. The raw long-read sequencing data (FASTQ files) can now be accessed through the European Nucleotide Archive (ENA) under accession number PRJEB89727, as part of a larger collection of 1019 sequenced probands from the 1000 Genomes Project (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The generated imputation panel and the structural variant calls, based on the 888 probands used in the present manuscript, remain freely available for download at https://opnme.com/genomiclens. We have now added the summary statistics of 32 SV-wide association studies to the same resource. In addition, we have submitted the imputed SV genotypes for UK Biobank participants to the UK Biobank; once processed by UK Biobank, these genotypes will be released via the UK Biobank Research Analysis Platform (RAP).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine mapping.

      The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

      We thank the reviewer for highlighting the value of our SV imputation reference panel.

      Weaknesses:

      (1) Although the authors employ state-of-the-art analytical approaches for the identification of SVs, the overall accuracy remains suboptimal, as indicated by an F1 score of 74.0%, particularly in tandem repeat regions. To enhance accuracy, it would be beneficial to explore alternative SV detection methods or develop novel approaches. Given the value of the reference panel and the fact that improved SV accuracy would lead to more precise SV imputation and GWAS results, investing effort in methodological refinement is highly encouraged.

      Accurate SV calling remains an active area of research and is beyond the scope of the present study. Tandem repeat regions are particularly challenging for standardised SV detection. We believe that achieving a benchmark for NA12878 of F1 = 74% on a genome-wide level and, notably, F1 = 91% when excluding longer tandem repeats, represents strong performance. This result is especially convincing when considering that our benchmarking compared the SV calls to data generated using a different sequencing technology and processed using different bioinformatics pipelines.

      (2) From the Methods section, it appears that the authors employed Beagle for both imputation and the UK Biobank imputation.

      (a) It would be better to explicitly clarify this in the Results section and provide a detailed description of the corresponding procedures and parameters in the Methods section for both analyses, as this represents a key aspect of the study.

      We thank the reviewer for these suggestions. Accordingly, we added the clarification to the Methods section that the leave-one-out imputation used exactly the same pipeline and settings as the UK Biobank imputation (page 14, section “Leave-one-out imputation performance”):

      “We excluded one individual from the panel and imputed SVs for this individual using the panel of the remaining 887 samples, applying exactly the same pipeline and settings as those later used for SV imputation into UK Biobank (see below).”

      (b) Additionally, Beagle is not specifically designed for SV imputation, the imputation quality of SVs is generally lower than that of SNVs. Exploring strategies to improve SV imputation, such as developing a novel method with reference panel data, may enhance performance.

      As stated in our manuscript (page 4), we believe that, in our study, the imputation quality of SVs is lower than of that of SNVs primarily because of a) the greater difficulty of SV calling compared to SNV genotype calling and b) the heterogeneity in SV representation across samples. Once SVs are encoded as bi-allelic markers in the reference panel, they can be imputed using the same LD/haplotype-based framework as any other variants. Accordingly, improving SV imputation is likely to benefit most from more accurate upstream SV calling and genotyping (e.g., through more robust multi-sample calling and harmonised variant representations) and not so much from improved or SV-specific imputation methods. While improved imputation is an important research direction, it is beyond the scope of the present manuscript.

      (c) It is also important to assess how this reduced imputation quality may influence GWAS results. For instance, it would be useful to examine whether associated SVs exhibit higher imputation quality and whether SVs with lower quality are less likely to achieve significant association signals. In addition, the lower imputation quality observed for INV, DUP, and BND variants (Figure 3) may be due to their greater lengths (Figure 2). It is better to investigate the relationship between SV length and imputation quality.

      We agree that imputation quality can influence GWAS results. For example, for the FEV1/FVC phenotype, SVs with INFO > 0.9 are almost twice as likely to reach genome-wide significance (p < 5e-8) compared with SVs with 0.7 < INFO < 0.9 (odds ratio 1.95; Fisher’s exact test p-value 2.5e-5). This is consistent with the intuitive notion (applicable to any variants, not only to SVs) that greater uncertainty in the imputed genotypes dilutes association signals and therefore reduces power. For a detailed discussion of the relationship between allele frequency, imputation accuracy, and GWAS association results, see Zhang et al., Human Molecular Genetics 31(1):146–155 (2022), https://doi.org/10.1093/hmg/ddab203

      We have now investigated the relationship between SV length and imputation quality (the new Supplementary Figure 6). The results suggest that the observed association between imputation quality and SV size is primarily driven by the SV-size–dependent minor allele frequency in the imputation panel.

      (3) All examples presented in the manuscript focus on SVs that overlap with genes. It may also be valuable to investigate SVs that do not overlap with genes but intersect with enhancer regions. SVs can contribute to disease by altering regulatory elements, such as enhancers, which play a crucial role in gene expression. Including such analyses would further demonstrate the utility of SV-GWAS and provide deeper insights into the functional impact of SVs.

      We agree with the reviewer that examining SVs intersecting with enhancer regions could be an interesting direction for future studies, as it would provide additional insights into regulatory mechanisms and disease associations. However, in the present proof-of-principle study, we prefer focusing on SVs overlapping with genes and have now highlighted this in additional detail in the revised manuscript (Discussion, page 7):

      “In the present proof-of-principle study, we focused on SVs overlapping with the coding sequence of genes. In future applications of our SV imputation panel, more refined gene mapping approaches could be employed, e.g., including SVs overlapping enhancer regions or epigenetic marks. Such an enhanced mapping would increase the number of identified associated genes and thus provide additional insights into regulatory mechanisms and disease biology.”

      (4) The data availability link currently provides only a VCF file ("sniffles2_joint_sv_calls.vcf.gz") containing the identified SVs.

      (a) It would be beneficial for the authors to make all raw sequencing data (FASTQ files) and key processed datasets (such as alignment results and merged SV and SNV files) available. Providing these resources would enable other researchers to develop improved SV detection and imputation methods or conduct further genetic analyses.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we made both the SV calls and the full and reduced SV imputation panel files freely available. We have now added SV summary statistics from 32 SVwide association studies to the same resource. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download, in the manuscript.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) Furthermore, establishing a dedicated website for data access, along with a genome browser for SV visualization, could significantly enhance the impact and accessibility of the study. Additionally, all code, particularly the SV imputation pipeline accompanied by a detailed tutorial, should be deposited in a public repository such as GitHub. This would support researchers in imputing SVs and conducting SV-GWAS on their own datasets.

      The Methods section provides a full and detailed description of the imputation pipeline and parameters in the section “Preprocessing and imputation of SVs into UK Biobank” on page 14 of the revised manuscript. Our data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) In the Results section, Figure 3b is mentioned before 3a, and it is better to switch them in the Figure.

      Thank you for highlighting this fact. We acknowledge that typically the sequence of sections matches exactly between text and figures. However, in this specific case, we would prefer to deviate from the norm: In our opinion, Figure 3 is easier to interpret in its current sequence. At the same time, the text flows more logically in its current sequence, describing 3b before 3a. We would therefore prefer to stick to the current order, even if it means that Fig. 3b is described before 3a in the text.

      (2) Page 10, "Figure 1e" -> "Figure 2e".

      Thank you, we corrected this issue.

      (3) Page 14, "Leave-one out" -> "Leave-one-out".

      Thank you, we corrected this mistake.

      (4) It is better not to use abbreviations in the subheadings, especially "UKB" (page 3).

      Thank you, we changed the acronym ‘UKB’ to ‘UK Biobank’ in all subheadings.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to develop a novel and efficient method for SV detection, utilizing data from the 1000 Genomes Project (1KGP) for modeling and calibration. This method was subsequently validated using UK population data and applied to identify structural variants associated with specific disease phenotypes.

      Strengths:

      Third-generation single-molecule sequencing data offers several advantages over traditional high-throughput sequencing methods, particularly due to its long-read lengths, which provide valuable insights into significant forms of genomic variation. The authors have developed an efficient method for detecting structural variations and optimizing the utilization of genomic data. We hope that this method will continue to be refined, enabling researchers to more effectively leverage long-read data, high-throughput data, or even a synergistic combination of both.

      Weaknesses:

      Although this research contributes to our ability to more effectively utilize long-length and high-throughput data, there are some key issues that need to be addressed in terms of analyzing the specific results as well as writing the article.

      Reviewer #2 (Recommendations for the authors):

      (1) How to discuss the lower detection rate of structural variations (SVs) in East Asian populations, it is worth considering whether the authors' training dataset, which may have been based on raw data with insufficient representation of East Asian individuals, could have introduced a bias favoring other populations. This potential bias might arise from the relatively limited data available for Asian ancestry. Alternatively, the observed differences could also be influenced by the role of natural selection, which may have shaped the genomic landscape of East Asian populations in distinct ways. Further investigation is needed to clarify these possibilities.

      Thank you for raising this important point. Although an interesting research direction, a detailed investigation of the factors affecting SV detection rates is beyond the scope of the present study. However, we do not think that the lower detection rate in East Asians is due to an underrepresentation of Asian ancestry in our dataset. To explain this to all readers, we have added the following explanation to page 7 of the Discussion:

      “In this context, we observed that the number of SVs detected per individual differed between superpopulations. We identified the highest average number of SVs in individuals of African descent and a slightly lower average in East Asians, compared to the other superpopulations. While we included a higher number of African ancestry individuals, the number of East Asian individuals included in our reference panel was comparable to the number of individuals from other, non-African ancestries. In fact, it was even larger than the number of European ancestry individuals (AFR n=241, SAS n=171; EAS n=168; EUR n=164; AMR n=144). Therefore, we do not expect a major bias from underrepresentation of any superpopulation in the training dataset. It is well established that African ancestry is more diverse than is the case for other superpopulations [32, 33] and previous studies indicate that East Asian populations tend to exhibit slightly lower genetic diversity compared to European populations [34], which is consistent with the lower observed SV counts per genome.”

      (2) The authors did not present the results of the detection of CNV.

      Copy number variations (CNVs) are considered a subclass of structural variants. In our analysis, we detected deletions and duplications, which represent the most common forms of CNVs. However, we did not specifically investigate high copy-number SVs, as these are often larger than what can be reliably detected using long-read sequencing. Large-scale CNVs are typically identified in biobank studies through analysis of intensity data from genotyping microarrays using tools like PennCNV, and there is extensive literature supporting the use of this microarray approach in UK Biobank and other genotyped cohorts, see for example Aguirre et al.: Phenomewide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019, Aug 1;105(2):373-383. doi:10.1016/j.ajhg.2019.07.001.

      (3) Multiple testing correction is essential for ensuring the validity of large-scale structural variation (SV) association analyses. It is strongly recommended that the statistical methods and correction strategies employed, such as Bonferroni correction or false discovery rate (FDR) control, be explicitly detailed to enhance the transparency and reliability of the findings.

      For genome-wide SV association analyses, we applied the commonly used genome-wide significance threshold of 5e-8, which is standard in genome-wide studies. Given that these were exploratory proof-of-principle analyses illustrating use cases for SV analyses, we decided not to correct on top of that for multiple testing for the number of traits (32) tested. For the pQTL analyses, we further adjusted this threshold using a Bonferroni-type correction based on the number of proteins tested (1,463), to account for the increased number of multiple comparisons.

      We have now added a more detailed description of this multiple testing procedure to the Methods subsection “SV-wide association studies in UK Biobank” on page 16 of the revised manuscript:

      “In the exploratory SV-WAS, we used the standard threshold for genome-wide significance of p < 5×10<sup>-8</sup>. For the pQTL analyses, we applied Bonferroni correction for multiple testing on top of that genome-wide threshold, correcting for the number of tested protein levels (n=1463): p < 5×10<sup>-8</sup>/ 1463 = 3.4×10<sup>-11</sup>.”

      (4) The study primarily relied on data from the 1000 Genomes Project (1KGP) and the UK Biobank; however, the UK Biobank cohort is predominantly composed of individuals of European ancestry, which may restrict the generalizability of the research findings to other populations.

      Our reference panel was constructed to cover multiple ancestries, enabling imputation for diverse populations. Thus, our imputation panel can be applied to biobanks around the world and is freely available for this purpose. As a proof of principle, we have demonstrated the feasibility and performance of SV imputation in UK Biobank as an example of a broadly accessible cohort. We are looking forward to biobanks from diverse ancestries downloading our imputation panel and applying it to their populations.

      (5) Although the study employed long-read sequencing technology, the validation of structural variation (SV) detection accuracy predominantly relied on internal data, such as 'leave-one-out' validation. To further strengthen the reliability of the SV detection methods, it is recommended to incorporate additional external independent datasets for validation.

      The leave-one-out procedure in our study was used to validate the imputation performance, not the accuracy of SV detection. To assess SV calling accuracy, we performed extensive benchmarking against external SV call datasets derived from PacBio long-read sequencing and Illumina short-read sequencing. These details are provided under the subheading ‘Structural variant calling and benchmarking’ in the Results section on page 2 of the manuscript.

      (6) Some of the SVs mentioned in the study overlap with disease association loci in the GWAS Catalog, but functional annotation and exploration of the biological mechanisms of these SVs are more limited. It is suggested that LD can be added to analyse whether there are SNP that are highly linked to them to further explore their functions.

      We thank the reviewer for this suggestion. We have actually conducted an analysis addressing exactly this question: We performed conditional association analyses of the SV signals with nearby short variants (SNPs and InDels) at the SV locus. Such a conditional analysis addresses whether the observed SV association is influenced by LD-correlated SNPs or not. The results of this analysis are reported in Supplementary Tables 16 and 17. These tables include both the conditional analysis results and the LD between each SV and the variant at the locus with the second-highest evidence for an association.

      Researchers interested in exploring the biological significance of the SV-WAS results in more detail can now download the full SV-WAS summary statistics from https://opnme.com/genomiclens.

      (7) The discussion section could be further expanded to explore the role of SV in complex diseases and its potential application in precision medicine. For example, it could discuss how SV information can be integrated into existing GWAS frameworks to enhance the accuracy of disease risk prediction.

      Thank you for the suggestion, we have now added the following sentences to the discussion (page 7/8):

      “Structural variants can influence complex disease biology through either the disruption of coding sequence or an altered regulation of gene expression. Such effects may not be well captured by short variants alone. Incorporating SVs into GWAS and follow-up analyses would thus provide more accurate disease risk prediction, uncover underlying pathomechanisms by highlighting actionable pathways and targets, and support precision medicine by providing biomarkers for patient stratification.”

      (8) The geographic labeling of certain samples in Figure 2 appears to contain inaccuracies. For instance, the CDX sample, which represents the Dai population from Xishuangbanna in China's Yunnan Province, is currently mislabeled as originating from China's Inner Mongolia. This discrepancy should be corrected to ensure the accuracy of the data representation.

      We apologise for the misunderstanding. The geographic map in Figure 2a serves as an illustrative mapping of the samples to countries. It is intended to provide readers with an overview of population coverage, rather than to indicate the precise geographic origins of individual populations. The populations CDX, CHB, and CHS are displayed within the outline of China in alphabetical order, without any intention to indicate their exact geographic origin. We changed the respective figure caption to make this clear (page 19 of the revised manuscript):

      “Map of the 888 samples from the 1000 Genomes project, mapping the samples to countries and not indicating detailed geographical origins of populations.”

      Reviewer #3 (Public review):

      This study successfully identified genetic loci associated with various traits by generating large-scale long-read sequencing data from a diverse set of samples. This study is significant because it not only produces large-scale long-read genome sequencing data but also demonstrates its application in actual genetics research. Given its potential utility in various fields, this study is expected to make a valuable contribution to the academic community and to this journal. However, there are several critical aspects that could be improved. Below are specific comments for consideration.

      Strengths:

      Producing high-quality, large-scale variant datasets and imputation datasets

      Weaknesses:

      (1) Data availability

      Currently, it appears that only the Genomic Lens SV Panel is available on the webpage described in the Data Availability section. It is unclear whether the authors intend to release the raw sequencing data. Since the study utilized samples from the 1000 Genomes Project, there should be no restriction on making the data publicly accessible. Given this, would the authors consider making the raw sequencing reads publicly available? If so, NCBI SRA or EBI ENA would be the most appropriate repositories for data deposition. I strongly encourage the authors to consider public data release. Additionally, accessing the Genomic Lens SV Panel data does not seem straightforward. The manuscript should provide a more detailed description of how researchers can access and utilize these data. In my opinion, the best approach would be to upload the variant data (VCF files) to a public database such as the European Variation Archive (EVA) hosted by EBI.

      I strongly request that the authors publicly deposit the variant data. At a minimum:

      (a) The joint genotype data for all 888 samples from the 1000 Genomes Project must be publicly available.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we make both the SV calls and the full and reduced SV imputation panels (provided as multi-sample VCF files) freely available. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) For the UK Biobank samples, at least allele frequency data should be disclosed.

      Supplementary Table 5 includes the allele frequencies of the SVs imputed into UK Biobank.

      (c) Since eLife has a well-established data-sharing policy, compliance with these guidelines is essential for publication in this journal.

      By sharing the FASTQ files, the SV calls, the SV imputation panels, the SV summary statistics, and (once processed by UK Biobank) the genotypes of SVs imputed into UK Biobank, we are providing all SV data generated in our study.

      (2) Long-read sequencing data quality

      While the manuscript presents N50 read length and mean or median read base quality for each sample in a table, it would be highly beneficial to visualize these data in figures as well. A violin plot or similar visualization summarizing these distributions would significantly improve data presentation.

      Notably, the base quality of ONT long-read sequencing data appears lower than expected. This may be attributed to the use of pore version 9.4.1, but the unexpectedly low base quality still warrants attention. It would be helpful to include a small figure within Figure 2 to illustrate this point. A visual representation of read length distribution and base quality distribution would strengthen the manuscript.

      We thank the reviewer for this suggestion. We have now included two violin plots (the new Supplementary Figure 1) to the revised manuscript, summarising a) the N50 read length per sequencing run and b) the median read quality per sequencing run. These plots provide a clearer visualisation of the underlying distributions. We do not consider the ONT base quality to be low. Importantly, structural variant detection is generally robust to modest variations of per-base quality. Therefore, we do not expect the observed base quality levels to significantly affect SV calling in this study.

      (3) Variant detection precision, recall, and F1 score

      This study focuses on insertions and deletions (indels) {greater than or equal to}50 bp, but it remains unclear how well variants <50 bp are detected. I am particularly interested in the precision, recall, and F1 score for variants between 5-49 bp.

      While ONT base quality is relatively low, single-base variants are challenging to analyze, but variants {greater than or equal to}5 bp should still be detectable as their read accuracy is still approximately 90%, making analysis feasible. Given that Sniffles supports the detection of variants as small as 1 bp, I strongly encourage the authors to conduct an additional analysis.

      A simple two-category classification (e.g., 5-49 bp and {greater than or equal to}50 bp) should suffice. Additionally, a comparative analysis with HiFi and short-read sequencing data would be highly valuable. If possible, I strongly recommend that all detected variants {greater than or equal to}5 bp be made publicly available as VCF files.

      Because short InDels are available from high-coverage Illumina sequencing data generated for the same individuals (i.e., the data referred to as the NYGC dataset in our manuscript), we decided against calling such short variants from our lower coverage Oxford Nanopore data and thus concentrated our efforts on reliably calling longer variants covering at least 50 bp, consistent with the conventional definition of structural variants.

      (4) Assembly-based methods

      Given the low read accuracy and low sequencing depth in this dataset, it is understandable that genome assembly is challenging. However, the latest high-quality human genome datasets-such as those produced by the Human Pangenome Reference Consortium (HPRC)demonstrate that assembly-based approaches provide significant advantages, particularly for resolving complex and long structural variants.

      Since HPRC data also utilize 1000 Genomes Project samples, it would be highly informative to compare the accuracy of ONT sequencing in this study with HPRC's assembly-based genome data. The recent publication on 47 HPRC samples provides a valuable reference for such a comparison. Given its relevance, the authors should consider providing a comparative analysis with HPRC data.

      The aim of the present study was to generate an SV reference panel that enables SV imputation for large biobanks. Detailed assessments of ONT sequencing quality in general and comparisons to other sequencing efforts and technologies are out of scope for the present manuscript. We invite the scientific community to use the FASTQ files provided at ENA for conducting such detailed assessments in follow-up studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      General comments:

      You will see that many of the reviewers’ comments overlap. From our discussion with them, we agree that several of these comments should be addressed in this study, particularly comments related to the interpretation of the effect of drugs acting on the cellular cytoskeleton (reviewers #1 and #2). We also agree that the comparison of isogenic cell lines such as the mcf10a series or the 4T1 series should address some of the concerns regarding the interpretation of the mechanical fingerprint (reviewer #3). Also, certain methodological aspects should be easily clarified (reviewers #1 and #3).

      We also agreed that other comments may be more difficult to address in the context of this study. This is the case for comments related to establishing a link between different mechanical signatures and different cellular functions/outcomes (Reviewers #1 and #3). One could test whether migration or proliferation is altered by changing the mechanical fingerprint, or you could simply discuss these aspects by carefully reviewing the literature to corroborate mechanical signatures with known cellular phenotypes (e.g. migration speed, adhesion, cell size...). This is also the case for comments on the influence of other cellular parameters such as molecular crowding and energy metabolism, which could be left for future work or where you could use a low dose of cycloheximide (below the level of deleterious effects) to address the effect of cytoplasmic proteins (reviewer #3).

      We thank the editor for providing this helpful overview of the requested revisions. We have carefully addressed these points throughout the revised manuscript. The only difficulty was to establish the isogenic cell lines as requested. It took us over 18 months to find a source of these cells in Europe, and since then we are trying hard, but not successful to get these cells stably growing in the condition necessary for the optical tweezers experiments. As we have now spent more than 2 years on this without success, we decided to resubmit the paper without this part to not further delay this manuscript. The additional experiments and revisions have substantially strengthened the manuscript. Especially, the addition of Latrunculin A as suggested was an excellent request, as now the results regarding actin depolymerization and mechanical properties are in excellent agreement with the expected effects, as Latrunculin A is much more efficient in depolymerizing actin than cytochalasin B. The major changes are summarized below, followed by a detailed point-by-point response to all reviewer comments.

      General changes

      (1) Repeated measurements on wild-type HeLa cells.

      (2) Repeated all Cytochalasin B and Nocodazole experiments and increased the number of analyzed cells to approximately 60 per condition.

      (3) Performed additional experiments using Latrunculin A and combined Latrunculin A + Nocodazole treatment.

      (4) Performed immunostainings for all cytoskeletal perturbation conditions (WT, Cytochalasin B, Latrunculin A, Nocodazole, Cytochalasin B + Nocodazole, and Latrunculin A + Nocodazole).

      (5) Refined the rheological analysis procedure and expanded the methodological description.

      (6) Revised the manuscript text throughout and expanded the discussion of limitations and biological interpretation.

      Public Reviews:

      Reviewer #1 (Public Review):

      A limit of the paper is that the biological mechanisms by which intracellular mechanics is modulated (e.g. among cell types) remains unexplored and only briefly discussed. Yet this limit is greatly offset by the rigor of the approach.

      We thank the reviewer for this positive assessment and agree that a more extensive discussion of the biological mechanisms underlying the observed mechanical fingerprints strengthens the manuscript. We have substantially expanded the Discussion and Conclusion sections to address potential contributions of cytoskeletal organization, intracellular transport, molecular crowding, and metabolic state. In addition, we now discuss the relationship between the identified mechanical phase space and known cellular phenotypes where appropriate, while explicitly outlining the limitations of the current study and the need for future investigations linking intracellular mechanics to cellular function.

      Reviewer #2 (Public Review):

      The most difficult part of the method is the part with actin polymerization inhibition with cytochalasin B. The data shows that viscoelastic parameters as well as active energy parameters are unaffected by cytochalasin B. It is reasonable to expect that elasticity will reduce and fluidity will increase upon application of such a drug. The stiffness-reducing effect was observed only when CB was used with nocodazole most likely because of phagocytosis of the bead, which is governed by microtubule. The use of other actin-depolymerizing drugs such as latrunculin A would be needed to test actin’s role in mechanical fingerprints. If actin’s role is only explained by accompanying microtubule inhibition, it is not a convenient system to directly test the mechano-adaptation process.

      We thank the reviewer for this important suggestion. To strengthen the interpretation of the actin perturbation experiments, we repeated the Cytochalasin B measurements with an increased number of cells and performed additional experiments using Latrunculin A, a mechanistically distinct and more potent actin-depolymerizing compound. Together with complementary immunostaining experiments, these additional data reveal distinct contributions of the two major cytoskeletal systems to the intracellular mechanical fingerprint. Whereas actin depolymerization primarily affects intracellular stiffness and fluidity, microtubule depolymerization has the strongest effect on intracellular activity while also contributing to cellular softening. These additional experiments provide a substantially clearer interpretation of the respective roles of actin filaments and microtubules in shaping the intracellular mechanical fingerprint.

      Depolymerization of MT with nocodazole did not reduce the solid-like property A. Adding discussion and comparison with other papers in the literature using nocodazole will be helpful in understanding why.

      We thank the reviewer for this suggestion. We have expanded the discussion and now compare our observations with previous AFM studies investigating Nocodazole treatment. While AFM measurements of cortical mechanics often report little change or even increased stiffness after microtubule depolymerization, our intracellular measurements reveal pronounced softening and strongly reduced intracellular activity. We now discuss that this difference likely reflects the distinct intracellular mechanical compartment probed by intracellular microrheology compared with cortical AFM measurements.

      Overall, the usefulness of the concept of mechanical fingerprints and comparisons with other cell mechanics studies (from other groups) will make this manuscript stronger.

      We thank the reviewer for this suggestion. Throughout the revised manuscript we have strengthened the comparison of the mechanical fingerprint with previous literature. In particular, we now discuss the cytoskeletal perturbation experiments in the context of published AFM studies, compare the observed mechanical differences between cell types with previous measurements where available, and expand the discussion of the biological interpretation and limitations of the proposed mechanical fingerprint.

      Reviewer #3 (Public Review):

      The importance of the mechanical fingerprint is diluted due to some missing controls needed for biological relevance.

      We thank the reviewer for raising this important point. To strengthen the biological interpretation of the mechanical fingerprint, we performed substantial additional experiments, including repeated cytoskeletal perturbation measurements with increased sample sizes, additional Latrunculin A experiments, and complementary immunostaining analyses. We also expanded the discussion to address the influence of factors beyond the cytoskeleton, including molecular crowding and metabolic state, and explored possible relationships between the proposed mechanical phase space and cellular phenotypes. While we agree that future studies using well-controlled isogenic model systems will be required to establish direct links between intracellular mechanics and biological function, we believe that the additional experiments and expanded discussion substantially strengthen the biological relevance of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      A caveat of the general methodology, which is partially acknowledged in the MS is that beads are endocytosed and likely end up in specific lysosomal compartments. Therefore, it is not clear whether the mechanical fingerprint fully represent the material properties of bulk cytoplasm, and not something more specific to lysosomal organelles. For instance, lysosome motion may be largely driven by motors moving along MT cytoskeletal track, and the extracted effective energy may as such not fully represent the crowding and effective active temperature of the cytoplasm. This limit certainly affect the interpretation of the results in other cell types, in which membrane trafficking and cytoskeletal organization may vary largely. I believe it would be very important to outline this limitation of the work and discuss it in light of the results obtained throughout.

      We thank the reviewer for pointing out this important limitation, which was not sufficiently addressed in the original manuscript. We have now acknowledged this issue throughout the manuscript and added a limitation section to the conclusion to clarify that our findings specifically relate to internalized objects surrounded by a membrane and therefore primarily reflect the properties of membrane-bound organelles in the 1 µm size regime, rather than the bulk cytoplasm as a whole.

      We consider this focus on membrane-enclosed intracellular objects to be biologically relevant and interesting in its own right. Alternative approaches for introducing tracer particles, such as microinjection or particle guns, are generally more invasive and less reproducible. We therefore deliberately focused on phagocytosed beads as a minimally perturbative and robust experimental system in this study.

      The evolution of the mechanics in Hela Cells using cytoskeletal drugs in interesting, but I was confused by the fact that authors interpret the effect of cytochalasin solely on the cortex. As they are probing intracellular rheology, variations (or lack thereof) may rather reflect bulk F-actin networks? Also the compensation mechanism is interested, but it would need to be strengthened by immunostaining for instance, to support the claim, that microtubule depolymerization enhances F-actin networks.

      We thank the reviewer for this important comment. To elaborate on the effect of cytoskeletal filaments, we extended our analysis by repeating the experiments, increasing the number of samples, and investigating the effect of an additional drug, Latrunculin A. Additionally, we conducted immunostaining with subsequent confocal imaging to deepen our understanding of the effect of the respective drugs. The additional experiments reveal that actin and microtubules contribute differently to the fingerprint. Actin depolymerization primarily affects intracellular stiffness and fluidity, whereas microtubule depolymerization has the strongest effect on both mechanics and intracellular activity. Combined perturbation produces the largest overall effect. Based on these additional data, we no longer invoke the compensation mechanism proposed in the original manuscript. While interactions between the actin and microtubule cytoskeleton have been reported previously, our immunostaining experiments do not provide evidence for a compensatory increase in actin organization following microtubule depolymerization. We have therefore removed this interpretation from the revised manuscript and replaced it with a discussion based on the newly acquired perturbation and imaging data.

      The final figure using principal component analysis is very interesting, but it would be important to link this to phenotypic signatures of the different cells. Could the authors try to link resistance, fluidity and activity to the different functions/behavior of cells? For instance, some of these cells are migratory but some may move much faster than others, and it would be very interesting to correlate the degree of activity or fluidity with speed of migration, or cell shape/size/contractile state for example.

      Indeed, this is an important point. Establishing direct links between the mechanical fingerprint and functional cellular properties such as migration, contractility, proliferation, or morphology would substantially strengthen the biological interpretation of the identified phase space. We carefully considered this suggestion and explored several approaches to relate the measured mechanical parameters to cellular phenotype. However, obtaining directly comparable quantitative functional data across all investigated cell types proved challenging. Parameters such as migration speed, adhesion, and contractility depend strongly on experimental conditions, including substrate properties, assay design, and culture conditions, making literature values difficult to compare across studies. To address the reviewer’s concern, we expanded the discussion and incorporated comparisons to available literature where appropriate. For example, previous studies have reported higher migration rates for HeLa cells compared with MCF7 cells, which is qualitatively consistent with the higher intracellular activity observed in HeLa cells. However, due to the limited comparability and availability of quantitative functional data across the investigated cell types, we refrained from performing a formal correlation analysis. In addition, we grouped the investigated cell lines according to several broad phenotypic classifications, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that could support definitive conclusions regarding structure–function relationships. We therefore agree with the reviewer that establishing direct links between intracellular mechanical fingerprints and cellular function represents an important next step. To this end, future studies will combine intracellular rheological measurements with independently quantified functional assays, ideally in well-controlled isogenic model systems.

      Reviewer #2 (Recommendations For The Authors):

      The study needs more thorough validation against known technology (such as AFM) or literature, e.g., rheological change upon the same drugs used in the current study.

      We thank the reviewer for this suggestion. We have expanded the discussion of the cytoskeletal perturbation experiments and now compare our observations to previous AFM studies and related literature on cytoskeletal mechanics. Consistent with AFM measurements of cortical mechanics, actin depolymerization using Cytochalasin B or Latrunculin A resulted in a reduction of cellular stiffness. In contrast, microtubule depolymerization produced effects that differ from many AFM studies, which report either no change or an increase in cortical stiffness following Nocodazole treatment. We now explicitly discuss that this discrepancy likely reflects the different mechanical compartments probed by the two techniques. AFM predominantly measures the actin-rich cell cortex, whereas our intracellular microrheology measurements probe the mechanical environment experienced by membrane-bound intracellular particles. We therefore interpret the differing response to microtubule depolymerization as evidence that intracellular active mechanics and cortical mechanics can be influenced by distinct physical mechanisms. These comparisons have been incorporated into the Results and Discussion sections of the revised manuscript.

      Page 8: Citation to Fig. 3a is missing before mentioning Fig. 3b.

      We revised the manuscript to ensure that all references are given in an appropriate order.

      Proper uses of hyphens are recommended to avoid confusion. For example, ’a yet not understood change’ can be written as ’ a yet-not-understood change’.

      We thank the reviewer for this suggestion. We carefully revised the manuscript to improve the use of hyphenation and compound modifiers throughout the text. The specific example highlighted by the reviewer, as well as similar constructions, have been corrected to improve readability and avoid ambiguity.

      Reviewer #3 (Recommendations For The Authors):

      As it reads, sinusoidal waves are applied sequentially from 1- 1024Hz. Please clarify if amplitude is the same for each frequency, also how many frequencies are used? On this point, due to perturbations due to alterations in pre-stress, are the orders of frequencies randomized?

      We thank the reviewer for pointing out this ambiguity. We have revised the manuscript to provide a more detailed description of the active microrheology protocol. Specifically, we now state that all measurements were performed using a constant trapping-laser oscillation amplitude of 200 nm and that the applied frequencies were 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, and 1024 Hz. The frequencies were applied sequentially in increasing order and were not randomized. This information has now been added to the manuscript.

      How many beads are probed in a given cell?

      We thank the reviewer for this question. We have clarified this point in the Methods section and now explicitly state that only a single phagocytosed probe particle was analyzed per cell. Of course, many different cells, and hence beads, have been analyzed per cell type.

      Is the graph in 1 c G’, G” per cell or average of many cells?

      We thank the reviewer for pointing out this ambiguity. In the original version of the manuscript, Figure 1b showed data from a representative cell, whereas Figure 1c displayed an average over multiple cells. To avoid confusion, we revised Figure 1 and now show representative data from a single measurement throughout the analysis workflow (Figure 1c,e,f).

      Figure 1e is quite nice, however, is there an equivalent performed in a nonlinear ECM such as collagen for comparison, in a similar vein can the equivalent be calculated for cells with/without treatment with low doses of cycloheximide to reduce protein synthesis? Yes, cytoskeletal elements are important for cell mechanics, but cytoplasm crowding is often an overlooked factor.

      We thank the reviewer for this important suggestion, and we are glad that the reviewer likes figure 1e. Regarding non-linear ECM, we have not done such experiments using optical tweezers. Collagen is a highly heterogeneous material and using the small deformations that we can obtain using the optical tweezers, our access to the non-linear contributions is rather limited.

      However, we agree that factors beyond the cytoskeleton, including molecular crowding and protein content, can make important contributions to intracellular mechanics. While investigating these effects experimentally, for example through cycloheximide treatment, would be highly interesting, such studies were beyond the scope of the present work.

      The primary focus of this study was to establish and validate a mechanical fingerprint for intracellular active microrheology and to investigate how this fingerprint responds to perturbations of the cytoskeleton. The additional experiments performed during revision therefore concentrated on strengthening the interpretation of the cytoskeletal contributions.

      At the same time, we agree that molecular crowding represents an important alternative mechanism influencing intracellular mechanics. We have therefore expanded the Discussion and Conclusion sections to explicitly acknowledge this limitation and now cite recent studies demonstrating strong effects of molecular crowding on intracellular rheology (Umeda et al,. 2023, Ebata et al., 2023). We further discuss that, besides cytoskeletal organization, metabolic state, intracellular transport, and molecular crowding are likely contributors to the observed mechanical fingerprint.

      The biggest issue is the interpretation of the different factors as each of these cells have different energetic needs. The comparison between cancer cells with different aggressiveness, immune and epithelial cells. For example, some types of cancer cells will be dominated by glycolysis vs oxphos, which will influence both the cytoplasmic and nuclear mechanics? It would be useful to carefully assess factors not restricted to

      (a) Cytoskeleton

      (b) Protein synthesis

      (c) Metabolic state

      For similar lines and/or cells where there are lineages that are either more metastatic in cancer, normal counterpart or drug resistant in an effort to link the fingerprint to a biological output. Specifically, is migration, proliferation, survival correlated with the measurements. The reviewer is sensitive to the technical difficulties of the experiments. However, the interpretation and importance of the mechanical fingerprinting requires additional work as mentioned above.

      We thank the reviewer for this thoughtful comment. We agree that intracellular mechanics is likely influenced by a broad range of biological factors beyond the cytoskeleton, including metabolic state, molecular crowding, intracellular transport processes, and protein synthesis. We also agree that the biological significance of the mechanical fingerprint would be strengthened by establishing direct links to functional cellular outputs such as migration, proliferation, or survival. To address the first point, we have expanded the Discussion and Conclusion sections of the manuscript to explicitly acknowledge that the observed fingerprint is unlikely to be determined solely by cytoskeletal organization. In particular, we now discuss the potential contributions of metabolic state, intracellular transport, and molecular crowding, and cite recent studies demonstrating the importance of these factors for intracellular mechanics. To address the second point, we explored several strategies to relate the measured mechanical fingerprints to cellular phenotype. We expanded the discussion of available literature, including examples where mechanical properties and migratory behavior appear qualitatively consistent. In addition, we grouped the investigated cell lines according to broad biological characteristics, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that would support definitive conclusions regarding structure–function relationships. We therefore agree that establishing direct links between intracellular mechanics and cellular function represents an important next step. Such studies will require quantitative functional assays performed under controlled and directly comparable conditions, ideally using well-defined isogenic model systems. We now discuss these limitations and future directions explicitly in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Reviewing Editor, the Senior Editor, and the three reviewers for their careful and constructive assessment of our manuscript. We were encouraged that the reviewers found the question timely and novel, the experimental design thoughtful and well-replicated, and the analyses diverse and informative. The reviewers also raised a number of valuable concerns, which clustered around three themes: (i) the framing of host-mediated selection as microbiome “engineering” versus a proof of concept; (ii) the interpretive challenges introduced by microbial dispersal and the resulting limits on the sterile-inoculated controls; and (iii) requests for clearer methodological detail and additional context from the recent literature. We have revised the manuscript to address these points through clearer framing, expanded discussion, and fuller methodological detail. Consistent with the nature of this long-term experiment, our revisions strengthen the interpretation and presentation of the existing dataset rather than adding new experiments.

      eLife Assessment

      The study has also shortcomings in that the rescuing effect is not benchmarked against healthy well-watered plants, the sterilized controls do not add much information, and the dispersal between inocula confounds the interpretation of the results… the presentation would overall benefit from more extensive consideration of recent developments in the field.

      We appreciate this balanced summary and have revised the manuscript accordingly. We have reframed the abstract and Introduction to present the study explicitly as a proof of concept rather than a completed engineering effort (ll. 27–31; ll. 96–101); we now address the well-watered benchmarking limitation and the limits of the sterile-inoculated controls directly in the Discussion (ll. 543–551); we discuss dispersal and its confounding effect on interpretation head-on, including an alternative hypothesis (ll. 546–551); and we have incorporated the recent studies suggested by Reviewer 3 (ll. 206–209, 442, 476–478). Each change is detailed in the point-by-point responses below.

      Reviewer #1 (Public Review):

      Weaknesses:

      The findings demonstrate the efficacy of host-mediated microbiome selection, but the engineering part for enhancing rice performance under drought-stress conditions has not been provided. The proposed mechanisms rely on correlations but not direct experimental proofs.

      We agree, and we have adopted this framing throughout. Our study demonstrates host-mediated selection as a discovery framework rather than a completed engineering pipeline, and we now say so explicitly: the abstract has been reframed (ll. 27–31) and a statement added at the end of the Introduction (ll. 96–101) clarifying that the work reproducibly enriches beneficial taxa and functions and yields simplified candidate communities, but does not yet benchmark those communities against single-isolate inoculants or test them in the field or against a resident native microbiome. We likewise agree that the functional inferences from our metagenome-assembled genomes (MAGs) are correlative; we now state this explicitly in the Methods and Discussion (ll. 776–778) and note that establishing causal roles for individual taxa or genes will require targeted isolation and genetic manipulation.

      Reviewer #1 (Recommendations For The Authors):

      The experimental design… could benefit from more detailed explanations. For instance, what are the criteria for choosing these soils and how are they relevant to rice growth phenotype? Also, the word ‘generation’ is misleading as it implies the use of seed-to-seed experiments… It would also be good to explain why the authors chose 6 generations for rice fields and 4 generations for deserts and serpentine seep. Importantly, the contribution of the rice seed microbiome… has not been considered and is also missing from… the discussion.

      We have addressed each part of this comment. Soil selection criteria: the Results section “Source inocula bacterial diversity” describes our rationale — we screened nine field soils in a pilot experiment, then selected the three that both supported rice growth and had negligible taxonomic overlap (providing three distinct starting points), with a stated per-soil expectation (rice-adapted, drought-adapted, and high-diversity). We are happy to expand this further if the reviewer feels additional detail is needed. “Generation”: we now define this term as a single 40-day selection cycle rather than a seed-to-seed generation (l. 109). Six vs. four generations: we explain in the Results (l. 309) that, having observed convergence of microbiome composition across soil treatments by the fourth selection generation, we concentrated resources on Rice Field and carried it through two additional cycles. Seed microbiome: we have added a note to the Discussion (ll. 444–446) that, although seeds were surface-sterilized before each generation, a residual contribution of seed-borne endophytes common to all treatments cannot be excluded.

      The authors stated that microbiomes were not selected for propagation into future generations in control lines. In this case, have the authors tested if the control LI microbiome in SG1 through SG6 did or did not significantly change in all the soil types?

      We have clarified the role of the live-inoculated (LI) lines in the text. LI lines were well-watered controls that were re-inoculated each generation with unsterilized selection-line material; they were included to identify drought-enriched taxa (by contrast with the droughted selection lines) and to test whether drought-optimized microbiomes were deleterious under well-watered conditions — not as an independently propagated selection line. Because LI communities were re-derived from selection-line inocula each generation, their composition necessarily tracked the changes occurring in the selection lines; this is the basis of the SL-versus-LI differential-abundance analysis (Figure 7B, Supplemental Figure 9). We note that comprehensive, temporally resolved 16S sequencing was performed for Rice Field, so we are appropriately cautious about extending LI comparisons across every soil type, and we have tempered our conclusions from the control lines accordingly (ll. 543–551).

      In Figure 3B (rice field), the tolerance in terms of AUC NDVI contrastingly increases to the biomass values in SG5 and SG6… the [NDVI] does not seem to be a good measure… It would be interesting to analyze these data sets under normal conditions… include representative pictures of all the ‘generations’… It will also be important to include the LI control data in Figures 3B and 3C.

      We appreciate these suggestions and respond to each. Our metric is biomass-adjusted AUC NDVI, which we use precisely to separate drought performance from plant size; NDVI itself was validated against shoot water content in preliminary experiments (R = 0.98; Supplemental Figure 2E), so we are confident it is an appropriate, validated proxy for drought status. We have substantially expanded the Methods to explain this adjustment and why the metric can diverge from raw biomass (ll. 666–669). Regarding the specific additions requested: analyzing the well-watered plants as a phenotypic dataset, adding representative images for every generation, and plotting LI data in Figure 3B/C would each require new analyses or figures that are outside the scope of this revision; moreover, LI plants were never droughted and therefore have no drought-response score comparable to the SL and SI lines, so they cannot be placed on the same axes. Representative images contrasting the first and last selection generations are already provided in Figure 3A. We have, however, added an explicit acknowledgement that our design does not quantify the absolute magnitude of drought rescue relative to well-watered performance (ll. 551–552).

      It is less clear how sterile soils acquired environmental taxa over time. Was this a seepage of microbes from inoculated samples to the calcinated clay, possibly via the water irrigation system? In this regard, four Venn diagrams representing all the generations… would be relevant.

      Each plant was grown in an individual container with its own separate water reservoir (Supplemental Figure 4), so shared irrigation was not a route of transfer; the most likely routes are airborne movement and handling within the growth chamber, together with within-treatment shuffling of plants. Our dispersal analysis (Figure 5) already traces the origins of taxa in each treatment, and Supplemental Figure 6 quantifies the ASVs shared among treatments over generations; we have added explicit criteria for these origin assignments (ll. 248–253). We therefore prefer to retain the existing Figure 5 / Supplemental Figure 6 presentation rather than add four separate Venn diagrams, which would convey the same information less quantitatively, but we are glad to reconsider if the editor feels a Venn representation would help readers.

      What is the logic behind the so-called ‘immigrating taxa’ in this study?

      “Immigrating” (dispersed) taxa are those that appear in a treatment despite not being attributable to that treatment’s own starting material — i.e., ASVs not detected in that treatment’s field soil or enrichment-generation inoculum, which must therefore have arrived by dispersal from other treatments or from the growth-chamber environment. We have made this definition explicit in the text (ll. 248–253).

      The decrease in alpha-diversity in subsequent generations… should be thoroughly discussed. Have authors tried to culture these few remaining taxa? If yes… tested for their individual drought tolerance supported by physiological assays… If no, is the microbiome of SG6 (and associated functions) ideal or sufficient to create drought tolerance in field conditions?

      We have expanded the discussion of the diversity decline. In addition to niche filtering along the soil-to-root gradient and dilution-to-extinction (already discussed), we now note that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent collapse in diversity may reflect enrichment for the taxa that were active in the original inoculum (ll. 206–209). We agree that culturing the remaining taxa and characterizing them with physiological assays (e.g., water potential, water-use efficiency, stomatal conductance) is a valuable next step; these experiments are outside the scope of the present study, which we have now framed explicitly as a proof of concept, and we identify field validation of selected communities as a key open question (ll. 96–101).

      The result that Ideonella was identified as the dominant taxa in all selection conditions is highly interesting… This… should have been followed up for isolating the strains and performing direct tests to test their importance for conveying drought stress.

      We agree that isolating and directly testing dominant taxa such as Ideonella is the logical next step, and we now emphasize that a central value of host-mediated selection is that it yields simplified communities from which such taxa can be more readily isolated (ll. 27–31). These isolation and functional-validation experiments are beyond the scope of the current study and we have framed them as future directions rather than undertaking them here.

      The MAGs shown in Figure 8 have apparently ‘been assigned to ASVs…’. These data are not shown anywhere… the MAG data only give correlations but not direct genetic proofs of the biological functions of the identified genes.

      We have expanded the Methods to describe how each MAG was matched to an ASV (by closest taxonomic assignment and by concordance of relative abundance across samples), and we now state explicitly that these assignments are approximate and that the functional inferences drawn from them are correlative rather than definitive (ll. 776–778). We would be glad to add a supplemental table listing the MAG-to-ASV assignments if the reviewer or editor would find it useful; because it reports assignments already in hand, it requires no new analysis.

      Reviewer #2 (Public Review):

      Strengths:

      I think this study examines an important and exciting topic in the area of plant microbiomes. I predict the findings of the experiments will inform a wide audience of researchers attempting similar studies and be helpful in their designs.

      We thank the reviewer for recognizing the novelty of this complex experiment as well as the effort we put into designing it. Like the reviewer, we hope that this manuscript can serve a wide audience and help inform subsequent experiments in this new topic area.

      Weaknesses:

      Although the controls were well designed, the dispersal of the microbiomes erased the utility of the sterile inoculated (SI) controls… the SI lines acquired microbes from the experiment and never appeared to significantly deviate from the SL plants. The dispersal of the microbes… also minimizes any conclusions that can be made about the different starting inocula and how prone to selection they may be.

      We agree that microbial dispersal confounded our ability to use the sterile-inoculated (SI) plants to account for batch variation between generations. By maintaining each plant as a spatially discrete unit (individual pots and watering reservoirs), we had originally intended SI plants simply to acquire a similar consortium of environmental microbiota each generation. Truly axenic SI plants would have been better suited to this purpose, but would have severely limited the number of replicates and replicate selection lines we could include. We have now addressed this limitation directly in the Discussion (ll. 543–551): we state that the SI lines cannot be treated as static, microbe-free baselines, that the batch-to-batch variation they were meant to capture is only partially controlled, and that dispersal limits the strength of the conclusions we can draw about differences between starting inocula. This shared trajectory of selection and intended control lines has been observed in other host-mediated selection studies but rarely discussed in detail, and we now foreground it as a lesson for experimental design.

      Reviewer #2 (Recommendations For The Authors):

      My first concern is the framing of the approach… the authors never show that this approach has better efficacy than single-isolate inoculates… The phase of the research is still proof of concept, understandably, but these caveats should be mentioned/addressed head-on in the Introduction and Discussion.

      We agree and have made these caveats explicit rather than implicit. The abstract now frames the work as identifying candidate taxa and communities rather than delivering a finished engineering solution (ll. 27–31); the Introduction now states plainly that this is a proof of concept that does not benchmark the passaged communities against single-isolate inoculants or evaluate them in the field or against a resident native microbiome (ll. 96–101); and the Discussion reiterates these limitations (ll. 543–551).

      I disagree with the authors that the selected microbiota better approximate field conditions (line 55) - because… the diversity of the microbiome is drastically reduced… it is likely that exclusion of taxa is just as important as the passaging of bacterial members to see the desired effect.

      We take this point and have revised the sentence at (former) line 55 accordingly (now l. 57): we now say that community-level screening more closely approximates field complexity than single-isolate screens only at the outset, and we no longer imply that the selected (diversity-reduced) communities better approximate the field. We agree that taxon exclusion may be as important as enrichment; this is consistent with our balance analyses, in which the denominator groups comprise taxa negatively associated with phenotype (Figure 7), and with the diminishing returns we observe as diversity collapses. We have also added an explicit sentence to the Discussion (l. 488) stating that the exclusion of detrimental taxa may be as important as the enrichment of beneficial ones, and that a microbiome’s finite membership may contribute to the diminishing returns of selection we observe over generations.

      (1) It is unclear what the reason (or methodology) for correcting NDVI by biomass. Much of the findings hinge on corrected NDVI values, so a more thorough explanation of the correction method… would benefit the reader.

      We have substantially expanded this explanation in the Methods (ll. 666–669). We now state that biomass and AUC NDVI were anti-correlated (Supplemental Figure 12) and that we adjusted for plant size by taking the residuals of a linear regression of AUC NDVI on shoot dry-weight biomass, using these biomass-adjusted values as our measure of drought performance so that selection would reflect drought tolerance rather than plant size alone.

      (2) Are data for panels B and C of Figure 3 scaled?… how can one have a negative area under the curve if all the NDVI values are positive? For panel B, the representative plant images are much larger than 0.8 grams.

      This is a helpful catch, and the confusion stems from our terse original description. The values plotted are the biomass-adjusted AUC NDVI (regression residuals), which are centered on zero by construction; negative values therefore indicate poorer-than-expected drought performance for a plant of a given size and do not reflect negative raw NDVI or a negative raw area under the curve. We now explain this explicitly (ll. 666–669). In panel C, shoot biomass is plotted as dry weight in grams; the representative plant images in panel A are qualitative illustrations and are not scaled to the biomass axis. We will make the axis labels and legend state the units and the residual nature of the adjusted metric explicitly (noted in our accompanying figure-revision guide).

      (3) The dispersal analysis… What are the criteria for classifying ASVs as specific to an input source? Was it that they were observed in all samples of field soil, i.e. was a prevalence threshold implemented? Could they be observed in any other soil at a smaller threshold?

      We have added the criteria explicitly (ll. 248–253). An ASV was attributed to a given soil treatment if it was detected (present/absent) in that treatment’s field-soil or enrichment-generation inoculum samples; ASVs detected in none of the field soils or source inocula were designated environmental in origin (“unk/env”), and ASVs meeting the criterion for more than one treatment were assigned to each. Assignments were thus based on detection in the source samples rather than on an abundance-prevalence threshold within later generations.

      This reviewer finds the results around [inoculum source] inconclusive… the serpentine seep microbiome appears to provide more benefit from the first round of selection than any other soil… The slope of improvement… is different between soils, but mainly because the serpentine microbiomes start out conveying greater benefits than the other soils.

      We agree the Serpentine Seep result is not clear-cut. The Discussion already presents inoculum provenance as one of several factors shaping the outcome rather than a decisive one, and we have now added an explicit acknowledgement that Serpentine Seep conferred comparatively large benefits in the earliest cycles before plateauing, so its weaker response to continued selection may reflect an early approach to a performance ceiling rather than an inherently poorer substrate for selection (l. 423). We have tempered our “source matters” language accordingly.

      Have the authors assessed the biomass and ndvi of the well-watered plants?… showing this data would allow the reader to assess the degree to which the microbiomes are rescuing the plant… and… whether tradeoffs exist… under fully watered conditions.

      We have added an explicit statement that our design does not pair each droughted line with a well-watered readout of the same phenotype, so we refrain from estimating the absolute magnitude of drought rescue (ll. 551–552). We note, however, that shoot biomass increased in parallel with drought performance across selection generations (Figure 3C), which provides no evidence that selection for drought tolerance came at a cost to growth under our conditions. Collecting matched well-watered phenotypes to quantify effect size and trade-offs is a worthwhile aim for future work but would constitute a new analysis beyond this revision.

      How can the authors exclude the possibility that environmental microbes pre-existing in the growth chamber taxonomically overlap with the field soil-specific microbes?… the alternative hypothesis should be mentioned… A clearer representation of the ASVs categorized as source soil-specific in Figure 5… would be useful and how many of these ASVs make up the bar plots.

      We now state this alternative hypothesis explicitly: because dispersed taxa came to dominate all treatments, we cannot fully exclude that taxa shared across treatments were recruited from a common growth-chamber pool rather than dispersing directly between soils (ll. 546–549). We note that the two processes are difficult to distinguish retrospectively, but that the bias of each SI line toward its own treatment’s native diversity (Figure 5) is more consistent with genuine cross-treatment dispersal. Regarding the figure, the number of ASVs underlying each origin category is available in Supplemental Figure 6; we describe in the accompanying figure-revision guide how the Figure 5 legend can be clarified to state the assignment criteria and the ASV counts.

      The sterile inoculated plants were a nice control in theory, but I question their utility… A contrast that should be made is the microbiomes of only SI plants. It is striking that sterilized controls assemble and retain more microbes from the unsterilized starting inoculum. I would expect everything to be acquired from dispersal.

      We agree, and we have foregrounded this in the Discussion (ll. 543–551). As the reviewer notes, SI communities were biased toward their own treatment’s native diversity rather than being assembled entirely from dispersal (Figure 5) — an informative observation, but one that also demonstrates why the SI lines cannot serve as the clean, microbe-free baseline we had intended. We now treat this as a key design lesson and note that a fully isolated (e.g., gnotobiotic) control would be required to separate these effects in future experiments (l. 560).

      Reviewer #3 (Public Review):

      Weaknesses:

      Sterile/non-inoculated calcined clay also tends to enrich similar microbes… In a future experiment, the work would benefit from including a truly sterile control… the reader may get to wonder whether these efforts are necessary at all… This is discussed across the paper but not directly addressed and I think the manuscript would benefit from a clear argument for or against this idea.

      We thank the reviewer for this insightful point and have made our argument explicit rather than leaving it implicit. First, we agree a fully isolated, truly sterile control would strengthen future iterations of this design; the manuscript notes that gnotobiotic plants would be the ideal (if costly) means of achieving this (l. 560). Second, on whether selection is necessary if plants recruit beneficial microbes from the environment: the phenotypic gains seen even in the sterile-inoculated lines do not indicate that selection was superfluous, but rather that those plants recruited from a metacommunity that was itself being optimized by selection in the neighboring selection lines each generation. In other words, environmental acquisition propagated the benefits of selection across the shared growth-chamber environment rather than replacing it. We have clarified this reasoning in the Discussion (ll. 543–551).

      Reviewer #3 (Recommendations For The Authors):

      It is mentioned multiple times… that host genotype is the driver of the microbiota selection… However, this is not the case [multiple lines] and therefore I don’t find that surprising that there is a convergence of the microbiota across soils and selection rounds.

      We agree and have added text making this explicit: all plants were a single, near-isogenic rice genotype, and because host genotype is itself a strong filter on microbiome composition, the use of one genotype — together with shared environmental conditions and selection criteria — makes convergence across lines an expected rather than a surprising outcome (ll. 438–441). We have softened language that could be read as attributing selection to host-genotype variation.

      Another possibility… is that those microbes that are found in the later generations are actually the ones that were active/alive in the initial inoculum. It is not possible to rule out that most of the sequenced microbes in the input were not actually dead. Similar observations were made… in Duran et al. 2022. New Phytol.

      We have added this possibility to the manuscript, noting that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent diversity decline may reflect enrichment for the subset of taxa that were active in the original inoculum, with reference to the transplantation work the reviewer cites (Durán et al. 2022; ll. 206–209).

      In the shotgun data, was there any observation of other microbes present (fungi, virus)? Did they follow the same trends as the bacterial communities?… I think addressing this will be very interesting and very novel.

      We agree this is an interesting question. Our shotgun workflow was designed and assembled specifically to recover high-quality bacterial and archaeal MAGs, and a rigorous cross-kingdom analysis (fungi, viruses) would require dedicated, eukaryote- and virus-specific assembly, binning, and reference databases — a substantial new analysis that lies outside the scope of this revision. We therefore flag cross-kingdom community dynamics as a promising direction for future work rather than presenting a new analysis here.

      Any interesting overlap with the results found in Karasov et al. 2022 (biorxiv)?

      We have added a comparison to drought-driven selection on host-associated microbiomes in Arabidopsis (Karasov et al. 2022) at the relevant point in the Discussion (l. 442).

      In Liu et al., 2024 Nat. Comms, the authors found Devosia as an interesting candidate for disease suppression (to add to the discussion?).

      Added — we now note that Devosia, one of the lesser-known genera enriched in our experiment, has recently been highlighted as a candidate mediator of disease suppression in the rhizosphere (Liu et al. 2024; l. 476).

      Lipids as a signal for host-microbe interaction: Rich et al., 2021 Science.

      Added — in the functional-enrichment discussion we now cite lipids as increasingly recognized central signaling molecules in host–microbe symbioses (Rich et al. 2021; l. 478).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Dwulet et al. combined experimental and modeling approaches to investigate how correlated spontaneous activity in the mouse's primary visual (V1) and primary somatosensory (S1) areas drives the development of multisensory integration in area RL. Notably, they focused on early developmental stages, before sensory experience occurs. Consistent with previous experimental findings, the authors first demonstrated that spontaneous activity becomes more sparse across development in all three areas, as measured by event amplitude, event duration, and participation ratio. Using a linear mixed model analysis to compare the maturation of this spontaneous activity, they found evidence that S1 matured the fastest. The authors then presented experimental evidence suggesting that these spontaneous events were moderately correlated both spatially and temporally.

      They hypothesized that activity-dependent mechanisms use these correlations to establish connectivity across these regions. To test this hypothesis, the authors modeled a feedforward network with connections from S1 to RL and from V1 to RL, where the strength of connections depended on a Hebbian term for potentiation and a heterosynaptic term for depression. By investigating different levels of V1-S1 correlations, they found that moderate levels of correlation led to the significant development of topographically organized connectivity while maintaining a mix of bimodal and unimodal cells in RL. Additionally, when simulating a network with a more mature S1, they observed that topographical maps improved not only between S1 and RL but also between V1 and RL. Finally, the authors use linear regression to suggest that the mixture of bimodal and unimodal cells in RL is optimal for encoding the maximum amount of information from both V1 and S1.

      However, there are significant gaps between the experimental data and the modeling setup, which weaken the paper's conclusions. Additionally, some key details are omitted, making it difficult to fully assess their analysis and interpret some of their figures.

      (1) Some of the statistical measures and techniques in Figure 1 could benefit from clearer definitions. While the thresholds for activation (peak with at least 5% dF/F0) and events (20% of recorded cells activated simultaneously) are provided, event duration and participation rate are not clearly defined. Based on this definition of event alone, it is unclear why the minimum participation rate in Figure 1F is not 20%. Additionally, the conclusion that S1 matures earlier than RL and V1 could be strengthened by including a direct comparison between S1 and RL, as the current analysis only compares these areas to V1.

      We thank the reviewer for this comment. We have now updated the Methods to include the event duration as time above half max, participation rate as % of cells out of total in that region active during an event. Also, the threshold of 20% recorded cells to identify an event was incorrectly stated, in fact the threshold was 5% consistent with what the reviewer observed in Figure 1F. This error has been corrected throughout the Methods. We chose 5% because spontaneous activity significantly sparsifies over development, with events involving far fewer cells, as previously shown by multiple studies (Golshani et al., 2009; Rochefort et al., 2009; Gribizis et al., 2019; Leighton et al., 2021; Murakami et al., 2022; Chini et al., 2022; reviewed in Lakhera et al., 2024).

      For the linear mixed model (LMM) analysis, we used V1 as a reference just for convenience, but this has no influence on the results. We now added a direct comparison using each area as reference in the LMMs. Several Supplementary Tables (S1-3) now show these results with coefficient estimates and stars showing statistical significance and are mentioned in the legend of Figure 1 and the main text.

      (2) The wide-field experiments in Figure 2 could be expanded to support the feedforward modeling assumptions. Currently, the spatial and temporal correlations presented leave open the possibility that these spontaneous events are traveling waves propagating from V1 to RL to S1 (or vice versa). This scenario would suggest a different connectivity scheme for the model. Clarifying this point with additional data analysis, specifically including temporal correlations involving RL, could provide stronger support for the model's assumptions.

      We agree with the reviewer that the correlation analyses shown in Figure 2 do not differentiate between two possibilities: activity that travels smoothly from one cortical area to another, thereby correlating correlations between these areas, versus activity that is spatially confined to individual areas but occurs near-synchronously across those areas. To address this point, we have revised Figure 2 in two ways.

      First, we added examples of spontaneous activity showing near-synchronous but spatially distinct activation of sub-areas in V1, RL and S1 (new Figure 2D). These examples show that localized activity can remain confined to individual sensory cortical areas and RL, while occurring at similar times across areas. Thus, the observed correlations are not simply due to single large events spreading continuously across the entire imaged field.

      Second, we added a lagged cross-correlation analysis between V1 and S1 activity (new Figure 2G). This analysis shows that the correlation between V1 and S1 peaks close to zero lag and decays for both positive and negative lags. This argues against a stereotyped travelling-wave-like propagation from V1 to S1 or from S1 to V1 with a fixed delay. The cross-correlation curves show a mild asymmetry, with somewhat higher correlations when S1 precedes V1. However, because the dominant peak is centered near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across sensory areas, rather than fixed directional propagation.

      Together, these two analyses support the modeling abstraction that V1 and S1 provide temporally correlated, spatially structured inputs to RL. We have added the new activity examples and the lagged cross-correlation analysis to Figure 2 and revised the Results accordingly. Although these analyses do not exclude all forms of propagating activity, they argue against the specific concern that the correlations are dominated by stereotyped traveling waves passing sequentially through V1, RL, and S1.

      (3) The functional correlation map in Figure 2D appears contradictory to the authors' modeling assumption that inputs are correlated spatially in V1 and S1. While V1 seed points align topographically with RL, this organization breaks down when extended into S1. In contrast, and in support of the modeling assumption, Figure 2E shows clearer topography across all three regions. A discussion of this discrepancy would be helpful, as it's a key conclusion of the figure. Additionally, it is unclear when this data was collected during development. Clarifying the developmental stage and analyzing how this map changes over time could strengthen the results.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This representation made it difficult to directly compare the V1- and S1-seeded maps and may have given the impression that topographic organization was preserved in one direction but not the other.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlations of these RL seeds with activity across the imaged cortical field. This allows us to ask directly whether different RL locations are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum-channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the thresholded representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend how the RGB maps are computed and how thresholded pixels are represented.

      The reviewer also asked about the developmental stage and progression of this phenomenon. The example shown in Figure 2 was recorded at PN9, and we now state this explicitly. In addition, we added examples from PN9–PN13 in Supplementary Figure S1, showing that similar functional correlation-map structure is present across the developmental period analyzed here. This is consistent with previous work showing that retinotopy-like patterns in higher visual areas can be recovered from functional-connectivity analysis of spontaneous activity before eye opening (Murakami et al., 2022), and with recent work showing that retinotopy-like and somatotopy-like patterns of ongoing activity, together with their rough topographic correspondence in RL, are already present before eye opening at PN10–11 (Matsumoto, Murakami & Ohki, 2025).

      (4) The modeling of spontaneous events with fixed amplitude and duration seems inconsistent with the experimental data in Figure 1, which shows variability in these parameters. This is particularly confusing in Figure 4, where S1 maturation is modeled as a stronger topographical alignment with RL, but the experimental data defines maturation based on amplitude, duration, and event rates. Justifying these modeling choices or adapting the model to reflect experimental variability would create a better connection between the theory and data.

      We agree with the reviewer that the original presentation did not sufficiently distinguish between the experimentally measured maturation of spontaneous activity and the way S1 maturation was implemented in the model. In the experiments (Figure 1), earlier maturation of S1 was reflected by lower event amplitudes, shorter durations, and higher event rates. In contrast, the original model explored the effect of a stronger or more spatially refined S1-to-RL projection (Figure 4). This modeling choice was motivated by pilot anatomical data suggesting that projections from S1 to RL become more elaborate earlier than projections from V1 to RL at comparable developmental ages. We include examples of these pilot data (Author response image 1), but we have not included them in the manuscript because the dataset is preliminary and does not yet allow for a sufficiently complete quantitative analysis.

      Author response image 1.

      Projections from V1 and S1 to RL at different developmental ages. Pilot anatomical data suggest that the S1 projection to RL becomes more elaborate and mature earlier than the V1 projection.

      To address the reviewer’s concern more directly, we have now extended the model to incorporate differences in the spontaneous activity patterns of V1 and S1, including the lower amplitude and higher frequency of S1 events. We then examined how these activity differences interact with different levels of initial connectivity bias between the primary sensory cortices and RL (Supplementary Figure S2). We also quantified the resulting topography, map alignment, and fraction of bimodal RL neurons as a function of the S1 bias and included these additional plots in Figure 4 (panels C-E).

      This analysis shows that incorporating the more mature S1-like activity patterns alone was not sufficient to generate the appropriate topographic and aligned maps. Rather, the model still required an initial connectivity bias, together with an appropriate level and structure of correlated activity. This is consistent with the results shown in Figure 3B,E,G–I and discussed in our response to Reviewer 2, point 3, where we show that the initial bias does not by itself determine the final map structure, but instead interacts with the level of V1–S1 correlation. We have added the new analysis to Supplementary Figure S2 and revised the text to clarify the interpretation. Rather than presenting the stronger S1 bias as a direct consequence of the more mature S1 activity dynamics revealed through the differences in amplitude, duration, and event rate, we now frame it as a model prediction: earlier S1 maturation may need to be accompanied by, or act through, a more advanced anatomical or functional S1-to-RL projection, whose refinement still depends on the temporal and spatial structure of spontaneous activity.

      The results suggest that differences in spontaneous activity dynamics and differences in projection maturity may act together during the emergence of topographically aligned multisensory maps, with neither component alone being sufficient to determine the final organization. Future experiments will be needed to establish whether such an S1-to-RL connectivity bias is present systematically, to quantify its developmental progression, and to disentangle the relative contributions of more mature spontaneous activity dynamics and more mature connectivity.

      (5) Several important details of the mathematical model are missing or unclear, partly due to typos. The Results section mentions the general framework of the input correlation matrix (e.g., "S1 and V1 neurons were driven by a combination of events, independent and shared in each V1 and S1" and "each independent event activated a randomly chosen, contiguous set of neurons"), but the specifics are not fully explained. Additionally, the caption of Figure 5 refers to a non-linear transfer function (a sigmoid), but these details are not provided in the Methods section, which instead suggests a linear model was used. A careful review of the main text and Methods section would help ensure that all the necessary details are included and that the story is both complete and accurate.

      We thank the reviewer for pointing out these missing details and inconsistencies. We have carefully revised the Results, figure captions, and Methods to make the model description more complete and internally consistent.

      First, we clarified how spontaneous input events were generated. Specifically, V1 and S1 activity was constructed from independent events in each area and shared events across the two areas. These event streams were generated using Poisson processes, with the rates chosen such that the total event rate was matched across simulations while varying the fraction of shared versus independent events. We also clarified that each event activated a spatially contiguous group of neurons, thereby implementing local spatial correlations within each primary sensory area, while shared events activated corresponding topographic locations in V1 and S1.

      Second, in the Methods we clarified the use of the nonlinear transfer function in the decoding analysis shown in Figure 5. The simulated RL activity was transformed with a sigmoid nonlinearity before performing the regression analysis, and we have now added the corresponding equation (15) to the Methods.

      Third, we clarified the distinction between the numerical decoding analysis and the analytical calculation of the optimal weight matrix. The decoding analysis uses the nonlinear transformation described above, whereas the analytical calculation uses a linearized version of the model to obtain a tractable closed-form solution. We now state this explicitly in the Methods to avoid the impression that two inconsistent models were used.

      Finally, we corrected several typographical errors and checked that the Results, Methods, and figure captions use consistent terminology for the input generation, correlation structure, and decoding analysis.

      (6) While Figure 5 supports the paper's conclusion that a mixture of unimodal and bimodal neurons in RL optimizes information encoding, the authors missed an opportunity to strengthen the connection between the model and experimental data. Specifically, they could apply this reconstruction method to the experimental data and examine how RL's ability to reconstruct V1/S1 activity changes across development. Their model predicts that this performance would improve over time, and if this trend is observed in the experimental data, it would provide strong validation that these feedforward connections are developing in line with the model's predictions.

      We agree with the reviewer that applying the reconstruction analysis directly to the experimental data would provide an important additional test of the model. However, the current experimental datasets are not well suited for this analysis. The two-photon recordings used to characterize spontaneous activity in V1, S1, and RL were acquired sequentially rather than simultaneously, and therefore cannot be used to reconstruct V1/S1 activity from RL activity. In principle, a related analysis could be attempted using the wide-field recordings, which are simultaneous across cortical areas. However, these data have lower spatial resolution, include movement-related variability, and do not provide cellular-resolution measurements of RL activity. We explored this possibility, but the resulting reconstructions were not sufficiently reliable or interpretable to include in the manuscript.

      We now state this explicitly as a limitation in the Discussion and identify simultaneous multiarea recordings at cellular resolution as an important future test of the model. Such experiments would make it possible to determine whether the ability of RL activity to reconstruct V1/S1 activity improves across development, as predicted by the model.

      Reviewer #2 (Public review):

      The authors aim to investigate the role of spontaneous activity in shaping the development of multisensory integration in the brain, specifically focusing on the connections between primary visual and somatosensory sensory areas (V1 and S1) and a higher-order cortical area rostrolateral to V1 (RL). They seek to understand how spontaneous activity guides the formation of aligned topographic maps and the emergence of bimodal neurons in RL.

      First, the authors found that spontaneous activity in all three areas sparsifies over time, but S1 exhibits more mature patterns earlier than V1 and RL. They claimed that correlated activity among neighboring regions of these areas during development carries topographic information. These data were used to implement a computational model that employed Hebbian rules of synaptic plasticity. The model indicated that correlated spontaneous activity can generate topographic connectivity between S1/V1 and RL and bimodal neurons in RL. The model suggested that the more mature spontaneous activity in S1 can guide map alignment between V1 and RL. In addition, the model also suggested that a mixture of bimodal and unimodal neurons in RL is optimal for decoding information from V1 and S1.

      While the data presented in the manuscript is promising and provides preliminary insights into the role of spontaneous activity in multisensory integration, it would be beneficial to strengthen the experimental foundation regarding the correlation between V1, S1, and RL. Incorporating more rigorous spatio-temporal analyses of spontaneous activity could enhance the robustness of these findings.

      Here are some important concerns:

      (1) The analysis of how spatial topography influences activity correlations in Figure 2 has several issues.

      (1a) While squares in V1 and S1 covered a small area of these sensory areas, the correlated territories in RL covered the entire area of RL. The topographic map in V1 continues caudally, so where is the rest of the map in RL? Something similar applies to the relationship between S1 and RL.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This made it difficult to directly compare the V1- and S1-seeded maps and could give the impression that the correlation structure extended differently across RL depending on the chosen seed area.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlation of each RL seed with activity across the imaged cortical field. This allows us to ask more directly whether different locations in RL are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the maximum-channel representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend and Methods how the RGB maps are computed, how the maximum-channel maps are generated, and how thresholded pixels are represented. In addition, we added Supplementary Figure S1 to show further functional-correlation-map examples across PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel.

      (1b) It is essential to know how areas were drawn. High precision is required.

      Consistent delineation of cortical areas is absolutely essential for interpreting the functional correlation maps. We have therefore expanded the Methods to describe how cortical areas were delineated from the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical-area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. This approach follows the procedure we previously validated for developmental wide-field recordings (Leighton et al., 2021).

      To make this transparent, we added Supplementary Figure S3, which illustrates how the reference-map-based outlines were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. We also clarified this in the Methods.

      (1c) It is not clear if correlated activity means different events in sync or large events that cover 2 or all 3 cortical areas of interest. The figure points to the second option, which contradicts the size of events at these stages, mainly in the oldest mice analyzed here.

      The reviewer asks whether the correlations reflect spatially confined events occurring near-synchronously in different cortical areas, or instead large events spanning V1, RL, and S1. To clarify this point, we revised Figure 2 to show representative activity traces and individual frames from the wide-field recordings (new Figure 2B–D). These examples show that activity can be localized to distinct subregions within V1, RL, and S1 while occurring at similar times across areas. Thus, the observed correlations are not well explained by single large events spreading continuously across the entire imaged field.

      We have revised the Results and Figure 2 to make this clearer. In addition, the lagged cross-correlation analysis in Figure 2G shows that V1–S1 correlations peak near zero lag and decay for both positive and negative lags, arguing against a stereotyped travelling-wave-like propagation between the two primary sensory cortices as the dominant explanation for the observed correlations.

      (1d) It is fundamental to know in detail and provide examples of how the detection of events was performed. For instance, could the dispersion of light from an event in V1 close to RL cause the detection of activity in RL?

      The reviewer asks how events were detected in the wide-field recordings and whether light dispersion could lead to false-positive correlations between neighboring areas. We have clarified this point in the Methods. For the functional correlation analyses shown in Figure 2, we did not perform event detection. Instead, the correlation maps were computed from the continuous fluorescence time courses by calculating Pearson correlations between seed region activity and the activity of every pixel in the field of view. Thus, the functional correlation maps do not depend on detecting or assigning individual events.

      To address the concern about whether correlations could reflect light spread from large events rather than genuine co-activity across areas, we revised Figure 2 to include representative activity traces and individual frames from the wide-field recordings. These examples show that activity can be spatially confined to distinct subregions in V1, RL, and S1 while occurring at similar times across areas. This argues against the interpretation that the correlations are simply caused by a single event spreading continuously across the imaged field or by light dispersion from one area into another. We have also described the area delineation procedure in more detail in the Methods and added Supplementary Figure S3 to illustrate how activity patterns and functional correlation maps were used to assign outlines of distinct cortical areas.

      Although wide-field imaging cannot completely exclude minor contributions from light scattering near area borders, the spatially localized activation patterns and the topographically ordered correlation maps support the interpretation that the correlations reflect genuine nearsynchronous co-activity across V1, RL, and S1.

      (2) For the correlations among V1, S1, and RL, it is crucial to have a consistent method to delineate the borders of cortical areas. The authors mention in one sentence that areas were drawn according to a reference map. More details are needed to convince the reader that the borders are accurate, especially because their shape and position change with age.

      As described in our response to point 1b, we have expanded the Methods to clarify how cortical-area borders were delineated in the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. We also added Supplementary Figure S3, which illustrates how the outlines based on reference maps were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. This makes the delineation procedure more transparent across animals and developmental ages.

      (3) The results from the model seem to be based on the initial bias in connectivity between neighboring cells from the different areas. Then, it seems straightforward that implementing correlated activity with Hebbian and synaptic depression rules will force the strengthening of connections between spatially close cells. Despite this apparent predisposition of the model towards a defined outcome, the flaws in the experimental data used prevent a rigorous interpretation of the computational model.

      We understand the reviewer’s concern that the initial topographic bias could predispose the model toward the emergence of topographic maps. However, the model results show that this bias is not by itself sufficient to determine the final organization (Figure 3B,E,G– I). When V1–S1 correlations are weak, many RL neurons decouple from the primary sensory inputs, resulting in poor topography and few bimodal neurons (Figure 3E,G–I). Conversely, when V1–S1 correlations are very strong, the two input maps become highly aligned, but topography is degraded because many RL neurons receive similar visual and somatosensory inputs at the same topographic location, thereby overriding the initial topographic bias (Figure 3E,G,H). Thus, the initial bias does not simply determine the final map structure. Rather, appropriate topography, map alignment, and the emergence of a mixture of unimodal and bimodal neurons require an intermediate level of correlated activity.

      We have revised the manuscript to make this interpretation clearer. We also strengthened the experimental basis for the activity structure used in the model by revising Figure 2 and the corresponding Results and Methods. The revised analyses now show near-synchronous but spatially distinct activation of V1, RL, and S1, a lagged cross-correlation analysis arguing against stereotyped travelling-wave-like propagation between V1 and S1, and functional correlation maps computed from common RL seed locations. Together, these additions clarify the spatial and temporal structure of the spontaneous activity used to motivate the model.

      Finally, as described in our response to Reviewer 1, point 4, we have extended the model to test the role of the initial bias more directly in combination with experimentally measured differences in V1 and S1 activity patterns. In this analysis, we incorporated these activity differences and examined how they interact with different levels of initial connectivity bias (Supplementary Figure S2). These simulations show that more mature S1-like activity patterns alone are not sufficient to generate the appropriate topographic and aligned maps, and that an initial connectivity bias is required. At the same time, consistent with Figure 3, this bias does not by itself determine the final organization; the outcome also depends on the temporal correlation structure of V1 and S1 activity.

      We agree that the initial topographic bias remains an important modeling assumption, consistent with the idea that coarse activity-independent mechanisms provide an initial scaffold for later activity-dependent refinement. We now present the model accordingly: not as showing that correlated activity alone creates topography from an entirely unstructured circuit, but as showing how structured spontaneous activity can refine an initially coarse topographic scaffold to produce aligned multisensory maps and a mixture of unimodal and bimodal RL neurons.

      (4) In the Introduction, the authors nicely and briefly explain the role of primary and higher order sensory cortices in information processing. They also explain how spontaneous activity during development helps to build these circuits by refining connections or establishing hierarchies. They continue explaining the relevance of aligning different topographic maps to allow multisensory integration. Then they provide some examples of sites of multisensory integration. This provides a general context for the data presented in the Results section; however, and importantly, there is no specific introduction of why they are interested in RL and its interaction with V1 and S1. The authors should introduce the RL area and explain why it is an interesting site for multisensory processing.

      We thank the reviewer for pointing this out. We have revised the Introduction to make the rationale for focusing on RL more explicit. Specifically, we now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices and contains overlapping visual and tactile representations. We also clarify that RL is a particularly relevant area for studying multisensory map alignment because corresponding locations in visual and whisker space can converge onto the same RL neurons, including bimodal neurons. Finally, we expanded the Introduction to explain that RL has been implicated in visually guided tactile behaviors and cross-modal generalization, making it an appropriate model system for studying how aligned multisensory representations emerge during development.

      (5) The results shown in Figure 1 corroborate published data from Golshani et al, Rochefort et al, Murakami et al. While the reproduction of data is more than welcome, the authors should specify which part of the data is completely new and acknowledge clearly the rest as corroboration of previous data. The sentence "As described in previous experiments ..." partially acknowledges this fact but is not clear enough. In addition, the transition between this part of the manuscript and the next data is not smooth. Data seems to be used to feed the model so perhaps the organization of the manuscript leaves room for improvement.

      We thank the reviewer for pointing this out. We have therefore revised the Results to clarify that the developmental sparsification of spontaneous activity in V1 is consistent with previous work, including Portera-Cailliau, Konnerth, Hanganu-Opatz, Crair and Ohki labs as well as our own (Siegel et al. 2021) and that similar developmental trends in S1 and RL corroborate and extend these observations across the sensory and higher-order cortical areas analyzed here.

      We also clarified what is new in the present analysis. Specifically, our contribution is not simply to reproduce previously described developmental sparsification, but to compare V1, S1, and RL within the same experimental and statistical framework, revealing that S1 exhibits more mature activity features earlier than V1 and RL. We also revised the transition to the next section to make clearer how these measurements motivate the subsequent analysis of temporally and spatially correlated spontaneous activity between V1, S1, and RL.

      Reviewer #3 (Public review):

      Summary:

      The study by Dwulet et al. explores how the development of spontaneous neural activity in primary sensory cortices influences the co-alignment of multiple sensory modalities in higher order brain areas (HOAs). To address this question, they focus on connectivity between the primary visual (V1) and somatosensory (S1) cortices and an associative cortical area (RL) in mice. The authors combine experimental (wide-field and two-photon calcium imaging) and computational approaches to show that spontaneous activity matures at a different pace across these brain regions. Their data indicate that S1 develops more rapidly than V1, which is possibly beneficial for RL's integration of visual and somatosensory inputs through correlated spontaneous activity. Using a computational model, they demonstrate that a moderate correlation between V1 and S1 activity can optimally guide the formation of bimodal neurons in RL, which are crucial for maximizing the decodability of multisensory stimuli. This finding highlights the role of correlated spontaneous activity in primary sensory cortices in establishing co-aligned topographic multimodal sensory representations in downstream circuits.

      Strengths:

      The manuscript is well written and it provides strong enough evidence to support the main claim of the authors. The insights on the role of correlated activity on instructing co-aligned multisensory maps in HOAs are not trivial and are an important advancement for the field.

      Weaknesses:

      In the opinion of this reviewer, the study has no major weaknesses. A drawback of the work is that none of the predictions of the computational modeling have been corroborated through mechanistic experimental manipulations of early brain activity.

      We thank the reviewer for their positive assessment of the manuscript and for highlighting the importance of the model predictions. We agree that a direct mechanistic perturbation of early spontaneous activity would provide an important future test of the model. Such experiments could, for example, perturb the temporal correlation structure between V1 and S1 during the relevant developmental window and then test whether this affects the alignment of V1/S1 maps in RL and the emergence of bimodal RL neurons.

      In the present study, we focused on identifying candidate features of spontaneous activity that could instruct multisensory map alignment and testing their sufficiency in a computational model. We now explicitly acknowledge in the Discussion that causal perturbations of early spontaneous activity will be needed to validate the model predictions experimentally. We believe this provides an important direction for future work while preserving the main conclusion of the current study: that structured, moderately correlated spontaneous activity provides a plausible developmental mechanism for refining aligned multisensory representations in higher-order cortex.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional comments/suggestions for the figures:

      (1) In Figure 1D-G, some of the dots lie almost directly on top of each other, essentially "hiding" certain data points. Using different shapes for each of the three regions might help alleviate this issue and make the data more visually distinct.

      We thank the reviewer for this suggestion. We have revised Figure 1D-G so that the three cortical regions are shown with different marker shapes. This should make overlapping data points easier to distinguish and clarify that each point corresponds to the average value for one animal and cortical region at the indicated postnatal age.

      (2) In Figure 2D-E, RGB color values are used to represent the highest correlation coefficient across the three seeded areas. It would be more informative if these also depicted the magnitude of the correlations, possibly through a color gradient. Additionally, the black regions in these panels are not currently defined and should be clarified.

      We have revised the functional correlation-map analysis and its presentation in Figure 2, as suggested by the reviewer. In the revised figure, the main correlation-map panels now use three seed locations in RL and show the resulting correlations across the imaged cortical field. We present the maps in two complementary ways. First, the raw RGB correlation map shows the correlation values for all three seed locations, with the intensity of each color channel reflecting the magnitude of the corresponding Pearson correlation coefficient (new Fig. 2E). Second, the maximum-channel representation assigns each pixel to the seed location with the strongest correlation, while still preserving correlation strength through pixel intensity (new Fig. 2F).

      We have also added color scales to relate pixel intensity to correlation magnitude and clarified that black pixels in the maximum-channel representation correspond to pixels below the correlation threshold used for visualization. The figure legend and Methods now describe how the RGB maps and maximum-channel maps were computed. Finally, we added Supplementary Figure S1 with additional examples from PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel. This illustrates that they are quite similar across the ages investigated here.

      (3) I found Figure 3F a bit difficult to interpret without referring to the Methods section for the definitions of Topography and Alignment. Since these definitions are relatively short and essential for understanding all the modeling figures, I suggest moving them into the main text where they are first introduced.

      The definitions of Topography and Alignment have been added to the text where they are introduced.

      (4) In Figures 3-5, it is unclear what causes the variability in the model’s responses, as there are two potential sources of randomness: the initial random connectivity matrix and the correlated inputs driving the system. Are either of these fixed? For example, is the distribution of dots along the y-axis in Figure 3G-H, which corresponds to zero correlation between V1 and S1, driven by variability in the initial connectivity matrix, the random timing of input events, or a combination of both? If it’s a combination, it would be interesting to tease this effect apart by fixing one form of randomness and recreating these plots.

      In the original simulations in Figures 3–5, neither source of randomness was fixed across runs: each point corresponds to an independent developmental realization with a newly sampled initial connectivity matrix and a newly sampled sequence of spontaneous input events. The initial connectivity was random but weakly biased toward matched topographic location, while spontaneous activity consisted of stochastic independent and shared events activating randomly chosen contiguous groups of neurons (as explained in the main text and Methods). Thus, for example, the spread of points at zero V1–S1 correlation in Figures 3G–H reflects a combination of variability in the initial connectivity and variability in the independent V1 and S1 event histories. At zero correlation, no shared V1–S1 events are present, so this spread does not reflect variability in correlated shared events, but rather run-to-run differences in the two independently refined maps.

      We have clarified this point in the text and figure legend. We agree that fixing one source of randomness while varying the other would be an interesting additional analysis to decompose the relative contribution of initial wiring versus input history. However, the goal of the present simulations was to characterize the ensemble of possible developmental outcomes when both initial connectivity and spontaneous activity vary, as expected biologically.

      This interpretation is also consistent with the earlier two-layer model from developmental refinements from retina/thalamus to V1 (Wosniack et al., eLife 2021) on which our model builds, where final receptive fields emerge from the interaction between weak biased initial connectivity and stochastic structured spontaneous activity. In the current three-layer extension, the same principle applies to two converging projections, from V1 to RL and from S1 to RL. The initial topographic bias constrains the possible map structure, while the spatiotemporal statistics of V1 and S1 activity determine whether the two maps remain separate, align, or collapse into overly bimodal representations.

      (5) The specific parameter values used to create the panels in the modeling figures (Figures 3 and 4) should be made clearer, at least in the figure captions. For example, in Figure 3E, the exact values for the “weak,” “medium,” and “strong” correlations should be provided. Additionally, Figure 4 does not mention the strength of the correlated input considered, which should be specified as well.

      The values for the weak, medium and strong correlations have been added to the figure caption of Figure 3. The input correlation for Figure 4 is also now specified in the figure caption.

      (6) There is an odd vertical line in Figure 3I that doesn’t appear to be discussed or defined. Its purpose should be clarified, or the line should be removed if it is unintentional.

      This line has been removed.

      (7) There is a typo in the caption for Figure 3. Panel 'K' should be panel 'J'.

      This typo has been corrected.

      (8) In the text, the authors write "With these connectivity refinements, the generated activity in RL became sparser in terms of amplitude and participation rate (Figure 3J)." While this appears to be the case for this single example, it is difficult to confirm without zooming in on the panel. These quantities should be computed across multiple instances, and a summary plot should be provided to support this statement.

      The experimentally measured developmental sparsification of RL activity is quantified (independent of the model) in Figure 1D–F.

      We see how the original wording placed too much weight on the illustrative example in Figure 3J. We have revised the text to clarify that Figure 3J shows a representative simulation illustrating how RL activity changes as V1/S1-to-RL connectivity refines, rather than a separate population-level quantification across model instances.

      At the same time, this example is not meant to introduce a new, unsupported mechanism. The model used here is an extension of our previous two-layer model of developmental refinements between retina/thalamus and V1, in which spontaneous activity refined feedforward receptive fields from thalamus to V1. In that study, we specifically quantified how receptive field refinement led to sparsification of cortical activity in V1 over development, including reduced event amplitude, reduced event size/participation, and reduced pairwise correlations (Wosniack et al., 2021). Thus, the example shown in Figure 3J is consistent with a mechanism that has already been systematically characterized in the simpler two-layer setting.

      In the present manuscript, the central modeling results concern the emergence of topography, alignment, and the balance of unimodal and bimodal RL neurons. We therefore have softened the corresponding statement and explicitly refer to Figure 3J as an illustrative example.

      (9) Figure 5C is a bit difficult to interpret. The corresponding text states, "However, when activity across V1 and S1 is moderately correlated, having some unimodal RL neurons can achieve a higher total maximum fraction of variance for both V1 and S1 compared to the purely bimodal case (Figure 5C)", from which I infer that these dots represent networks resulting from "moderate correlations." However, the exact range of correlations considered should be mentioned in the text or figure caption. Additionally, I find it unusual that some networks with close to 0% bimodal cells perform quite well in reconstructing both S1 and V1. Many data points overlap, but I notice quite a few pale dots in the upper right of the plot. I believe this should be addressed in the main text.

      We thank the reviewer for this helpful comment. We have added the correlation values used for the simulations in Figure 5C to the figure caption and clarified the interpretation in the Results. The high reconstruction performance for some networks with relatively few bimodal cells arises because, when V1 and S1 activity are not perfectly correlated, unimodal RL neurons can provide unambiguous information about activity in one sensory area. In contrast, a purely bimodal population can make it more difficult to distinguish whether one or both primary sensory cortices were active. Thus, for moderately correlated inputs, a mixture of unimodal and bimodal RL neurons can reconstruct both sensory areas better than a population composed entirely of bimodal neurons. We have revised the main text to make this point explicit.

      (10) The network schematics in Figures 3A and 5A could be improved to better illustrate the network setup using a similar approach as the one used by this research group in Wosniack et al. (2021). Adding arrowheads to the lines from V1/S1 to RL would clarify that these are purely feedforward inputs. It would also be helpful to depict that V1 and S1 are driven by correlated events that are spatially structured.

      We thank the reviewer for this helpful suggestion. We have revised the schematics in Figures 3A and 5A to make the feedforward nature of the model clearer by adding arrowheads to the projections from V1 and S1 to RL. We have also clarified the depiction and description of the input activity. Specifically, Figure 3C shows the spontaneous events driving V1 and S1 in the model, including shared events that are both temporally correlated and spatially structured across corresponding topographic locations in the two primary sensory areas. These shared events activate matched contiguous groups of neurons in V1 and S1, while independent events activate randomly chosen contiguous groups within each area. We have clarified this point in the Results and Methods.

      General comments regarding the text (including typos):

      (1) In Statistical analysis, "In wide-field calcium imaging (we re-analyzed data from [46] (Figure 1))..." should be referencing Figure 2.

      Typo fixed.

      (2) Right before Table 1, the authors mention that they ran the simulations for 500,000 milliseconds, which is 500 seconds. This doesn't seem long enough for the weights to approach their steady-state values given the inter-event interval. Since the example simulations in Figure 3 are 1,000 seconds long, I'm guessing this is a typo.

      Typo fixed. Indeed the simulations in Figure 3 were 1,000 ms (1 s) long.

      (3) The specific time step used for the simulations should be specified. Currently, the text only mentions "sufficiently small time steps".

      We have now specified the simulation time step in the Methods.

      (4) In the Rate-based network model section, you write "These biased weights decay with a Gaussian profile with increasing distance (Figure 3)), with amplitude a and spread s", but Table 1 denotes these parameters differently.

      We have corrected the notation so that the parameter names are consistent between the Methods and Table 1.

      (5) Currently, all differential equations are written as 1/tau*df/dt. Based on the units of your time constants (seconds), I believe these equations should be tau*df/dt.

      We have corrected the differential-equation notation.

      (6) Equations 5-6 and 8 should be differential equations.

      We have corrected these equations so that they are written as differential equations. These mistakes happened because we changed formats between from Word to Latex.

      (7) The expectation in Equation 8 is not clearly defined and I would think here that the W_ij's should be within expectations. In the next paragraph, the authors specify that they are interested in a specific case of W_ij's, but this condition has not been introduced yet.

      We thank the reviewer for pointing out this ambiguity. We have revised the text around Equation 8 to define the expectation more clearly and to introduce the specific steady-state connectivity configuration before it is used. Because the expectation is taken over the input activity statistics at steady state, the weights are fixed quantities in this calculation. Including W_ij inside the expectation would therefore not change the result, but we have revised the notation and explanatory text to make this clearer.

      (8) The expectation in Equation 8 is not clearly defined, and I believe that the W_ij’s should be included within the expectations (in the following paragraph, the authors mention that they are interested in a specific case of W_ij’s, but this condition has not yet been introduced).

      This comment is the same as the one above. Please see the point above for the reply.

      (9) At the start of "Optimal weight matrix for correlated input populations", you write that the vector X is M x 1. If that is the case X'X would be a 1x1 matrix. I'm not sure if you meant to write X as 1 x M or to examine XX'.

      We thank the reviewer for pointing out this dimensional inconsistency. We have corrected the notation in the Methods. The concatenated input vector X=[v; s] has size M x 1, so the relevant input covariance matrix is X X^T not X^T X. This covariance matrix has size M x M, as required for the eigenvector analysis. We revised the corresponding equations and explanatory text accordingly.

      (10) Equation 11 has an s_i on the right-hand side that should be a \mu_s.

      Typo fixed.

      Reviewer #2 (Recommendations for the authors):

      Some sentences may require more scientific rigor. For instance: "We found that activity between the visual and the somatosensory cortex is often, but not always, temporally synchronized.

      We have revised the Results to state the quantitative observations more explicitly. Specifically for this example, we now report that the average activity in V1 and S1 across PN9PN12 animals showed a range of Pearson correlation coefficients with a mean of approximately 0.5. We also describe the examples in Figure 2B-D as near-synchronous but spatially distinct activation of subregions in V1, RL, and S1, and we use the lagged cross-correlation analysis in Figure 2G to support the conclusion that V1-S1 correlations peak near zero lag rather than reflecting stereotyped propagation with a fixed delay.

      Reviewer #3 (Recommendations for the authors):

      Minor suggestions on how to improve some specific aspects of the manuscript.

      Introduction:

      (1) What do the authors mean when they write "Higher-order areas (HOAs) situated between primary sensory areas"? This sentence might need some editing.

      We have revised the sentence to clarify that we are referring to higher-order cortical areas that receive and combine inputs from multiple primary sensory areas. We now also state explicitly that some of these areas, including RL, are anatomically positioned between the primary sensory cortices whose inputs they integrate.

      (2) In later portions of the manuscript, it becomes clear what the authors mean when they write “whereby sensory neurons converge onto higher-order cortex while preserving space”, but I think that it would be beneficial if this statement would be better explained also in the introduction.

      This has been clarified in the introduction. Specifically, we now clarify that topographic convergence means that neurons representing corresponding regions of sensory space in different primary sensory areas can project to overlapping or nearby locations in higher-order cortex. In the case of RL, this means that visual and tactile representations with corresponding spatial organization can converge onto RL neurons, including bimodal neurons.

      (3) Could the authors provide some more information about RL and the rationale as to why it was chosen as the HOA that they investigated in the study?

      We have expanded the Introduction to make the rationale for focusing on RL more explicit. We now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices. We also explain that RL contains overlapping visual and tactile representations, including bimodal neurons, and that corresponding locations in visual and whisker space can converge in RL. In addition, we now note that RL has been implicated in visually guided tactile behavior and cross-modal generalization. These anatomical and functional properties make RL a particularly suitable model system for studying how aligned multisensory representations emerge.

      Results:

      (1) "RL was found to slightly lag behind V1 and S1". On what evidence is this statement based upon? As far as I can understand, there are no significant differences between V1 and RL besides amplitudes being higher in RL, which I don't think can be univocally interpreted as a sign that RL lags behind V1 in the developmental profile.

      The evidence for a delayed RL maturation relative to V1 and S1 is limited and comes from the pattern of coefficient estimates in the linear mixed models, now shown in Supplementary Tables S1-S3, rather than from a robust difference across all measured activity features. We have therefore revised the Results to state more conservatively that RL and V1 develop more similarly during the second postnatal week, while S1 shows more mature activity features earlier in development. The full linear mixed-model comparisons using V1, S1, and RL as reference areas are provided in Supplementary Tables S1-S3.

      (2) Figure 1H is very hard to read.

      (a) The slopes and the intercepts have values that differ by orders of magnitude, so the slopes get squeezed and become invisible. Further, the different parts of the plots (e.g. the one of amplitude and duration) are almost overlapping, which is a bit confusing. Slopes and intercepts should also have different units of measure (see Equation 3), so I wonder how they can lie on the same axis. Can the authors try to plot the data in a manner that is easier to visually inspect?

      (b) Including the "reference" (V1) intercept in H is also a bit misleading, as one might intuitively interpret it as a difference between V1 and other brain areas. Perhaps the overall differences between brain areas (regardless of age) might be best represented in a plot without age on the x-axis (only brain area). Alternatively, one might point them out directly on the plots in DG.

      (c) In D-G, what do the individual dots represent? The legend states N=10 animals, but I only see ~6 dots per plot.

      We thank the reviewer for these helpful points. We have revised the caption of Figure 1H and added Supplementary Tables S1-S3, which provide the full linear mixed-model estimates for each choice of reference area. These tables report the intercepts, slopes, interaction terms, confidence intervals, and significance levels in a format that avoids placing quantities with different units and scales on the same visual axis.

      For the caption of Figure 1H: The V1 value corresponds to the model intercept at PN8, whereas the age coefficient corresponds to the slope for V1. The S1, RL, Age: S1, and Age: RL terms represent differences relative to this reference model. To avoid the impression that the V1 intercept represents a difference between areas, we now explicitly state that the coefficients in Figure 1H are interpreted relative to V1 at PN8, and that the complete comparisons using S1 and RL as reference areas are provided in Supplementary Tables S2 and S3.

      Finally, we clarified that the individual points in Figure 1D–G represent animal-level averages for each cortical area at the indicated age. The value N = 9 refers to the total number of animals included across the dataset, not to the number of animals at each postnatal age. Because recordings were distributed across ages and some points overlap visually, fewer points are visible in individual panels than the total N.

      (3) Figure 2B-C: at which lag does this correlation peak? Is it at 0ms? Or does one brain area precede/follow the other one?

      We thank the reviewer for this comment. We have revised Figure 2 to include a lagged V1–S1 cross-correlation analysis. The V1–S1 correlation peaks close to zero lag and decreases for both positive and negative lags, indicating that the dominant temporal relationship is near-synchronous rather than consistent with fixed-delay propagation from one primary sensory cortex to the other. The curves show a mild asymmetry, with somewhat stronger correlations when S1 precedes V1, but because the dominant peak is near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across areas rather than stereotyped travelling-wave propagation. We have added this interpretation to the Results and clarified the temporal-lag convention in the Figure 2 legend.

      (4) Figure 2D-E: in the methods section the authors report that "The actual color of each pixel represents the highest coefficient of correlation value across the three channels." I think that this important information should be included in the main text or the legend of the figure.

      We have changed Fig. 2 now to clarify the quantification of the functional correlation maps and also added the information requested by the reviewer to the figure legend.

      (5) Figure 3D: I think that it would be beneficial if the authors would highlight directly in the figure that those connectivity matrices are between V1/S1 and RL.

      This information has been added to the figure.

      (6) Figure 3I: does the vertical line correspond to the "critical amount of temporal correlation" (eq. 2)? If so, could the authors provide this information in the figure or the figure legend?

      This line was unintentional and has been removed.

      (7) It would be nice if the data that was generated for this study (and the data that has already been published and was used to generate Figure 2) would be made publicly available on an open-access repository.

      We agree that open data sharing is important. We have made the code used for the model and figure generation available in the repository listed in the Data and Code Availability section. At present, we are not able to deposit the complete raw imaging datasets in an open repository because the wide-field and two-photon imaging files are very large, amounting to multiple terabytes, and we do not currently have a sustainable hosting solution for these raw data. We will share data upon request, and we will deposit the raw imaging datasets in an appropriate open repository if a feasible long-term hosting solution becomes available.

      References

      M. Chini, T. Pfeffer, and I. Hanganu-Opatz. An increase of inhibition drives the developmental decorrelation of neural activity. eLife, 11:e78811, 2022.

      P. Golshani, J. T. Gonçalves, S. Khoshkhoo, R. Mostany, S. Smirnakis, and C. PorteraCailliau. Internally mediated developmental desynchronization of neocortical network activity. Journal of Neuroscience, 29(35):10890–10899, 2009.

      A. Gribizis, X. Ge, T. L. Daigle, J. B. Ackman, H. Zeng, D. Lee, and M. C. Crair. Visual cortex gains independence from peripheral drive before eye opening. Neuron, 104(4):711–723.e3, 2019.

      S. Lakhera, E. Herbert, and J. Gjorgjieva. Modeling the emergence of circuit organization and function during development. Cold Spring Harbor Perspectives in Biology, 17(2):a041511, 2025.

      A. H. Leighton, J. E. Cheyne, G. J. Houwen, P. P. Maldonado, F. De Winter, C. N. Levelt, and C. Lohmann. Somatostatin interneurons restrict cell recruitment to retinally driven spontaneous activity in the developing cortex. Cell Reports, 36(1):109316, 2021.

      H. Matsumoto, T. Murakami, and K. Ohki. Topographic correspondence between retinotopic and whisker somatosensory map in mouse higher visual area and its development. Frontiers in Neural Circuits, 19:1552130, 2025.

      T. Murakami, T. Matsui, M. Uemura, and K. Ohki. Modular strategy for development of the hierarchical visual network in mice. Nature, 608:578–585, 2022.

      N. L. Rochefort, O. Garaschuk, R.-I. Milos, M. Narushima, N. Marandi, B. Pichler, Y. Kovalchuk, and A. Konnerth. Sparsification of neuronal activity in the visual cortex at eyeopening. Proceedings of the National Academy of Sciences of the United States of America, 106(35):15049–15054, 2009.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds on earlier work showing that early-life odor exposure can trigger glial-mediated pruning of specific olfactory neuron terminals in Drosophila. Moving from indirect to direct functional imaging, the authors show that pruning during a narrow developmental window leads to long-lasting suppression of odor responses in one neuron type (Or42a) but not another (Or43b). The combination of calcium and voltage imaging with connectomic analysis is a strength, though the voltage imaging results are less straightforward to interpret and may not reflect synaptic output changes alone.

      Strengths:

      Biologically, one of the main strengths of this work is the direct comparison between two odor-responsive OSN types that differ in their long-term adaptation to early-life odor exposure. While Or42a OSNs undergo pruning and remain persistently suppressed into late adulthood, Or43b OSNs, which also respond to the same odor, show little lasting change. This contrast not only underscores the cell-type specificity of critical-period plasticity but also points to a potential role of inhibitory network architecture in determining susceptibility. The persistence of the Or42a suppression well beyond the developmental window provides compelling evidence that early glia-mediated pruning can imprint a stable, life-long functional state on selected sensory channels. By situating these functional outcomes within the context of detailed connectomic data, the study offers a framework for linking structural connectivity to long-term sensory coding stability or vulnerability.

      Weaknesses:

      The narrative begins with the absence of changes in PN dendrites and axons. While this establishes specificity, it is a relatively weak starting point compared to the novel OSN functional results.

      We agree that switching the order of Figures 1 and 2 recontextualizes the negative PN morphology findings to make their significance more clear, especially with the addition of PN odour-evoked activity data (see Figures 2A, B of the revised manuscript).

      Calcium imaging with GCaMP, though widely used, is an indirect measure of synaptic function, and reduced signals could reflect changes in non-synaptic calcium influx as well as release probability. The interpretation of the voltage imaging results is also unclear: if suppression were solely due to impaired synaptic release, one might expect action potential-evoked voltage signals to remain unchanged. The reported changes raise the possibility of deficits in action potential initiation or propagation, which would shift the mechanistic explanation.

      Although it is true that non-synaptic Ca<sup>2+</sup<> influx could contribute to odour-evoked signals in OSN axon terminals, it seems likely to be a relatively small contribution when compared to Ca<sup>2+</sup> influx via voltage-gated Ca<sup>2+</sup> channels at the active zone. Given the observation that synaptic markers are eliminated during this form of critical period plasticity and remain decreased even after OSNs regrow their terminals days later (consistent with our observed continued decrease in odour-evoked responses), the most parsimonious explanation is that we are seeing a reduction in synaptic Ca<sup>2+</sup> influx. We cannot dismiss the possibility that there is a decreased voltage signal arising from fewer action potentials being elicited by the odour stimulation. However, the reduction in voltage signal must arise at least in part from the observed reduction in Ca<sup>2+</sup> influx. We have therefore provided additional text to this effect in the results section.

      The difference between Or42a and Or43b OSNs is attributed to varying inhibitory input densities from connectome data, but this remains speculative without functional tests such as manipulating GABA receptor expression in OSNs. In Or43b, there is essentially no strong phenotype, making it premature to ascribe the absence of suppression solely to inhibitory connectivity.

      We have tempered our conclusions to posit additional mechanisms that could explain the more mild pruning that occurs for Or43b OSNs. While the pruning phenotype for Or43b OSNs is not as strong as Or42a, it is not absent. To further explore the contribution of inhibition as a candidate mechanism underlying differences in susceptibility of Or42a and Or43b to this form of critical period plasticity we compared the relative impact of knocking down expression of GABA-A receptor (called “rdl”) in Or42a and Or43b OSNs. Consistent with the degree of pruning being regulated inhibition, knocking down expression of rdl enhanced pruning for both Or42a and Or43b OSNs. However, because the magnitude of the enhancement was similar between both OSN types, we agree with the reviewer that inhibitory connectivity cannot be the sole mechanism that explains the difference and have therefore tempered our language appropriately.

      Finally, the study does not connect circuit-level changes to behavioral outcomes; assays of odor-guided attraction or discrimination could place the findings in an organismal context.

      We agree that behavioral assays will be a critical component for understanding the functional consequences of this form of critical period plasticity. However, the goal of this study was to extend our prior work to determine the longevity and selectivity of the critical period pruning. Behavioral assays testing the consequences of this form of early life plasticity will be a component of future studies.

      Some introduction material overlaps with the authors' 2024 paper, and the novelty of the present study could be signposted more clearly.

      We have included text to highlight the novelty of the present study.

      Reviewer #2 (Public review):

      Recent work from the authors identified the synaptic changes and glial reaction that occur during exposure of a Drosophila odorant receptor neuron population to continued exposure of a stimulating odorant. This work markedly advanced our understanding of cellular response to critical periods. This current Advance manuscript carries that work forward and examines the non-autonomous responses to constant odorant exposure. The authors discover that the changes to ORN populations are not accompanied by changes to either PN dendrite or PN axon volume, nor are they concurrent with changes in postsynaptic PN structures. These changes are, however, notable, accompanied by changes in Ca2+ and voltage responses in ORNs. Importantly, this set of responses is specific to the Or42a ORNs (that are highly sensitive to the odorant in question, ethyl butyrate) and not the Or43b ORNs (which respond to ethyl butyrate, but not as drastically). Finally, the authors include connectomics analyses showing that Or43b and Or42a ORNs differ in their synaptic input/output relationships.

      This is an excellent use of the Advance mechanism for the journal, as these are important follow-up findings for the parent story. The non-autonomous effects (or lack thereof) on PNs is an important part of the story, as is the functional response of Or42a ORNs and the differing response of similarly (but not identically) sensitive Or43b ORNs. The experiments are well-conceived, controlled, and conducted. Where the story falters a bit, though, is with the connectomics analysis. The authors show distinct differences between Or43b and Or42b ORN input-output relationships, and suggest that those differences may underlie the differences observed in their response to ethyl butyrate exposure during the critical period. This is certainly a possibility, but as it stands now, it is too disconnected to offer significant proof. There would have to be additional experiments to address this. Right now, the inclusion of the connectomics work feels like a distraction at best, and a complete non sequitur at worst. To be clear, the connectomics work is well done and I have no issues with its validity, but it is not helpful to the central thesis of the work. I would suggest the authors either remove it entirely or strongly rethink how it fits into the paper.

      We have tempered our stated interpretations of the connectivity analysis and include new experiments examining the impact of GABA signaling on pruning. We have therefore opted to retain the connectivity analysis as we feel that it has been better integrated into the overall narrative of the paper.

      Major Concerns:

      (1) The examination of PN axon terminals in the MB and LH is interesting, but it is only one possibility. Oftentimes, the volume of neurons remains constant with perturbation, while the synapse number is affected. Figure 1C and E would be greatly helped by examining synapse number (via Brp or Brp-Short) in the PN axons.

      We agree that the counting synapse number would provide greater resolution information about synapse function relative to axon volume and have added this analysis to what is now Figure 2.

      (2) The use of dlg1[4K] is a strong use of a new tool, but the result is surprising. The presynaptic ORN synapse number onto the PNs is notably changed, but that is not reflected in a postsynaptic PSD-95 change. That suggests a compensatory mechanism that the authors might explore. A good proportion of PN puncta should be postsynaptic to those ORNs, so why aren't they adjusted?

      We agree that this result suggests that a compensatory mechanism may be present. We have therefore added new text to point out this observation and potential explanation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The interpretation of the voltage imaging results would benefit from clarification. If these signals are reduced because of upstream action potential changes rather than synaptic release, this should be explicitly discussed and illustrated with representative raw traces for both OSN types. The proposed link between inhibitory connectivity and selective vulnerability could be tested more directly, for example, by manipulating GABA receptor function in OSNs.

      We have now tested the link between inhibitory connectivity and susceptibility to glial pruning by testing the effects of GABA receptor knockdown in either Or42a or Or43b OSNs (fully described above).

      Adding an intermediate post-exposure time point for Or42a responses could help resolve whether suppression is immediate or develops over time.

      The suppression of Or42a odour-evoked responses is present immediately after the 2 day exposure period and responses remain suppressed until 25 days post-eclosion, indicating that the suppression is immediate and sustained. We therefore respectfully disagree that another physiological time point will help resolve whether the suppression is immediate or develops over time.

      In terms of presentation, the introduction could be tightened to reduce overlap with the 2024 paper, figures should have clear axis labels and consistent terminology for neuron types and glomeruli, and a schematic summarising key inhibitory connections for Or42a vs. Or43b would aid clarity

      We have now streamlined the introduction, improved clarity on axis labels and checked for consistency of terminology.

      Minor Concerns:

      (1) The dlg1[4K] is made with a V5 epitope but the authors have it labeled mCD8::GFP in Figure 1F. This is likely a typo and should be corrected.

      This typo has now been corrected.

      (2) Can the responses be separated in Figures 2A, C, and E? It is difficult to see the differences in oil and EB exposure. This would make it much more straightforward to tell the difference if both traces were clearly visible.

      Overlaying the averaged response traces for in Figure 2C, E and G (now Figures 1C, E and G) enables the reader to make direct visual comparisons between the responses of OSNs from flies in each condition to both mineral oil and ethylbutyrate. Separating the individual traces would make it much more difficult to make these comparisons.

    1. Author response:

      We would like to thank all the reviewers and the editors for their considerate evaluation of our study.

      We are pleased that overall the reviewers were positive about the bulk of our study establishing a role of tissue macrophage programming/specialisation in regulating the macrophage lipidome, in the peritoneum, including the exemplar sphingolipid class. The reviewers raise understandable issues about the specificity of the available inhibitory compounds, such as zileuton meaning that conclusive statements about the role of LTE4 are not possible.

      In a revised manuscript, we will address all points but predominantly focus on the second aspect of the study, ensuring that reviewers comments are addressed appropriately, detailing and weaknesses, or ambiguities, with our study. This will include, but will not be limited to:

      - Further commentary on the regulation of eosinophil numbers within the tissue;

      - Addressing the specificity of zileuton and the implications of this for interpretation of our results with respect to eosinophil biology;

      - More careful framing of the transcellular biosynthesis potential;

      - A detailed discussion of sex dependency with regard to eosinophil numbers in general and any potential effect on the reported Gata6-dependent phenomenon;

      We are grateful for the constructive comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper describes an interesting phenotype of C. elegans lite-1 mutants. Previous work showed that lite-1 mutants lose a violet/blue light avoidance response. The authors show here that lite-1 mutants also show a defect in negative diacetyl chemotaxis. While wild-type worms avoid diacetyl at high concentrations, lite-1 mutants are instead *attracted* to it. The authors go on to perform Ca2+ imaging in sensory neurons and find that ADL and ASK neurons show altered Ca2+ responses to diacetyl in lite-1 mutants, suggesting LITE-1 is required for these responses. As unc-13 mutants with defective synaptic transmission show similar diacetyl Ca2+ responses as wild-type, this suggests these neurons respond cell autonomously to diacetyl. However, whether lite-1 also acts cell-autonomously is not discussed. Indeed, because unc-13 and lite-1 mutants show different ADL and ASK Ca2+ responses, it seems the diacetyl response regulated by LITE-1 is likely acting outside of those cells. An interesting result that is not commented on is the switching of the valence of the ASK Ca2+ response in lite-1 mutants. ASK neurons still respond to diacetyl, but instead of a strong increase in Ca2+, diacetyl appears to drive it strongly lower. This may be consistent with the switch in valence in the diacetyl chemotaxis assay. It also argues against the idea that LITE-1 is a low-affinity diacetyl receptor that drives avoidance or the Ca2+ responses in ASK, since it is still present in lite-1 mutants. The authors then use a strain that expresses LITE-1 in the body wall muscles and show this expression is sufficient to engender them with sensitivity to diacetyl, as measured through altered swimming and hypercontractility. The authors interpret this result as LITE-1 may act as a diacetyl receptor. The authors test whether a structurally similar molecule, 2,3-pentanedione, shows similar effects, and they find it does. Alpha-fold modeling and molecular docking analysis show where diacetyl might bind to the LITE-1 protein. They then test whether lite-1 mutants show chemotaxis defects to other molecules, as seen with diacetyl. Generally, they find that the observed diacetyl responses are unique, although lite-1 mutants do lose their avoidance response to 2,3-pentanedione. However, unlike the acquisition of diacetyl attraction in lite-1 mutants, 2,3 pentanedione avoidance is *lost*; it is not switched to attraction. Overall, I felt the description of the results and their implications could have been more in-depth. Further, the evidence that LITE-1 is a chemoreceptor itself, rather than acting in some way to shape chemoreceptor responses (via light or otherwise), remains unclear, as conceded by the authors.

      Strengths:

      Overall, the study follows up on an interesting and useful result. The experiments as presented are generally well-conceived and performed. The authors use a variety of behavioral and imaging approaches to test how LITE-1 mediates diacetyl avoidance.

      Weaknesses:

      The study is missing experiments needed to resolve whether LITE-1 is doing what they propose. The evidence that LITE-1 is a diacetyl receptor is lacking support since lite-1 mutants have their avoidance and calcium responses flipped, which would not be expected if it were acting solely as an avoidance receptor. Presumably, the authors are concluding that the attractive response that is left in the lite-1 mutant is mediated by ODR-10, but that experiment is not shown.

      We interpret the shift from avoidance to attraction in lite-1 mutants as consistent with the loss of an aversive sensory component in the presence of an underlying attractive response to diacetyl. We initially hypothesised that this residual attraction was mediated predominantly by ODR-10. To test this, we now generated and analysed lite-1; odr-10 double mutants. The double mutants retained an attractive response to diacetyl, indicating that ODR-10 alone does not account for the attraction observed in the absence of LITE-1 and that additional receptors or sensory pathways are likely to contribute. This finding is consistent with previous studies where loss of ODR-10 did not lead to a complete loss of diacetyl responsiveness.

      Similarly, the authors concede that "the use of lite-1 point mutants that affect specific LITE-1 function, such as light sensing, channel gating, or binding pocket, could further elucidate LITE-1 mechanisms." This reviewer agrees, and such experiments designed to localize diacetyl binding site(s) would be necessary to conclude definitively that LITE-1 is a diacetyl receptor. The body wall muscle assay used or some other heterologous experimental system could work for such a structure-function analysis. A concern is whether the extensive number of LITE-1 point mutants described in the literature affect cell surface expression vs. receptor function, which might complicate the interpretation of a result showing loss of diacetyl responses.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and is an important future direction for research, which we have included in the discussion.

      Reviewer #2 (Public review):

      Summary:

      Koh and colleagues investigate the broader sensory role of LITE-1, a gustatory receptor previously linked to UV light detection in C. elegans. Their study explores whether LITE-1 also mediates avoidance of specific chemical stimuli-namely, high concentrations of diacetyl and 2,3-pentanedione. They show that LITE-1 is required in the ADL and ASK neurons for calcium responses to diacetyl, and that its expression in body-wall muscles is sufficient to trigger hypercontraction upon odorant exposure. Molecular docking suggests both odorants may directly bind to LITE-1 with micromolar affinity. These findings suggest LITE-1 may act as a multimodal receptor for both light and chemical stimuli.

      Strengths:

      (1) Methodological Precision: The study is technically strong, with well-executed calcium imaging and quantitative behavioral assays that clearly show neural and muscular responses to chemical stimuli.

      (2) Novelty and Scope: The work presents a compelling case for LITE-1 functioning as a multimodal sensor, which is an intriguing expansion of its known role.

      (3) Potential Impact: If validated, the findings could significantly advance the understanding of sensory integration in C. elegans, and the tools developed may be broadly useful to the research community.

      (4) Relevance to the Field: The study adds to evidence that C. elegans uses non-canonical sensory pathways and may inspire further exploration of multimodal receptor functions in other systems.

      Weaknesses:

      (1) Lack of Rescue Experiments: The absence of rescue experiments makes it difficult to definitively link the observed phenotypes to loss of lite-1.

      We have now performed the rescue experiment expressing lite-1 in ADL, and showed that LITE-1 in ADL is sufficient for avoidance, although it is not a complete rescue to wild-type levels.

      (2) Single Loss-of-Function Approach: The reliance on a single genetic mutant limits interpretability. Additional strategies such as RNAi (e.g., neuron-specific knockdown) would provide stronger evidence.

      We observed the loss of avoidance in three independent lite-1 alleles. Combined with the new cell-specific rescue experiment, we think this provides sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Unclear Neuronal Contribution: While calcium responses in ADL and ASK are reduced, it's unclear which neuron(s) are necessary for behavioral avoidance. Cell-specific rescue or knockdown experiments are needed.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      (4) Unvalidated Docking Data: The molecular docking predictions lack experimental validation. Site-directed mutagenesis would be needed to support claims of direct interaction.

      We agree that the docking data does not in itself establish direct binding (we think the muscle expression and paralysis provides stronger evidence). Based on previously reported docking experiments, we wanted to check if diacetyl could occupy the same binding pocket. We have now also included docking data of the other odorants from the chemotaxis assays in the manuscript.

      (5) Limited Odorant Specificity Testing: Docking analysis does not include non-binding odorants, making it difficult to assess binding specificity.

      We agree and have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Incomplete Quantification: Some calcium imaging results (e.g., in AWA neurons of unc-13 mutants) lack statistical comparisons, which limits their interpretive value.

      We have generated the scatter plots of calcium imaging responses across the different sensory neurons, and the statistical significance was assessed using two-sided t-tests with FDR correction, which is now included in the manuscript.

      Reviewer #3 (Public review):

      In this work, Brown and colleagues report that the photosensor protein LITE-1 of the nematode C. elegans may also be a chemosensor that can be activated by high concentrations of the compound diacetyl. LITE-1 was described as a putative ion channel of the gustatory receptor family, which is mainly constituted by insect odorant receptors. These form tetrameric ion channels that can be activated by odorants. Specificity is achieved by forming heteromeric channels from three copies of the odorant receptor co-receptor (ORCO) and another subunit that resembles ORCO in the pore-forming C-terminus, but brings in a binding site for the respective odorant. LITE-1 has a very similar structure, according to Alphafold3 predictions, and also carries a binding pocket. In LITE-1, this was proposed to be occupied by a light-absorbing molecule that activates the channel when a photon is absorbed. Alternatively, compounds generated by absorption of high-energy photons may be formed in vivo and bound by the LITE-1 binding pocket. Koh et al. now demonstrate that another, non-light-activated compound, diacetyl, at high concentrations, can activate cells expressing LITE-1. Such (chemosensory) cells are also responsible for the avoidance of high concentrations of diacetyl. LITE-1 activation in excitable cells, i.e, muscles, causes strong body contraction and paralysis, and the authors show that this is also the case when diacetyl is presented. The authors further present molecular docking studies showing that diacetyl could occupy the binding pocket of LITE-1. Last, they show that another compound chemically resembling diacetyl, i.e., 2,3-pentanedione, can also induce avoidance in a LITE-1 dependent manner, though not as potently.

      The data are intriguing, and the demonstration of LITE-1 being a diacetyl chemosensor is interesting. Yet, there are a few questions arising that the authors should address.

      The authors identified mutants lacking diacetyl responses. In their chemotaxis assay (Figures 1A, B), they show that lite-1 mutants do not avoid high concentrations of diacetyl. However, the animals actually showed attraction, as the chemotaxis index was positive. If the lite-1 animals were insensitive, they should be indifferent, and the chemotaxis index should be close to zero. This means, other neurons contribute to the diacetyl response, and the result of these neurons being activated means/remains attraction? If so, the authors need to rule out any effects of these neurons on the effects they attribute to LITE-1 in the other assays.

      We have tested tax-4 mutants in the chemotaxis assay and found that, contrary to the predicted chemotaxis index of zero, these animals retained strong avoidance of high concentrations of diacetyl. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4 independent sensory pathways contribute to high diacetyl avoidance. We agree that these experiments cannot completely rule out indirect neuronal effects. We have therefore revised the text to acknowledge this limitation. Nevertheless, the rapid paralysis and contraction observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background support the idea that LITE-1 is sufficient to confer a diacetyl-evoked response in these cells.

      The effect of diacetyl on muscle cells (Figure 3C) is pretty rapid, i.e., already during 1 minute after application, the animals are almost maximally contracted. How fast is it really? Can the authors provide a time course with more time points during the first minute? This is a relevant question, as the compound would have to either pass the worm cuticle or enter through the gut and diffuse through the body to reach the muscle cells. Can one expect this to occur within (less than) a minute? In this context, the authors need to rule out that other mechanisms may be at play. E.g., diacetyl may be immediately sensed by ciliated chemosensory neurons that might release a signaling molecule that leads to activation of LITE-1 in muscles, or that sensitizes it somehow, responding to light used for filming animals. The authors should repeat this assay in a lite-1 mutant background.

      We repeated the paralysis assays under red-filtered illumination to minimise potential effects of light, with animals maintained in darkness from hatching to adulthood. We also included lite-1 mutants to assess whether neuronal LITE-1 contributed to the paralysis response. In addition, the assay was repeated with more frequent time points, revealing that paralysis and body contraction occurred within 10 s and neuronal LITE-1 does not contribute to the effect.

      Furthermore, the authors tested unc-13 mutants to rule out indirect effects on the neurons recorded. Likewise, they should eliminate neuropeptide signaling via unc-31 mutants (a recent paper cited by the authors showed involvement of neuropeptide signaling in LITE-1-mediated light avoidance behavior).

      We agreed and have acknowledged and discuss in the manuscript that contributions from gap junction-mediated communication, neuropeptide signalling and other chemosensory pathways cannot be excluded.

      Last, to demonstrate that effects are not indirect in response to chemosensory neurons, the authors should repeat the contraction or swimming assay in a tax-4 mutant, which largely lacks chemosensation. This also applies to the chemotaxis assay. Animals should exhibit a chemotaxis index to diacetyl of zero, then.

      We have tested tax-4 mutants, and like wild-type animals, retained strong avoidance of high concentrations of diacetyl, indicating that TAX-4-independent sensory pathways contribute to this response. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4-independent sensory pathways contribute to high diacetyl avoidance. Therefore, repeating the contraction or swimming assay in a tax-4 background would not completely exclude indirect input from other chemosensory neurons. In addition, rapid paralysis and contraction were observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background, and together with the calcium imaging and rescue data, they support a role for ADL and ASK in mediating high diacetyl avoidance. The tax-4 chemotaxis data is now included in the manuscript.

      Does diacetyl activate other neurons expressing LITE-1? A number of cells express LITE-1 at high levels, which the authors have not tested (they restricted their analyses to chemosensory neurons). This is important to address because it leaves the possibility that LITE-1 requires a specific partner only present in these chemosensory neurons to detect diacetyl. This partner would have to be present also in muscles, where diacetyl could activate ectopically expressed LITE-1. According to CeNGEN scRNAseq data, cells expressing LITE-1 can be identified. The ADL and ASH neurons actually come up only at the lowest threshold, so some of the other cells showing much higher levels of LITE-1 mRNAs, i.e., AVG, ALM, PLM, ASG, PHA, PHB, AVM, RIF, or some pharyngeal neurons, should be tested. ASG was among the cells the authors recorded from, but this neuron did not show a response.

      We have acknowledged and discuss in the manuscript that other non-sensory neurons may contribute to the avoidance behavioural, and which should be the future direction for investigation.

      The authors need to show that diacetyl responses of ADL and/or ASK can be rescued by expressing LITE-1 specifically in these neurons in a lite-1 mutant background.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      Molecular docking studies are not described in detail. How was this done?

      Molecular docking was performed in two stages. First, diacetyl was docked to the tetrameric LITE-1 model using DynamicBind without a predefined binding pocket. The generated complexes were ranked using the DynamicBind confidence score, and the highest ranked poses were used to identify the candidate binding site. The top DynamicBind pose was then used to define the box region for redocking with Gnina. Gnina poses were ranked using the CNN score. A more detailed molecular docking procedure has now been updated in the Methods section.

      Diacetyl is a very small molecule. How well can docking algorithms assess this at all?

      We agree that the small size of diacetyl limits the precision of docking scores because it forms relatively few protein contacts. However, its small size and limited conformational flexibility also simplify pose sampling. To increase robustness, we used two conceptually different docking approaches. First, DynamicBind was used without a predefined binding pocket to identify candidate binding regions while allowing ligand-associated protein conformational adjustments. Second, the resulting pocket was subjected to focused redocking and CNN-based pose ranking with Gnina. The results are interpreted as a structural hypothesis for the probable binding site and relative affinity ranking, rather than as definitive proof of binding or an accurate quantitative affinity measurement.

      Did the authors preselect the binding pocket, or did the algorithm sample the entire molecular surface of the LITE-1 model and end up with the binding pocket?

      The binding pocket was not predefined. Diacetyl was first docked to the tetrameric LITE-1 model using DynamicBind without specifying pocket residues or grid coordinates. DynamicBind therefore performed global, pocket-agnostic docking. The highest-ranked poses identified a candidate pocket, which was then used for focused redocking with Gnina.

      The latter would be very convincing. The authors should provide control docking experiments with other molecules that caused avoidance in their hands (i.e. benzaldehyde, 2,4,5,trimethlythiazole, isoamyl alcohol, nonanone, octanone), but did not activate LITE-1. Also, they should try docking molecules related to diacetyl, and if there are some that do not dock under the same conditions, such molecules should be used in a behavioral experiment. Ideally, they should also not activate LITE-1. Examples could be, e.g., diacetyl monoxime or 2,4-pentanedione.

      We have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      Last, the authors should provide a PDB file with the docked diacetyl to allow readers to assess the binding for themselves. Since a large number of mutations of LITE-1 have been reported, it may be that amino acids shown to be essential for LITE-1 function are also required for diacetyl binding. If so, this could be backed up with an experiment.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and which we have highlighted in the discussion as an important future direction for research. Additionally, we have now provided the PDB file containing a representative DynamicBind derived docking pose of diacetyl within the LITE-1 binding pocket.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) lite-1 mutant animals, as described, fail to avoid' high concentrations of diacetyl, but isn't it more accurate to say that the valence of the response is changed from avoidance to attraction? Do the authors believe that attraction is mediated by ODR-10? Can you build an odr-10; lite-1 double mutant and determine if they lose this attraction to high concentrations of diacetyl and/or 2,3 pentanedione?

      Our initial interpretation was that, in the absence of LITE-1-mediated avoidance, attraction to high concentrations of diacetyl is driven by the low-concentration receptor ODR-10. Interestingly, however, the lite-1; odr-10 double mutants remained strongly attracted to high concentrations of diacetyl, suggesting that this phenotype is independent of ODR-10 and may instead be mediated by other, less specific odorant receptors. This is consistence with the odr-10 mutants not completely losing attraction to low concentration of diacetyl, suggesting the involvement of other potential/putative receptors (Sengupta et al., 1996; Taniguchi et al., 2015). The lite-1; odr-10 double mutant data is now added in the results section (lines: 82 to 90; Figure S1B).

      (2) For ADL, yes, it seems like lite-1 mutants have a reduced diacetyl response, but the ASK response seems... different. While it goes up (slowly) in wild-type, cellular Ca2+ levels in ASK (and maybe ADL) are *reduced* by diacetyl in lite-1 mutants. Can the authors comment on this, and the behavioral responses change in valence?

      ADL and ASK are involved in both attractive and aversive responses so it is possible that the attractive component of the diacetyl response suppresses their activity. In the absence of LITE-1 activation, the observed decrease in calcium responses may reflect the unopposed inhibitory input. The slower decay of the calcium signal in ASK neurons of unc-13 mutants further supports the presence of additional inhibitory signals influencing their activity. This explanation is now included in the results section (lines: 101 to 117).

      (3) ADL and ASK calcium traces in unc-13 mutants look generally similar and lack the effects seen in lite-1 mutants. Does that mean the lite-1 effect is in cells other than ADL or ASK? Can the authors spend more time discussing these differences?

      We acknowledge that LITE-1 is expressed in multiple cell types beyond chemosensory neurons, and that non-chemosensory neurons may also contribute to the observed phenotype. In the previous version, we had highlighted the interneuron AVG as a potential contributor, given its role in light-induced escape. In the revised discussion, we have now expanded this section to include additional possible contributors such as the LITE-1-expressing phasmid neuron PHA and the pharyngeal interneurons I2, both of which have been implicated in hydrogen peroxide sensing (lines: 177 to 181).

      (4) Chemotaxis responses to diacetyl and 2,3-pentanedione in lite-1 mutants are rather different. Diacetyl switches from repulsive (CI < 0) to *attractive* (CI > 0), which is not what would be expected for mutations that eliminate a receptor. In contrast, the 2,3-butanedione responses are more what would be predicted: diacetyl goes from inhibitor to no effect (CI ~0). Again, if the authors feel that this is because of the ODR-10 function, can they discuss whether 2,3-butanedione is predicted to bind ODR-10 like diacetyl?

      We think a switch to attraction is expected for the removal of a receptor for an aversive signal, in an attractive background signal. We did indeed think that this attraction was mediated by odr-10, but the double mutant results now show other receptors must be involved. This is also not wholly unexpected since previous work has identified other receptors of high-concentration diacetyl and the original odr-10 paper didn’t report a complete absence of diacetyl response, suggesting the presence of other receptors that mediate attraction to diacetyl (Sengupta et al., 1996; Taniguchi et al., 2014). The new data is now added in the results section (lines: 82 to 90; Figure S1B; Supplementary video 1 and 2).

      (5) Were the behavior experiments performed in the dark? I realize the calcium imaging experiments and some of the video behavior recordings are not possible in complete darkness, but maybe the authors made efforts to exclude visible light effects (e.g., infrared illumination, etc.) in some assays that might help determine whether light plays *no* role in the effects observed. Alternatively, the authors could try repeating their chemotaxis experiments in the dark or at least communicate in the methods that this was not viewed as a concern (and why). As the authors propose and discuss LITE-1 modulating diacetyl responses via light sensation as a possibility, it is incumbent upon them to communicate the steps they took to overcome this concern for themselves.

      We acknowledge this and have performed a chemotaxis experiment to compare assays performed under dark and ambient light conditions, and no significant differences were observed (results section: lines 75 to 80; Figure S1A, material and methods section: 241 to 246). Therefore, subsequent chemotaxis assays were carried out under ambient light while avoiding exposure to strong illumination.

      Paralysis assays were repeated under red-filtered illumination to minimise light effects, with animals maintained in darkness from hatching to adulthood. Additionally, the assay was expanded to include lite-1 mutants, ruling out contributions from neuronal LITE-1 to paralysis. The new data is now incorporated into the result section (diacetyl; lines: 135 to 137, figures 4 and S3; 2,3-pentanedione; lines: 152 to 154, figures 4 and S6), materials and methods section have been updated to reflect these changes (lines: 250 to 253 and 257 to 260).

      Reviewer #2 (Recommendations for the authors):

      Minor Issues:

      (1) Pmyo-3::LITE-1 worms shrink in the absence of odorants (Figures 3C, 4D); possible effects of ambient light should be discussed.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis.

      To minimise potential light-induced effects, the paralysis assays were repeated with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) The title is misleading, as ASH does not show altered activity in lite-1 mutants and should be removed from the claim.

      ASH has been removed from the title (line: 97).

      (3) Specific Kd values should be provided for the reported micromolar binding affinities.

      The values from DynamicBind and Gnina are provided in Figure S5.

      Recommendations:

      (1) LITE-1, a member of the gustatory receptor family, was previously shown to mediate UV light responses in C. elegans. In this study, Koh and colleagues demonstrate that LITE-1 is also required for the nematode's avoidance of high concentrations of diacetyl - an odorant that is attractive at low levels but aversive at higher concentrations. Using calcium imaging, the authors show that LITE-1 is necessary in the sensory neurons ADL and ASK for calcium transients in response to high concentrations of diacetyl. Additionally, they find that expressing LITE-1 in body-wall muscles causes hypercontraction upon diacetyl exposure. Similar LITE-1-dependent responses were observed for 2,3-pentanedione, another structurally related odorant. Molecular docking analyses suggest that both diacetyl and 2,3-pentanedione directly bind to LITE-1 with micromolar affinity.

      These findings are intriguing and have the potential to significantly advance our understanding of LITE-1 as a multimodal sensory receptor. However, several major issues need to be addressed to support the authors' conclusions:

      (1) Rescue experiments are missing. The authors should rescue at least one lite-1 mutant to confirm that the observed avoidance defects are specifically due to loss of lite-1.

      Because the avoidance defect was observed in three independent lite-1 alleles, we think background mutations are unlikely to be causal. We have now also performed a rescue experiment with lite-1 expressed in ADL neurons. In this strain, attraction is restored. The new data have been incorporated into the results section (lines: 119 to 121; Fig. 2D).

      (2) Alternative loss-of-function approach. To strengthen the findings, the authors should use a different method to disrupt lite-1 function-such as RNAi by feeding or cell-specific RNAi driven by the lite-1 promoter (see PMID: 17459615).

      We believe the multiple alleles (Fig. 1B) and new cell-specific rescue experiment (Fig. 2D) provide sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Clarify the role of ADL and ASK neurons. While calcium imaging data show reduced activity in these neurons in lite-1 mutants, it remains unclear whether lite-1 is required in ADL, ASK, or both for avoidance behavior. Cell-specific rescue or RNAi experiments, along with additional calcium imaging, are needed to determine the contribution of each neuron.

      We agree that more in-depth work will be required to dissect the neuronal pathways involved in LITE-1-mediated diacetyl avoidance. We have avoided making specific comments on how exactly the observed imaging results relate to the behavioural phenotype. The new rescue experiment expressing lite-1 in ADL does at least show that LITE-1 in ADL is sufficient for avoidance, although it’s not a complete rescue to wild-type levels (lines: 119 to 121; Fig. 2D).

      (4) Validation of molecular docking results. While molecular docking suggests direct binding of odorants to LITE-1, experimental validation is needed. Mutations that reduce predicted binding affinity (engineered in transgenes or via CRISPR) should be tested for functional impact on avoidance behavior.

      We agree the docking does not in itself establish direct binding (we think the muscle expression and paralysis provides much stronger evidence). Based on previously reported docking experiments, we were simply curious whether diacetyl would be predicted to occupy the same binding pocket. We have now updated the discussion in the use of lite-1 mutants to test for impact on diacetyl avoidance (lines: 188 to 191).

      (5) Include analysis of non-binding odorants. Docking results should also be presented for odorants that did not elicit LITE-1-dependent avoidance, to help establish specificity.

      We have now included docking data of odorants from the chemotaxis assays, with the corresponding docking values shown in Supplementary Figure 5. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Figure 2C concerns. In neurons such as AWA, calcium transients in unc-13 mutants appear reduced compared to wild-type. A statistical comparison for all the neurons should be included to assess significance.

      Scatter plots of calcium imaging responses across the different sensory neurons were generated, and statistical significance was assessed using two-sided t-tests with FDR correction. Only ADL and ASK neurons showed significant differences between lite-1 mutants and wild-type animals. No significant differences were observed in any neuronal pairs between unc-13 mutants and wild-type, including AWA neurons, although the difference is close to significant (p = 0.07). The scatter plots were now included as Supplementary Figure 3.

      Minor points:

      (1) Figures 3C and 4D: Pmyo-3::LITE-1 worms appear to shrink even without diacetyl or 2,3-pentanedione. Could this be due to ambient light? The authors should discuss this possibility.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis. We repeated the paralysis assays with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, to minimise potential light-induced effects, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) Title revision needed: The title "Chemosensory neurons ADL, ASK, and ASH are involved in avoidance of diacetyl" is misleading, as calcium transients in ASH appear unaffected in lite-1 mutants. The title should reflect the actual data.

      ASH have been removed from the title (line: 97).

      (3) Binding affinity clarification: The authors report micromolar binding affinity for LITE-1 but should provide specific dissociation constants for clarity and completeness.

      The values from DynamicBind and Gnina are now provided in Supplementary Figure 5.

      Reviewer #3 (Recommendations for the authors):

      How did the authors measure body length if the animals were swimming in the diacetyl solution? Standard 6-well plates have an area of roughly 10 cm², meaning that if one adds 1 ml, the liquid level should be 1 mm. The animals would be able to move in 3D, so it is likely that animals swim up and down and do not move in one flat plane, i.e., head and tail would be out of focus, and only a projection image would be recorded that would lead to an underestimation of actual worm length.

      We acknowledge that this issue may led to an underestimation of worm length in the previous assay. However, every effort was made to exclude worms that moved out of focus. The paralysis assay has since been modified to include spreading a thin layer of solution across the worms, which keeps them mostly in focus, particularly those that are paralysed. Worms that were partially out of focus were excluded from the analysis.

      We have incorporated the new data into the results section, reflected in the updated Figures 3C, 4E, S4C, and S6C. Corresponding revisions have also been made in the materials and methods (lines: 250 to 253 and 257 to 260).

      Could diacetyl be a compound that results from UV absorption in cells? This may be worth discussing. What could be the precursor molecule?

      We are not aware of such a precursor, but we cannot rule it out. Even if UV absorption in cells leads to diacetyl production, it is likely that the resulting diacetyl levels are insufficient to activate LITE-1, as our data suggest that LITE-1 functions as a receptor for high concentrations of diacetyl. We have updated the discussion accordingly (lines: 169 to 173).

      In lines 57-61, the references to Edwards 2008 and Ward 2008 do not seem to fit the statements made in this sentence.

      The inclusion of Edwards et al., 2008 was an error, and it has now been removed. Ward et al., 2008 demonstrated that ASJ phototransduction requires cGMP and CNG channels, stating that “Our studies indicate that C. elegans photoreceptor cells also employ CNG channels and the second messenger cGMP for phototransduction.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper examines whether humans use protracted temporal integration in a noise-free, deferred-response contrast discrimination task, using a covert evidence-duration manipulation combined with EEG (SSVEP, CPP, Mu/Beta). The key finding is that evidence for protracted sampling is behaviorally and neurally supported, but even joint CPP + behaviour fitting cannot fully discriminate a standard integration (DDM) model from a novel "extremum-flagging" non-integration model. The paper is transparent about this outcome.

      Strengths:

      This is a well-conducted and well-written study that makes a genuine contribution to the perceptual decision-making literature by introducing a clean experimental design for probing temporal integration without participants adapting their strategy and demonstrating for the first time that a non-integration model (extremum-flagging) can replicate CPP waveform dynamics that have long been considered hallmarks of evidence accumulation. The transparent treatment of equivocal modelling outcomes is commendable.

      Weaknesses:

      My main concerns relate to statistical power, the under-specification of the and the extremum-flagging mechanism. Addressing these would greatly strengthen the paper.

      (1) The sample of 16 participants (15, after the exclusion of one participant) is described as "close to similar EEG studies" with no formal power analysis. Given that the paper's core claim rests on subtle quantitative differences between two model classes - differences that are, by the authors' own admission, not sufficient to declare a winner - even a modest increase in sample size might yield a more decisive outcome. At a minimum, the authors should report a sensitivity analysis or post-hoc power calculation to indicate what effect sizes the current N could reliably detect, particularly for the rmANOVA comparisons and the neural constraint fitting.

      We appreciate the reviewer’s concern regarding sample size and statistical sensitivity. To address statistical robustness throughout the paper, we have now reported effect sizes for our statistical tests (e.g. η2 for rmANOVA; including Tables S1 and S2), and we provide error-shading around the ERP waveforms to indicate the reliability of the key patterns our models are aimed at capturing (i.e. the dramatically higher and earlier CPP peak for high-contrast, and very little systematic differences across the four low-contrast durations - see revised Figure 3). We also conducted an indicative post-hoc power analysis using the G*Power software based on the behavioural data. Using the observed partial η2 = 0.44 for the test of duration effect on accuracy among only the low-contrast conditions, and the final sample of 15 participants, this amounts to a statistical power of 0.998.

      On the model comparison, while we agree that larger sample sizes are generally beneficial for population-level inferences, we respectfully maintain that our current sample size is sufficient to support the core claim that qualitative dynamics of neural signatures of decision formation, usually assumed to reflect temporal integration, can be successfully reproduced using non-integration models in the delayed-response task conditions we examine here. Any marginal changes in quantitative fit resulting from having a higher N contribute to grand averages are unlikely to substantively alter this conclusion of the model comparison. The statistical reliability of the data to which our models are fitted is also bolstered by the number of trials (about 256 per condition per participant). It is common in behavioural modelling studies for data to be collected from a much smaller sample (e.g., fewer than 10 subjects) but with a high trial yield - a relevant precedent for us being Stine et al., (2020), who provided a compelling demonstration of similar model fits for Extrema and Integration models using only 6 subjects. In sum, the key qualitative data patterns of accuracy improvements with duration and broader, lower and duration-invariant low-contrast CPPs are statistically robust and provide a strong basis to reveal the fundamental principle that the Extremum-flagging and Integration models are both able to produce these key qualitative dynamics.

      (2) The Extremum-flagging model is the paper's most novel contribution, yet its physiological basis is underspecified. The model posits that each decision-terminating bound-crossing triggers a stereotyped, half-sine-shaped centroparietal signal, but no neural circuit or computational mechanism is proposed for how the brain could detect the first bound-crossing event in a non-accumulating evidence stream or generate a temporally precise, fixed-amplitude signal in response. Possible connections to P3b theories of context updating and response facilitation are acknowledged, but these are vague functional descriptions rather than mechanistic accounts. I think the discussion should engage more directly with potential neural substrates that could generate this flagging signal, and whether these are consistent with the known generators of the CPP/P3b. Without this, the extremum-flagging model risks being viewed as a mathematical convenience rather than a biologically plausible alternative.

      We thank the reviewer for this constructive comment. While the focus of this paper was indeed on simple mathematical descriptions in the spirit of classical cognitive modelling, we agree that expanding on the potential neural substrates of the Extremum-flagging model strengthens its utility as an alternative framework. We have revised the Discussion to engage with potential biological mechanisms, particularly those previously proposed to underlie the P300/P3b, such as Nieuwenhuis’ (2005) proposal that it reflects a phasic arousal response mediated by the LC/NE system that serves to activate task-relevant areas following completion of a decision. We agree that aside from this, accounts of ERP component functions over the years have often been vague and non-mechanistic, but the idea that they reflect discrete neural activations marking an internal cognitive event in a stereotyped way persists, and remains a basic assumption of several new and influential ERP signal analysis toolboxes (e.g. Ehinger 2019; Weindel 2024). If the flagging signal’s fixed amplitude seems physiologically implausible, all-or-nothing neural activation events are not generally unheard of in neurophysiology, and, again, we are taking an approach favouring parsimony in the spirit of cognitive modelling, and we found that we did not need to assume any variation in the amplitude of the flagging signal in order to capture the key decision signal dynamics alongside behavioural accuracies in this particular case.

      We also discuss the study of Latimer et al. (2015), who demonstrated that discrete, step-function state transitions that on single trials may mark extrema detection events, can produce ramp-like signals when trial-averaged. While the biological plausibility of such step-function dynamics remains a subject of debate, it serves as another example of how continuous evidence integration is not the only way to reproduce the ramping neural signals traditionally observed in grand-average neural signals.

      (3) The Integration model at the preferred neural weighting estimates a high-to-low contrast drift rate ratio of 8.7, whereas the empirical Mu/Beta lateralization slopes suggest a ratio of approximately 3.5. The authors attribute this discrepancy to the nonlinear contrast response function of early visual cortex and the salience of the high-contrast evidence onset, but these explanations are speculative. These outcomes are arguably the most quantitatively damaging result for the integration model, so they deserve more than a brief discussion. I would recommend that the authors (a) estimate what range of contrast response nonlinearities would be required to close this gap, (b) test whether an alternative drift rate parameterization (e.g., scaling drift rates directly by SSVEP amplitude rather than contrast) reduces the discrepancy, or (c) be more explicit about treating this as a point against the Integration account.

      We agree that the quantitative discrepancy we demonstrated between the empirically observed buildup rate ratio in motor preparation signals (3.5) and the greater drift-rate ratio (8.7) required by the Integration model to fit the CPP waveforms is an important one that should be emphasised and discussed with greater depth and clarity. As we said, nonlinear contrast response functions and a boosting effect of the salient high-contrast step-change are two plausible ways that a drift rate might scale disproportionately more steeply with contrast, but in principle, assuming straightforward transmission of evidence accumulation to the motor level, Mu/Beta lateralization slopes should then reflect this steeper drift rate scaling, or at least approach it even when allowing for some temporal blurring. We have thus put more emphasis on the discrepancy by confirming that if we constrain the drift rates to be directly proportional to contrast, the Integration model is indeed significantly hampered in its ability to produce the much steeper CPP buildup for higher-contrast trials, much more so than the Extremum-flagging model (Figure 4 - Supplement 7). We have also applied a temporal blurring equivalent to the short-time Fourier Transform to the simulated motor preparation waveforms (convolving with a boxcar of the same duration as the Fourier window) in Figure 4N-P so that the real and simulated traces are on an equal footing in this respect. We have also revised the Discussion to elaborate on how this quantitative discrepancy represents a point against the Integration account, and possible ways it might be reconciled with an Integration account. One reason, for example, why the relative steepness of the Centroparietal ERP in the high-contrast condition so far exceeds that of Mu/Beta might be that additional processes are evoked by the very salient step-change, which may make a positive-polarity contribution to the centroparietal ERP waveform and hence cause overestimation of how early and steeply the underlying, high-contrast CPP decision signal rises. We looked into this by carefully examining time courses and topographies through the initial period of buildup, with no additional smoothing low-pass filter applied, now presented in Figure 3 - Supplementary Figure 1. While the smoothed waveforms that we show in the main paper and to which we fit models could be seen to have a brief inflection during the main buildup for the high-contrast condition, removing the smoothing shows that this arises not from random noise but from a distinct bimodal morphology, with a distinct early peak and lull during the buildup, which temporally coincides with a very strong bilateral occipital N2 (associated with a low-level evidence-onset detection or ‘target selection’ process - Loughnane et al 2016), in a way that suggests that the positive tail-end of the dipolar neural generators of the N2 may contribute to the initial part of the positive centro-parietal buildup. It is difficult to estimate the extent to which the neurally-constrained model estimate of high-contrast drift rate is inflated by this initial overlapping potential, because we can’t precisely know the ground truth of the N2 tail’s contribution, but this analysis provides a potential explanation that can be explored in future (e.g. through softening the strong-evidence onset with a ramp or use of auditory evidence). We thank the reviewer for raising this as we feel that this extra discussion positively adds to the theme of the paper to highlight methodological challenges with neurally-constrained modelling. In the process, we have updated the methods section to present in full detail the centroparietal electrode selection and waveform smoothing that was applied to provide the models with a relatively uninterrupted buildup signal to capture, which is important for readers to appraise the potential impact of this overlapping potential.

      (4) The sensitivity analysis over neural constraint weightings (w = 0.1 to 1000) is thoughtful, but the paper ultimately acknowledges that the preferred weighting is w=10, chosen because it achieves "a good fit to CPP dynamics without substantively sacrificing behavioral fit" - a qualitative criterion. No principled statistical framework is used to select the optimal weighting or to compare models at a given weighting. A Bayesian model comparison could provide a more formal framework for combining behavioral and neural fit components, and would allow a clearer statement about the relative posterior probability of each model.

      We agree with the reviewer that theoretically, the Bayesian framework provides a principled way to combine behavioural and neural evidence by weighting each source according to its statistical reliability. However, a Bayesian formulation typically quantifies reliability through across-trial variance, which applies quite differently for accuracy and EEG data. While the precision of EEG measurements can be estimated empirically (e.g., from noise characteristics), we currently lack a formal measure of uncertainty for the linking function itself, that is, the theoretical mapping between neural signatures and latent decision processes. This represents an unresolved methodological issue rather than a straightforward parameter estimation problem. Second, although hierarchical Bayesian approaches are well established for standard diffusion models, the mechanisms examined here for extremum flagging do not currently have tractable closed-form formulations suitable for Bayesian integration. Developing a dedicated hierarchical Bayesian framework for these non-standard mechanisms would require substantial methodological work and is beyond the scope of this research.

      Thus, rather than imposing a single assumed reliability relationship between neural and behavioural data, we chose to perform a systematic sweep across weighting values. We view this approach as a transparent sensitivity analysis that accommodates different scientific priors regarding the relative contribution of neural versus behavioural constraints. By presenting the full range of w (including in supplemental tables and figures), readers can directly evaluate how model behaviour changes when emphasis is shifted between behavioural data and neural data, transparently revealing how the behavioural and neural signal fits can trade against one another.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Hajimohammadi, Mohr, O'Connell and Kelly is intended to demonstrate that participants integrate evidence over time to make a decision, even in a noise-free, static decision context. This is validated by the observation that (1) participant accuracy improves with increased exposure to the stimulus; and (2) there is a correlation between participant accuracy and a neural index of evidence accumulation, as measured by centro-parietal positivity (CPP).

      Strengths:

      (1) Joint modelling of accuracy and CPP dynamics is a significant achievement, as behaviour alone often cannot distinguish between competing theories of decision-making. In the case of protracted sampling in particular, the absence of reaction times (RT) due to the delayed nature of the response makes this method highly appealing.

      (2) The experimental manipulations and the method used to extract the different neural indices are well chosen, enabling the mapping of putative cognitive processes such as evidence accumulation and motor preparation onto the recorded EEG with clarity.

      (3) The in-depth discussion of the results clearly articulates those reported by the authors and in previous works.

      Weaknesses:

      (1) One main issue to support the interpretation of the authors toward the need for protracted sampling is the timing of the evidence. By design, participants believe that the signal is present for 1.6 seconds (reinforced by the fact that easy trials were displayed for 1.6 seconds). However, the difference in stimuli is turned off either 1.4, 1.2, 0.8 or 0 seconds before the cue to respond. While this makes sense in the context of the authors' question, it also raises the possibility that participants will focus on the last samples before answering. Even if participants apply equal weighting, this still favours them delaying evidence accumulation until they are sufficiently certain that the evidence should be present (e.g. participants might start accumulating after the stimulus has disappeared in the 0.2 condition). I do not see an easy way to test these alternative explanations outside of running a study in which the evidence is always offset before the go cue.

      This is a reasonable question about the design - if participants were under the impression that they had a whole 1.6 sec of stimulation, couldn’t they afford to wait until later into the stimulus to start sampling? However, the task was designed to be so difficult that participants would be deterred from ignoring any initial evidence, and the fixed and explicitly instructed lead-in period as well as the interleaved easy trials, would have continually reinforced their ability to time their sampling onset quite precisely. Indeed, key aspects of the data confirm they did not appreciably delay sampling. First, accuracy in even the shortest (0.2 s) condition was reliably above chance (t(15) = 2.60, p = 0.0201) and improved steadily across evidence durations (Figure 1B). This places an upper bound on the accumulation onset: participants cannot have delayed accumulation until after the evidence disappeared and still achieve above-chance performance; if they only used the ‘last samples,’ at the end of the stimulus, they would have performed at chance level for all durations except 1.6 sec. Second, we fit a model that allowed for such a delayed sampling onset, captured in the parameter ‘sampT,’ which, across the range of neural weightings (Tables S3, S5-8), consistently landed within a few tens of msec of evidence onset (often slightly before rather than delayed), and improved the overall model fit very little relative to the addition of starting point variabilities or collapsing bound. The Methods section now addresses these aspects of task design.

      (2) Regarding the behavioural models, are these identifiable based on accuracy data alone? This should be addressed using a parameter recovery study, in which a set of parameters is used to generate data, and the same fitting routine used for the real data is used to estimate the parameters. This would enable us to determine what can be inferred from the model comparison presented. This is not a serious problem for the manuscript, as it specifically aims to go beyond behaviour. It is, however, worth noting that such a parameter recovery addition could be used to demonstrate the need for a joint modelling framework to answer the question of protracted sampling on delayed response times (RT).

      As the reviewer notes, we did have the specific aim of going beyond behaviour, and the need to do so is demonstrated in the inability to adjudicate between the alternative models based on behaviour alone. We took this as sufficient justification without a formal parameter recovery test to assess the degree to which behaviour-only models could accurately estimate parameter values. Still, we agree that it is valuable to address parameter identifiability in some way. Since a full parameter recovery covering the full possible parameter space for each of the many models would be too great in volume to add to this paper, we can instead address identifiability somewhat indirectly through parameter estimate consistency across the 10 fits we conducted with different instantiations of noise; we now provide the standard deviations alongside the mean parameter values for the D1, D2 and B parameters of each of the behaviour-only models in Table 1 - Table Supplement 1, which indicates that the parameter estimates were reliable across 10 different instantiations. 

      Minor comments:

      (1) I would advise authors to fix the D1 parameter and use it as a scaling parameter across all models. Currently, as I understand it, the models are scale-free, meaning the same fit is achieved by multiplying all parameters by two, for example. This makes the fit more complex (bounds on parameter values are required) and means that the models are less comparable in terms of their estimates. Perhaps I'm missing something, but I would have thought that fixing D1 (the common parameter across all models) would solve these issues.

      The models are not scale-free because they are constrained relative to a fixed sampling noise parameter value of s = 0.1; All tables in the main text have now been updated to make this more immediately clear. Aside from this being standard in diffusion modelling (Ratcliff & Smith, 2004), this enabled us to replicate the observation by Stine et al., (2020) that since the non-integration models depend on the magnitude of individual evidence samples rather than an integration of many, the drift rate values must be set much higher to achieve the same choice accuracy as the integration models (Table 1).

      (2) Why is the snapshot model so bad despite being a good model in Stine et al 2020? Can the authors speculate in the discussion?

      We thank the reviewer for querying this. We had originally thought that the poor performance of the snapshot model made sense because the continued presentation of zero contrast difference for short-evidence trials renders it a bad strategy. Because our main purpose was to briefly substantiate the principle that accuracies alone are an insufficient basis for model comparison and move on to the main goal of jointly modelling accuracies and CPP dynamics, we did not take the same level of care to ensure we attained the very best fit of the behaviour-only models, as we did for the neurally-constrained models. In the neurally-constrained modeling, we took care to check for every parameter whether the range of allowed values (Table S4) was narrow enough to avoid the optimisation algorithm getting lost in untenable parts of parameter space, yet wide enough to include the optimum point, and wherever we saw parameter values landing at or near the edge of the allowed range we expanded that range and re-ran the model fit. Applying these same checks to the behaviour-only fitting, we found that the SnapShot model needed a wider range on drift rate and when we applied this, the fit was much more competitive, in line with Stine et al., (2020), though it remained the worst-fitting model among all two-drift-rate behaviour-only models (see updated Table 1). We similarly conducted these checks across all behaviour-only models and re-ran them. The extrema detection model with last-sample default when no bound is hit also improved its fit, though again it did not fit better than the version with guess default. Thus, the point we were making with this section, that behaviour alone can be captured competitively by a range of integration and non-integration models, is bolstered by the updated model fits. Since the last-sample default was competitive in the behaviour-only fits, we also ran a version of the Extremum-flagging model jointly fit to accuracies and CPP dynamics with a last-sample rather than random guess default when a bound was not reached, and show in new Figure 4 - Figure Supplement 8 that the conclusions are the same. Again, thank you for prompting us to look back at those fits.

      (3) The meaning of the flag width is unclear. Figure 4 provides the reader with an intuitive understanding of the model that the authors have in mind. However, the tables in the appendices report values between 0.2 and 0.9. I understand that these values represent the width of the half-sine in seconds. This suggests that the actual estimated values for these flag events are much broader than those displayed in Figure 4. While this is probably fine for most models, it can be problematic for the extremum-flagging model, as it means that the rise to the peak takes between 0.1 and 0.45 seconds. While strictly speaking, this is still a 'flag' model, such a slow rise to the peak, given the usual expectation of evidence accumulation, would place this model closer to a smooth integration model than to a boundary-crossing flagging mechanism.

      We thank the reviewer for raising this about the flag width parameter. In so doing, they enabled us to catch that our schematic depiction of the model in Figure 4 was misleading, and have now revised it to make clear that the flag signal is a post-decision one triggered by the bound crossing, and we have updated explanations accordingly (in ‘Neurally-constrained models’ and Discussion). The reviewer is correct that the reported values in the supplemental materials (approximately 0.2–0.9 s) correspond to the width of the half-sine kernel used to model the post-decision flag event. However, the flag signal is stereotyped, evidence-independent, and is triggered once the decision threshold has already been crossed, so it does not share the key characteristics of evidence integration, regardless of how wide the model estimates it. In the extremum-flagging model, the boundary crossing remains a discrete event. The width parameter instead captures the temporal extent of the neural process that follows this commitment event. Such a post-decision neural process unfolding over several hundred milliseconds is in line with some classic theories of the centroparietal P300/P3b component, and we now expand our discussion point on this to address proposed neural substrates (e.g. Nieuwenhuis et al’s (2005) implication of a phasic noradrenaline system response).

      (4) In the modelling section, it is not clear overall (i.e. for G<sup>2</sup> and R<sup>2</sup>) how the participant dimension is taken into account. Are these individually fitted models, and if so, how are the secondary statistics generated from the individual estimates? Or were these fitted over all participants?

      All models were fitted to the grand-average neural and behavioural data across participants, rather than to individual participant data. We chose this approach as the CPP signal at the individual level is highly noisy, which can introduce substantial instability and noise into the model fitting procedure. We have revised the Modelling section to explicitly state that the reported G<sup>2</sup> and R<sup>2</sup> values are derived from models fitted to the grand-average data, and in the revised discussion acknowledged this as a limitation of the current modelling framework.

      (5) On page 7, in the last sentence of the first paragraph of the section titled 'Decision-Related Neural Signals', the authors state that 'this stable contrast-difference encoding suggests that a constant (i.e. non-adapting) drift rate is a reasonable simplifying model assumption'. However, I am not sure how this is true given that SSVEP quantifies encoding, yet the drift rate can vary due to non-sensory aspects (e.g. attention).

      The reviewer makes a good point - even if sensory encoding is stable, non-sensory factors like attention could cause dynamic changes in the effective drift rate independently of the sensory representation itself. However, our point in that section, which we have revised to put more clearly, was to test for one particular well-known time-varying effect that could impact drift rate, namely sensory adaptation, a well-established phenomenon behaviorally and at the level of sensory neuronal responses, where prolonged stimulation produces reductions over time in sensory neural activity. If strong adaptation were present in the sensory evidence representation indexed by the SSVEP, we would expect corresponding temporal changes in the signal. The absence of such changes lends support to the simplifying assumption (as in most accumulation models) that the drift rate is approximately stationary over time, even if we cannot be sure there isn’t a time-varying effect downstream.

      (6) The mu/beta lateralisation does indeed favor the integration model more, but in terms of boundary estimation and starting-point analyses, both models are pretty far apart. Providing an interpretation of this observation, e.g. regarding alternative linking functions for mu/beta, would add to the manuscript.

      In response to this comment, we revised the manuscript in the Discussion to say that in the current analyses, we implicitly assume an approximately linear mapping between Mu/Beta amplitude and decision units. However, the true relationship may instead reflect another monotonic transformation (e.g., involving power rather than amplitude, logarithmic scaling such as dB units, or a nonlinear saturating function). This uncertainty could affect the apparent correspondence between the neural signal and the model-derived estimates of boundary position or urgency dynamics. While our analyses support a close relationship between Mu/Beta lateralisation and the evolving decision process, the precise quantitative mapping remains uncertain. One possibility is that urgency itself evolves nonlinearly (e.g., decelerating over time), even if the measured neural trajectory appears approximately linear under the current transformation assumptions.

      Reviewer #3 (Public review):

      Summary:

      The authors aim to compare proposal models of perceptual decision making using a joint modeling approach, where they fit models to both behavioral outcomes as well as CPP. Most notably, they compare a standard evidence accumulation model with models that track the evidence without integrating it over time (extrema detection). The authors report that the joint CPP-behavioral data do not discriminate between two of their proposals.

      Strengths:

      This is an interesting finding that reinforces the idea that what we believe to see based on aggregation over trials may not be what happens on every single trial. The models are creative, and the simulations are convincing, relating the models to multiple neural markers of decision formation. These include the CPP but also mu/beta power spectra.

      Weaknesses:

      The paper makes some strong points, and the work seems generally well-executed. The weaknesses that I identified are twofold:

      (1) Embedding in the literature/exposition of the main argument.

      The focus in the introduction is on the noise-free nature of the stimulus and the prolonged presentation time. However, after reading the paper, I felt these were mostly experimental design choices that enable comparison of the different models using the CPP. Perhaps my misreading of the goals of the paper stems from two other observations:

      (a) The fact that the stimulus is noise-free does not entail that perception is noise-free. Thus, the argument that using a noise-free stimulus precludes the necessity of temporal integration seems not completely valid. Of course, one could argue that noise is limited in this case, but that makes a noise-free stimulus more of a design choice.

      (b) The focus on prolonged stimulus presentation, but at the same time the contrast with expanded judgement, did not make sense to me. Perhaps, as a non-native speaker, I am misreading the subtle difference between "protracted sampling" and "longer sampling", but again, the longer duration seems mostly a design choice.

      We thank the reviewer for this impression, which has helped us revise the introduction to more clearly motivate the paradigm as an interesting case for close examination. The primary driver of our choice of stimulus and task parameters was not to enable model comparison using the CPP; it was to examine a decision scenario that exists in everyday life but that has not been examined in terms of underlying decision mechanisms because it offers only sparse behavioural data - the scenario in which plainly visible objects (without noise or stochasticity, as in daylight conditions) need to be examined for a subtle feature difference to guide a later action. The reviewer echoes our point in the Intro, that despite the absence of physical noise in the stimulus, perceptual processing itself is not noise-free. Therefore, temporal integration is certainly not precluded, but its benefit is minimised and less obvious to the decision maker. Given the examples we raise where integration was found to not be employed to its optimal extent (e.g. bound setting foregoing accuracy improvements with duration), and the various theoretical accounts citing energy costs associated with integration and the fleeting nature of many natural environments where prolonged deliberation about a static stimulus is not the norm (e.g. Uchida et al 2006), it is quite hard to guess a priori whether humans will engage in protracted sampling and integration in this case, in practice, even if it is optimal under basic assumptions. As we make clear in our revised Intro, this theoretical interest in the uncertain case of long, noise-free stimuli where perfect, unbounded integration may be optimal but seems doubtful given extant empirical findings, is coupled with a methodological interest in the extent to which neural signatures of decision formation can ‘come to the rescue’ and provide grounds for reliable adjudication between competing mathematical models, when behavioural data fall short.

      More could be said about the optimality of the extrema detection methods. In particular, decades of work (centuries?) have shown that evidence integration is an optimal decision-making procedure: For example, the Sequential Probability Ratio Test is Bayes-optimal wrt mean RT (Wald, 1946); evidence accumulation together with collapsing threshold serves to maximize rewards in repeated choices (e.g., Bogacz et al., PsychRev, 2006; Boehm et al. APP, 2020). Given all this work, why would the brain have evolved to adopt a different mechanism? I realize that the paper is not about optimal decision making, but some discussion of this point seems warranted.

      We had a similar impression initially when reading Stine et al., (2020) where extrema-detection was pitted against integration - is extrema detection so suboptimal that it is too implausible to even consider? Ditterich (2006) argued that signal-to-noise ratio would have to be implausibly high for extrema-detection to produce the behaviour observed on typical decision tasks. However, the fact is, we do not know the effective signal-to-noise ratio, nor can we precisely quantify the costs associated with prolonged evidence accumulation, such as attentional or energetic costs (Drugowitsch et. al., 2012). Even if the extrema detection strategy appears implausibly suboptimal, it is an important principle to demonstrate how not only behavioural but also neural decision signal dynamics can be so nicely consistent with integration yet technically can be quantitatively captured with non-integration mechanisms.

      (2) Modeling choices.

      The authors introduce a parameter, sampT, that represents uncertainty in the sampling onset time. It was not clear to me whether this parameter represented an offset of all trials, or a distribution (probably the latter). I wonder how exactly this parameter was integrated into the models, and in particular, if and how it interacts with the starting-point parameters. My intuition is that on a single-trial, IF early sampling occurs, you can model that with either a negative sampT and z at 0, or with sampT at 0 but a shift in z. This would suggest trade-offs between these parameters, making them hard to estimate independently. Since the paper does not depend on the identification of parameter estimates, this may not be a huge problem, but nevertheless it is good to explore the consequences.

      We thank the reviewer for raising an important question regarding the relationship between sampT and starting-point variability (sz). Mechanistically, early accumulation onset can indeed generate effects that resemble starting-point variability: if accumulation begins during a period containing only zero-mean noise, then by the time informative evidence appears, the decision variable will already have diffused away randomly from zero. In this sense, negative sampT can induce variability in the state of the accumulator at evidence onset. However, the two mechanisms are not mathematically equivalent. The sz parameter assumes a uniform distribution over starting points, whereas the variability induced by early accumulation onset would instead reflect the distribution resulting from integrating zero-mean Gaussian noise over variable durations. Aside from this distinction between distribution shapes, the reviewer is correct that these parameters could partially trade off with one another when sampT takes negative values. In our model, however, sampT was allowed to take either positive or negative values. Positive values delay the onset of evidence integration relative to the evidence, thereby ignoring the first samples, very different from the effect of starting point variability. Nevertheless, to the extent that they can partially trade off each other to some degree, the consequent problem this might cause to accurately estimating both parameters is part of the reason we do not fit a model that includes both simultaneously.

      The way the Bounded Integration model (BIntg) is formulated seems very close to the EZ-diffusion model (Wagenmakers et al., PBR, 2007). This model states that the proportion of correct responses Pc = 1/(1+exp(-B*D/s^2), with B and D the bound and drift rate parameters, respectively. However, filling in the numbers for the high contrast condition from Table 2, and assuming that s=2 (because the model description states that dt=2, with s undefined), I get a Pc of 80% for the 1.6H condition. This seems substantially less than what Figure 2 suggests.

      As we had stated in the Methods section, the model used “a standard deviation of 0.1 arbitrary units” for the evidence. We now make it more explicitly clear that this corresponds to setting within-trial noise s = 0.1 as the scaling parameter (the first paragraph of ‘Model Fits to behaviour only’ and the first paragraph of ‘Integration models’ in Methods). Replacing (s = 0.1) in the suggested calculation yields a predicted accuracy close to 1 for the high-contrast condition, consistent with both the behavioural data and the model predictions shown in Figure 2. We have now clarified this parameter explicitly in the revised Methods and updated Table 1 and Table 2 to avoid confusion.

      On some occasions, it is unclear to me what modeling choices are being made:

      (a) It seems as if the models are fit on accuracy data alone (before introducing the neural data). This seems suboptimal given that the authors do report differences in RT.

      Because of the delayed-report feature of the task, RTs do not directly reflect decision termination time, which is the basis of the use of RT in cognitive modelling normally. Here, whether the decision process has concluded during the stimulus or not, indeterminate response-cue detection and motor execution processes intervene between the stimulus and RT, which would necessitate complicating the models with additional mechanisms, of which there are several possibilities as reflected in the response to the reviewer’s final comment about the RT effects below. This could potentially obscure the core mechanisms of decision formation during the stimulus itself, which was the focus of the study.

      (b) Are the models fit on all data combined, or on the data of individual participants? Fitting individual participant data is preferred, as combined or aggregated data may be distorted by individual differences.

      Because of the noise and variability of EEG data at the single-participant level, we model data averaged across participants, which we have ensured is clear in the revised paper. We provide individual accuracy trends in Figure 1, to verify that the accuracy improvements with increasing evidence duration seen on average are representative of the vast majority of individual subjects. We also added a comment on the limitation this incurs regarding individual difference analysis in the revised discussion.

      (c) The authors seem to suggest that the diffusion coefficient s is estimated (in the section "Integration models"). Most likely, however, this is set to a fixed value. Obviously, it matters for the model comparison using AIC whether this parameter was freely estimated or not.

      As noted in our response to an earlier Comment, the diffusion coefficient was fixed at s=0.1, and to make this explicit, we have entered it for all models in the revised Table 1 and Table 2.

      Not really a weakness, but I wondered about the effect of stimulus duration on RT. In particular, what hypothesis (or post hoc explanation) do the authors have for these RT effects? I could think of at least three hypotheses that are consistent with the behavioral data:

      (a) H1: The shorter the evidence duration, the more likely participants are to require a double-check before response execution, reflecting their uncertainty about their decision.

      (b) H2: There is a collapsing threshold that initiates at stimulus offset, leading to quicker responses on trials where there is more evidence.

      (c) H3: motor preparation is correlated with the evidence signal, which leads to faster responses on trials with more evidence.

      We thank the reviewer for these hypotheses. We agree that the RT effects admit multiple possible interpretations, and while we are cautious not to overinterpret them mechanistically in the paper given our focus on the decision process during the stimulus preceding these response-cue-triggered responses, we do take them to signify that the decision process has not always fully completed and been transformed to a finalised action plan by the time of response cue (start of Results section). To consider these interesting possibilities further:

      We agree that the longer RTs for shorter duration, more uncertain trials could reflect a “double-checking” process (H1), but it could alternatively reflect the fact that if a bound has already been reached during the stimulus, this commitment can be translated to fully-selected action plan that only needs triggering, whereas if a bound has not been reached by stimulus offset, which would occur more often for shorter evidence-duration trials, more of the motor action-selection process would yet need to be completed to initiate the action, causing the slight delay in RT. In other words, on longer-duration trials, the accumulated evidence is more likely to have already reached the bound before the response cue, allowing participants to both commit to a choice and prepare the associated motor response in advance. Therefore, RTs would be shorter as the remaining processes after the cue primarily involve cue detection and motor execution.

      This interpretation is broadly compatible with the reviewer’s H3 account, in the sense that motor preparation may track the evolving decision variable/evidence state. It is also possible that collapsing bounds are set on a post-stimulus, cue-evoked process (H2), which is not mutually exclusive with the above possibilities. It would be hard to determine whether such a process is primarily a response cue-detection decision process that is modulated by uncertainty state at stimulus offset, or a cue-triggered “double-check” process that perhaps operates on the iconic memory of the evidence, and our paradigm does not allow these alternatives to be cleanly dissociated.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see, the reviewers are positive about the work and highlight several strengths, while at the same time offering recommendations for improvement. Once these points are satisfactorily addressed, this may also lead to a revision of the eLife assessment below.

      Reviewer #1 (Recommendations for the authors):

      As outlined in my public review, my most important recommendations relate to sample size and possible neural bases of the extremum-flagging model.

      Minor Comments:

      (1) p. 4 rmANOVA statistic is listed as 12.35.31.

      This has been corrected, with thanks for spotting it.

      (2) The stable d-SSVEP amplitude during the evidence period is used to justify a constant (non-adapting) drift rate assumption. This is a reasonable inference, but the SSVEP reflects early sensory encoding rather than the decision variable per se. Neural adaptation or gain changes at later processing stages could still produce a non-constant effective drift rate even with a stable sensory representation. This inference should be qualified.

      This is true. We have clarified that these checks for one important potential source of a time-varying drift rate, namely adaptation at the level of early sensory representation, but admit that other effects may happen downstream.

      (3) The Methods describe a leaky accumulation extension; the Results note leak ≈ 0.0002 at w=10, which is effectively zero. This is a positive result (evidence against leaky integration) that should be stated more explicitly in the Results or Discussion rather than appearing only in supplementary tables.

      We thank the reviewer for highlighting this. We have pointed to this result now in the first paragraph of the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) The panels in Figure 4 H, I, J are not discussed in the Neurally-constrained models Section, while I believe they are probably more informative than Figure 4E alone.

      The manuscript has been revised to explain Figure 4H,I, J fully under ‘Neurally-constrained models’.

      (2) In the method section, we don't know how many trials were rejected based on the chosen threshold.

      The Method section is updated with the rejected trials after preprocessing. It now reads, “After preprocessing, 12017 trials remained across all conditions, with an average rejection rate of 15% (± 12.9%) across participants.”

      (3) Figure 1: For b and c, data are mean {plus minus} s.e.m. after between-participant variance was factored out. -> reference or detailed method.

      We have now explained in the caption that this is done by subtracting the overall mean of each individual from their data and adding back the grand mean, retaining the between-condition differences - that is, we remove the component of variance that repeated-measures tests ignore.

      (4) Figure 2: Shouldn't the snapshot model only feature one sample, as in Stine et al. 2020? This figure and others would also benefit from a better resolution.

      The reviewer is correct regarding the schematic of the snapshot model presented in Stine et al., (2020). We used small, light orange dots to represent the evidence samples presumably being encoded and a larger orange dot to indicate the single randomly chosen sample used as evidence for the decision. In our revised figure, we have increased the visual distinction between the evidence dots and the selected dot, and pointed this out in the caption. We have also improved the figure resolutions.

      (5) Typo:

      - Semi-saturation Table S4.

      - pi missing in the text of the G^2 equation.

      Thank you for catching these.

      Reviewer #3 (Recommendations for the authors):

      Small, random points:

      (1) How was it ensured that participants indeed did not detect the change in contrast throughout the 1.6s interval? In previous work (Winkel et al., PBR, 2014) we did something similar in a random-dot motion task, but observed that participants always observed the change, if we did not slowly change the coherence of the stimulus (unfortunately, Winkel et al., 2014 is not explicit about the exact parameters of the change, but Figure 2 suggests that the change in coherence lasted 50ms, independent of stimulus strength).

      It is true that abrupt changes in stimulus strength are more salient and detectable than a ramped change. The experimenters tested the stimulus during the task design subjectively, to satisfy themselves that they could not tell when the contrast stepped back to baseline, but this was not verified systematically with psychometrics, nor can we be sure that a more sensitive observer couldn’t sometimes detect the change. However, based on the task design, stimulus properties, and both behavioural and neural data, we are confident that participants are very unlikely to have been sufficiently confident in detecting the step-back in contrast to ceasing their contrast-comparison decision process at that point:

      (1) Participants were naive to the underlying manipulation. They were informed that trials would naturally vary in difficulty. It was normal for them to perceive some trials as harder than others without suspecting a mid-trial structural change.

      (2) The contrast difference in the hard condition was very subtle, and though it is possible that the step-down in contrast could be detected with above-chance accuracy if instructed to do so, given there was no instruction on whether and when the step-down would happen, it is very unlikely they could be detected with sufficient confidence to be certain there is no remaining evidence in the stimulus. Furthermore, the rapid, flickering nature of the stimulus would have helped to mask the transition point, in comparison with a sudden change in a continuously-playing random-dot motion stimulus (as in Winkel et al., 2014).

      (3) Our behavioural post-cue RT data imply the participants did not cease decision formation at evidence offset. If they had, then they would have been afforded the most time to prepare their chosen action in advance of the response cue in the case of the earlier evidence offset (i.e. shorter durations), yet these were the conditions with the longest, not the shortest post-cue RTs.

      (4) The low-contrast CPP traces remained elevated for the full 1600 ms interval, unperturbed by the evidence offsets. If participants were explicitly detecting a sudden change in contrast, we would expect to see a transient evoked response locked to that change, marking that detection. Instead, the sustained elevation of the CPP is characteristic of a continuation of the decision process, uninterrupted, through the subtle offsets.

      (2) Could you include the regression coefficients of the statistical modeling of the behavioral data?

      The regression coefficient is now added to the second paragraph of the Results section.

      (3) I felt Figure 5B was a bit confusing: Are the dashed lines here the ipsilateral sides or the non-linear bounds? This was confusing because "data" only has a solid line in the legend.

      We agree that the subtle nonlinearity, which appears visually close to the linear model, may have caused this confusion. To improve clarity, we have added arrows to Figure 5B to explicitly indicate which legend refers to which panel, and distinguish the dashed nonlinear bounds from the other traces.

      (4) "amplitude variations [...] to be used as an independent evaluation of model fit": Could you refer to where these model predictions are presented? I think this is in Figure 4 - Sup 5?

      We thank the reviewer for raising this. The model-predicted waveforms showing amplitude variations across durations for all neural weightings are presented in Figure 4 - Supplementary Figures 2 and 5. We have clarified this in the revised manuscript.

      References

      Ditterich, J. (2006). Evidence for time‐variant decision making. European Journal of Neuroscience, 24(12), 3628–3641. https://doi.org/10.1111/j.1460-9568.2006.05221.x

      Drugowitsch, J., Moreno-Bote, R., Churchland, A. K., Shadlen, M. N., & Pouget, A. (2012). The cost of accumulating evidence in perceptual decision making. Journal of Neuroscience, 32(11), 3612–3628.

      Ehinger, B. V., & Dimigen, O. (2019). Unfold: An integrated toolbox for overlap correction, non-linear modeling, and regression-based EEG analysis. PeerJ, 7, e7838.

      Latimer, K. W., Yates, J. L., Meister, M. L. R., Huk, A. C., & Pillow, J. W. (2015). Single-trial spike trains in parietal cortex reveal discrete steps during decision-making. Science, 349(6244), 184–187. https://doi.org/10.1126/science.aaa4056

      Loughnane, G. M., Newman, D. P., Bellgrove, M. A., Lalor, E. C., Kelly, S. P., & O’Connell, R. G. (2016). Target selection signals influence perceptual decisions by modulating the onset and rate of evidence accumulation. Current Biology, 26(4), 496–502.

      Nieuwenhuis, S., Aston-Jones, G., & Cohen, J. D. (2005). Decision making, the P3, and the locus coeruleus–norepinephrine system. Psychological Bulletin, 131(4), 510.

      Stine, G. M., Zylberberg, A., Ditterich, J., & Shadlen, M. N. (2020). Differentiating between integration and non-integration strategies in perceptual decision making. Elife, 9, e55365.

      Uchida, N., Kepecs, A., & Mainen, Z. F. (2006). Seeing at a glance, smelling in a whiff: Rapid forms of perceptual decision making. Nature Reviews Neuroscience, 7(6), 485–491.

      Weindel, G., van Maanen, L., & Borst, J. P. (2024). Trial-by-trial detection of cognitive events in neural time-series. Imaging Neuroscience, 2, imag–2.

      Winkel, J., Keuken, M. C., Van Maanen, L., Wagenmakers, E.-J., & Forstmann, B. U. (2014). Early evidence affects later decisions: Why evidence accumulation is required to explain response time data. Psychonomic Bulletin & Review, 21(3), 777–784. https://doi.org/10.3758/s13423-013-0551-8

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers and editors for their time and valuable input into improving our manuscript. The reviewers recognized the value of this work while identifying places where further explanation and/or additional experiments could strengthen the manuscript. We greatly appreciate this feedback and have addressed reviewer comments through additional experiments and necessary textual edits.

      In response to the reviewer comments, our principal data-driven changes to revise the manuscript include: 1) testing how stabilizing HIF-1 in the ADF and NSM neurons affects healthspan; 2) measuring for interactions between HIF-1 stabilization in serotonergic neurons and the mt-UPR; and 3) measuring nlp-17 expression downstream of HIF-1 stabilization in serotonergic neurons. We also made textual changes including: 1) clarifying that our study focuses specifically on the vhl-1-mediated genetic hypoxic response; 2) summarizing which parts of the working model are experimentally validated vs. speculative; and 3) expanding our discussion of future directions needed to understand the epistasis of the many signals acting in this circuit. Our responses to each reviewer comment are below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study by Kitto et al., the authors set out to identify specific signaling components regulating the hypoxic response from the neurons to the periphery and which components are required for lifespan extension. Their previous work had shown that expression of a stabilized HIF-1 mutant in the nervous system extends lifespan through the serotonin receptor SER-7 and leads to the induction of fmo-2 in the intestine. In the current study, they mapped the precise neural circuits required for this response, as well as the signaling mediators. Their work reveals that neurotransmitters GABA and tyramine, and the neuropeptide NLP-17, act downstream of neuronal HIF-1 to convey a "hypoxic signal" to peripheral tissues. Through cell-type-specific expression studies, targeted knockouts, and comprehensive lifespan analysis, the authors provide robust evidence to support their conclusions. The insights gained from the study are both moving the field forward as they advance our understanding of neuro-peripheral hypoxic signaling, but they also lay the groundwork for potential therapeutic strategies aimed at the modulation of such signaling pathways.

      We appreciate the reviewer’s positive assessment of the topic and general interest in this work.

      Strengths:

      (1) This study provides new evidence further delineating signaling components required for hypoxic signaling-mediated longevity, from the nervous system to the periphery. Using a rigorous approach where they express stabilized HIF-1 mutant selectively in ADF, NSM, and HSN serotonergic neurons, followed by cell-type-specific tph-1 knockouts to pinpoint ADF-dependent serotonin signaling as essential for both lifespan extension and intestinal fmo-2 induction.

      This was followed by generating 11 transgenic lines that drive SER-7 expression under distinct neuron-specific promoters, to systematically tease out in which of 27 candidate neurons SER-7 functions to mediate hypoxia-induced longevity. This ultimately highlighted the RIS interneuron as the required signaling hub.

      (2) As the intestine lacks direct neuronal innervation, the authors employ neuron-specific RNAi (TU3311 strain) and dense core vesicle analyses to identify that the neuropeptide NLP-17 is required to transmit the hypoxic signal from RIS to induce fmo-2 in the intestine.

      (3) Overall, the paper is very well written. The experiments were carried out carefully and thoroughly, and the conclusions drawn are also well supported by the results they are showing.

      Weaknesses:

      Overall, I don't see many weaknesses. One point relates to their read-outs, which rely heavily on lifespan measurements and fmo-2 induction without evaluating other physiological processes that serotonin or NLP-17 might affect. For translational relevance, it would be valuable to assess or mention potential adverse effects, such as changes in reproduction, pharyngeal pumping, or proteostasis capacity (proteostasis capacity specifically in the tissue showing fmo-2 upregulation).

      We thank the reviewer for the positive review and fully agree and acknowledge that the primary readouts used in this work, fmo-2 induction and lifespan, may not reflect other important elements of health and physiology. To address this weakness, we have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We focused on examining these HIF-1 stabilized strains rather than our nlp-17 knockout animals because stabilizing HIF-1 in the NSM or ADF neurons is sufficient to extend lifespan. nlp-17 is necessary, but it remains unclear whether nlp-17 signaling is sufficient to extend lifespan. Our new data show that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly in the new thrashing data, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). Lastly, our new data show that ADF and NSM HIF-1 stabilization had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      While we did not examine reproductive capacity in these strains, we have expanded our discussion section to mention the importance of fully characterizing other elements of physiology, health, and behavior in future work.

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [8,80], longevity manipulations can also have adverse effects on reproduction [81,82] and behavior [83,84]. In this study, we find that HIF-1 stabilization in the ADF neurons does not prevent the deleterious effects of hif-1 knockout on mobility, but that HIF-1 stabilization in the NSM neurons may attenuate age-related decline in pumping and thrashing (Fig. S1). However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [41,42,39,40,43].”

      While lifespan assays and fmo-2 expression do provide strong evidence, incorporating additional markers of stress resistance could strengthen the link between hypoxic signaling and organismal health as well.

      We also measured hsp-6 expression via qPCR in the serotonergic neuron-specific HIF-1 stabilized strains to determine whether these conditions that lead to upregulated fmo-2 may also affect the mt-UPR. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G). This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion to emphasize that further investigation of cell- and tissue-specificity within the hypoxic response should be a focus of future work.

      Finally, we agree that it is important to examine whether activating the hypoxic response promotes stress resistance in addition to longevity. We did not focus on this element of the hypoxic response in this work because HIF activity is known to promote adaptive stress-responses to some stressors like infection [7,8] and oxidative stress [9,10], while simultaneously impairing the proteosasis stress response [11]. We have added this important information to our introduction section, and have expanded our discussion section to emphasize that a key future direction will be to test which specific components of this pathway also facilitate stress resistance.

      Introduction Section Modification:

      “However, the physiological changes induced by the hypoxic response are broad and involve adaptations such as increased vascularization, metabolic rewiring, and changes in cell survival pathways. In mammals, some of these same adaptations can be detrimental, as mutations in components of the hypoxic response have been linked to conditions like cancer and cardiovascular disease [12-15]. Additionally, HIF activity is essential to promote some forms of stress resistance to infection [9,10] and oxidative stressors [7,8] but can have detrimental effects on proteostasis [11].”

      Discussion Section Modification:

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [1,2], longevity manipulations can also have adverse effects on reproduction [3,4] and behavior [5,6]. In this study, we find that HIF-1 stabilization in the ADF or NSM neurons has little effect on mobility in young and aged animals. However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [7,8,9-11].”

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to identify the specific neurons, neurotransmitters, and neuropeptides that mediate the longevity effects of the hypoxic response in C. elegans. By genetically dissecting the pathway downstream of HIF-1, they define a neural circuit involving ADF serotonergic neurons, the SER-7 receptor in the RIS interneuron, tyraminergic signaling from RIM, and neuropeptide NLP-17, ultimately linking neuronal hypoxic sensing to pro-longevity signaling in the intestine.

      Strengths:

      The study employs a diverse genetic toolkit, including neuron-specific transgenes, tissue-specific knockouts and rescues, RNAi knockdowns, allowing the authors to pinpoint causality, sufficiency, and necessity with high resolution. The comprehensive mapping of cell-nonautonomous signaling adds depth to our understanding of how HIF and serotonin signaling interface with aging pathways. The conclusions are supported by consistent survival assays and fmo-2 gene expression analyses.

      Weaknesses:

      A key limitation is the lack of clear evidence showing epistasis of so many identified molecular/neuronal components downstream of HIF-1 and serotonin. Thus, the mechanisms of how a diverse set of molecules/neurons coordinate and mediate neuronal HIF-1 effects on intestinal fmo-2 and longevity remain murky.

      We thank the reviewer for these important points. We agree that the epistatic relationships between ADF serotonin, RIM tyramine, RIS GABA, and neuropeptide NLP-17 signaling remain unclear within this pathway. Determining the epistatic relationships of each signal within this complex pathway will require: 1) generating genetic manipulations to each identified signaling component that may mimic vhl-1 knockout to promote longevity; 2) crossing these new strains into multiple genetic knockouts we identified as required for vhl-1 mediated longevity; and 3) measuring the lifespans of each double and triple mutant. We are very interested in testing the epistatic relationships between all of these molecules and neurons, and believe this extensive follow-up exploration will generate significant future results.

      To better address these limitations of the current work, we have added a summary table to Fig. 7 as well as a paragraph to our discussion section. This table and paragraph better explain which components of our working model have been tested for necessity, sufficiency, and epistasis, and which components of this model remain unclear (Fig. 7). See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      Some rescue strategies may inadvertently cause non-physiological expression.

      This is a great point. We bring attention to this limitation in the discussion section. Our cell-specific knockouts (ADF tph-1 KO, Fig. 1E) or ablations (RIS ablation, Fig. 2C) data showed that the neurons identified using rescue strains (i.e., tph-1 in the ADF and ser-7 in RIS) are likely not false positives. However, we did not generate a RIM-specific knockout or ablation strain to confirm our tdc-1 results and have added this limitation to the discussion of these results.

      Discussion of rescue strategy limitations and the need for a RIM-specific knockout:

      “The circuit-mapping approaches employed in this work are also impacted by limitations in cell-specific genetic modifications and in the use of RNAi knockdown. For example, cell-specific rescue constructs can sometimes lead to unintended rescues in other cell types due to cell-nonautonomous signaling. Because all serotonin-producing neurons also express the serotonin reuptake transporter mod-5, serotonin produced by one cell in our rescue strains could be taken up by other serotonin-producing neurons, leading to unintended signaling effects. This may also be true of the uv1 and RIM tyraminergic rescue strains, although little is known about tyramine reuptake in C. elegans. Two cells identified via tissue-specific rescue experiments (ADF and RIS, Fig. 1G and Fig. 2D) were also found to be necessary via cell-specific knockout (ADF tph-1 KO, Fig. 1E; RIS ablation, Fig. 2C), decreasing the likelihood of a false positive from the rescue strain technique. However, the role of the RIM neuron was identified via a tdc-1 rescue strain and was not validated using a RIM-specific knockout (Fig. 7B). Therefore, the contribution of RIM signaling to this circuit is less well-validated, and a RIM-specific tdc-1 knockout strain should be examined in future work.”

      Additionally, environmental hypoxia was not tested in parallel, so the claim on "hypoxia response" throughout the manuscript is not justified by genetic manipulation alone, and the translational relevance of the genetic manipulations remains somewhat uncertain.

      We thank the reviewer for identifying the need for additional specificity in our language. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in Discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Reviewer #3 (Public review):

      Summary:

      This study found that ADF serotonergic neurons have a significant role in extending lifespan mediated by HIF-1, as well as serotonin receptor SER-7 in the GABAergic RIS interneurons. The author focuses on the sufficiency and necessity of components from the central nervous system and how they contribute to aging upon hypoxia.

      Previous work from the lab has identified that the stabilization of HIF-1 in neurons is sufficient to extend lifespan through the serotonin receptor, SER-7, which subsequently activates fmo-2 in the intestine and leads to lifespan extension. Building on this, the author sought to determine which serotonergic neurons are involved and found that serotonin signaling in ADF neurons is required for lifespan extension mediated by HIF-1.

      The author next tested which subset of neurons requires Ser-7 expression to rescue hypoxic response. They found that ser-7 expression in multiple neurons is sufficient to induce fmo-2, with the top candidate being the RIS neuron. Ablation of the RIS neuron did not extend lifespan, suggesting that ser-7 expression in the RIS neuron is required for lifespan extension, positioning it as a key component in the longevity signaling pathway.

      The author also investigated neurotransmitters and found that GABA and tyramine are important components in this circuit. They showed that the tyramine receptor called tyra-3 is required for vhl-1-mediated longevity. Given that tyra-3 is expressed in oxygen- and carbon dioxide-sensing neurons, the author demonstrated that these sensing neurons work downstream of serotonin signaling. Lastly, the author screened neuropeptide/receptor binding pairs and identified NLP-17 as playing a role in hypoxia-mediated longevity.

      Originality and Significance:

      This research is significant in that it uncovers components that are sufficient and necessary for lifespan extension via the hypoxic response. It provides comprehensive data supporting longevity induced by HIF-1-mediated hypoxic response, in conjunction with fmo-2, a longevity gene, as demonstrated in previous work from the lab. Moreover, it provides a number of new transgenic worm tools for C. elegans and aging communities.

      We thank the reviewer for the positive assessment of the manuscript. We appreciate all suggestions for further improving the work and have made changes based on these suggestions (see details below).

      Data and Methodology:

      (1) The experiments were thoroughly conducted, especially the generations of strains using different neuron-type promoters and crossing into mutant strains to demonstrate sufficiency and necessity.

      (2) Some figure legends from the text do not match what the data show. (Figure 6E, F, G).

      We have made changes to the legends accordingly to make sure figures and legends are consistent.

      (3) The lifespan graph legends are confusing and could use some revamping for better clarification.

      We have updated the lifespan graph legends to provide the statistics in the same section as the labels, indicating which line is which condition (see updated Figure 1E). We hope this change better clarifies the lifespan graph legends.

      Conclusions:

      This study provides insights into how hypoxic response regulates aging in a cell non-autonomous manner, outlining a potential circuit involving neurons, neurotransmitters, and neuropeptides.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As suggested by the Reviewers 2 & 3, including environmental hypoxia will broaden the impact of the study. If not, the authors should consider clarifying their response as "genetic activation of the hypoxic response" (see Reviewer 2 below).

      We thank the editor and reviewers 2 and 3 for this clarification. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We have also expanded our discussion of this important future direction.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Suggested minor changes by Reviewer 3 should also be made to improve the readability of the paper.

      We have made changes suggested by Reviewer 3 to improve the readability of the paper.

      Reviewer #2 (Recommendations for the authors):

      (1) Suggestions for additional experiments or analyses:

      (a) To clarify the hierarchical relationships among the components of the identified circuit, epistasis experiments between serotonin, tyramine, NLP-17, and oxygen-sensing neurons (e.g., double mutants or sequential rescues) would strengthen the proposed model and help determine whether these signals act in parallel or downstream of each other.

      We thank the reviewer for identifying this important caveat. As described in the public review response, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. We hope to address these questions about epistasis and interactions between different signals in upcoming projects.

      To clarify that the order of many of these signals remains speculative in our working model, we have added a table to Fig. 7, that summarizes what is known and unknown about the epistasis of these pathway components. We have also expanded our discussion of this limitation. See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron we found to be required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant.”

      (b) Testing whether environmental hypoxia (e.g., 0.5-1% O₂ exposure) elicits similar neuronal requirements and fmo-2 induction as the genetic HIF-1 stabilization would validate that the described pathway is relevant to the actual hypoxic response and improve translational relevance.

      We thank the reviewer for identifying the need for clarification. As also discussed in the public review section, we agree that it is important to emphasize that this paper exclusively examines genetic mimetics of hypoxia, rather than actual exposure to a hypoxic environment. To clarify this point, we have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (ie “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      (c) Functional readouts beyond lifespan and fmo-2 induction (e.g., neuronal activity monitoring or optogenetic modulation of key neurons) could help clarify how information flows through the circuit.

      We thank the reviewer for this valuable suggestion. We are hoping to be able to implement these experimental techniques in future work. We agree that in combination with genetic epistasis analyses, direct measurements of neuronal signaling will greatly improve our understanding of the directionality and interactions between signals in this circuit. Our discussion section recommends employing these techniques in future studies.

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (2) Recommendations for improving the writing and presentation:

      (a) The manuscript is well written overall, but clarity would be improved by explicitly stating in the abstract and introduction that the study is based on genetic activation of the hypoxic response, rather than environmental hypoxia.

      We appreciate this valuable suggestion and have updated the abstract and introduction to clarify that this work examines genetic activation of the hypoxic response.

      Updated sentences from the abstract:

      “Here, we interrogate the cell-nonautonomous signaling pathway downstream of genetic activation of the hypoxic response.” 

      “Together, these insights develop a circuit for how genetic induction of the hypoxic response cell-nonautonomously modulates ageing and suggests valuable targets for modulating ageing in mammals.”

      Updated sentences from the introduction:

      “In this study, we uncover key neural components of the longevity circuit initiated by genetic induction of the hypoxic response. Within this circuit, we identify individual cells, signals, and receptors necessary and/or sufficient to extend lifespan downstream of genetic activation of the hypoxic response. More specifically, we find serotonin signaling in the ADF serotonergic neurons is both necessary and sufficient to extend lifespan through genetic activation of the hypoxic response. This pathway signals through the serotonin receptor SER-7 in the RIS interneuron. We further demonstrate additional neurotransmitters (GABA and tyramine), and a neuropeptide (NLP-17) are critical for mediating these longevity effects. Finally, we identify that oxygen sensing neurons (URX, AQR, PQR and BAG) act downstream of neuronal HIF-1 in this circuit. Our insights into this longevity pathway provide a mechanistic understanding of how genetic activation of the hypoxic response delays aging and improves health.”

      (b) In the discussion, clearly delineating which parts of the proposed pathway are firmly established versus inferred would aid interpretation.

      We thank the reviewer for this idea, and have added the following text to the discussion:

      “Evidence for the necessity, sufficiency, and epistatic relationships between each signal are summarized in new Fig. 7B. In brief, all signaling molecules presented in this work are necessary for vhl-1 to extend lifespan. Rescuing ADF serotonin production, RIS ser-7 expression, and RIM tyramine production is sufficient for vhl-1 to extend lifespan. The sufficiency of oxygen sensing neurons BAG and UPA/PQR/AQR, and the neuronal signals of GABA, NLP-17, and TYRA-3 to restore vhl-1 mediated longevity remains unclear. All signals act downstream of vhl-1. ADF HIF-1 stabilization and SER-7 signaling act upstream of fmo-2 induction, and the oxygen sensing neurons (BAG, UPA/PQR/AQR) act upstream of or in parallel to ADF HIF-1 stabilization.”

      (c) Adding a summary table or schematic that visually distinguishes necessity vs. sufficiency for each component (e.g., ADF, RIS, RIM, NLP-17) would make the overall model more accessible.

      We thank the reviewer for this great suggestion and have added a table to new Fig. 7B that summarizes what is known about necessity for vhl-1, sufficiency for vhl-1, and sufficiency to extend lifespan independently of vhl-1 for each signal (table included in response to Reviewer 2, comment 1a).

      (3) Minor corrections and clarifications:

      (a) Define or replace "hypoxic response" with "genetically induced hypoxic response" where appropriate to avoid conflating genetic manipulations with actual environmental hypoxia.

      We appreciate this valuable suggestion and have replaced “the hypoxic response” and with “genetic activation of the hypoxic response” or “genetically induced hypoxic response” throughout the manuscript to clarify this point.

      (b) All genes and alleles should be italicized per worm nomenclature.

      We thank the reviewer for this comment and have reviewed the manuscript to italicize all gene names and alleles. In some locations, the protein is referred to instead of the gene using the conventional uppercase non-italicized format.

      Reviewer #3 (Recommendations for the authors):

      (1) Using vhl-1 RNAi as the sole approach to demonstrate hypoxic response appears somewhat limited, as vhl-1 is also involved in HIF-1 independent processes that can influence lifespan in C. elegans. Including additional downstream effectors of HIF-1, such as egl-9, or having HIF-1 nondegradable strain as validation could strengthen the findings.

      We thank the reviewer for this valuable comment and agree that an important next step is to test whether these signals are also required for other genetic (HIF-1 stabilized, egl-9) and environmental activators of the hypoxic response to extend lifespan. We have worked to clarify that this paper focuses primarily on vhl-1 mediated longevity throughout the text and have also included this limitation in our discussion section (for details, please see response to Reviewer 2 public review).

      (2) Investigating how the healthspan is affected by serotonergic neuron-specific hypoxic responses would be interesting and could enhance understanding of the physiological mechanisms underlying lifespan extension. 

      We appreciate this suggestion, and have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We found that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). ADF and NSM HIF-1 stabilization also had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      (3) While the experiments were thoroughly performed, the connections between components such as NLP-17, GABA, and tyramine in regulating aging appear critical for establishing a "cell non-autonomous circuit." Additionally, how the potentially antagonistic roles of RIM and RIS neurons influence this axis could be interesting to further explore.

      We thank the reviewer for identifying this important caveat. As described in the public review response to Reviewer 2, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. Our current data showed that intestinal fmo-2 is required for neuronal HIF-1 stabilization to extend lifespan, indicating information must be communicated between the nervous system and the intestine through serotonin, tyramine, NLP-17 and responsible neurons using a “cell non-autonomous circuit” [22]. However, we will continue to address questions about epistasis and interactions between different signals in upcoming projects to fully establish the circuit.

      With respect to RIM and RIS, we value this suggestion and agree that there could be interesting signaling occurring between RIM and RIS in this circuit, as is observed in initiation of reversal behaviors [16,21]. We have updated the discussion section to mention this interesting antagonistic relationship between RIM and RIS signaling in the context of reversal behaviors:

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (4) Does serotonergic neuron-specific rescue impact the mitochondrial unfolded protein response (mtUPR), given that serotonin signaling has been shown to modulate mtUPR?

      We appreciate this question and suggestion. To determine whether activating the hypoxic response in serotonergic neurons modifies the mt-UPR, we measured hsp-6 expression via qPCR in the ADF and NSM-specific HIF-1 stabilized strains. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G. This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion of these results to emphasize that further investigation of this response should be a focus of future work.

      (5) To remain consistent with the flow of Figure 1, the authors should include fmo-2 expression in NSM:HIF-1S in addition to the ADF:HIF-1S (Figure 1I).

      We appreciate this suggestion and have added NSM::HIF-1S data to Fig. 1I. We found that stabilizing HIF-1 in either the ADF or the NSM has a similar effect on fmo-2 induction.

      (6) It would greatly strengthen the RIS observation if the authors demonstrated that ablation of another neural subtype from their screen does not abolish lifespan extension by vhl-1 RNAi. This would be a good supplemental figure, but it is not necessary for the overall story.

      We thank the reviewer for this suggestion and agree that ablating a ser-7 expressing neuron that was not a hit from our screen would be an excellent additional control. We did find that ablating a non-ser-7-expressing neuron (the RIC, Fig. 3E) did not affect vhl-1 mediated longevity, suggesting that impairing the signaling of any interneuron is not sufficient to disrupt the phenotype. In addition, ablating a ser-7 expressing neuron other than RIS would provide much stronger support for this finding. We have suggested this approach to further validate this working model in the future directions section of our discussion.

      “Finally, additional genetic controls could better support the role of RIS-specific ser-7 expression in genetic activation of the hypoxic response. For example, a ser-7 expressing neuron that was not a hit in our screen could also be ablated and tested for necessity in vhl-1 mediated longevity. This experiment would test whether the ability of RIS ablation to attenuate vhl-1-mediated longevity is not a false positive driven by any disruption to ser-7 expression.”

      (7) For consistency with the rest of the manuscript, it would strengthen the hypothesis if modulating the expression of nlp-17 or its receptors impacted the intestinal activation of fmo-2 transcription.

      This is a great point. We attempted this experiment, but were unable to achieve consistent results (see Author response image 1). This result could be due to indirect effects of the overexpression of nlp-17 signaling modulating fmo-2 induction in a complicated circuit, variability in expression of its receptor, or other complexities within the circuit.

      Author response image 1.

      (8) It would strengthen the manuscript to determine whether serotonergic, GABA, and/or tyramine signaling activate the expression or secretion of this nlp-17 neuropeptide.

      We thank the reviewer for this great idea of experiment. To address this suggestion, we performed qPCR to measure nlp-17 mRNA in WT, hif-1 KO, ADF HIF-1 stabilized, and NSM HIF-1 stabilized strains. Compared to WT and hif-1 KO controls, we observed no change in nlp-17 expression when HIF-1 was stabilized in the ADF or NSM serotonergic neurons (new Fig. S5F). This could suggest either that nlp-17 signaling acts in parallel to serotonergic signaling following genetic activation of the hypoxic response. Alternatively, neuronal HIF-1 stabilization may modify nlp-17 splicing or translation without resulting in detectable differences in mRNA levels. Together, these data indicate NLP-17 signaling is required for longevity following genetic activation of the hypoxic response, although whether this peptide is synthesized or released in response to hypoxic response remains unclear.

      We agree that it is also important to connect nlp-17 expression and/or secretion to other components of this pathway. However, we believe the most effective experiment to confirm a connection between GABA and tyramine signaling and nlp-17 in the context of hypoxia would be to measure nlp-17 expression in strains that manipulate GABA and/or tyramine signaling in a manner that mimics vhl-1 knockdown and extends lifespan. Because we have not yet validated hypoxic-response mimetics for these specific signals, we hope to first generate these strains and then measure their effect on nlp-17 expression in future work. The importance of identifying manipulations to GABA and tyramine signaling that promote longevity has been added to our discussion section. 

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      References:

      (1) Calabrese, E. J., Dhawan, G., Kapoor, R., Iavicoli, I. & Calabrese, V. What is hormesis and its relevance to healthy aging and longevity? Biogerontology 16, 693-707 (2015). https://doi.org/10.1007/s10522-015-9601-0

      (2) Zhou, I. K., Pincus, Z. & Slack, J. F. Longevity and stress in Caenorhabditis elegans. Aging 3, 733-753 (2011). https://doi.org/10.18632/aging.100367

      (3) Yuan, R., Hascup, E., Hascup, K. & Bartke, A. Relationships among Development, Growth, Body Size, Reproduction, Aging, and Longevity - Trade-Offs and Pace-Of-Life. Biochemistry (Mosc) 88, 1692-1703 (2023). https://doi.org/10.1134/S0006297923110020

      (4) Mautz, B. S., Lind, M. I. & Maklakov, A. A. Dietary Restriction Improves Fitness of Aging Parents But Reduces Fitness of Their Offspring in Nematodes. J Gerontol A Biol Sci Med Sci 75, 843-848 (2020). https://doi.org/10.1093/gerona/glz276

      (5) Duric, V., Clayton, S., Leong, L. M. & Yuan, L.-L. Comorbidity Factors and Brain Mechanisms Linking Chronic Stress and Systemic Illness. Neural Plasticity 2016, 1-16 (2016). https://doi.org/https://doi.org/10.1155/2016/5460732

      (6) Mariotti, A. The Effects of Chronic Stress On Health: New Insights Into the Molecular Mechanisms of Brain–Body Communication. Future Science OA 1 (2015). https://doi.org/https://doi.org/10.4155/fso.15.21

      (7) Bellier, A., Chen, C.-S., Kao, C.-Y., Cinar, H. N. & Aroian, R. V. Hypoxia and the Hypoxic Response Pathway Protect against Pore-Forming Toxins in C. elegans. PLoS Pathog 5, e1000689 (2009).

      (8) Palazon, A., Goldrath, W. A., Nizet, V. & Johnson, S. R. HIF Transcription Factors, Inflammation, and Immunity. Immunity 41, 518-528 (2014). https://doi.org/https://doi.org/10.1016/j.immuni.2014.09.008

      (9) Vora, M. et al. The hypoxia response pathway promotes PEP carboxykinase and gluconeogenesis in C. elegans. Nature Communications 13 (2022). https://doi.org/https://doi.org/10.1038/s41467-022-33849-x

      (10) Nakazawa, S. M., Keith, B. & Simon, C. M. Oxygen availability and metabolic adaptations. Nature Reviews Cancer 16, 663-673 (2016). https://doi.org/https://doi.org/10.1038/nrc.2016.84

      (11) Fawcett, M. E., Hoyt, M. J., Johnson, K. J. & Miller, L. D. Hypoxia disrupts proteostasis in Caenorhabditis elegans. Aging Cell 14, 92-101 (2015). https://doi.org/https://doi.org/10.1111/acel.12301

      (12) Ohh, M., Taber, C. C., Ferens, F. G. & Tarade, D. Hypoxia-inducible factor underlies von Hippel-Lindau disease stigmata. Elife 11 (2022). https://doi.org/10.7554/eLife.80774

      (13) Wind, J. J. & Lonser, R. R. Management of von Hippel-Lindau disease-associated CNS lesions. Expert Rev Neurother 11, 1433-1441 (2011). https://doi.org/10.1586/ern.11.124

      (14) Lonser, R. R. et al. von Hippel-Lindau disease. Lancet 361, 2059-2067 (2003). https://doi.org/10.1016/S0140-6736(03)13643-4

      (15) Kaelin, W. G. Molecular basis of the VHL hereditary cancer syndrome. Nat Rev Cancer 2, 673-682 (2002).

      (16) Costa, S. W. et al. A GABAergic and peptidergic sleep neuron as a locomotion stop neuron with compartmentalized Ca2+ dynamics. Nature Communications 10 (2019). https://doi.org/10.1038/s41467-019-12098-5

      (17) Wu, Y., Masurat, F., Preis, J. & Bringmann, H. Sleep Counteracts Aging Phenotypes to Survive Starvation-Induced Developmental Arrest in C. elegans. Curr Biol 28, 3610-3624.e3618 (2018). https://doi.org/10.1016/j.cub.2018.10.009

      (18) Busack, I. & Bringmann, H. A sleep-active neuron can promote survival while sleep behavior is disturbed. PLOS Genetics 19, e1010665 (2023). https://doi.org/10.1371/journal.pgen.1010665

      (19) Powell-Coffman, J. A. & Coffman, C. R. Apoptosis: Lack of oxygen aids cell survival. Nature 465, 554-555 (2010). https://doi.org/10.1038/465554a

      (20) Kruempel, J. C. P. et al. Hypoxic response regulators RHY-1 and EGL-9/PHD promote longevity through a VHL-1-independent transcriptional response. Geroscience 42, 1621-1633 (2020). https://doi.org/10.1007/s11357-020-00194-0

      (21) Bach, M., Bergs, A., Mulcahy, B., Zhen, M. & Gottschalk, A. (2023).

      (22) Leiser, S. F. et al. Cell nonautonomous activation of flavin-containing monooxygenase promotes longevity and health span. Science (2015). https://doi.org/10.1126/science.aac9257

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      We sincerely apologize for the insufficient image clarity and anatomical orientation in the original Figure 3I. To resolve this issue, we have performed the following revisions in the revised Figure 3I:

      (1) Replaced the original panels with the high-resolution confocal images showing clear c-FOS immunofluorescence in the LPGi.

      (2) Included explicit anatomical orientation indicators (Bregma −6.75 mm) to clearly demarcate the boundaries of the LPGi.

      (3) Added ROI outlines surrounding the LPGi region.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      We thank the reviewer pointing out this oversight. We have updated the Methods section under "Animals" and specific procedure subsections to clearly state the exact age of animals.

      (1) For retrograde trans-synaptic PRV tracing, 8-week-old mice received intrahepatic viral injections and were sacrificed 5 days post-injection.

      (2) For chemogenetic and optogenetic manipulations, stereotaxic AAV injections were performed at 8 weeks of age. Mice were allowed 4 weeks for viral expression and recovery before undergoing metabolic tests or light stimulation at 12 weeks of age.

      (3) For chemical denervation (6-OHDA), 8-week-old mice were injected into targeted lobes and examined 7 days post-denervation.

      (4) For postnatal innervation mapping, neonatal mice at postnatal week 0 (P0), week 1 (P7), and week 2 (P14) were harvested for tissue clearing.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      We appreciate the reviewer for this feedback. We agree that our original description lacked precise anatomical localization regarding the pre- and postganglionic neurons, and it was inaccurate to state that descending fibers pass through the sympathetic chain (SyC).

      Based on our whole-mount tissue clearing data, we observed that the preganglionic neurons of the brain-liver sympathetic circuit are primarily located in the T6–T12 segments of the thoracic spinal cord. Accordingly, we have revised the text in Results 4 to specify these exact locations.

      Manuscript Revision (Results 4):

      "Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the thoracic spinal cord (T6–T12) send descending fibers via the splanchnic nerves to innervate postganglionic neurons in the CG-SMG (Figure 4A)."

      Furthermore, following your valuable suggestion to compare our findings with existing literature, we reviewed a recent study published in Nature Communications (Harima, Yukiko et al. Parallel labeled-line organization of sympathetic outflow for selective organ regulation in mice. Nat Commun. 2024;15(1):10478). In that study, researchers injected retrogradely transducible AAVs directly into the CG-SMG and traced the preganglionic neurons predominantly to the T8–T13 segments. Their results are largely consistent with our findings. Interestingly, the broader anatomical range observed in our trans-synaptic liver-to-brain mapping (T6–T12) compared to their CG-SMG-specific tracing (T8–T13) reveals a slight discrepancy. This observation suggests an intriguing anatomical hypothesis: a subset of sympathetic preganglionic nerves may bypass the CG-SMG relay entirely and project directly to the liver.

      (4) Figure legends should be revised and matched with the text.

      We apologize for the oversight. We have conducted a comprehensive audit of all figure and legends to ensure precise matching between the main text and the figures.

      Specifically, we have corrected a typographical error in the Figure 1 Legend where panel (C) was mistakenly labeled as a second panel (B), and we fixed a spelling error ("LPG" corrected to "LPGi"). Additionally, we corrected a miscitation in Results (Section 3) regarding Figure 3. In the original text, Figure 3C was incorrectly grouped with blood glucose data, whereas it actually displays the Western blot validation of sympathetic denervation.

      We have revised the corresponding sections in the manuscript as follows:

      Manuscript Revision (Figure 1 Legend):

      “(C) Quantification of PRV-labeled neurons in left and right LPGi across different hepatic lobes: left lateral, median, right posterior, right anterior, caudate, and porta hepatis (n = 3).

      (D) Sankey diagram showing projection patterns from left and right LPGi to individual hepatic lobes. (E and F) Representative slices of EGFP+ and mRFP+ neurons in left (top) and right (bottom) LPGi following PRV-EGFP (right anterior lobe) and PRV-mRFP (median lobe) injections. Proportions of EGFP+, mRFP+, and co-labeled neurons in left and right LPGi (F, n = 3). Scale bars, 100 μm.”

      Manuscript Revision (Results 3):

      “Despite the absence of directly sympathetic input to denervated lobes, systemic blood glucose levels were unchanged compared with controls (Figures 3A-3B, Figure S5A), indicating functional compensation through the remaining intact liver.”

      Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      We sincerely thank the reviewer for their time and evaluation. We respectfully note that these comments mirror those raised during the previous round of review. We would like to kindly direct the reviewer to the extensive revisions we implemented in our previous resubmission, which directly and comprehensively addressed these exact concerns. These revisions remain intact in the current version of the manuscript. Below, we briefly summarize how each point was previously addressed for your convenience.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      As addressed in our previous revision, we fully agree with this suggestion. We updated the title of the manuscript to explicitly include the species: "Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice." We also clarified the species used throughout the main text to ensure accuracy.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      As detailed in our previous response, we acknowledge this important limitation. In our prior revision, we addressed this concern through both additional data analysis and text revisions:

      (1) We provided SyGlass 3D reconstruction data demonstrating that the TH-positive nerve fibers originate from the celiac-superior mesenteric ganglia (CG-SMG), a well-established sympathetic ganglion (Figure S5F).

      (2) In parallel, we collected dorsal root ganglia (DRG) from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While the T7-12 DRG segments are historically known to contain the sensory neurons that innervate the liver (Anat Rec A Discov Mol Cell Evol Biol. 2004; Auton Neurosci. 2024), we detected only a remarkably sparse number of PRV-positive neurons in these segments. This effectively functionally distinguishes this efferent pathway from primary sensory afferents (Supplementary figure B).

      (3) We explicitly added this methodological limitation to the Discussion section (paragraph 6) of the current manuscript, noting that more selective approaches, such as genetic targeting of sympathetic lineages, will be important for future validation."

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      As outlined in our previous response, we incorporated this crucial context into our revised manuscript. Specifically, we expanded the Discussion section (paragraph 3) to contrast our precise cell-type-specific chemogenetic and optogenetic approaches with historical studies that relied on coarse electrical stimulation. This addition highlights how our current methodology reveals the contralateral and lobe-specific architecture of brain-liver sympathetic control that was previously obscured.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We appreciate this insightful physiological perspective, which we addressed comprehensively in our previous revision. As we previously agreed, the central nervous system integrates a broad range of afferent signals, and compensatory sensory or parasympathetic mechanisms likely contribute to the observed LPGi activation following hepatic sympathetic denervation.

      To address this, we significantly expanded our Discussion section (paragraph 4) in the prior revision. We explicitly proposed a model wherein hepatic glucose production is regulated by an integrated afferent-central-efferent loop, acknowledging that our current study primarily resolves the efferent component. We clearly noted the lack of direct assessment of sensory or parasympathetic innervation as a limitation and highlighted this dynamic crosstalk as a critical avenue for future investigation.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

      Recommendations for the authors:

      Reviewer #4 (Recommendations for the authors):

      No further recommendations aside from tempering the CGRP language, as marking all sensory fibers.

      We appreciate the reviewer for pointing out this important anatomical distinction. We entirely agree that CGRP specifically labels peptidergic sensory afferents and does not represent the entirety of the sensory nervous system.

      We have carefully reviewed the entire manuscript and tempered our language accordingly. Wherever CGRP is mentioned, we have clarified that it serves as a marker for peptidergic sensory fibers, rather than functioning as a pan-sensory marker.

      Manuscript Revision (Results 1):

      "Unlike the NTS, a well-established hepatic sensory center served here as a positive control, the LPGi contained few CGRP-positive cell bodies (Figure S1G), indicating a lack of peptidergic sensory projections."

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are grateful that both reviewers found the study to provide a strong behavioral framework for defining sleep in Aedes aegypti and appreciated the breadth of the behavioral and genetic approaches used. We also appreciate the reviewers’ careful identification of several issues requiring clarification, particularly regarding the interpretation of post-blood-meal sleep, the support for the 10-min sleep threshold, possible nutritional confounds in the BSA experiments, the framing of the host-seeking model, and the description of statistical analyses. In the revised manuscript, we have addressed these concerns by clarifying our rationale, tempering several conclusions, revising the statistical reporting and methods, explicitly stating sample sizes, and expanding the Discussion to better acknowledge limitations and alternative interpretations. Where appropriate, we have also revised the text to distinguish more clearly between increased sleep and reduced locomotion, and to frame mechanistic conclusions more cautiously.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Conventionally, a coincidence of sleep increase and locomotion reduction would weaken the certainty of a sleep increase assessment. The authors implied this concurrence observed after blood meal is derived from internal "drowsy" neural state instead of physical "cripple", but they did not use their two high-resolution video tracking velocity or pDoze/Wake to clarify this.

      Thank you for addressing this point. We understand the need to validate locomotion when used as a readout of sleep. We note that analysis of waking activity is normalized to time spent awake, and therefore should be separate from the time spent inactive that is classified as sleep. Based on the reviewers’ suggestions we have reanalyzed some data and revised the relevant sections in include this analysis.: In brief we performed pDoze/pWake analyses on the two high-resolution tracking video from EthoVision XT system. A velocity threshold of 0.4 mm/s was used, with velocities above 0.4 mm/s defined as wake/activity and velocities below 0.4 mm/s defined as doze/sleep state. pWake and pDoze were defined as proportional time metrics of wake/active (velocity > 0.4 mm/s) and doze/sleep (velocity > 0.4 mm/s) within each LD cycle. The conclusion that sleep is increased following blood feeding is supported by these data. We also note (as described in response to Reviewer 2, that this paper represents a step towards describing sleep in mosquitoes. We hope that future application of approaches used in Drosophila, such as brain imaging and indirect calorimetry will further refine our understanding. Along these lines, we have also included a section in the Discussion about how additional measures, including systems like FlyVista might be applied in the future.

      (2) The major molecular component underlying blood meal effect on sleep/locomotion is less certain, because the BSA solution used for feeding contains ATP, which itself is able to enter haemolymph and potentially exerts sleep/locomotion effect. Additionally, the basal or control sleep recording is done after sucrose feeding. It is, however, unclear from the method if this is 10% too? And if the observed sleep level increase after a blood meal is a result of sugar level reduction in the blood (~0.1%).

      We thank the reviewer for raising this important issue. We think it is unlikely that the small amount of ATP used for feeding is driving the sleep phenotype. We have now included this point as a caveat within the discussion, and explained its inclusion.

      (2) Sucrose concentration in controls

      We apologize that this was not clearly stated. Yes, the control mosquitoes were maintained on 10% sucrose, and we have now clarified this explicitly in the Methods and figure legends where relevant.

      (3) Could the effect reflect reduced sugar intake rather than blood/protein?

      We think this is unlikely, however it cannot be ruled out based on the experiments we have run. We have added discussion of this point. However, we note in fruit flies, this has been studied extensively, and loss of sugar under certain contexts reduces sleep. The points above highlight the need for systematic analysis of the dietary components that contribute to sleep in mosquitoes. While we regret being unable to include them in this manuscript, we note that many of these experiments are challenging (with many controls) and have been ongoing for over a decade (with contributions from many labs) in Drosophila.

      Reviewer #2 (Public review):

      (1) The authors settle on a 10-minute immobility threshold, but their own data do not convincingly support this choice… A 15-minute threshold would be better supported by the data as presented.

      We appreciate this evaluation of the sleep threshold. We chose 10 minutes because the first significance in arousal threshold is at the time-point of 10-15 minutes. Therefore, we believe that sleep bouts longer than 10 minutes should be qualified as sleep. We are particularly interested in why arousal threshold continues to increas at 15 minutes. This is either incomplete sleep between minutes 10 and 15 or the presence of multiple sleep states. We have established a new system in the lab using Zantiks that we believe will allow for simultaneous recording of posture and arousal threshold. We now explicitly comment on this in the discussion, and the need for further analysis of the timeframe for which sleep is defined. Nevertheless, we believe we have honed in on a period of 10-15 minutes that serves as a good proxy for sleep regulation. We hope that this initial description of sleep in mosquitoes provides an initial step towards defining sleep, and that future studies that include techniques applied in Drosophila including brain imaging, indirect calorimetry and additional videography will define more nuanced changes in sleep. We have written in limitations and future opportunities to better define sleep throughout the manuscript.

      (2) The primary experimental paradigm measures sleep beginning at Day 4 post-blood feeding, immediately after oviposition... what is being measured as ‘sleep’ could reflect post-reproductive quiescence or recovery rather than diet-induced sleep per se. The BSA experiment partially addresses this, but since BSA also triggers vitellogenesis and egg production, the confound persists.

      We agree this is an important concern. Our intent in measuring sleep after oviposition was to isolate prolonged post-feeding effects from the well-established transient suppression of host-seeking that occurs during the first ~72 h after blood feeding. However, as the reviewer notes, this design does not by itself distinguish post-feeding sleep from other physiological processes associated with reproduction, including vitellogenesis, oviposition, or post-reproductive recovery. To address this issue, we included the experiment measuring sleep immediately after blood feeding, before oviposition. We agree, however, that this rationale should have been stated more clearly and that the limitation remains relevant, particularly because BSA can also support egg development. In the revised manuscript, we have therefore: In the current version we have clarified more explicitly that the immediate post-blood-meal recording was included to show that the sleep increase begins before oviposition; We have also tempered our interpretation of the Day 4–5 phenotype to avoid implying that it is purely diet-driven and fully independent of reproductive state; and expanded the Discussion to acknowledge that blood feeding, protein feeding, and reproductive physiology are closely linked in female mosquitoes and that our current experiments do not fully disentangle these processes. These changes frame the data more cautiously: blood/protein feeding is sufficient to induce a sleep-promoting state that begins immediately after feeding and persists into the post-oviposition period, but the relative contributions of nutrient sensing, egg development, and reproductive recovery remain to be determined.

      (3) The opportunistic vs. determined host-seeking hypothesis… requires actual measurement of host-seeking alongside sleep to be substantiated, or at least the caveats need to be discussed more explicitly.

      We agree with the reviewer. Our intention was to present this as a conceptual model motivated by the temporal dissociation between published host-seeking recovery and the prolonged sleep phenotype observed here, not as a demonstrated behavioral framework directly tested in this study. In the revised manuscript, we have substantially softened this section by clarifying that we did not directly measure host-seeking behavior in the current study; adding explicit caveats that the proposed framework remains speculative until sleep and hostseeking are measured simultaneously in the same animals across the same post-feeding time course. We appreciate this comment and agree that the distinction should be presented as a model for future testing rather than as a central conclusion established by the current data.

      (4) The methods describe ‘one-way ANOVA, followed by Mann-Whitney tests with Welch’s correction,’ which is an internally inconsistent combination…

      We thank the reviewer for catching this lack of clarity. We apologize for this inconsistency. We have fixed this error. In the revised manuscript, we have carefully rewritten the statistical analysis section to specify: which datasets were analyzed using parametric tests (e.g., ANOVA, with appropriate post hoc comparisons where assumptions were met), which datasets were analyzed using non-parametric tests (e.g., Mann-Whitney), and where Welch’s correction was applied, specifically for unequal-variance t-tests, not Mann-Whitney tests. We have also revised Methods, Figure legends and reporting throughout to ensure that the statistical test named in the text matches the reported test statistics. The changes include statistical methods rewritten for consistency and accuracy, and updated figure legends that include exact sample sizes. In addition, one summary spreadsheet of statistical analysis throughout this study is provided and will be submitted as a supplementary file.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear whether there are any systematic changes in preferences over the course of testing that could explain the observed changes in correlation with neural responses, such as changes due to learning (e.g., flavor nutrient conditioning, relief of neophobia), changes in deprivation state, or habituation to/proficiency with the BAT setup.

      For the revision, we have added analysis, including a new figure (Figure 3) between what are now Figures 2 & 4, testing the hypothesis that preference changes across testing days are non-random in direction (e.g., that they reflect attenuation of neophobia). This new analysis failed to reveal evidence supporting the hypotheses that: 1) preference for palatable tastes increases with experience (a result that would make sense given research on neophobia; 2) the preference for aversive tastes decrease with experience; or 3) absolute consumption of any particular taste changes in a reliable direction from session to session (lines 142-157 and new Figure 3).

      A secondary point is whether any changes in preference are attributed to internal individual versus external contextual factors. Both types of variation (i.e., across individuals and across time within an individual) are mentioned in the introduction, but it is not clear what the authors believe about the nature or neural representation of these sources of variation.

      While we assume that differences between rats are due to internal factors (given the controlled home-cage environment), we can’t be sure that some subtle, subthreshold (for us as observers) factor impacts taste preferences. Similarly, while changes across time within an individual is categorically within the individual, we cannot be sure whether some subtle facet of their experiences determines how preferences change (as opposed to it being purely internal). We have added prose to the Discussion session on this topic—including citation of Hilary Schiff’s recent work showing nurture-related preference changes as part of this new prose (lines 387-398).

      With respect to neural data analysis, no individual animal/day data are shown, making it difficult to assess the extent to which differences in correlation match individual differences in preferences and/or changes in preference with time within individuals.

      The revision now explicitly includes Figure panels (with analysis) showing the relationships between individual neural responses and consumption in the first and last BAT tests for a representative rat (lines 172-198; Figures 4A and 4D). As requested in the non-public comments, we have also added waveforms recorded for the representative neuron in an inset to Figure 4B.

      The correlation analysis is also lacking control for the fact that there is a certain degree of "chance" associated with behavioral and neural measures having matching ranks.

      Certainly chance cannot explain our results, which consist centrally of within-rat differences in match (that is, regardless of chance match levels, what we observed was specifically an enhancement of that match for the most recent behavioral assessment compared to an earlier assessment in the same rat)—a finding that is all the more surprising given that: 1) 2 weeks separate that behavior test and the electrophysiology session; and that 2) that gap between the ephys test and the (less well-matched) first behavioral test is only 1-3 days longer. Nonetheless, in appreciation of Reviewer 1’s concern, we have added an independent, convergent analysis to the revision, testing whether the observed pattern vanishes when we shuffle the preference ranks between tastes with neighboring ranks in the behavioral data (a more conservative test than complete shuffles among tastes). The results of this analysis, which are in the new Figure 5, provide further proof that our result is not based on chance—that they specifically reflect a match between neuronal activity and behavior (lines 242-251).

      Finally, …it is unclear to what extent changes in correlation may be attributed to overall changes in responsiveness of the neural population.

      We include several new analyses in the revision that test the hypothesis that the reduction in match between behavioral rankings and neural responses in the second electrophysiology sessions reflects spontaneous or taste-driven changes in neural excitability. These additional analyses reveal no clear between-session differences in baseline and/or taste-evoked responses, or in the percentages of neurons that are taste responsive and/or palatability-related (lines 292-309; Figure 7).

      Reviewer #2 (Public review):

      The manuscript could use additional corollary analyses to provide a more complete picture of the phenomenon. For instance, how many neurons (per animal and in total) have significant correlations with the final BAT patterns? And with the first BAT? Can a time course of such counts be provided? Can some decoding analyses be performed at a single session level to reconstruct a rat's behavioral preference pattern from its neural activity?

      These are all really good ideas. As noted in our response to Reviewer 1, we have implemented all but the last of the suggested analyses, which did not produce evidence suggesting that our results can be explained by changes in neuronal properties between the two recording sessions (lines 292-309; Figure 7). We have also made attempts to apply the decoding analysis; unfortunately, we don’t have large enough samples to obtain stable results such a subtle decoding task (reflecting the last BAT session’s preference pattern is significantly better than the first session’s pattern).

      The manuscript could benefit from additional polishing, both in the text as well as in the figures.

      An extensive holistic edit has been done, starting with suggestions made by Reviewer 2 in the non-public comments.

      Reviewer #3 (Public review):

      Without a behavioral measure collected after recording day 1 intraoral exposure, it is not possible to determine whether taste preference was altered by that experience…The authors' conclusion would be strengthened by adding an intervening brief access test between recording days 1 and 2.

      We very much appreciate Reviewer 3’s suggestion. Alas, the primary authors involved in data collection on this project have moved on, and we won’t be able to collect the additional dataset that would be required. Instead, we have softened the conclusion that we reached in the last section, and suggested the proposed experiment as a future direction (lines 366-374).

      The current experimental design exposes animals to 3 distinct sets of substances … [that] differ in identity … and concentration. Because palatability is known to be comparative depending on the other substances available and concentration-dependent, this introduces challenges to interpretation, [and] without more clarity, it is difficult to evaluate whether the interaction of different tastes within the sets of stimuli biases the main conclusions.”

      This is an interesting point. Analyzing each set of batteries separately and performing between-battery comparisons would require a larger number of experimental subjects then we have in our current sample size. That said, while we acknowledge that taste preference ranking is relative, we believe the ranking system used here deviates little, if any, from the 'true' ranking (and is therefore significantly relevant to gustatory activity). This is supported by our newly obtained result in response to Reviewer 1 & Reviewer 2 (see above), where an ancillary shuffle analysis (Figure 5C) showed that swapping adjacent preference orders eliminated the experimental effects across all batteries.

      Responses to sweet tastes are not reported in the electrophysiology data. This is seemingly the case because rats given set 1 received no sweet stimulus while rats given set 2 received to 2 distinct sweet tastes. Finally, rats given set 3 did not receive quinine, yet quinine is reported in electrophysiology data.

      We are unsure of the source of this confusion—in every case, the rat received the same tastes in the electrophysiology sessions that were delivered in the BAT preference tests—but in appreciation of Reviewer 2’s concern, we have modified the text and table to ensure: 1) that panels reflecting data from single example rats (panels that therefore necessarily include only a subset of possible tastes) are clearly marked as such; and 2) that the nature of which taste batteries were delivered is more explicit (lines 104-112; 172-178).

      The choice of reporting average lick cluster size is problematic because the authors use thirsty rats with 10-second-long trials. Thirsty rats are likely to lick in relatively long clusters, especially for neutral and palatable tastes. If the rat is mid-cluster when the trial ends, the final cluster would be cut off prematurely, resulting in shorter overall average lick cluster size, disproportionately affecting neutral and palatable tastes over aversive tastes.

      We have ourselves been deeply concerned with this issue, and in fact have recently published a paper that includes within it a direct test demonstrating that calculations of lick bout lengths from 10-sec BAT trials result in taste palatability estimates that are identical to (and less noisy than) those generated from more classically-used 15-min ad lib licking. We now cite this paper (Stone, Lin, et al., 2026) in the Methods section, along with text clarifying how we calculated lick clusters. We also conducted an additional analysis that estimates taste preference after removing these “prematurely ended bouts” without changing the observed pattern of results (lines 494-510).

      Of course, even if this last analysis had changed things, the result of clusters being cut short by the end of a trial would be an underestimation of the preference for the palatable tastes (which drive far more licking than aversive tastes and are therefore more likely to be mid-bout at the end of a trial). Such an underestimation would in turn be expected to reduce the observed neural-behavioral correlation. This fact highlights the robustness of our findings.

      Canonical palatability rankings may not apply to the concentrations selected in every stimulus set. This is particularly true for set 1, which included two concentrations of citric acid and quinine for the behavior. It is also not clear which concentrations are reported in Figures 3A2 and 3B2. Meanwhile, the concentrations of quinine and citric acid used for electrophysiology are quite low.

      In the revised Methods section, we explicitly motivate our reasoning (including citations) behind canonical rankings for each taste battery used (lines 513-522). Every taste used was of agreed-upon preference levels, and in the rare case that two concentrations of the same taste were used, both were known to have distinct palatabilities (e.g., 0.1M NaCl is preferred to 0.05M NaCl). This careful selection of tastes ensured that it was trivial to avoid misordering of canonical palatability rankings.

      And even if mistakes in canonical rankings had been made, the impact of these inaccuracies in these rankings would have been minimal. Our findings are primarily driven by high levels of inter-individual (between different rats) and intra-individual (day-to-day fluctuations within the same rat) preference differences. Given this variability, the fact that the brain-behavior correlation were consistently worse using these rankings almost certainly means that the neural activity matches preference behavior—our thesis. This conclusion is further supported by our shuffle analysis, which demonstrated that randomizing the order did not yield superior correlations between taste ranking and GC activity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [<sup>3</sup>H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we have also explicitly stated this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we mention and describe alternative approaches for assessing parasite viability. This also includes the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We have revised the text and now explicitly mention the use of dual staining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al., 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non-commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We have clarified this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We have added this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post-antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We have revised the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      The authors observed differences between their i-lacZ assay and conventional PRR assays attributable to the accumulation of lacZ enzyme at higher levels of surviving parasites, followed by degradation. Have the authors tested how long lacZ remains stable in standard or overgrown parasite cultures?

      We thank the reviewer for this important question. We assessed the stability of β-galactosidase activity in parasite lysates stored under different conditions and observed that the enzymatic activity remained stable for up to 21 days when lysates were stored at either -20°C or 37°C (Hellingman et al., 2024).

      We have not systematically characterized β-galactosidase stability in standard or overgrown parasite cultures. However, in experiments involving overgrown cultures, we observed that the β-galactosidase-derived signal decreased rapidly in overgrown culture settings, suggesting that enzyme stability in overgrown cultures is lower than in standard cultures and parasite lysates.

      At what parasitemia were the counts reported in Figure 1D conducted at?

      We thank the reviewer for this clarification request. The measurements shown in Figure 1D were performed at approximately 3% parasitemia, assuming an erythrocyte infection rate of 10-fold within 48 hours as parasite cultures were initiated at 0.3% parasitemia and incubated for 48 hours under rapamycin before the measurements were performed.

      Figure 2: Is the increasing background in DMSO-treated parasites attributable to leakage of the di-cre system contributing to a background level of LacZ induction? To what extent would this impact results in the PRR assay format?

      We thank the reviewer for this important observation. We agree that low-level leakage of the loxP-DiCre system may contribute to the increased lacZ signal observed in DMSO-treated parasites under overgrowth conditions. However, this effect was only observed when parasites were allowed to proliferate extensively in the absence of effective drug pressure.

      In the context of the MULT-i<sup>2</sup> assay, these conditions correspond to compound concentrations that do not affect parasite survival or replication. Such concentrations are outside the range of interest for evaluating antimalarial activity, as they represent inactive treatment conditions. Therefore, although reporter leakage may contribute to background signal under extreme overgrowth conditions, we expect this effect to have a negligible impact on the interpretation of MULT-i<sup>2</sup> assay results.

      Please define the abbreviations used (e.g., NONMEM and DV).

      We thank the reviewer for pointing this out. We have revised the manuscript to define all abbreviations at their first occurrence in the text and have added the relevant terms to the abbreviation list.

      Line 297: cPRR assay - give citation.

      We thank the reviewer for pointing this out. We have added the appropriate citation for the cPRR assay at the indicated location in the revised manuscript.

      The 2 in MULT-i<sup>2</sup> is not always superscripted.

      We thank the reviewer for pointing this out. We have corrected the formatting throughout the manuscript to ensure that the “2” in MULT-i<sup>2</sup> is consistently presented as a superscript where appropriate.

      Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the MULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We have added more explanations to the Discussion including the strengths and limitations.

      Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The data are high-quality and broadly consistent with the interpretation, but alternative mechanisms for the defects are not considered; additional work should rule out vesicular trafficking, discuss other mechanisms, and address methodological issues.

      We thank the editor and reviewers for assessing our work. We have now included additional experiments in a new main Figure 2, which directly rule out any general or Orai1 plasma membrane trafficking defects in Syntaxin11-depleted cells. There are additional experiments and/or analysis in many other figures, throughout the paper. We have included new and missing methods, quantifications and calibrations, and provided response to each of the reviewer’s comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Patients with STX11 mutations develop familial hemophagocytic lymphohistiocytosis Type 4, a fatal immune disorder marked by defective T and NK cell cytotoxicity and cytokine storm. The conventional explanation attributes this to impaired cytotoxic granule release, but this has never fully accounted for the broader disease picture. This study proposes an alternative mechanism. The authors show that STX11 is required for store-operated calcium entry through ORAI1 channels, which are essential for both cytotoxic killing and NFAT-driven gene expression in T cells. In STX11-deficient cells, ORAI1 currents drop, NFAT nuclear translocation fails, IL-2 expression is suppressed, and degranulation is impaired. These defects are largely rescued by ionomycin or a constitutively active ORAI1 mutant, placing the primary lesion at calcium signaling rather than the fusion machinery. Mechanistically, STX11 binds the C-terminal tail of ORAI1 via its Habc domain and maintains ORAI1 in a state competent for productive assembly prior to STIM1-dependent gating, a step the authors call "priming."

      Strengths:

      The paper identifies a novel and disease-relevant role for STX11 in calcium channel regulation and raises the possibility of using channel agonists as a therapeutic strategy in the disease. The biochemical and functional data are of high quality and generally consistent with the interpretation. The proposal that a non-conventional syntaxin directly interacts with ion channels to prime its activation is novel and interesting.

      Weaknesses:

      For readers to appreciate the value of patient experiments derived from a single individual, the authors should quote prior studies showing that STX11 protein levels are abolished in all known human STX11 mutations. The priming model, while functionally well-supported, rests on indirect structural evidence, and the precise conformational transition involved remains to be defined. These are acknowledged limitations, but alternate mechanisms have not been explored and formally excluded. More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional ORAI1 channels from vesicles.

      In the revised version, we have included references for all those prior STX11 human mutations that have been biochemically characterized till date. The reviewer has correctly pointed out that STX11 protein levels were almost abolished in almost all previously reported mutations. See line 168-173. Therefore, the prior STX11 patient mutations are essentially comparable to the frameshift mutation characterized in this study, in terms of STX11 protein depletion and, therefore, the mechanisms underlying the phenotypic defects reported here as well as earlier. We, therefore, believe that our data from even a single FHLH4 patient, with severely depleted STX11 levels, and additional knockdown studies across three different cell lines, are representative of majority of STX11 mutant FHLH4 patients that have been previously characterized.

      Regarding the Reviewers’ concern that absence of STX11 as a conventional SNARE could affect Orai1 channel delivery from intracellular vesicles. We would like to point out the following:

      (1) In Miao et al. 2013 (1), Figure 3C-D, we showed that expression of a dominant-negative mutant of NSF, a non-redundant protein in vesicle trafficking, impaired vesicle trafficking but did not affect SOCE. This experiment had essentially ruled out a role for vesicle trafficking in SOCE. In the same paper, we had also shown that Orai1 levels in the PM do not increase post-store depletion (Figure 3-figure supplement 2).

      (2) SNAP23/25 form a four helical bundle with R- and Q-SNAREs in orchestrating vesicle fusion. In this paper, we have ruled out a direct role for SNAP23, SNAP25 and SNAP29 in SOCE (Figure 2-figure supplement 3).

      (3) In v1 of this manuscript, we had shown that U2OS cells stably expressing Orai1-BBS-YFP have identical levels of Orai1 in the PM with and without STX11 depletion (Supplementary Figure 3B). This showed that the biosynthesis or delivery of Orai1 to the PM is not affected by STX11 depletion. The levels were also assessed in store-depleted U2OS cells but not included because in Miao et al. 2013 we had already established that levels of PM Orai1 remain essentially equal in resting versus store-depleted cells.

      In the revised version, we have included the data from store-depleted cells in U2OS and also done quantification of PM Orai1 in HEK293 and Jurkat T cells. In addition, we have added three independent membrane trafficking/ vesicle secretion assays performed in STX11-depleted cells (new Figure 2 and associated supplements). In all cases, we find no evidence of Orai1 in intracellular vesicles or a general defect in membrane trafficking/ secretion in STX11-depleted cells. Orai1 is constitutively and stably expressed in the PM in resting as well as store-depleted cells in three different cell lines.

      (4) Most importantly, in Figure 7I-J of this manuscript, we showed that calcium influx from a constitutively active mutant Orai1 (Orai H134S) is identical between STX11-depleted and scramble control cells. If wildtype Orai1 was indeed stuck in vesicles in STX11-depleted cells, then how would mutant H134S Orai1 be able to rescue the defect in SOCE? We have included the quantification of PM levels of Orai1 mutants w.r.t WT Orai1 in new Figure 8-figure supplement 3B and 3D.

      In summary, we have now done several new experiments to directly measure Orai1 levels in the PM and general vesicle trafficking assays in HEK293 and Jurkat T cells and have found no defects in these upon STX11 depletion.

      Regarding STX11 induced precise conformational transition, we are trying to setup collaborations with scientists who might be able to visualize this in situ. Please note that while purification of isolated pore subunits of ion channels followed by crystallization or expression in synthetic membranes for cryo-EM is currently considered a gold standard in the analysis of ion channel pore subunits, we have shown that ion channels are dynamic macromolecular complexes, in vivo (2), where synaptic proteins dynamically bind to induce conformational changes and affect their stoichiometry (2). Please also see (3) and (4). More advanced approaches, therefore, need to be developed to enable visualization of the dynamics of ion channel macromolecular complexes in their native environment in situ. In the absence of such approaches, the structural insights obtained from detergent-purified isolated subunits will remain incomplete.

      Reviewer #2 (Public review):

      Summary:

      Vig's lab delineates a critical role for STX11 in CRAC channel function, particularly in the context of the fatal immune disorder familial hemophagocytic lymphohistiocytosis type 4 (FHL4). They demonstrate that Syntaxin 11 directly binds and regulates Orai1, and that STX11 depletion abolishes CRAC currents and downstream signaling. Loss of STX11 reduces IL2 gene expression and impairs degranulation, both of which are rescued by the constitutively active Orai1 mutant H134S, whereas a gain‑of‑function mutant targeting the C‑terminus fails to restore these defects. The authors conclude that STX11 primes Orai1 for optimal local assembly that is independent of STIM1 yet required for CRAC channel gating.

      Strengths:

      This study is firmly grounded in disease biology and demonstrates that STX11 downregulation leads to profound functional defects. Using a comprehensive suite of methods and analyses, the authors interrogate the co-regulation of STX11 and Orai1 and present a near-complete view of STX11's modulatory role in CRAC channel function and downstream signaling pathways. The figures are clear, and the statistical analyses are rigorous and convincing.

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach, including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

      We thank the reviewer for acknowledging that analysis of the dynamic binding of STX11 will require standardization of advanced techniques which are beyond the scope of the present study. We plan to continue developing methods that will allow us to visualize the binding and unbinding of STX11 to Orai1 in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Mechanistic issues:

      (1) More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional functional channels, stored in secretory vesicles or recycling endosomes, to the membrane. A significant fraction of the ORAI1 channel is in vesicles, and the mobilization of this intracellular pool regulates the rates of calcium fluxes in HEK-293 cells (PMID 26116575) and effector T cells (PMID: 35217583). Mobilization of this pool could account for part or all of the functional effects reported here. The only evidence that STX11 depletion does not impact the plasma membrane availability of the channel relies on one single flow cytometry profile (Supplementary Figure 3). This experiment is performed in U2OS cells stably expressing a fusion protein containing an extracellular bungarotoxin site. This cellular system is only used here; the other data are obtained either in HEK-293 cells or in Jurkat T cells. The level of endogenous STX11 in U2OS cells is unknown, and the efficiency of protein depletion has not been assessed. The efficiency of depletion should be shown, and the total number of channels assessed by comparing the expression levels of permeabilized and non-permeabilized cells. A flow cytometry profile of cells treated with thapsigargin should be included to match the experimental conditions of functional recordings. It would be valuable to repeat this experiment in Jurkat T cells by expressing ectopically the channel tagged on the extracellular side. To further exclude the involvement of vesicular trafficking, the authors should also evaluate the contribution of VAMP8, as this R-SNARE has been proposed to interact with STX11 to regulate the exocytosis of specialized granules in cytotoxic T cells (PMID: 26124288).

      We have not observed significant levels of intracellular Orai1, as described in the PMID 26116575 paper. Multiple technical reasons could explain the artefactual appearance of intracellular Orai1. HEK293 is an embryonic kidney cell line, with most cells showing a distinct spindle shape and filopodia as shown on the ATCC website (https://www.atcc.org/products/crl-1573). The HEK cells shown throughout the Hodeify et al. 2015 paper (PMID 26116575) lack this typical morphology that majority of the HEK cells should show. Transfecting cells with high amounts of DNA and liposomes can severely affect the health and morphology of the cells leading to artifacts where PM proteins appear to be stuck intracellularly. The plasma membrane of the cell in movie 1 of PMID 26116575, for instance, also shows membrane blebs or membrane ruffles. The authors should have used a PM marker such as WGA (wheat germ agglutinin) to distinguish PM Orai1 from any intracellular Orai1 to establish whether what appear as intracellular Orai1 vesicles are not blebs of PM Orai1. Similarly, no endo/exocytotic vesicle marker is used to distinguish them from Orai1’s presence in apoptotic vesicles of unhealthy cells. In view of the overall abnormal morphology and any PM or intracellular vesicle marker, the claim that Orai1 resides in intracellular vesicles is unfounded.

      Similarly, the PMID: 35217583, quoted by reviewer #1, is about in vitro differentiated primary mouse T cells. The study lacks any detail of the generation of the HA-tagged mouse Orai1 plasmid, used in the study, or even a back reference to show how this plasmid was validated for normal expression in primary T cells earlier. Ectopic expression of CMV promoter-driven plasmids in mouse primary T cells is extremely challenging and often results in poor cell health and incomplete and selective expression in only 5-10% of cells. It is unclear if the same HA-tagged human Orai1 plasmid that was used in PMID 26116575 is being used in this study to express in primary mouse cells. In the absence of all this information, it is unclear whether there was an issue with the generation of a new HA-tagged mouse Orai1 construct, which made the protein get stuck intracellularly. Or potentially the expression of a human protein in mouse primary T cells is the problem. The functional verification of the construct by showing rescue of SOCE in Orai1-deficient primary mouse cells is an absolutely essential control but is missing. In the absence of these controls, one cannot disregard decades of robust data on the localization of Orai1 in the PM from multiple labs and papers that have established its PM localization conclusively (5) (6).

      Most importantly, in both the PMID 26116575 and PMID: 35217583, HA tagged-Orai1 is detected using a bivalent antibody, followed by secondary antibody. It is well known that cross-linking of cell surface receptors or proteins using bivalent antibodies has a major caveat that involves antibody-mediated clustering and capping, especially in lymphocytes, which typically induces rapid internalization of the entire antigen-antibody complex. The paper by Sekine-Aizawa et al., in 2004, showed that tagging of receptors/ channels with bungarotoxin binding site (BBS) followed by labelling with bungarotoxin (BTX) bypasses this confounding factor and, therefore, allows accurate estimation and localization of PM versus intracellular proteins. The artefactual bivalent antibody-induced endocytosis continues even when cells are incubated on ice as endocytosis is only slowed but not stopped on ice.

      Due to this challenge, we have generated BBS tagged Orai1-YFP. We use flow cytometry to show PM Orai1 because it is an unbiased and quantitative way of showing surface Orai1 expression (estimated by measuring the intensity of surface-bound BTX), simultaneously, in thousands of cells with no potential for visual bias in the selection and imaging of cells. BTX labelling is done on ice and cells are washed and fixed right after labelling with BTX-A647 to stop endocytosis. In the revised version, we have used both complementary approaches of flow cytometry and microscopy, showing representative images of cells, alongside quantifications. All new experiments were performed in HEK293 and Jurkat T cells to estimate surface versus total expression in a new main Figure 2. The post-store-depletion data have been added to the existing U2OS experiment in new Figure 2-figure supplement 2A-E.

      Our BBS-tagged Orai1 construct also has a YFP tag at the C-terminus. Since the flow cytometry experiment involves gating on YFP-positive cells, the total number of channels (biosynthesis) can be compared by looking at the YFP intensities in the scramble versus STX11-depleted groups. These data have now been included in the previous and new experiments. In none of the cases could we detect any difference in YFP or BTX-A647 intensities, pre- or post-store-depletion in scr or STX11-depleted cells which would indicate defects in biosynthesis or Orai1’s presence in vesicles in any group.

      Regarding a role for VAMP8, in Miao et al. eLIFE 2013 (1), Figure 3C-D, we showed that expression of a dominant negative mutant of NSF, a non-redundant protein in the vesicle trafficking pathway, impaired Transferrin receptor recycling within 20 hours but did not affect SOCE at all. This experiment had conclusively ruled out any role for vesicle trafficking in SOCE and therefore assessment of the role of each of the individual proteins involved in membrane trafficking becomes redundant. We have additionally ruled out a role for SNAP23/ SNAP25/ SNAP29 (old supplementary Figure 12, new Figure 2-figure supplement 3A-C). SNAP23/25 form a four helical bundle with most R-SNARE and Q-SNAREs to orchestrate vesicle fusion. R-SNAREs, typically, need SNAP23/25 to interact with Q-SNAREs. Since a role for these non-redundant proteins has also been ruled out by us, it is unlikely that VAMP8 plays a role in modulating the effects of STX11 in SOCE.

      (2) Both ORAI1 and STX11 are S-Acylated on cysteine residues, and this post-translational modification promotes their recruitment to the immune synapse (PMID: 24910990, 34913437). One possibility that should be discussed is that STX11 could enhance the recruitment of the ORAI1 channel into lipid domains rich in cholesterol, thereby favoring its activation. This type of priming would still require direct interaction between the two proteins but involve a different mechanism than the one discussed by the authors. This could be experimentally tested by expressing a STX11 mutant lacking the cysteine residues required for its S-Acylation. It would also be interesting to test whether depletion of STX11 impairs the recruitment of ORAI1 to the immune synapse forming between Jurkat T cells and antigen-presenting cells.

      In our experiments reported in this paper, we have used soluble anti-CD3 as well as plate-coated anti-CD3 in combination with soluble anti-CD28 to stimulate Jurkat and primary T cells. These antibodies are routinely used to stimulate T cells and this type of stimulus doesn’t depend on the formation of a classical immune synapse with an antigen-presenting cell (APC). Despite the absence of synapse, the T cells get fully activated and functional, as seen by NFAT translocation and secretion of cytokines, such as IL-2 in new Figure 4. Therefore, whether there is a defect in the recruitment of Orai1, or STX11, to T cell synapse formed with an APC is not within the scope of this study. Furthermore, accurate analysis of protein localization within the immune synapse requires a dedicated study employing sub-diffraction resolution microscopy approaches.

      Regarding PMID: 24910990, and the mechanism of recruitment of STX11 to the membranes. We believe this remains unknown. The frameshift mutant used in our study lacked all terminal cysteines, which have been earlier proposed to be crucial for membrane targeting, as well as a terminal part of the SNARE domain and yet it localized to the PM just as well as wild-type STX11 (See new Figure 5E). We have added this result in the text line 302-305 and removed the line stating that post-translational modifications of the terminal cysteines target STX11 to the PM as previously claimed in PMID: 24910990 and mentioned in v1 of this paper. In view of these new data, it currently remains unknown whether potential attachment to PIP2 in PM via various basic residues, spread throughout the sequence (7, 8), or binding to another protein targets STX11 to the PM (PMID: 26771955). There is no obvious poly-basic stretch in STX11 sequence, therefore, a systematic and focused deletion and mutagenesis study will be needed to individually assess the above possibilities which is outside the scope of the present study.

      Methodological issues

      (3) Since STX11 colocalize with Orai1 already in basal conditions, independently of STIM1, it could influence basal calcium levels. This cannot be appreciated from the data presented, because all the SOCE protocols start in calcium-free conditions, preventing baseline comparison between WT and STX11-deficient cells. A potential difference in basal calcium levels should be explored, and the impact of STX11 depletion on basal calcium fluxes should be documented by calcium shifts (2 mM → 0 mM → 2 mM) or manganese quenching approaches.

      The cells are typically loaded with Fura2 in 2mM calcium containing Ringer’s buffer. We switch the cells to 0mM right at the start of the SOCE protocol and start imaging within 5-10 seconds. Therefore, in our experience, the baselines of scramble versus STX11-depleted cells should show a difference even in the SOCE protocol if the basal calcium levels are affected because Fura2 is already present and bound to basal calcium present in the cytosol at the start of the protocol. Still, we have done the experiment suggested by the reviewer as shown in Author response image 1. The assay started with cells in 2 mM extracellular Ca<sup>2+</sup> followed by addition of 10 mM EGTA, which according to the following equation quenches the 2 mM extracellular Ca<sup>2+</sup> (https://somapp.ucdmc.ucdavis.edu/pharmacology/bers/maxchelator/CaEGTA-TS.htm).

      where, [Ca2+]<sub>Free</sub> is the free/unbound Ca2+, [Ca2+]<sub>Total</sub> is the total Ca2+, [EGTA]<sub>Total</sub> is the total EGTA concentration and Kd is the dissociation constant between Ca2+ and EGTA at 37°C and pH 7.4.

      To confirm complete sequestration of the extracellular Ca<sup>2+</sup>, we also repeated the assay with 20 mM EGTA but did not notice any difference between the 10 mM and 20 mM EGTA conditions. We do not see any differences in basal calcium, under any condition, between Scr and STX11 shRNA treated cells HEK or Jurkat T cells.

      Author response image 1.

      Representative Fura-2 traces of Scr (black) and STX11 (red) shRNA-treated HEK293 (A) and Jurkat (B) cells, where the cells were incubated with 2 mM Ca<sup>2+</sup> followed by addition of 10 mM EGTA to quench the 2 mM Ca<sup>2+</sup>. Since we do not have access to a perfusion system, we could not test Fura2 response after re-addition of 2mM calcium to the existing EGTA and Ca<sup>2+</sup> mixture. However, the transition from 0mM to 2mM is already shown in Figure 8.

      (4) The quantification of the calcium imaging data is problematic and requires clarification. In most figures, the data are shown normalized to the control condition. According to the method section (lines 721-724), 100% is the maximum value of the scramble shRNA group (amongst the three experiments). But what was measured here? The slope during calcium readmission? The peak amplitude after calcium readmission? Expressed as a ratio or as calcium values? The recordings are presented in micromolar calcium concentration. This implies calibration, but a calibration procedure is not mentioned. Please clarify. For the recordings of constitutive calcium entry in Figure 7, this normalization is not performed, and the data are expressed as ratio values. Here, it looks like the parameter quantified and compared is the absolute ratio value after calcium readmission. This is inappropriate. The trace in Figure 7G shows that the basal levels differ by more than two ratio units between control and STX11-depleted cells. Normalizing the data in Figure 7G to the basal ratio value would show no difference in the peak response amplitude between the two conditions. These data should be re-analyzed, and both the slope and the amplitude of the response should be presented, with statistics performed on independent recordings, not on individual cells pooled from different experiments. Cells from the same recording are not experimentally independent samples but rather replicates of the same experiment.

      We have updated the relevant method section with more details in lines 823-858. The peak amplitude after calcium readmission was measured and compared to the baseline as described in the updated methods. The Fura 2 calibration method has also been added to the revised version, we apologize for this omission earlier. Separate calibration was done for Fura 2 experiments in all figures except Figure 7 from version 1 (new Figure 8). The experiments in new figure 8 were done using a different objective (20X, water) and, therefore, although a separate calibration was done for these experiments it was not applied to the data. We apologise for this omission on our part and have now applied the respective calibration to these experiments.

      The trace in Figure 7G of version 1 should look different even in 0mM calcium in our opinion. The reason is the same as that explained in point 3 above; CAD-mediated constitutive activation of Orai1 is one of the strongest. The cells, even when they are being loaded with Fura2 in Ringer’s buffer with 2mM calcium, are constitutively recruiting calcium ions. This should result in a shift in Fura2 excitation due to higher levels of basal calcium. When the cells are switched to 0mM calcium and imaged within 4-5 seconds, the intracellular Fura 2 is still bound to all this extra calcium in the cytosol and therefore the baselines should show a significant difference. If the cells were imaged for several minutes in 0mM calcium, we might have seen the difference in basal ratios slowly reducing. However, we switched to 2mM calcium within 120sec. At this point any free Fura2 would be expected to bind incoming calcium again. For the same reason, the cytosol of Scr cells which would still have relatively higher levels of intracellular calcium concentration compared to STX11-depleted cells will show a smaller further increase in 2mM due to calcium-dependent inhibition of CRAC currents within the time frames we have measured. We have now added the Fura-2-calibrated response in the new Figure 8G-H. We have shown below the baseline-subtracted (normalized) response for Figure 8G. As you can see, there is still a significant difference in 2mM calcium between Scr and STX11-depleted cells but in this representation the important difference at 0mM is masked, we have therefore chosen to retain the original figure with Fura-calibrated values at 0 as well as 2mM calcium in Figure 8G-H.

      The cells shown in the old Figure 7G of version 1 were not from the same recording but from three different experiments. Because the cells were imaged with 20X objective in these experiments to allow selection of Orai1-CFP or mutant Orai1-CFP and CAD-YFP double-positive cells, the number of cells analyzed per experiment was less compared to other experiments. We have now shown Fura 2 calibrated values in the new Figure 8G-L. We have also re-done the statistical analysis on three independent experiments from each. As shown in Author response image 2, the difference is still statistically significant whether we show merged cells from all three experiments or single representative experiment out of three repeats. We believe merged cells have more information to offer and therefore have retained the same figures with the original analysis in the main Figure 8G-L.

      Author response image 2.

      Box plots representing quantification of individual repeats of constitutive calcium influx in Scr and STX11 shRNA-treated HEK293 cells expressing YFP-CAD and Orai1-CFP without (A) and with baseline subtraction (B). (C-D) Quantification of individual repeats of Scr and STX11 shRNA-treated HEK293 cells expressing Orai1-H134S (C) and Orai1-ANSGA (D) mutants.

      (5) Quantification of pull-down experiments. The binding data in Figures 4F, 5F, and 5J are presented largely qualitatively. Densitometric quantification with statistical comparisons across wild-type and mutant conditions would make these results more convincing, particularly given that the authors themselves acknowledge the interaction appears relatively weak in vitro.

      This has been done and included alongside the respective panels in the new Figure 6G and 6L (for old Figure 5F and 5J of version 1, where differences appeared relatively small in some experiments). The differences across lanes in both figures and their repeats were statistically significant.

      Figure 4F showed a clear and visually significant difference in binding across lanes and repeats and therefore no quantification is needed for these experiments in our opinion.

      Limitations of the study and mechanistic inferences.

      (7) Interpretation of the ORAI:ORAI FRET and crosslinking data. STX11 depletion increases basal ORAI:ORAI FRET (Figure 7A-C) and shifts crosslinked species toward higher molecular weights (Figure 7D-E). The authors interpret this as ORAI1 being trapped in an unprimed state, but higher FRET and higher-order species would conventionally suggest increased rather than decreased assembly. The paper needs a clearer mechanistic explanation of what "unprimed" looks like structurally. Is this aberrant crowding, non-productive oligomerization, or something else? The distinction between a change in intermolecular distance within existing oligomers versus an increase in oligomer density matters here and should be addressed.

      Higher ORAI: ORAI FRET and a shift in the size of crosslinked Orai1 oligomers, when analyzed together, suggests formation of ‘non-functional’ higher-order oligomers. Higher order does not necessarily translate to better function in the case of ion channels, it can also lead to non-selectivity or formation of ‘non-productive’ oligomers, as mentioned by the reviewer. It was shown by us earlier in Li et al. (2016) (2) that bigger oligomer size revealed by higher number of photobleaching steps of Orai1 did not translate to better function but led to non-selectivity.

      In new Figure 8, crosslinking with BS3, which has a spacer arm and working distance of ~11 Å, very likely reflects a change in the number of subunits within individual oligomers and not crosslinking of independent existing oligomers. This is because we show that neither total Orai1 expression nor Orai1 expression in the PM change in any group in new Figure 2. FRET works best within 1 to 10 nm distance, and therefore, in theory, can lead to energy transfer between neighbouring Orai1 oligomers in high Orai1-expressing cells. However, because there was no change in Orai1 abundance in the PM (new Figure 2) or distribution within PM (new Figure 7E,F,J,L) of any group, FRET changes also likely reflect intra-oligomer changes rather than inter-oligomer interactions. FRET changes can also arise from a change in the respective orientation of fluorophore pairs but when analyzed together with crosslinking studies, changes in pore assembly likely coincide with conformational shifts in Orai1 protomers. Furthermore, FRET has been used earlier to show shifts in conformation of other ion channels (9). Therefore, we believe that, when used together, these two approaches strongly suggest an intermediate conformational state along with a change in number of Orai1 subunits per channel since there was no evidence of overcrowding in the PM or obvious segregation of Orai1 in specific regions of PM in new Figures 2 and 7E,F,J,L.

      We could not assess whether the oligomers of Orai1 formed in the absence of STX11 possess an intact pore. The presence or absence of pore in STX11-depleted cells will require extraction of Orai1 oligomers from native membranes and performing systematic structural analysis using cryo-EM or related approaches which is outside the scope of this study.

      (8) The ANSGA versus H134S discrepancy. H134S ORAI1 rescues calcium influx in STX11-depleted cells (Figure 7I-J), but the ANSGA mutant does not (Figure 7K-L). The authors conclude from this that STX11 induces molecular shifts within ORAI1 transmembrane helices, and not in its C-terminal tail. This is an important mechanistic inference that needs more discussion. What does this imply about the conformational state of primed ORAI1? And why is straightening of the tails not sufficient for full opening without the correct TM helix arrangement? This distinction has implications for how the STX11-ORAI1 interaction should be modelled and should be engaged with more thoroughly.

      There is no discrepancy here, please also see our response to reviewer 2’s comment #4. The experiment implies that the conformational state of primed Orai1 involves shifts in the TM region of Orai1 and is different from unprimed state. The structural similarities between H134S and ANSGA Orai1 mutants have not been formally established. Unlike H134S, no structure exists for the ANSGA mutant. In the absence of this, it is impossible to comment on whether the two constitutively active mutants are structurally comparable or whether there are multiple ways to stabilize open states of CRAC channel pore, especially when using TM mutants of Orai1.

      The goal of this experiment was to determine what kinds of structural shits STX11 potentially induces in native Orai1. Using previously characterized constitutively active mutants and fusion proteins from the CRAC field, we have ruled out a potential role for STX11 in simply changing the orientation of Orai1 C-terminal tails. A discussion on the topic of why tail straightening of Orai1 is insufficient to open Orai1 is outside the scope of this paper. As pointed by reviewer 2, it is possible that C-term tails already exist pointing towards the cytosol in native, resting Orai1, although this has not been shown in any study using structure of full-length WT Orai1 and is purely speculative at this point. We prefer to not engage in speculative structural insights.

      Other points

      (9) Figures 1G and 1H. The patient-derived mutant STX11 band runs at approximately 37 kDa rather than the predicted 39.5 kDa. The authors suggest instability or reduced antibody reactivity, but premature translation termination is also a possibility that should be acknowledged.

      We have added this point in line 168.

      (10) Figure 2B. The traces and current voltage relationships should be rescaled to show the rectification and inactivation profile of the current in cells depleted of STX11.

      This has been done and modified in new Figure 3 (Figure 2 of version 1).

      While preparing source data files for all figures, we noticed an error in the value of the SE in the STX11-depleted group of old Figure 2C, which has now been corrected. The SE value in the older version was erroneously pasted from an adjacent data column.

      Similarly, in old Figure 3 (version 1), new Figure 4C, we noticed that some data points in the STX11 group were pasted twice in the same excel column. These cells were removed and additional cells were analyzed from the same experiment and added to this group. The overall result remains the same but the distribution of data points looks a bit different.

      (11) Figure 4B: This experiment should be repeated in cells treated with thapsigargin to deplete intracellular calcium stores, and the extent of colocalization quantified by measuring the Pearson's correlation coefficient.

      This has been done. Pearson’s correlation coefficient is included in new Figure 5D.

      (12) Figure 4F. Why is there no detectable band in the input lane of the left blot?

      Western blots show relative intensities of bands of proteins across lanes. A faint band in the input lane of old Figure 4F suggests that the IP/ co-IP/ pull down was robust. If we increase the exposure, the input band would become stronger but the pull-down band would become over-saturated and the difference in the intensities would not be linear. The faint non-specific bands in other lanes represent a fraction of soluble STX11 that tends to crash out of solution over time and gets spun down with the beads. See lines 535-541 explaining this.

      (13) Figure 5. Immunofluorescence data showing the membrane staining of the mutated syntaxin and channel should be included, as well as calcium recordings of cells expressing YFP-CAD with WT and mutated ORAI1.

      In version 2 Figure 6A, we have now also shown co-localization of mutant STX11 with Orai1-YFP in resting and store-depleted cells, in addition to WGA. Pearson’s correlation (not shown) did not show any significant difference in the localization of mutant synatxin 11 w.r.t Orai1. Calcium recordings of CAD-induced constitutive calcium influx from wild-type versus mutant Orai1 are now shown in new Figure 6O-P.

      (14) Figure 6B. A clear colocalization of CFP-O1 and STIM1-YFP is visible on the images, yet the authors conclude from morphometric analysis that the channel is not recruited into ER-PM clusters. Please show the difference in colocalization quantified by measuring the Pearson's correlation coefficient. Pictures should also be provided with the C-terminally tagged construct.

      The quantification of CFP-Orai1 localization inside Stim1-YFP puncta was already shown in old Figure 6E and F. The residence of Orai1 inside STIM1 puncta versus total Orai1 in the PM of STX11-depleted groups was clearly reduced. We have now also shown Pearson’s correlation coefficient for Stim Orai co-localization inside puncta in new Figure 7F.

      TIRF microscopy images of C-terminally tagged Orai1 were already included in Supplementary Figure 10. No defect in co-clustering of C-terminally tagged Orai1-YFP and N-terminally tagged CFP-Stim1 was seen and yet SOCE was inhibited. Therefore, we never concluded from Figure 6 that Orai1 and Stim1 fail to co-localize. We said, they fail to form ‘functional’ clusters. We have now moved the representative TIRF images from supplementary figure 10 to the new main Figure 7G. The Pearson’s correlation coefficient for Stim Orai1 co-localization inside puncta is shown in new Figure 7L.

      (15) Figure 6E and 6F show the same data.

      Figure 6E showed fraction of Orai1 inside Stim1 puncta divided by total Orai1, and 6F showed fraction of Orai1 outside puncta divided by total Orai1. The plots are different but we agree that the data are coming from same cells. We have removed old panel 6F and replaced it with Pearson’s correlation coefficient of Stim1:Orai1 colocalization in puncta in new Figure 7F.

      (16) Figure 7G-L. The difference in constitutive calcium fluxes should be confirmed by Manganese quench recordings. The surface expression of the Orai1 mutants should be shown.

      We have now shown the quantification of surface expression of Orai1 mutants for each respective mutant in the new Figure 8-figure supplement 3B and 3D. The Orai1 mutants we have used in this paper are well established in the literature, they showed clear surface localization and the differences in calcium influx between Scr and STX11 treated cells upon overexpression of Orai1 mutants in HEK are robust. Therefore, we do not see any compelling reason for repeating all of the experiments from Figure 7G to 7L to also show manganese quench recordings, as suggested by the reviewer. We have applied Fura 2 calibration done for these experiments to calculate intracellular calcium. These have been shown in the revised and new Figure 8G-L, where F340/380 ratios of representative calcium assays have also been replaced with the calibrated intracellular calcium concentration.

      (17) Supplementary Figure 12. The recordings show a very large variability between experiments. The different SNAREs that are depleted here could compensate for each other, accounting for this variability. It would be interesting to show the effect of the combined silencing of all the SNARES tested here. The efficiency of the protein depletion should also be documented.

      Genome-wide high- or medium-throughput screens are inherently noisy. None of the genome-wide high- or medium-throughput screens show evidence of protein depletion for each gene in any of the published screens to our knowledge. We chose to only characterize the candidates that reproducibly showed > 70% inhibition of SOCE, others were ignored as noise.

      Silencing of all SNAREs together will definitely lead to loss of morphology and early lethality as all membrane trafficking will be stopped. We never analyze cells that do not show normal morphology and have compromised viability for ablation of SOCE.

      (18) Lines 236-238. The authors note that STX11 harbors a stretch of C-terminal cysteines proposed to be essential for its membrane localization, but do not elaborate on the underlying mechanism. It would strengthen the discussion to explicitly acknowledge that this membrane anchoring is mediated by S-acylation of these cysteines PMID: 24910990 and to connect this to the known enrichment of Orai1 in lipid rafts and the immune synapse PMID 34913437. Both observations are relevant to understanding how STX11 and Orai1 are brought into proximity at the plasma membrane, and their omission leaves an explanatory gap in the proposed interaction model.

      Please see our response to point #2 above. We do not think C-terminal cysteines target STX11 to the PM. We have corrected this claim based on an earlier study, PMID: 24910990, in the revised version of this paper. Analysis of immune synapse and lipid rafts are outside the scope of this paper. The mechanism of PM targeting of STX11 is currently unestablished and will require a systematic and focused mutational analysis which is outside the scope and main focus of this paper.

      (19) Line 351. The statement that syntaxin depletion does not alter the structure or proximity of junctional ER to the plasma membrane is not supported by data. Neither electron microscopy nor TIRF imaging has been performed, which would be required to back up this claim.

      Because Stim1 itself can be used as a marker of ER-PM junctions, this statement was supported by data shown in Figure 6C, D, G, H of version 1 of this paper where the intensity and area of Stim1 clusters was assessed using TIRF microscopy and found to be indistinguishable between STX11 and scramble control cells. The imaging done in Figure 6G, H was TIRF imaging and this was already specified in the legend. We have now also done TIRF imaging of GFP-Mapper-expressing scr and STX11-depleted cells. Mapper is a genetically encoded fluorescent protein that was previously shown to mark ER-PM junctions (10). We found no significant difference in the area or intensity of GFP-Mapper puncta (new Figure 7O-Q), just like Stim1 puncta didn’t show any defect in STX11-depleted cells. Please see modified text from 416-423.

      (20) The molecular dynamics methods need more detail: force field, simulation length, water box dimensions, and convergence criteria should all be specified to allow replication. The supplementary RMSD plots (Supplementary Figure 5B) should also show individual replicate trajectories rather than averages only.

      We had already mentioned the force field (OPLS4) and simulation length (500ns) in the methods section. Also, the RMSD plots in Supplementary Figure 5B already showed individual replicates in version 1.

      We have now updated the methods with following additions:

      The OPLS4 force field was used for all 500 ns simulations in an orthorhombic water box with a buffer distance of 10 Å beyond the solute in each direction. Simulation stability was assessed based on the protein backbone RMSD over simulation time.

      Trajectory clustering was performed using the trajectory clustering tool in Schrödinger, which applies affinity propagation to the pairwise backbone RMSD-based similarity matrix. Within each affinity propagation run, convergence was defined as no change in the set of exemplar frames for 15 consecutive iterations, with a maximum of 400 iterations per run. If convergence was not reached, the damping factor was increased from 0.5 in increments of 0.01 until convergence.

      Reviewer #2 (Recommendations for the authors):

      Overall, this is a timely and impactful study supported by a broad set of methods and cell types. Before publication, the manuscript should address the following points.

      Major:

      (1) The authors note that STX11 contains cysteine residues that enable membrane association. What is the specific mechanism of membrane attachment? Could it occur via S-acylation (palmitoylation)? Both Orai1 and STIM1 are known to undergo S-acylation, which raises the possibility that this modification might also facilitate STX11 membrane anchoring and/or co-residence with Orai1. Is STX11 constitutively membrane-associated, or does it show preferential localization to specific membrane subdomains, particularly in proximity to Orai1?

      STX11 is constitutively membrane-associated and does not show any preferential localization to specific membrane subdomains in confocal images. Figure 4, panel A and B from version1 clearly showed this. In an earlier paper by Hellewell et al. 2014, PMID: 24910990, S-acylation of terminal cysteines of STX-11 was proposed to be crucial for membrane attachment of STX11 and its recruitment to the immune synapse. However, please see our response to reviewer 1’s comment #2 and a new Figure 5E for the localization of the frameshift FHLH4 mutant characterized in this paper. The frameshift mutant that we have characterized lacked all terminal cysteines as well as a short terminal part of the SNARE domain. Cloning and ectopic expression of this mutant still showed constitutive localization to PM and did not show preferential distribution to any specific regions. Therefore, we do not think that terminal cysteines of STX11 contribute to its membrane attachment, we have accordingly modified lines 291-293, 302-305, 540 in the revised version. Also see our response to your point#6 below.

      (2) Is there a possibility to monitor a dynamic change in STX11 co-localization from before to after store-depletion?

      We did not observe any change in the overall distribution of STX11 in cells expressing STX11 alone or co-expressing Orai1 with STX11, pre- or post-store-depletion (please see new figure 5B-C). In cells co-expressing ORAI1, STX11 and STIM1 (see new Figure 5M-N), we could not capture the dynamic segregation of STX11 into regions of PM devoid of STIM:ORAI puncta and therefore have only pre- or post-store-depletion images. Dynamic change in STX11 distribution would require live, multi-colour, high-resolution imaging of diffraction-limited ER-PM junctions and adjacent regions which is technically extremely challenging, and especially due to our inability to tag STX11 with a fluorescent tag without disrupting its localization. Also see our response to your point#6 below.

      (3) The authors use CAD to prove that the interaction with the R289A_E272A_E275A_E278A mutant is normal as for the wild-type. Does this also hold for STIM1 wild-type full-length?

      This is also true for full-length STIM1. The data have now been added to the new Figure 7-figure supplement1.

      (4) The authors state that STIM1 binds both the N- and C-termini of Orai1. While STIM1 binding to the Orai1 C-terminus is well established, the nature of its interaction with the N-terminus remains debated. Fragment-based assays suggest direct binding to the N-terminus; however, direct interaction with full-length Orai1 has not been conclusively demonstrated. This point should be phrased more cautiously to reflect the current uncertainty.

      We have re-phrased the sentence as follows in line 543: “The individual relevance of Orai1 N- versus C-terminus in the trapping versus gating of Orai1 remains unclear”

      (4) In the discussion, the authors report: "Though crucial for trapping and gating, Orai1 tails were missing from early structures of Drosophila Orai [28]". A previous NMR structure suggested that the C-terminal tails of two adjacent Orai1 subunits bend and pair with each other in an antiparallel fashion, and sit closely apposed to PM [37]. However, in recent structures of constitutively active H134 mutant Orai, the C-terminal tails were found to orient away from the membrane [28]. In STX11-depleted cells, switching the CFP-tag from the Orai1 N- to the C-terminus could rescue its clustering but not gating by Stim1. Furthermore, STX11 depletion inhibited the constitutively active ANSGA mutant of Orai1 [29], where the tails of Orai1 are proposed to be constitutively unlatched. These data essentially reinforce our conclusions that STX11 induced molecular shifts encompass Orai1 transmembranes.' However, the information provided here is not fully correct. The early Drosophila Orai structure lacks the full N-terminus but retains most of the C-terminus. It was the X‑ray, not cryo‑EM, structure that suggested an antiparallel arrangement of the Orai1 C-termini. Although the "open" X‑ray structure shows unlatching and straightening of TM4-C-termini, it remains uncertain whether these features reflect physiological gating or crystallization artifacts. It is also unclear whether the Orai1 ANSGA gain‑of‑function mutant adopts a similar unlatching; however, prior work indicates that ANSGA impairs proper coupling to the C‑terminal binding interface (in contrast to Orai1 H134S, which maintains effective coupling). This raises the key question: why do H134S and ANSGA respond differently to STX11 depletion? One possibility is that these mutants stabilize distinct conformations of the TM4-C-termini ("latched" vs "unlatched" states) that differentially dictate the requirement for STX11 in channel assembly or gating. We recommend refining the discussion

      We have changed the word ‘missing’ to ‘truncated’ in line 545 and 546.

      We agree that there is no evidence in literature that establishes similarity between H134S and ANSGA mutation-induced conformations of Orai1. It is, however, implied in most previous studies of mutant Orai1s that there is only one possible open state/conformation. We have added the suggested point and modified the discussion in line 551-556.

      (5) The authors highlight: 'A major problem with this interpretation is that even though amplification of CRAC currents was shown, none of the previous patch clamp studies established whether the higher currents resulted from a greater number of active channels or unchecked conductance per channel by performing single channel recordings.' It should be noted that CRAC channels have extremely low single‑channel conductance, making direct single‑channel recordings challenging. As a result, estimates of open probability and channel number typically rely on fluctuation (noise) analysis rather than direct measurements of single‑channel events (see https://doi.org/10.1085/jgp.200609588). We suggest acknowledging this limitation in the discussion to contextualize the interpretation of gating and channel density.

      We acknowledge how challenging it is to record the single-channel conductance from CRAC channels. We have added this fact to the discussion and the reference that the reviewer has suggested in line 570-572.

      (6) STX11 appears to shift Orai1 localization into puncta. Activated STIM1 is known to engage plasma membrane PIP2 to facilitate Orai1 coupling. How, if at all, is STX11 linked to PIP2 or PIP2-rich microdomains? Is there evidence for direct PIP2 binding by STX11, or for indirect recruitment via PIP2-binding partners? Any available data on STX11's lipid interactions or its enrichment within PIP2-enriched regions would help clarify this mechanism.

      We have not claimed that STX11 shifts Orai1 into puncta. We already showed in old supplementary figure 10 and Figure 6G-J of version1 (v1) of this paper that the C-terminally tagged Orai1-CFP can very well form puncta and co-localize with STIM1 in STX11-depleted cells. To avoid this confusion, we have moved the old Supplementary Figure 10 from v1 to the main figure in revised version, see new Figure 7 panel G. Despite the presence of ORAI1-CFP in puncta with YFP-Stim1, the SOCE was inhibited in STX11-depleted cells. Please also see new Pearson’s correlation coefficient for Stim1 and Orai1 colocalization in Figure 7 panel L. Therefore we concluded that, Orai1 forms ‘nonfunctional’ clusters with Stim1 in STX11 depleted cells, please see modified lines 408-412, clearly explaining this.

      Although syntaxin 1A has been shown to interact with cholesterol (11) as well as PIP2 (7, 8) using either a stretch of polybasic residues or basic residues spread throughout several domains. To our knowledge, these have not been proposed to recruit or segregate STX11 in membranes. STX11 doesn’t contain an obvious stretch of poly-basic residues in its sequence, either, to quickly mutate and address this question. Please also see our response to your point #1 and #2 above. Answering this question will require a systematic and dedicated mutagenesis study.

      Minor:

      (1) Please indicate in Figure 1 in the respective graphs in which cell type the Ca2+ imaging studies have been performed.

      Done.

      (2) Figure 4C, D: Why are the input bands so weak?

      Please see our response to reviewer #1’s similar comment 12 above.

      (3) Figure 5A: Please clarify what WGA is.

      WGA is wheat germ agglutinin which is used to mark PM in imaging experiments. It binds to N-acetyl-D-glucosamine and sialic acid residues found in mammalian cell membranes and glycoproteins. We have added the explanation to the new Figure 6A legend.

      The authors state that "the constitutively active ANSGA (261-265) mutant of Orai1 (Supplementary Figure 11G), which harbors 4 consecutive mutations in the Orai1 C-terminus ...". Please clearly state this is the nexus region connecting the C-terminus with TM4. The 5 aa stretch is not the C-terminus; it is just close to the C-terminus.

      We have modified this, as suggested, in line 470-471.

      References:

      (1) Miao Y, Miner C, Zhang L, Hanson PI, Dani A, Vig M. An essential and NSF independent role for alpha-SNAP in store-operated calcium entry. Elife. 2013;2:e00802.

      (2) Li P, Miao Y, Dani A, Vig M. alpha-SNAP regulates dynamic, on-site assembly and calcium selectivity of Orai1 channels. Mol Biol Cell. 2016;27(16):2542-53.

      (3) Chorev DS, Baker LA, Wu D, Beilsten-Edmands V, Rouse SL, Zeev-Ben-Mordehai T, et al. Protein assemblies ejected directly from native membranes yield complexes for mass spectrometry. Science. 2018;362(6416):829-34.

      (4) Dorwart MR, Wray R, Brautigam CA, Jiang Y, Blount P. S. aureus MscL is a pentamer in vivo but of variable stoichiometries in vitro: implications for detergent-solubilized membrane proteins. PLoS Biol. 2010;8(12):e1000555.

      (5) Vig M, Peinelt C, Beck A, Koomoa DL, Rabah D, Koblan-Huberson M, et al. CRACM1 is a plasma membrane protein essential for store-operated Ca2+ entry. Science. 2006;312(5777):1220-3.

      (6) Feske S, Gwack Y, Prakriya M, Srikanth S, Puppel SH, Tanasa B, et al. A mutation in Orai1 causes immune deficiency by abrogating CRAC channel function. Nature. 2006;441(7090):179-85.

      (7) Murray DH, Tamm LK. Clustering of syntaxin-1A in model membranes is modulated by phosphatidylinositol 4,5-bisphosphate and cholesterol. Biochemistry. 2009;48(21):4617-25.

      (8) van den Bogaart G, Meyenberg K, Risselada HJ, Amin H, Willig KI, Hubrich BE, et al. Membrane protein sequestering by ionic protein-lipid interactions. Nature. 2011;479(7374):552-5.

      (9) Miranda P, Contreras JE, Plested AJ, Sigworth FJ, Holmgren M, Giraldez T. State-dependent FRET reports calcium- and voltage-dependent gating-ring motions in BK channels. Proc Natl Acad Sci U S A. 2013;110(13):5217-22.

      (10) Chang CL, Chen YJ, Liou J. ER-plasma membrane junctions: Why and how do we study them? Biochim Biophys Acta Mol Cell Res. 2017;1864(9):1494-506.

      (11) Lang T, Bruns D, Wenzel D, Riedel D, Holroyd P, Thiele C, et al. SNAREs are concentrated in cholesterol-dependent clusters that define docking and fusion sites for exocytosis. EMBO J. 2001;20(9):2202-13.

    1. Author response:

      eLife Assessment

      This valuable study provides insights into the role of steroid signaling during tumorigenesis in the adult male drosophila accessory gland (functional equivalent of the prostate gland in mammals), hinting at a possible counterintuitive anti-tumoral role of sex hormones during prostate cancer in certain patients. While the Drosophila model provides an elegant way to study the hypothesis derived from the Cancer Atlas analysis, the analyses of public prostate cancer expression data are incomplete and critical knowledge on patients' treatment modalities and normalization across different datasets is missing. This work would be of interest to prostate cancer researchers as it suggests that the absence of androgen receptor signaling in humans could constitute a mechanism promoting tumor escape.

      We thank the reviewers for the time they have spent on the manuscript, the production of a public review and their useful recommendations. As a general goal for the corrected version, we will try to provide more insight on the data (especially the human data), and more controls, to strengthen our conclusions. We are aware of the lack of a definitive proof of the role of the apparent decrease in AR signaling on tumour progression, but hope that this manuscript will encourage medical scientists to test/challenge its counterintuitive results in large cohorts of tissues and mouse/human models.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this article, Vialat and his colleagues examine the early stages - which remain largely unknown - of the tumor escape process, particularly the basal extrusion of tumor cells following endocrine therapy for prostate cancer.

      They first used the "Prostate Cancer Atlas" database, which provides access to a vast amount of transcriptomic data, to perform high-throughput analyses. Interestingly, analyzing a series of Androgen Receptor (AR) target genes in castration-resistant prostate cancers, they concluded that the loss of the canonical AR signaling pathway may contribute to tumor resistance.

      Using a well-established model in Drosophila, they then replicated in vivo an endocrine therapy targeting the accessory gland by genetically inhibiting the expression of ecdysone, the only sex steroid present in Drosophila. These experiments induced basal extrusion similar to the mechanism observed in tumor escape in humans.

      These results suggest that the deprivation of sex steroids may play an important role in tumor progression.

      However, although the data from the "Prostate Cancer Atlas" constitutes a powerful tool that serves as the basis for this new concept, clinical validation using carefully selected human tumor samples would help strengthen the authors' conclusions.

      Strengths:

      (1) The Prostate Cancer Atlas is a comprehensive collection of clinical data derived from RNA sequencing and serves as a powerful tool for conducting high-throughput analyses in this paper.

      (2) The Drosophila model used in this article is well established and has already been the subject of publications by the team. In addition to being an in vivo model, Drosophila offers a threefold advantage for this study: i) the presence of an accessory gland, similar to the prostate, which allows for the simulation of tumor formation and, in particular, extrusion mechanisms; ii) its regulation by a single sex steroid, ecdysone; iii) the genetic ability to modulate or inactivate ecdysone expression, which allows for a parallel to be drawn with hormonal deprivation in humans.

      (3) This study presents interesting and original findings. The data are, for the most part, of high quality.

      Weaknesses:

      (1) The Prostate Cancer Atlas, which is an essential tool in this study, was described only briefly - if at all - in both the introduction and the "Materials and Methods" section. The selection criteria used to distinguish CRPC or NEPC from adenocarcinoma in the Atlas or as determined by the authors, as well as the analytical methods, were not specified. It is therefore difficult to be convinced by the results, particularly those presented in Figures 1 and 2.

      As the tool has been published in different articles, we chose to limit its description. However, we agree that explanation are necessary, that will be added in the new version. First, we initially used here just basic categories, in order to avoid any possible bias; so mCRPC includes rare DNPC and NECP patients. We will also put the data with true ARPC, with essentially the same results.

      For human data, we also expect to use transcriptomic data from an independent cohort to check whether the same loss of AR signaling occurs during progression. Furthermore, we consider to add the data showing that decrease in canonical AR signaling also (logically with the previous results) correlates with castration status or ADT exposure (these info are available on ProstateCancerAtlas). Interestingly, and this can be put in supplementary data, prostatecanceratlas detects changes in EMT genes or proliferation genes that are coherent with what is known about cancer progression, indicating that the apparent decrease in AR signaling should correspond to a real phenomenon.

      (2) Although the hypothesis put forward by the authors - that the deprivation of sex hormones contributes to tumor progression - is strongly supported by the Drosophila model and by the in silico analysis of transcriptomic data from the Atlas, this concept still needs to be clinically validated by analyzing a series of prostate cancer samples, either through transcriptomic analysis or by tracking gene expression in histological sections.

      There are many indirect evidences that loss of AR signaling induces tumor progression in mouse (as stated in the intro or the discussion of the manuscript). However, as suggested in the introduction of the letter, we believe that medical scientists are the most qualified to prove that sex steroid deprivation indeed induces tumor progression in human. We will add in any case data to at least reinforce this puzzling finding of a decrease in AR canonical signaling during progression.

      (3) With regard to the cells responsible for tumor escape, stem cells have been described as "candidates for the initiating resistant tumor growth" (lanes 50-55), but it is also essential to address the recent concept of "persistent cells". Indeed, these cells have been primarily associated with their tolerance to treatment (chemotherapy) and are referred to as "drug-tolerant cells". However, persistent cells could also correspond to cells that evade hormone therapy in the case of prostate cancer. This possibility should be discussed in the article.

      This is an interesting suggestion, which can be discussed: on the one hand, intrabasal cells may not have accumulated mutations to survive the loss of EcR signaling, as would do persistent cells. On the other hand, they strongly proliferate, and show no sign of senescence, behaving more like resistant cells. So, it does not look to us that we induced the appearance of persistent cells in the Drosophila accessory gland, except if these cells are quickly reactivating to give rise to intrabasal cells.

      Reviewer #2 (Public review):

      Summary:

      In this study, Vialat and collaborators study the role of steroid hormone signalling on the development of prostate cancer (patients) and of accessory gland tumours in Drosophila, a tissue functionally equivalent to the prostate. Mining publicly available prostate cancer expression data and using gene expression signatures, they uncover that androgen signalling is actually down-regulated in castration resistant prostate cancers (CRPC) compared to "primary" cancers, leading the authors to wonder whether down-regulation of canonical androgen signalling could represent an important event increasing tumour aggressiveness. They then take advantage of their recently published tumour model in the accessory gland of Drosophila adult males, in which cells are primed for tumorigenesis by the constitutive activation of the EGFR receptor, to test directly this hypothesis. They show that the genetic invalidation of ecdysone reception and signalling increases the aggressiveness of the "pre-cancerous" lesions, and that ecdysone-insensitive tumours present higher proliferation and initiate basal extrusion.

      Strengths:

      The authors bring original observations on the role of ecdysone signalling to prevent male accessory gland tumour development in Drosophila

      Weaknesses:

      (1) The link between the human data mining and Drosophila model is not straightforward.

      (2) Important information, in particular clinical information, is missing in the presentation of the cancer patients' data, making it complicated to grasp the solidity of the claims.

      (3) Data-mining insights should be validated by orthogonal approaches.

      (4) Ecdysone signalling activity should be monitored.

      While the two parts of the study both investigate the role of steroid signalling on tumour growth, the link remains slightly artificial. I think starting with Drosophila and then opening with some patient data would be better suited to the level of proof reached here, implying that the anti-tumoral role of steroids observed experimentally in the fly might be conserved based on data mining in patients, rather than trying to prove in the fly the hints gained from public data mining. Indeed, there are many important differences between the mammalian prostate and the fly accessory gland, as well as between sex hormone androgen signalling and developmental timing ecdysone signalling.

      This is an interesting suggestion. Actually, we first wrote the manuscript by starting with Drosophila data and then going to patients data, and previous reviewers said that this was not possible to directly go from Drosophila to human. So, we suppose that the real way to solve this will be by the validation or refutation of the data by other teams in different models.

      The prostate cancer data mining and re-evaluation brings some interesting observations that appear to challenge the androgen driver, contrary to the vast amount of literature. Indeed, the authors observe an apparent decrease in androgen signalling in the more advanced states of the disease, in particular CRPC. In order to better evaluate its clinical relevance, more background on the tumours analysed should be provided.

      This point is also important to Reviewer 1, and will be implemented.

      What treatments were received by the patients? Hormonotherapy? LH/RH analogues? +/- anti-androgens? Are these treatments still given when CRPC emerge and tissues were banked? Metastatic disease? Are these only primary tumours in situ? Are there metastases included in the analyses?

      Most CRPC come from metastatic sites. Most of the CRPC were treated by ADT. We chose to have an approach that included all the samples; but we will provide insight, whenever available, on these absolutely relevant questions.

      Frequently, castration resistance is associated with alternatively spliced variants of the AR (AR-V7) that become constitutive and could bind to new AR-sensitive enhancers, even in the absence of androgen. Is the splice variant status of patients known, or could it be inferred from the expression data? Would there be different responses according to AR-V7 status?

      As a first approach, from the cohorts that were used, it seems that in the PCA patients, there are around or less than 15% of patients harboring the AR-V7 driver. It could be of interest to test their behavior regarding the same set of genes, and it will be done if we can identify the patients.

      Regarding the signature used. Why not monitor PSMA, one of the major prostate cancer markers, which is regulated by AR?

      It will be done. PSMA behaves as the others, even though the drop between primary samples and ARPC samples is very limited and just statistically significant.

      Finally, to consolidate the surprising observation that AR signalling is repressed in CRPCs, the authors should back these in silico predictions with orthogonal approaches such as histochemistry on patients' TMA or tissues from mouse models, monitoring AR activity.

      As said previously, we believe that this specific work will be better done by medical scientists.

      Regarding the fly experiments, the observation that ecdysone signalling depletion cooperates with EGFR-lambda activation to generate big overgrowths that delaminate basally without passing through the muscular sheet is interesting. However, several important controls need to be provided in order to support the claims:

      Considering the comments regarding the fly experiments, we agree that, if experiences are taken individually, controls are lacking. However, we have to explain our strategy and why the results taken in their entirety have a significance. In our model of epithelial tumorigenesis, we have started to explore the EcR pathway after years of work on other pathways. At the first experiment (with the EcR RNAi line), we were struck by the intrabasal phenotype that did not occurred in our previous experiments, and especially for the 14 RNAi lines that we published in two independent articles on Ras/MAPK, Pi3K/Akt pathways and cholesterol metabolism. As justly said in the review, many unexpected effects can happen, so we decided to explore the role of five other genes of the same pathway to be sure of the reproducibility of the phenotype when we block the EcR pathway. The odds of having the same specific phenotype for 6 lines of the EcR pathway when there is always another phenotype for 14 lines targeting other pathways can be calculated: p = 0.00000494. So, the best control we offer, and it is largely significant, is the repetition of the experiments intended to downregulate the EcR pathway, that produce the same phenotypes independently of the target.

      Furthermore, all lines we used were previously tested, validated and most of the time published in other scientific works. This is essential to us, as one complexity of doing rare clones in an otherwise normal tissue is that decreasing an mRNA in less than 5% of the cells of course difficultly leads to a detectable drop in overall expression in the whole gland. This is also the reason why we always tried pairs of fly lines to block the receptor activity itself (RNAi EcR, RNAi Shd), the receptor's downstream targets (RNAi HR3, RNAi HR4), and the production of ecdysone (RNAi Sad, RNAi Phtm). As we validated RNAi Sad, we can try anyway to validate at least another RNAi of another category. Furthermore, we did use a RNAi White control: it behaves in the same way as the GFP control. We will put the results comparing the two lines in supplementary data.

      (a) The authors should use an ecdysone reporter (ERE-LacZ, ERE-GFP...) to monitor and show that Ecdysone signalling is indeed lower in the tumours after genetic manipulations, or that it is higher in EGFR-lambda small clones.

      This would be of interest to validate that the 6 lines are behaving in the same way (at least, they give similar phenotypes). However, in the adult accessory gland, ERE activity is largely lower than during development (DOI: 10.1016/j.jinsphys.2011.03.027), and to be able to decrease it, authors had to express notoriously strong dominant-negative EcR-DN. We can try the experiment but are really not persuaded that we will be able to see a drop of activity with only a decrease of expression of the gene. If we can think of another solution that could be more efficient, we will try it as the idea is of course interesting.

      (b) EcR is normally a repressor, which is turned into an activator in the presence of 20-hydroxyecdysone. The removal of EcR could lead to de-repression of genes and thus slightly activate the pathway. Monitoring ecdysone signalling activity is thus critical.

      Actually, there are different EcR isoforms. EcR-B1 is generally considered as the main activator of the pathway, as EcR-A is a repressor of the pathway. The EcR RNAi line which was used does not target a specific isoform.

      (c) The authors should also monitor the expression of Phantom, Shadow, Shade, and EcR in the different accessory glands (wild-type, EGFR-lambda, EGFR-lambda & EcR-RNAi). It is extremely surprising that systemic ecdysone has so little role since Phantom, Shadow, or Shade RNAi appear as potent as EcR-RNAi. This quantification has actually been performed for Sad in Figure S5, which is not even mentioned in the text. It should be done for Phtm.

      The levels of ecdysone are tenths of times lower in adult compared to the peaks during embryogenesis or metamorphosis. And one source of production is the epithelial cells of the accessory glands themselves. Considering that EcR is expressed in all the cells of the accessory gland (epithelial cells and muscle cells), it seems plausible that there is only a very little amount of ecdysone that can in fact be available for the other epithelial cells.

      A UAS-yellow-RNAi (or similarly irrelevant RNAi) rather than UAS-GFP should be used as a control for the EcR, Phtm, Sad, Shd, and Tub RNAi. Indeed, loading the RNAi machinery could have some unexpected effects not controlled by the UAS-GFP.

      The NLSGFP line we used here is the one we already published twice (and we compared it to RNAi lines), and this is the reason why we used this already validated control. However, we tested a RNAi White line, and it behaves in the same way.

      The authors should not use the term "sex steroid" when referring to ecdysone. It is a steroid hormone important for developmental timing and rate of growth, but is not a sex hormone, as sex is cell autonomously genetically determined in the fly.

      In human, sex hormones control sexual differentiation (up to adult characteristics) and sexual reproduction. In Drosophila, ecdysone controls sexual reproduction in both sexes and sexual differentiation at least in the female (DOI: 10.1007/s004270050186). From these results, we do not think that saying it is a sex steroid (not a sex hormone) is ill suited. We intend to precise what we put in this term in the introduction to avoid overinterpretation from our part.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, they generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.

      The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      A major question that remains, is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      The authors need to discuss this issue explicitly to make it clear that conservation of the mechanism is not complete and that other interpretations are possible.

      The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the the acidic residues and their mutants are not shown.

      This has been addressed.

      There are some issues with figure S2B and S5.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

      We are grateful to the Editors and reviewers for their careful assessment of our revised manuscript and for identifying the remaining points that required clarification. We have addressed each of these comments in the present revision. We believe that these revisions have resolved the remaining concerns and have further improved the clarity and accuracy of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please discuss the large difference in the effect of the mutants on activity explicitly

      According to the reviewer’s suggestion, we have added the following discussion to the revised manuscript:

      “Although mutation of the two acidic residues impaired the ATPase activity in both aqMutL and aqGyrB, the magnitude of the effects differed substantially, with much greater reduction in the catalytic efficiency in aqMutL than in aqGyrB (Table 1). The molecular basis for this quantitative difference is currently unclear. One possible explanation is that subtle differences in the active-site architecture and surrounding residues alter the relative contribution of each acidic residue to catalysis, allowing aqGyrB to tolerate perturbation of either residue more effectively than aqMutL.” (p.6 line 257-262 in the revised manuscript)

      (2) Figure S2B is a completely different view from the other panels, please provide the correct one.

      We have revised Supplementary Fig. S2B so that the E48A structure is now shown from a viewpoint as similar as possible to those used in the other panels. We note, however, that the E48A structure cannot appear completely identical to the other panels because the E48A mutant does not bind the ATP analog and therefore does not undergo the nucleotide-binding-associated conformational changes observed in the other structures.

      (3) S5 : It is not clear to me what is meant by " The scale bar indicates the number of amino acid substitutions per site." : there is a 'Tree Scale 1" but no other numbers in my version.

      We thank the reviewer for pointing out that the scale bar in Supplementary Fig. S5 was insufficiently explained. The value “1” in the tree scale corresponds to a branch length of one amino acid substitution per site. To avoid ambiguity, we have revised the scale-bar label in Supplementary Fig. S5 to explicitly indicate “1 substitution/site” and have clarified its meaning in the figure legend:

      “Branch lengths are proportional to the evolutionary distances inferred by IQ-TREE. The scale bar represents an evolutionary distance of one amino acid substitution per site.” (p. 22 line 767-769 in the revised manuscript)

      Reviewer #2 (Recommendations for the authors):

      (1) P. 9, in the "Data Accessibility Statement", all three PDB codes (23UX, 23UY, and 23UZ) should be listed.

      We thank the reviewer for this comment. We carefully rechecked the Data Accessibility Statement and confirmed that all three PDB accession codes (23UX, 23UY, and 23UZ) are included in the statement.

      (2) Supplementary Figures S2 and S3. The authors have written "Asn33" and "Asn52", instead of "Glu32" and "Asp51" in both the figure and figure legend of Supplementary Figure S3. They have also written "TND" instead of "NTD" in the figure legend.

      “Asn33” and “Asn52” in Supplementary Fig. S3 are not typographical errors. Asn33 in aqMutL and Asn52 in aqGyrB are the residues that directly coordinate the Mg<sup>2+</sup> ion and are distinct from the acidic residues discussed in this paper. To avoid confusion, we have added the following sentence to the legend of Supplementary Fig. S3:

      “These Mg<sup>2+</sup>-coordinating asparagine residues are adjacent to, but distinct from, the second acidic residues Glu32 in aqMutL and Asp51 in aqGyrB examined in this study.” (p. 21 line 749-751 in the revised manuscript)

      We have also corrected the typographical error “TND” to “NTD” in the figure legend. (p. 21 line 749 in the revised manuscript)

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Bajohr and colleagues propose a transcription factor-driven approach to generating bonafide oligodendrocyte lineage cells (OLCs) from primary mouse astrocytes. Ectopic expression of Olig2, Sox10, or Nkx6.2 in isolated astrocytes produced a range of OLC-like cell states, with Sox10 emerging from lineage tracing and single cell RNA sequencing experiments as the most successful transcription factor in driving direct lineage reprogramming. The authors strengthened their claims with an unbiased, deep learning perturbation model to predict genetic drivers of the astrocyte cluster to OLC cluster transition observed in their scRNA seq dataset. Here, Sox10 surfaced in the top ten correlated genes, and the top transcription factor, mediating this fate shift. Altogether, this paper presents an interesting approach to generate OLCs, a cell type historically difficult to procure, from primary mouse astrocytes to study this lineage in development and disease and perhaps repopulate it in dysmyelinating conditions. While this certainly addresses a technical gap in the field, authors defined iOLCs as ones with lineage-specific gene expression and morphological characteristics, lacking any functional analysis to assess the reprogrammed cells' capacity to myelinate. This comment and other critiques are discussed below.

      While Sox10 and Mbp expression in iOLCs, as confirmed by IHC, is a promising result suggesting that ectopic Sox10 instructs transduced cells to develop into cells of myelinating potential, functional confirmation is essential. As mentioned in the discussion, the absence of a substrate for myelination may have also contributed to the low DLR efficiency. Co-culturing Sox10 iOLCs with primary neurons and examining the cells' potential to engage and enwrap axons would greatly strengthen the authors' claim that this could be an effective therapeutic approach to myelin regeneration in vivo, or even a technical approach to studying myelin dynamics in vitro.

      In Figure 1B, it appears that Mbp expression in tdTomato+ cells decreases in Sox10 transduced iOLs during the observed time period. Can the authors elaborate on this result, given that MBP expression is crucial for myelination and should, if anything, increase with time?

      The authors acknowledge that there is a conversion of tdTomato- zsGreen+ cells with an astrocyte-like morphology to OLC cells expressing Mbp following Sox10 induction (Supplementary figure 5C,D). While they note the diversity of the astrocyte lineage in the discussion, further analysis should be applied to this subset of cells to confirm the subset of astrocyte or progenitor-like cell type that gives rise to their cell endpoint of interest (Sox10-driven Mbp+ iOLs).

      Finally, ectopic expression of Olig2 and Sox10 in primary astrocytes resulted in very different OLC subtypes, as evidenced by OLC marker expression seen in IHC and the subclustering of these cell types in scRNA seq. Although this diversity in OLC type and generation efficiency follows with previous reports showing that these two transcription factors vary in effect, might the authors further discuss this discrepancy given that the two transcription factors regulate one another (as mentioned in the introduction) and should theoretically give rise to more similar cells? Perhaps due to the lower specificity of Olig2 in marking a pure OLC population relative to Sox10?

      We thank the editor and reviewers for their additional comments, which have significantly improved our manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors use Aldh1l1+ astrocyte with a GFAP promoter linked to TF expression, claiming that the converting cells are cortical astrocytes. However, during early mouse development, radial glia that can give rise to multiple cell types, including oligodendrocytes, also express Aldh1l1 and low levels of GFAP. Therefore, it is not proven whether the resulting iOLs came from mature astrocyte or a radial glia population. Especially since in Figure 3B,G it is shown that a majority of the D0 population expresses high amounts of Vimentin and Nestin, both markers associated with radial glia and immature astrocytes. It would be beneficial for the authors to confirm that either there are no contaminating radial glia or that the radial glia don't express the TF, especially since the iOL population is so small.

      Thank you for this comment. We agree that previous studies have demonstrated Aldh1l1 expression in radial glia cells [1]. However, our dissection protocol to obtain the postnatal astrocytes is cortex-specific and does not take portions of the VZ/SVZ, preventing radial glia contamination.

      Nevertheless, to confirm the astrocytic identity of our Aldh1l1+ starting cells we used AUCell enrichment scoring [2]. First, Aldh1l1+ cells were subclustered from our starting culture single cell dataset (Author response image 1A). We then defined two gene signature modules: an “astrocyte” module, comprised of canonical astrocyte markers (Aqp4, Gja1, Slc1a2, Glul, Aldoc, S100b, Nfia, Thbs1, Cst3, Clu), and a “radial glia” module, containing common radial glia and progenitor markers (Pax6, Fabp7, Sox2, Hes1, Prom1, Top2a, Mki67, Ube2c, Cdk6, Mcm2). When we scored all Aldh1l1+ cells (n=869) for enrichment of each signature, no cells were classified as radial glia (Author response image 1B). Instead, the astrocyte signature predominated, with 68.3% classified as astrocytes (Author response image 1B,C). The remaining cells (36.7%), were classified as transitional, reflecting substantial expression of genes from both modules (Author response image 1B,C). Therefore, although the Aldh1l1 astrocytes do express genes common to radial glia, there are no cells that express only progenitor markers. This is consistent with literature showing that many radial glia genes are commonly found in astrocytes [3], [4], [5].

      Taken together, our stringent dissection protocol and bioinformatic profiling of our starting Aldh1l1 cells suggests that the resulting Aldh1l1+iOLs are originating from astrocytes, rather than radial glia.

      In Figure 2D, Sox10 and Nkx6.2 have an n=4 while the Cre control has an n=3. Why is this the case? Was the 4th point excluded? Similar attention should be given to other panels in Figure 2 for consistency.

      We thank the reviewer for highlighting this. No outliers or datapoints were excluded from the analysis. Rather, the fourth culture for our control treated cells was not viable for analysis.

      Author response image 1.

      Astrocyte gene expression in Aldh1l1+ cells. (A) UMAP clustering of Aldh1l1+ cells in our starting cultures (0DPT, non-transduced). (B) UMAP clustering from (A) overlayed with cell classification based on AUCell gene signature expression scoring. (C) Feature plots showing AUCell enrichment scores (darker purple indicates higher enrichment) for the astrocyte signature (left), radial glia signature (middle), and the differential enrichment score (right, Astro_AUCell - RG_AUCell) (darker purple scores indicate higher astrocyte signature and negative scores (gray) indicate higher radial glia signature).

      Representative images in Figure 2 do not convincingly support the argument by the authors. It appears that some of the cells highlighted by the arrows are just background (e.g. PDGFRa and td Tomato in Figure 2E, or zsGreen in Figure 2F). Additionally, the authors should show a different representative image depicting astrocyte morphology in Figure 2G 7DPT.

      Thank you to the reviewer for this comment. We have replaced the images in Figure 2E,F to better represent our findings (updated manuscript Figure 2E,F). We have also adjusted the representative image in Figure 2G 7DPT to better visualize the astrocyte morphology (updated manuscript Figure 2G) as well as included as supplementary additional examples of pre-conversion astrocyte morphology to supplement our morphology analysis (Author response image 2).

      Author response image 2.

      Lineage tracing confirms true conversion of astrocytes to oligodendrocyte lineage cells. Representative images of astrocyte morphology observed prior to cell conversion (arrow indicates converting cells, scale bar =50um).

      References

      (1) L. C. Foo and J. D. Dougherty, “Aldh1L1 is expressed by postnatal neural stem cells in vivo,” Glia, vol. 61, no. 9, pp. 1533–1541, Sep. 2013, doi: 10.1002/glia.22539.

      (2) S. Aibar et al., “SCENIC: Single-cell regulatory network inference and clustering,” Nat Methods, vol. 14, no. 11, pp. 1083–1086, Nov. 2017, doi: 10.1038/nmeth.4463.

      (3) M. Götz and Y.-A. Barde, “Radial Glial Cells: Defined and MajorIntermediates between EmbryonicStem Cells and CNS Neurons,” Neuron, vol. 46, no. 3, pp. 369– 372, May 2005, doi: 10.1016/j.neuron.2005.04.012.

      (4) P. Malatesta, I. Appolloni, and F. Calzolari, “Radial glia and neural stem cells,” Cell and Tissue Research, vol. 331, no. 1, pp. 165–178, 2008, doi: 10.1007/s00441-0070481-8.

      (5) S. Clavreul, L. Dumas, and K. Loulier, “Astrocyte development in the cerebral cortex: Complexity of their origin, genesis, and maturation,” Front Neurosci, vol. 16, p. 916055, Sep. 2022, doi: 10.3389/fnins.2022.916055.

    1. Author response:

      eLife Assessment

      This study presents a useful database resource containing protein conformations generated through molecular dynamics simulations, with extensive quality evaluation and benchmarking. While the database is well-constructed and professionally organized, the evidence supporting its claimed representation of protein conformational landscapes is incomplete, as the short simulation times and starting structure bias prevent true Boltzmann sampling of the conformational space.

      We thank the editors for recognizing the usefulness of ProteinConformers and the value of its quality evaluation and benchmarking. We will revise the manuscript to clarify that ProteinConformers provides large-scale, energetically profiled descriptions of protein conformational landscapes, with broad coverage of locally stereochemically valid and energetic compatible structures from non-native to near-native regions, rather than a complete equilibrium sampling of all conformational states. These revisions will better define the scope of the resource while preserving its intended use for benchmarking and data-driven studies of protein conformational variability.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a new database that rigorously explores protein conformations.

      Strengths:

      It is extremely well done, using state-of-the-art tools by a group at the top of the field of structural modeling. The evaluation of qualities and the benchmarking of the structures are outstanding, and it is expected that the new database will have a significant impact on the field.

      We thank Reviewer #1 for the positive evaluation of our work and for recognizing the potential impact of the ProteinConformers resource.

      Weaknesses:

      The authors are using MD simulation to generate some of the structure, and therefore should have access to standard MD energies. I am surprised that no evaluation is provided based on these energies that can be extended to free energies.

      We thank the reviewer for this helpful suggestion. We reprocessed the original MD energy files and extracted three MD-derived energy terms, including total energy, system potential energy, and protein-only potential energy. These energy terms have been added to the ProteinConformers resource and the web portal. We will update the manuscript to describe these additional energetic annotations.

      Reviewer #2 (Public review):

      Summary:

      The authors developed a dataset of protein conformations by running molecular dynamics simulations starting from both native and decoy conformations for a large number of proteins. These conformations were put together as a dataset for querying and downloading, along with their energies under different force fields. The authors suggest that such conformations represent the proteins' conformational landscape, so that they will be useful for evaluating methods generating multiple conformations of proteins.

      Strengths:

      The dataset is online and working. It has good documentation for others to use.

      We appreciate Reviewer #2’s positive assessment of the online resource and documentation.

      Weaknesses:

      The biggest weakness is that the collected conformations very likely do not represent the true conformational landscape. To represent the conformational landscape, the structures need to be sampled based on the Boltzmann distribution. However, in this study, conformations are generated by running very short (125ps to 375ps) MD simulations starting from near-native conformations and decoys. Such short simulations will produce small fluctuations around the starting conformations, so the distribution of conformations is largely dominated by the distribution of the initial conformations, which by one means are Boltzmann distributed. A conformation might be physically plausible, but it might have very small weight in the Boltzmann distribution. On the other hand, conformations with large weights might not be in the dataset.

      We thank the reviewer for this important and constructive comment. We agree that the conformations in ProteinConformers should not be interpreted as an equilibrium ensemble sampled according to the Boltzmann distribution. Because the MD simulations used here are short, the resulting snapshots around each seed mainly reflect local relaxation and limited thermal fluctuation from that seed, rather than exhaustive equilibrium sampling. Therefore, the relative population of conformations in our dataset should not be interpreted as a Boltzmann weight, and some thermodynamically important states may be underrepresented or absent.

      Our goal in this work is different from conventional long-timescale MD studies that aim to estimate equilibrium populations from one or a few initial structures. ProteinConformers was designed as a large-scale, multi-seed, MD-refined conformer resource. The broad structural coverage comes primarily from initiating simulations from many diverse seed decoys for each protein, while the short all-atom MD protocol is used to relax structures under a molecular mechanics force field, remove structures that fail to converge, reduce steric clashes and unrealistic local geometries, and generate energetically annotated conformers. We will revise the manuscript to make this distinction clearer and to avoid implying that ProteinConformers provides a rigorous Boltzmann-sampled representation of the underlying thermodynamic landscape.

      We also agree that longer simulations and enhanced sampling methods, such as replica-exchange MD, metadynamics, or umbrella sampling, would be necessary to estimate equilibrium populations and improve sampling of rare but thermodynamically relevant states. We will add this point as a limitation and future direction in the revised manuscript. Thus, ProteinConformers should be viewed as a broad, energetically annotated, MDrefined conformer library for benchmarking, data-driven modeling, and descriptions of protein conformational landscapes, rather than as a complete equilibrium ensemble itself.

      Reviewer #3 (Public review):

      Summary:

      This manuscript describes a web-based tool that allows researchers to compare large numbers of representative ("plausible") conformations of proteins. It also includes energetic analysis from multiple widely used structure-prediction methods.

      Strengths:

      This tool will likely be useful for students who want to learn more about the ensemble properties of proteins. The resource is well organized and it represents a large amount of computing resources.

      We thank Reviewer #3 for the positive assessment of the ProteinConformers resource and for recognizing its potential value for community education.

      Weaknesses:

      It is not entirely clear how the database may be utilized by other groups to advance research. It could be helpful if the authors add a short section that provides example use cases that illustrate how this database can support new strategies for studying protein dynamics.

      We thank the reviewer for this constructive suggestion. We agree that the manuscript should more explicitly explain how other groups can use ProteinConformers to advance research. In the revised Discussion, we will add a short section describing concrete use cases of the database. In particular, we will emphasize that the benchmark analysis already presented in this manuscript provides a worked example of how ProteinConformers can be used by other groups. ProteinConformers-lite, together with the released evaluation metrics and codes, can serve as a standardized reference set for testing new multi-conformation or protein ensemble generation methods. Other groups can generate conformational ensembles for the same targets, compare their coverage of low-energy regions using the diversity metrics reported in Table S1, and evaluate the agreement of residue-pair geometric statistics using the plausibility metrics reported in Table S2.

      We will also describe additional use cases enabled by the full ProteinConformers resource. The dataset can be used as a training, validation, or pretraining resource for conformation generators, energy-aware ranking models, and model quality assessment methods, because each conformer is paired with structural similarity annotations and multiple energetic scores. The broad coverage from non-native to near-native conformations also enables systematic analysis of how local stereochemical validity, global structural similarity, and energetic evaluations co-vary across diverse conformational perturbations, and may provide useful structural proxies or starting points for modeling flexible or disordered-like protein states, where experimentally resolved structural data are often limited. In addition, the interactive portal allows users to filter protein-specific conformers by structural similarity, energetic annotations, and secondary-structure features, making it possible to construct customized subsets for downstream biomolecular modeling, hypothesis generation, and educational exploration.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Hall et al provides a novel and important new finding about communication between the anterior cingulate cortex (ACC) and the CA1 region of the dorsal hippocampus: there is a clear ability of ACC to predict CA1 activity, and that is modulated by learning/experience. Furthermore, they have some evidence that the modulation differs by whether the CA1 neurons were in the deep versus superficial sub-layer of CA1. The evidence is suggestive of new and exciting findings, but some gaps and weaknesses remain to be addressed before I believe all of the authors' claims can be supported. The figures also need to be slightly better organized, and the discussion is missing a major dimension in my opinion. Overall, this is a strong submission, but with some gaps to fill.

      Strengths:

      (1) This is a well-written manuscript - the introduction was especially clear, well-cited, and motivating.

      (2) The sub-layer specific communication between ACC and CA1 represents the discovery of a novel and functionally impactful piece of neurobiology.

      (3) Optogenetics was an important verification of ACC-CA1 communication, as was the analysis of neurons by waveform type.

      Weaknesses:

      (1) Figure 2: Why are the data separated into two groups from the outset? If all data are combined, is there a general drop in prediction gain from pre to post?

      Thank you for bringing this to our attention. In Figure 1F, all data is combined for GLM decoding. We found a significant pre-to-post decrease in prediction gain specifically using the –200 to 0 ms window to predict CA1 spiking during ripples. Figure 3 builds upon these findings to examine how prediction gain changes relate to task engagement.

      (2) 2b and 2c are important since they are complementary means to show the same thing, and it is important that they cross-validate each other, especially since the non-significant task active neuron difference in 2b appears to be nearly as strong as the significant difference to its left. A more holistic analysis can be done to compare these dimensions.

      We appreciate this feedback. In light of this comment, as well as similar concerns raised by other reviewers regarding the binary classification of neurons as task-active or task-inactive (modulation index > 0 or < 0), we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement. However, we were unable to replicate modulation index changes for neurons split by prediction gain score percentile and removed those figures from the manuscript. We have adjusted the text accordingly to account for this new effect.

      (3) Sup vs deep neuron definition: Did the authors have any means to validate this anatomical separation using histology or otherwise? I don't believe they described anything like that, and instead use physiology to infer anatomical location. I understand anatomy-based methods may be practically impossible with tetrodes, but this limitation should at least be mentioned, and it should be explained that without something like silicon probes or histological validation, anatomy had to be inferred from physiology.

      We think this is an important limitation to address and thank you for bringing this to our attention. Given our technical restraints, we only validated radial position through physiological properties. This remains an outstanding limitation of this study. We have since added text into the main discussion bringing attention to this limitation. See below:

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (4) Superficial vs deep differences in firing rate ratio based on PG: there are many fewer CAdeep neurons, but in 4c, the trends appear to be the same pre-training, top PG lower than others. It seems the lack of difference in CA1deep in 4c may be due to the much lower power/n. This should be discussed or addressed.

      We appreciate this feedback and since have added more recordings to address these lower Ns (CA1sup: Previous N = 71, Revisions N = 89; CA1deep: Previous N = 21, Revisions N = 61). Notably, we find that previous firing rate ratio (now called modulation index) is no longer significant with the inclusion of more neurons and is reported accordingly (Figure 2—figure supplement 1).

      (5) In Figure 5, the term "firing rate ratio" is used, and it sounds the same as in previous figures, but this is a different ratio (based on modulation by opto stim, not task).

      To improve clarity and avoid confusion with task-related modulation metric used across the paper, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion.

      (6) I would like to learn more about these v-type neurons. I understand we do not yet know about their molecular or morphologic correlate, but more analysis can be done with the current data.

      We thank you for this feedback. We have included further analysis into V-type properties. Namely, we performed autocorrelegrams, theta phase modulation, and burst index analyses. See Figure 6 and Figure 6—figure supplement 1.

      Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlogram analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type interneurons and CA1 pyramidal neurons.

      (7) I would like more discussion of ACC-CA1 connectivity.

      We have since added greater discussion of ACC-CA1 connectivity into the discussion section. See below:

      “Finally, an important caveat to mention is that it remains an ongoing debate whether ACC directly projects to CA1 (Andrianova et al., 2023; Rajasethupathy et al., 2015; Shi et al., 2022). One lab reported clear monosynaptic ACC-to-CA1 connection (Rajasethupathy et al., 2015), while another lab replicated those same experiments and were unable to come to the same conclusions (Andrianova et al., 2023). Further studies report no direct connection (Shi et al., 2022). The contention in connectivity may arise from differences in targeting strategies, injection coordinates, or viruses used. Our findings reported an excitatory response in CA1 V-type interneurons in response to ACC stimulations, proposing another possibility for ACCàCA1 connectivity. Interestingly V-type interneurons responded with extremely low-latency as fast as 4.2 ms after stimulations, compatible with monosynaptic timing (Cho et al., 2013; Petreanu et al., 2007; Wang et al., 2009). If the ACC→V-Type connection was monosynaptic pathway, it could help explain discrepancies in the field, as the relative sparsity of V-type interneurons may reduce the likelihood of detecting ACC→CA1 connectivity. However, future anatomical studies are needed to conclusively determine connectivity.

      Alternatively, ACC→CA1 communication may be mediated by multiple intermediate structures (Behzadi et al., 1990; Oh et al., 2014; Shi et al., 2022; Souza et al., 2022). The ACC sends monosynaptic projections to the nucleus reuniens (RE) and median raphe (MnR), both of which project directly to CA1 (Oh et al., 2014; Shi et al., 2022). The RE has a known role in contextual discrimination learning and memory specificity (Ramanathan & Maren, 2019; Ramanathan et al., 2018; Ratigan et al., 2023; Silva et al., 2021; Xu & Südhof, 2013). Interestingly, RE→CA1 activity tuned to immobility (freezing) emerges only after shocks are presented, suggesting a learning-induced modification between regions, similar to that seen in our ACC–CA1 data (Ratigan et al., 2023). As for MnR, it receives dense inputs from the ACC (Behzadi et al., 1990; Souza et al., 2022), and its projections to the CA1 are predominantly glutamatergic (Jackson et al., 2009; Senft et al., 2021; Szonyi et al., 2016). Notably, these glutamatergic MnR inputs directly target CA1 interneurons, including CCK basket cells (Miettinen & Freund, 1992; Morales & Bloom, 1997; Senft et al., 2021), while avoiding PV interneurons and pyramidal neurons (Acsady et al., 1993; Freund et al., 1990; Halasy et al., 1992; Miettinen & Freund, 1992; Papp et al., 1999; Turi et al., 2019). This connectivity suggests that the ACC may indirectly modulate CA1 activity through the MnR, potentially suppressing PV interneuron and pyramidal neuron activity via local inhibitory circuits, thereby contributing to the regulation of hippocampal oscillations and memory consolidation (Huang et al., 2022; Wang et al., 2015). Ultimately, future experiments combining pathway-specific manipulations with simultaneous recordings will be necessary to distinguish direct from polysynaptic mechanisms.”

      (8) Some elements may be missing from the discussion, relating baseline functioning versus post-learning function.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have added these discussions into the main text. See below:

      “Alternatively, ACC→CA1 communication may contribute to the homeostatic downscaling of memory-unrelated synapses during sleep. Evidence finds that slow-wave sleep is strongly linked to downscaling of non-learning related neuron activity (Gulati et al., 2017; Liu et al., 2010; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Slow-wave sleep ripples in particular depotentiate memory-unrelated synapses (Gulati et al., 2017; Norimoto et al., 2018). In our study, we find that learning-related reduction in communication between ACC and CA1 were selective for task-inactive neurons. Therefore, ACC→CA1sup communication may not simply weaken following learning but rather becomes selectively disengaged from task-inactive neurons, enabling homeostatic downscaling while preserving behaviorally relevant synapses (Liu et al., 2010; Norimoto et al., 2018; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Still, behavioral recruitment alone cannot account for the observed remodeling of ACC→CA1 communication, as CA1sup and CA1deep neurons did not display significant differences in task-related activity (Figure 2—figure supplemental 2). Instead, these learning-related changes of task-inactive neurons appear sublayer-specific.”

      Reviewer #2 (Public review):

      Summary:

      This study uncovers an inhibitory pathway from the anterior cingulate cortex (ACC) to pyramidal cells in the superficial sublayer of hippocampal area CA1 (CA1sup). As ACC neuron spiking tends to precede hippocampal ripples, this presents the intriguing possibility that ACC inputs are selectively inhibiting particular CA1sup neurons, which could play a role in the reactivation of task-related ensembles known to take place during hippocampal ripples. Indeed, through a generalized linear model (GLM) analysis, the authors demonstrate that the ACC activity within the 200ms immediately preceding the ripple is predictive of the ripple content.

      Strengths:

      The biggest strength of the work is the optogenetic manipulation experiments, which convincingly demonstrate that stimulation of ACC pyramidal neurons activates an interneuron population with symmetric spike waveforms, and inhibits parvalbumin interneurons and pyramidal cells in CA1sup but not CA1deep sublayer.

      An additional strength in the GLM analysis which consistently shows that ACC activity preceding the ripple is predictive of hippocampal activity during the ripple considerably more than in shuffled data for all cells and periods tested.

      Weaknesses:

      The major weakness of this work is that the link with learning and memory is not very well supported.

      The only evidence of rebalancing and reorganization appears to be a single statistical test (the test in Figure 1f, p=0.013) demonstrating a decrease of the GLM prediction gain from pre-task sleep to post-task sleep; the same test is repeated for subsets of the data in the rest of the figures. As the idea of rebalancing and reorganization is central to the paper as currently written, exploring it through another measure, independent of the GLM prediction gain, should be expected. The notion that this pathway is suppressed in sleep following learning can be supported by demonstrating a decrease in any of the following measures: ACC spike-triggered average CA1sup responses, cross-covariances (Wierzynski et al 2009) between ACC and CA1sup cells in post-task sleep, or ripple-triggered cross-correlations (Sirota et al. 2009).

      We thank the reviewer for this helpful feedback. We have added an additional analysis the reviewer pointed out to address this concern. Specifically, we performed an ACC spike‑triggered analysis. The ACC spike‑triggered average further supported the learning‑related decrease in ACC-to-CA1 activity (see Figure 1—figure supplement 1). We did not include a separate cross-covariance analysis because the ACC spike-triggered average captures essentially the same temporal relationship between ACC and CA1 activity.

      The differences between task-active and task-inactive neurons are not convincing. The separation between task-active and task-inactive neurons is to divide a distribution that is far from bimodal into what appears to be two arbitrary groups. Similarly, the authors divide cells relative to their prediction gain ("Top PG" and "Bottom PG" in Figure 2c), which fails to select for the population of significantly predicted cells (relative to the shuffle). Within CA1sup cells, after learning, there is a significant decrease in the prediction gain for "task-inactive" cells but not "task-active" cells, but it is important to keep in mind that the "task-active" group contains only 24 neurons, and there was no difference between the two groups of cells ("task-active" vs "task-inactive") when directly compared.

      We agree with this concern. To address this, we removed conclusion based on those arbitrary criteria instead opting for a more continuous approach. Specifically, we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement.

      Finally, it is not clear whether the identity of the pathway-responsive CA1sup neurons is fixed or whether it may change with learning. A deeper analysis into the cell pair cross-correlations or the weights of the GLM analysis may reveal whether there is a reorganization of CA1sup responses (some cells that were inhibited are no longer inhibited, and vice versa) or a dampening (the same CA1sup cells are inhibited in both cases, but the inhibition is less-pronounced in post-task sleep). The possibility of a rigid circuit dampened immediately following fear conditioning, is not discussed by the authors.

      We appreciate this feedback. To address this concern without weight analysis, we examined the stability of prediction gain scores between pre- and post-training sleep. Preservation of neuronal prediction-gain rankings would suggest that learning weakens existing predictive communication while maintaining the relative contribution of individual neurons, consistent with a dampening response. In contrast, poor preservation of prediction-gain rankings would be more indicative of a reorganization of predictive relationships across the population. This led to interesting findings regarding sublayer differences: ACC→CA1sup communication is more dynamic and evolving following learning, whereas ACC→CA1deep communication remains comparatively stable (See Figure 2A&B; Figure 3 C–F).

      Reviewer #3 (Public review):

      Summary:

      In this study, Hall and colleagues investigate how the coupling of activity from ACC to CA1is altered by fear learning, showing that during sleep immediately before learning, there is evidence for increased coupling of ACC activity with neurons that will subsequently be inhibited during the learning process. They go on to show that this effect seems to be mediated most by a subpopulation of neurons in the superficial layer of CA1. This fits with previous reports suggesting that these superficial neurons are key for the flexible updating of memory. The authors then go on to show that artificial activation of ACC using optogenetics results in varied effects in CA1, including a subtle decrease in activity of superficial neurons that lasts longer than the stimulus itself. Finally, the authors present some preliminary data suggesting that different interneurons may be recruited by this optogenetic stimulation in different ways and at different times.

      Overall, this is an interesting paper, but much of the analysis is very preliminary, and much of the crucial data about the learning effects and alterations to cell firing are not presented clearly and fully. This is further confounded by a rather opaque description of the results and analysis in the text. Overall, there is something very interesting here, but there needs to be a substantial series of extra analyses to clearly say what this is. In many cases, more robust analysis may render the results underpowered, which could dramatically change the conclusions of the paper.

      Strengths:

      The authors performed difficult, dual-location recordings across a multi-day learning paradigm, which seems like it could be a really nice dataset. They delve into the circuit basis of an interesting finding regarding ACC to CA1 connectivity and how this changes before and after fear conditioning. They provide data to suggest this connectivity may be through specific and distinct subcircuits in CA1.

      Weaknesses:

      (1) There is essentially no information in the text or figures about what the actual learning was, how it was done, how individual animals performed, and how any of these metrics related to learning. Looking at the methods, the authors did a number of things never mentioned anywhere in the text or figures, including novel arena exposure, contextual reexposure in extinction after learning, etc. It seems that this is a very rich dataset that has not been presented at all. I would recommend at the very least:

      We appreciate the reviewers’ feedback and have worked to address these concerns. See below our response.

      (a) Plot all of the behavioural training data, and how each mouse relates to one another - did the mice learn? At this stage, we don't know!

      We have now plotted all contextual fear conditioning behavioral data for each mouse (Figure 2—figure supplement 1). All mice exhibited high level of freezing during the contextual fear test, suggesting successful learning of the context–shock association.

      (b) Explain in the text in detail exactly what was done and why, and what this tells us about the neuronal activity.

      We have now added text to describe in detail the behavioral results and how that may relate to our GLM analyses. We have also more clearly detailed our experimental objectives (what was done and why) utilizing contextual fear conditioning,

      “In this approach, we were able to examine ACC–CA1 communication prior, during, and after learning, enabling us to examine how this communication evolves across fear learning. Specifically, we emphasized investigation into communication changes between pre- and post-training sleep to understand whether functional connectivity undergoes learning-related reorganization.”

      “Lastly, we examined whether PG scores correlated with the freezing response in mice during recall. Across all mice, freezing was significantly higher during recall than pre-shock baseline during training (Figure 2—figure supplement 2A). Overall, we found no correlation between PG and freezing (Figure 2—figure supplement 2B–D). However, there was a trend for a positive correlation (p = .07) between ΔPG and freezing percentage. An important consideration is that the uniformly high levels of freezing in mice limited behavioral variability, potentially reducing our ability to detect relationships between ACC–CA1 communication decoding and behavioral differences.”

      (c) If there is variance in learning and or conditioning, does this relate to features in the analysis, such as the GLM result.

      We examined this question by first investigating whether prediction gain scores in pre-training, post-training, or overall ΔPG correlated with freezing percentage. We found no significant correlation between any of the variables (Figure 2—figure supplement 1). We speculate this may be a result of a relatively robust freezing response limiting the ability for our fine-grained GLM decoding analyses to detect those differences. We have added this consideration to the main text.

      (2) Along similar lines, a key metric for most of the paper is that neurons most coupled with ACC are more likely to be inhibited during training. However, there is nothing anywhere in the paper showing these data. How do neurons in general respond to contextual shocks? The methods describe this as the average firing rate during training, normalised to pre-sleep activity. This metric seems a bit coarse and may obscure really important task-relevant dynamics. Are the neurons active at specific times, are they tuned to relevant parts of the task, and do any of these features of the cell activity also relate to the coupling with ACC? Similarly, how did the authors mitigate the influence of electrical artefacts caused by the foot shock in their recordings? Again, there is a huge amount of data here that is not being described, and likely holds very valuable information about what is actually happening. The paper would really benefit from the inclusion of these data in an accessible form, such as heatmaps of spiking, how these patterns change over time, and around e.g., foot shock, etc. Also key is how these features are altered by the variability of learning across subjects.

      We thank reviewer for this feedback. As pointed out, electrical artifacts caused by the footshocks prevents our ability to examine, with temporal sensitivity, neurons’ responses to footshocks. Therefore, we are left to examine activity changes across longer timescales. We acknowledge that our current modulation index analysis is a bit coarse. One reason is that our preliminary analyses using more temporally sensitive approaches did not reveal robust CA1 activity associated with specific behaviors, such as freezing or transitions between mobility and immobility. Thus, we chose a more holistic approach looking at the full CFC session to include all components that CA1 may be encoding during the training session. For example, although the pre-shock baseline period does not contain any footshock stimuli, it serves a key part in the process as mice begin to encode their environment around them. Nevertheless, we have added an additional analysis to examine how modulation changes across the pre-shock versus post-shock window (Figure 5—figure supplement 1). Although this provides greater insight into how ACC and CA1 activity changes across different dimensions in the task, further investigation utilizing casual manipulations is necessary to elucidate which phase ACC→CA1 activity is most involved.

      (3) A number of the effects are presented by comparing a statistically significant effect to a non-statistically significant effect (e.g. in Figure 2b, Figure 2d, Figure 4 b,c, and others). This isn't really valid - the key test that the two groups are different is either with a direct test of the difference or an interaction term in an e.g., ANOVA test. In some places, I am not sure the same conclusions will be drawn from the data with these tests.

      We want to thank the reviewer for this critical feedback. We have since added the appropriate statistical measure including linear mixed-effects models and ANOVA tests and for our analysis to avoid our previous statistical errors.

      (4) To what extent is defining superficial and deep CA1 neurons solely by ripple waveform an accepted method? Of the two papers referenced for this approach, one is a 2-photon calcium imaging paper that does not do electrical recordings (as far as I am aware), and the second uses this as a descriptor after defining the positions of units on an array. It would be good to clarify how accepted this is, and also how robust this is. At the very least, some kind of metric or walkthrough in the supplement as to how this was done, and how well each cell was classified and with what confidence, or some metric of how distinct and separate the two populations were (or was it just a smudge).

      We appreciate this feedback. While the Berndt 2023 paper implemented 2-photon calcium imaging, they also used tetrode classifications for radial axes in that paper which help informed our approach. Ultimately, our tetrode classification remains an outstanding limitation which we have since added to the main text (See below).

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (5) In the optogenetic experiment in Figure 5, the effect on the CA1 sup neurons seems to be driven by changes in a small subpopulation of this group, with no change in the others. Related to point 2, is there anything else in the data that can pull out what these cells are? More detailed analysis of the firing of these neurons might pull out something really interesting.

      We thank the reviewer for this feedback. Firstly, we want to clarify that optogenetic experiments were performed in a separate cohort of mice that did not undergo contextual fear conditioning. We have adjusted the text accordingly to make this distinction clearer. We have also added a per-animal separation of CA1 heatmap responses to ACC stimulations to demonstrate suppression is preserved across animals (Figure 5—figure supplement 3; Figure 6—figure supplement 2). Finally, our optogenetic experiments were primarily focused on understanding anatomical connectivity. Consequently, common waking behaviors (e.g., exploration and feeding) were not standardized across animals, and our analyses were therefore restricted to comparisons between slow-wave sleep and wakefulness more broadly.

      (6) Related to this - a number of comparisons simply pool neurons across mice and analyse them as if independent. This is done a lot in the past, but it would be better if an approach that included the interdependence of neurons recorded from the same mouse at the same time were used (such as a hierarchical model). While this is complex, a simpler approach would just be to plot the summary data also per mouse. For example, in Figure 5, how do the neurons inhibited by ACC activation spread across the different mice? Is the level of inhibition related to how well the mice learned the CS-US association?

      For dual-site analysis we have now added animal-level comparison for some key analyses (See Figure—figure supplement 1&2). As for optogenetic experiments, we have added per-animal heatmaps for ACC stimulation response (Figure 5—figure supplement 3; Figure 6—figure supplement 2).

      (7) Figure 6 is interesting, but very preliminary. None of the effects are quantified, and one of the cell types is not identified. I think some proper analysis needs to be done, again across mice, to be able to draw conclusions from these data.

      We thank the reviewer for this feedback. Reviewer 1 had a similar concern, and we have since added additional analyses to the revisions. Specifically, we performed autocorrelegrams, theta phase modulation analyses and a burst index analysis for V-Type interneurons. Importantly, however, these approaches still collapse neurons across mice. Given the sparse nature of V-type and PV interneurons, files often contain 2 or fewer interneurons making within-animal comparison difficult. That said, we have added supplemental figures displaying per-animal changes in response to ACC stimulations (Figure 6—figure supplement 1).

      (8) Finally, in general, I felt that the way the paper was written was very hard to follow, often relying on very processed levels of analysis that were hard to relate back to the raw traces and their biological meaning. In general taking more words to really simply and fully explain each analysis, and taking the words and figures to walk through how each analysis was done and what it tells us about the neuronal data/biology would be really beneficial, especially to someone who is not an extracellular electrophysiologist or immersed in the immediate field.

      We thank the reviewer for this feedback. Throughout the manuscript, we have revised the text to improve clarity in explaining our approaches and their results.

      In summary, while this manuscript explores an intriguing hypothesis about pre-learning circuit dynamics, it is currently held back by insufficient clarity in behavioural analysis, data presentation, and statistical quantification. Addressing these core issues would greatly improve interpretability and confidence in the findings.

      Additional comment:

      For the optogenetic experiments, we reprocessed and resorted the neuronal dataset to ensure accurate cell classification. Following this re-analysis, the principal findings remained unchanged. However, we found that sublayer-specific differences in response to ACC stimulation were restricted to the first second following ACC stimulation. Consequently, we removed the previous Figure 5E, which examined firing rate changes across successive 1-sec time bins, as the additional time windows did not provide further evidence of sublayer-specific effects.

      Acsady, L., Halasy, K., & Freund, T. F. (1993). Calretinin is present in non-pyramidal cells of the rat hippocampus--III. Their inputs from the median raphe and medial septal nuclei. Neuroscience, 52(4), 829-841. https://doi.org/10.1016/0306-4522(93)90532-k

      Andrianova, L., Yanakieva, S., Margetts-Smith, G., Kohli, S., Brady, E. S., Aggleton, J. P., & Craig, M. T. (2023). No evidence from complementary data sources of a direct glutamatergic projection from the mouse anterior cingulate area to the hippocampal formation. eLife, 12, e77364. https://doi.org/10.7554/eLife.77364

      Behzadi, G., Kalén, P., Parvopassu, F., & Wiklund, L. (1990). Afferents to the median raphe nucleus of the rat: Retrograde cholera toxin and wheat germ conjugated horseradish peroxidase tracing, and selective<span class="small">d</span>-[<sup>3</sup>H]aspartate labelling of possible excitatory amino acid inputs. Neuroscience, 37(1), 77-100. https://doi.org/10.1016/0306-4522(90)90194-9

      Berndt, M., Trusel, M., Roberts, T. F., Pfeiffer, B. E., & Volk, L. J. (2023). Bidirectional synaptic changes in deep and superficial hippocampal neurons following in vivo activity. Neuron, 111(19), 2984-2994.e2984. https://doi.org/10.1016/j.neuron.2023.08.014

      Cho, J. H., Deisseroth, K., & Bolshakov, V. Y. (2013). Synaptic encoding of fear extinction in mPFC-amygdala circuits. Neuron, 80(6), 1491-1507. https://doi.org/10.1016/j.neuron.2013.09.025

      Freund, T. F., Gulyas, A. I., Acsady, L., Gorcs, T., & Toth, K. (1990). Serotonergic control of the hippocampus via local inhibitory interneurons. Proc Natl Acad Sci U S A, 87(21), 8501-8505. https://doi.org/10.1073/pnas.87.21.8501

      Gulati, T., Guo, L., Ramanathan, D. S., Bodepudi, A., & Ganguly, K. (2017). Neural reactivations during sleep determine network credit assignment. Nat Neurosci, 20(9), 1277-1284. https://doi.org/10.1038/nn.4601

      Halasy, K., Miettinen, R., Szabat, E., & Freund, T. F. (1992). GABAergic Interneurons are the Major Postsynaptic Targets of Median Raphe Afferents in the Rat Dentate Gyrus. Eur J Neurosci, 4(2), 144-153. https://doi.org/10.1111/j.1460-9568.1992.tb00861.x

      Huang, W., Ikemoto, S., & Wang, D. V. (2022). Median Raphe Nonserotonergic Neurons Modulate Hippocampal Theta Oscillations. J Neurosci, 42(10), 1987-1998. https://doi.org/10.1523/JNEUROSCI.1536-21.2022

      Jackson, J., Bland, B. H., & Antle, M. C. (2009). Nonserotonergic projection neurons in the midbrain raphe nuclei contain the vesicular glutamate transporter VGLUT3. Synapse, 63(1), 31-41. https://doi.org/10.1002/syn.20581

      Liu, Z.-W., Faraguna, U., Cirelli, C., Tononi, G., & Gao, X.-B. (2010). Direct Evidence for Wake-Related Increases and Sleep-Related Decreases in Synaptic Strength in Rodent Cortex. The Journal of Neuroscience, 30(25), 8671. https://doi.org/10.1523/JNEUROSCI.1409-10.2010

      Miettinen, R., & Freund, T. F. (1992). Convergence and segregation of septal and median raphe inputs onto different subsets of hippocampal inhibitory interneurons. Brain Res, 594(2), 263-272. https://doi.org/10.1016/0006-8993(92)91133-y

      Mizuseki, K., Diba, K., Pastalkova, E., & Buzsáki, G. (2011). Hippocampal CA1 pyramidal cells form functionally distinct sublayers. Nature Neuroscience, 14(9), 1174-1181. https://doi.org/10.1038/nn.2894

      Morales, M., & Bloom, F. E. (1997). The 5-HT3 receptor is present in different subpopulations of GABAergic neurons in the rat telencephalon. J Neurosci, 17(9), 3157-3167. https://doi.org/10.1523/JNEUROSCI.17-09-03157.1997

      Norimoto, H., Makino, K., Gao, M., Shikano, Y., Okamoto, K., Ishikawa, T., Sasaki, T., Hioki, H., Fujisawa, S., & Ikegaya, Y. (2018). Hippocampal ripples down-regulate synapses. Science, 359(6383), 1524-1527. https://doi.org/10.1126/science.aao0702

      Oh, S. W., Harris, J. A., Ng, L., Winslow, B., Cain, N., Mihalas, S., Wang, Q., Lau, C., Kuan, L., Henry, A. M., Mortrud, M. T., Ouellette, B., Nguyen, T. N., Sorensen, S. A., Slaughterbeck, C. R., Wakeman, W., Li, Y., Feng, D., Ho, A., . . . Zeng, H. (2014). A mesoscale connectome of the mouse brain. Nature, 508(7495), 207-214. https://doi.org/10.1038/nature13186

      Papp, E. C., Hajos, N., Acsady, L., & Freund, T. F. (1999). Medial septal and median raphe innervation of vasoactive intestinal polypeptide-containing interneurons in the hippocampus. Neuroscience, 90(2), 369-382. https://doi.org/10.1016/s0306-4522(98)00455-2

      Petreanu, L., Huber, D., Sobczyk, A., & Svoboda, K. (2007). Channelrhodopsin-2–assisted circuit mapping of long-range callosal projections. Nature Neuroscience, 10(5), 663-668. https://doi.org/10.1038/nn1891

      Rajasethupathy, P., Sankaran, S., Marshel, J. H., Kim, C. K., Ferenczi, E., Lee, S. Y., Berndt, A., Ramakrishnan, C., Jaffe, A., Lo, M., Liston, C., & Deisseroth, K. (2015). Projections from neocortex mediate top-down control of memory retrieval. Nature, 526(7575), 653-659. https://doi.org/10.1038/nature15389

      Ramanathan, K. R., & Maren, S. (2019). Nucleus reuniens mediates the extinction of contextual fear conditioning. Behavioural brain research, 374, 112114. https://doi.org/https://doi.org/10.1016/j.bbr.2019.112114

      Ramanathan, K. R., Ressler, R. L., Jin, J., & Maren, S. (2018). Nucleus Reuniens Is Required for Encoding and Retrieving Precise, Hippocampal-Dependent Contextual Fear Memories in Rats. The Journal of Neuroscience, 38(46), 9925. https://doi.org/10.1523/JNEUROSCI.1429-18.2018

      Ratigan, H. C., Krishnan, S., Smith, S., & Sheffield, M. E. J. (2023). A thalamic-hippocampal CA1 signal for contextual fear memory suppression, extinction, and discrimination. Nat Commun, 14(1), 6758. https://doi.org/10.1038/s41467-023-42429-6

      Senft, R. A., Freret, M. E., Sturrock, N., & Dymecki, S. M. (2021). Neurochemically and Hodologically Distinct Ascending VGLUT3 versus Serotonin Subsystems Comprise the r2-Pet1 Median Raphe. J Neurosci, 41(12), 2581-2600. https://doi.org/10.1523/JNEUROSCI.1667-20.2021

      Shi, W., Xue, M., Wu, F., Fan, K., Chen, Q. Y., Xu, F., Li, X. H., Bi, G. Q., Lu, J. S., & Zhuo, M. (2022). Whole-brain mapping of efferent projections of the anterior cingulate cortex in adult male mice. Mol Pain, 18, 17448069221094529. https://doi.org/10.1177/17448069221094529

      Silva, B. A., Astori, S., Burns, A. M., Heiser, H., van den Heuvel, L., Santoni, G., Martinez-Reza, M. F., Sandi, C., & Gräff, J. (2021). A thalamo-amygdalar circuit underlying the extinction of remote fear memories. Nature Neuroscience, 24(7), 964-974. https://doi.org/10.1038/s41593-021-00856-y

      Souza, R., Bueno, D., Lima, L. B., Muchon, M. J., Gonçalves, L., Donato, J., Jr., Shammah-Lagnado, S. J., & Metzger, M. (2022). Top-down projections of the prefrontal cortex to the ventral tegmental area, laterodorsal tegmental nucleus, and median raphe nucleus. Brain Struct Funct, 227(7), 2465-2487. https://doi.org/10.1007/s00429-022-02538-2

      Szonyi, A., Mayer, M. I., Cserep, C., Takacs, V. T., Watanabe, M., Freund, T. F., & Nyiri, G. (2016). The ascending median raphe projections are mainly glutamatergic in the mouse forebrain. Brain Struct Funct, 221(2), 735-751. https://doi.org/10.1007/s00429-014-0935-1

      Tononi, G., & Cirelli, C. (2003). Sleep and synaptic homeostasis: a hypothesis. Brain Research Bulletin, 62(2), 143-150. https://doi.org/https://doi.org/10.1016/j.brainresbull.2003.09.004

      Tononi, G., & Cirelli, C. (2006). Sleep function and synaptic homeostasis. Sleep Med Rev, 10(1), 49-62. https://doi.org/10.1016/j.smrv.2005.05.002

      Turi, G. F., Li, W. K., Chavlis, S., Pandi, I., O'Hare, J., Priestley, J. B., Grosmark, A. D., Liao, Z., Ladow, M., Zhang, J. F., Zemelman, B. V., Poirazi, P., & Losonczy, A. (2019). Vasoactive Intestinal Polypeptide-Expressing Interneurons in the Hippocampus Support Goal-Oriented Spatial Learning. Neuron, 101(6), 1150-1165 e1158. https://doi.org/10.1016/j.neuron.2019.01.009

      Wang, D. V., Yau, H.-J., Broker, C. J., Tsou, J.-H., Bonci, A., & Ikemoto, S. (2015). Mesopontine median raphe regulates hippocampal ripple oscillation and memory consolidation. Nature Neuroscience, 18(5), 728-735. https://doi.org/10.1038/nn.3998

      Wang, J., Hasan, M. T., & Seung, H. S. (2009). Laser-evoked synaptic transmission in cultured hippocampal neurons expressing channelrhodopsin-2 delivered by adeno-associated virus. J Neurosci Methods, 183(2), 165-175. https://doi.org/10.1016/j.jneumeth.2009.06.024

      Watson, Brendon O., Levenstein, D., Greene, J. P., Gelinas, Jennifer N., & Buzsáki, G. (2016). Network Homeostasis and State Dynamics of Neocortical Sleep. Neuron, 90(4), 839-852. https://doi.org/https://doi.org/10.1016/j.neuron.2016.03.036

      Xu, W., & Südhof, T. C. (2013). A Neural Circuit for Memory Specificity and Generalization. Science, 339(6125), 1290-1295. https://doi.org/doi:10.1126/science.1229534

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggested fixes (these correlate with the numbered points in the Weaknesses section of the Public Review):

      (1) That seems like the larger finding to start with, before dissecting by Firing Activity Index. I'd first show the overall finding, then dissect it.

      We thank the reviewer for this recommendation. In our manuscript, figure 1F examines the overall pre-to-post prediction gain changes prior to any neuron separation. Further separation of neurons’ characteristics and classification occurs in subsequent figures.

      (2) The overall picture painted by Figure 2 suggests a correlation analysis should be carried out to search for a general property: prediction score gain vs firing activity index. Are those two variables considered significant by Pearson correlation? Did the authors try that and it didn't work, so they did these analyses? If there is no significant correlation, what does a more detailed look at those two variables on an x-y plot teach us? I suggest considering showing such a plot to readers, at least in a Supplement.

      We thank the reviewer for this feedback and have incorporated this approach in our revised manuscript. Specifically, we performed a correlation analysis between prediction gain change and modulation index (formerly called firing activity index). We uncovered a significant positive correlation between variables, suggesting that task engagement modifies ACC→CA1 communication (Figure 2C).

      (3) (No additional comments).

      (4) This weakness may be able to be addressed by doing a correlation of depth (LFP amplitude of sharp wave) versus the firing rate ratio. For example, the threshold used for deep may have been such that it reduced the number of detected deep neurons, but if a general relationship between depth and degree of FRR is found, it can remove issues from this difference in statistical power.

      We thank the reviewer for this comment. We were unable to perform a reliable link between correlation depth and modulation index. LFP amplitude can vary substantially between tetrodes due to differences in electrode impedance, placement, and recording conditions, making direct comparisons across animals difficult. While within-animal analyses could largely circumvent these issues, many recording sessions did not include tetrodes spanning the full superficial-to-deep CA1 axis, preventing a reliable assessment of this relationship.

      (5) I would give it a different name - "opto-modulation index" or "opto firing rate ratio" perhaps. This would make it clear that you are not measuring task-based modulation of firing.

      We have since modified our wording to improve clarity. Specifically, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion

      (6) Specifically: can the post-opto lag of v-type versus wide-waveform and PV-type neurons be analyzed? Are the V-type neurons increasing firing before the others decrease? What about cross correlograms between v-type and pyramidal neurons, either at baseline or post-stim?

      We appreciate this feedback. We have added a figure showing differences in response lags to the optostimulation. We demonstrate that V-Type interneurons clearly fire prior to PV and pyramidal cells (Figure 6—figure supplement 1E&F). Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlation analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type and pyramidal neurons.

      (7) Can the authors discuss the candidate pathways for connectivity from ACC to CA1?

      We thank the reviewer for this feedback. We have since added discussions on ACC-to-CA1 connectivity and discussed possible relay brain regions between ACC and CA1.

      (8) There is mounting evidence about the role of sleep oscillatory events playing homeostatic roles, not only memory-based. The authors bring this up, but do not offer it as an explanation for their findings, but I believe they probably should. For example, Norimoto et al 2018 cited by the authors. Also, Gulati/Gunguly et al 2017 Nature Neuroscience suggests downscaling as a default NonREM activity. Gulati and also Roux/Buzsaki NatNeuro 2017 show that certain privileged or tagged neurons can be protected from this. This therefore reflects that default activity in nonREM may have a homeostatic role, but then learning may alter that default. I believe this should be discussed as a possible reason for the dissociation between ACC and CA1, the authors observe after CFC.

      In more detail, the authors state that CFC worsened ACC ability to predict ripple spike rate vectors. The authors suggest this may reflect "worsened" communication from ACC to HPC. It could also reflect a SHIFT (not worsening) in the information state of the hippocampus, where ripples reflect novel information and/or information coming from other brain regions. Essentially, the novel information may out-compete usual information flows. For example, ACC may be a default "feeder" into ripples (for example as part of default mode network) when there was no recent highly salient information, but under non-default conditions such as after CFC, ripple content may be fed from other sources (be they internal or external to the HPC). I believe this should be discussed.

      For example, were CA1 sup task inactive neurons basically DMN-active neurons? Figure 2 shows neurons with the highest pre-training ACC prediction were the ones that dropped the most in training - again suggesting these neurons may be tuned to internal or default dynamics rather than CFC (or other novel experiences).

      This shift from a default communication mode to a more experience-based one should probably be discussed as an alternative explanation, rather than simply "worsening" of communication.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have incorporated many of the listed citations and ideas they discussed into our discussion section proposing homeostatic downscaling and a shift in the default mode network as possible explanations for our results.

      Minor Weaknesses:

      (1) Introduction Line 52: "during replays" should probably be "during replay events".

      Changed.

      (2) Introduction Line 76: "how communications" should be "how communication"

      Changed.

      (3) 200-0, 400-200, 600-400 time bins are a bit unclear in Figure 1f. Are they really negative times, rather than positive? Perhaps negative signs could be put into the legend of 1f, or the time windows can be shown on the left side of 1e. Or potentially 1f could use the same -0.6, -0.4, etc as 1e so readers understand they are linked (if I understand correctly).

      They are negative in the sense the occur before the ripple event. Figure 1e now displays the negative signs.

      (4) I don't believe the methods describe how many tetrodes are put into the ACC. It would seem this should be put in the ACC portion of the "Stereotaxic surgery" section.

      We now clearly explain the number of tetrodes (8) in the "Stereotaxic surgery" section.

      (5) Results line 125: "learning induced" should be "learning-induced".

      Changed.

      (6) In terms of display, deep and sup are swapped in the various figures in terms of which is shown first/second (at least for readers assuming left is first). I suggest putting deep first or sup first in all figures. To me, sup first seems more natural, but homogeneity seems best regardless. This will help readers easily track results.

      We adjusted the figures so that superficial is typically displayed first with some exceptions. For example, in Figure 3A, CA1deep is shown first to preserve the anatomical (dorsal-ventral) relationship.

      (7) Results line 178: "optogenetics stimulation" should be "optogenetic stimulation".

      Changed.

      (8) Results line 180: "upon stimulations" should be "upon stimulation".

      Changed.

      (9) Results line 180: "to different capacities" could be "to different degrees" or "in different manners".

      Changed.

      Reviewer #2 (Recommendations for the authors):

      The sleep scoring procedure is not described clearly. The text references delta waves and ripple oscillations, but the accompanying citation (Wang et al. 2015) does not use such a procedure. Since the post-task rest sessions are called "sleep sessions", there is some confusion about whether the data was restricted to slow wave sleep or not. If data from each sleep session were taken without restricting to actual sleep, that would be problematic because animals may be less likely to sleep immediately following fear conditioning, which could introduce some sleep/wake bias into the comparisons. In particular, the relationship between cortex and ripple activity has been reported to dramatically change between awake and sleep states (Tang & Jadhav, 2019). I am not including this point in the public review in case the data was in fact restricted for sleep, and it simply needs to be clarified in the text.

      Thank you for this feedback. The recordings were in fact restricted to sleep. We have added text to the manuscript to make this clearer. Moreover, we added further discussion on how sleep was calculated.

      There is a puzzling paragraph in the discussion, arguing that "Here, we add to this understanding with CA1sup neurons having a diminished role in fear memory formation". Sparse task-related activity in CA1sup does not imply that CA1sup is not involved in memory. Indeed, while the median of CA1sup neurons' firing rate ratio was below zero, there is a substantial proportion of neurons that are recruited, and these could be extremely important for memory. Most studies on reactivation and replay would only concentrate on cells sufficiently active in the task, and observe whether these patterns of activity are enhanced in post-task sleep.

      We thank the reviewer for this feedback and have removed the text claiming sublayer difference in fear conditioning.

      If the ACC is indeed inhibiting the CA1sup pyramidal cells through V-type interneurons, then one would expect the average GLM weights predicting the activity of those best-predicted CA1sup cells to be negative. If that is true, that could nicely tie the prediction effect to the optogenetic results, demonstrating that ACC's relationship to pyramidal cells is inhibitory in natural conditions as well.

      We thank the reviewer for this suggestion. Unfortunately, our primary GLM analysis did not properly save weight coefficients to perform such analyses. To address this as closely as possible, we performed a preliminary analysis using a modified version of our GLM to examine whether the coefficients predicting CA1sup pyramidal neuron activity exhibited a bias toward negative weights. While this modified analysis did not generate coefficients directly comparable to the prediction gain values reported in the manuscript, it allowed us to assess whether an overall difference in coefficient sign was evident between CA1 sublayers. We found no significant bias toward negative coefficients and no clear differences between CA1sup and CA1deep neurons. Although this result does not provide additional support for an inhibitory relationship under natural conditions, it does not necessarily contradict our optogenetic findings. GLM coefficients quantify statistical dependencies between neural activities and reflect not only direct interactions but also indirect network effects, shared inputs, and the model structure. Consequently, the sign of a GLM coefficient should not be interpreted as a direct measure of whether the underlying synaptic relationship is excitatory or inhibitory.

      Figure 1f is strangely missing comparisons for positive delays. If such windows were to be included and if the reactivation gain is lower for them, that could really drive home the point that communication takes place in the ACC->CA1 direction more than in the CA1->ACC direction.

      Our goal of this study was to examine how incoming information from the cortex may differentially drive CA1 sublayer activity. While we think examining the reverse direction offers a compelling future direction, it was beyond the scope of our present manuscript.

      There is some confusion about the N-s. There's a total of 190 CA1 neurons (Figure 2a legend). 21 of them are deep, and 77 are sup (Figure 3b legend), so presumably 92 would be neither. The legend of Figure 4b agrees with this: 24 task-active CA1sup and 53 task-inactive CA1sup cells, while in CA1deep, there were 14 task-active and 7 task-inactive cells, but in Figure 3c, there is a comparison of N=24 CA1deep cells and n=94 CA1sup cells (so 72 neither).

      For one dual-site animal, the CFC recording file was corrupted, while the pre- and post-training recordings remained intact. As a result, this animal was included only in the GLM analyses, leading to slight differences in sample size across analyses. We have clarified these sample sizes in the revised manuscript and highlighted this discrepancy in the Methods section.

      There appears to be a typo on line 213, the reference should be to Figure 6f.

      Changed.

      Reviewer #3 (Recommendations for the authors):

      To what extent do the authors think that this is learning dependent, as opposed to stress dependent? Not that I want an experiment here, but it is important to note that both of these regions are very much involved in stress responses. Do you think you would get the same result with a purely appetitive learning paradigm? Or is this specific to stress? It might be nice to add this to the discussion.

      We thank the reviewer for this raising this point of discussion. Future experiments utilizing appetitive behavioral tasks could help address these important questions. While we have avoided speculating too much, we agree it is valuable to acknowledge this possibility. Accordingly, we have added a brief discussion point to at least call attention to this possibility for the reader.

      “Another caveat to mention is that both the ACC and CA1 are involved in the stress response (Kim et al., 2015; Lamotte et al., 2021). Future experiments utilizing appetitive learning paradigms, rather than the aversive contextual fear conditioning used here, will help disentangle learning-related remodeling of ACC→CA1 communication from changes driven by stress.”

      There are a number of typos that confuse the message - for example, in the abstract, the authors say that ACC suppresses superficial CA1 interneurons. This seems most likely an error - I think the authors mean superficial CA1 neurons? Or PV interneurons? Similar errors exist throughout, as well as odd combinations of bold and italics across and within words, etc. Overall, especially in consideration of my final main point above regarding clarity of the text, it would be good to have a proper proofread to make sure the text is as clear as possible

      We thank the reviewer for identifying these issues. The specific error in the abstract has been corrected, and we have since carefully proofread the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Okazaki et al. showed flickering stimuli to patients with unilateral spatial neglect (USN) and measured EEG responses. They compared this with another patient group (post-stroke, but no USN) and healthy controls. The author's rationale was to entrain intrinsic brain rhythms using the flicker of different frequencies (3-30 Hz). Effects found unique to the 9-Hz stimulation condition differentiate USN patients from the other groups, leading them to conclude that USN can be characterized by increased hemispheric alpha asymmetry, driven by a relatively increased response in the intact hemisphere.

      Strengths:

      This study is principled empirical work that benefits from access to special patient groups of considerable size (about 60 stroke patients in total, and 20 USN). The authors use state-of-the-art established methods to (1) deliver and (2) quantify the responses to the flicker stimulation in the EEG recordings. In addition, they use phase-coupling measures to investigate cross-frequency coupling (here: alphagamma) and a measure of directed connectivity between brain areas, transfer entropy. The results are supported by means of simulations using a coupled oscillators model.

      Weaknesses:

      In my eyes, the major conceptual weakness of the study is that the authors make the a priori assumption that the flicker stimulation entrains intrinsic brain rhythms, especially alpha (9 Hz). To date, there is no direct (and only equivocal indirect) evidence that alpha rhythms can be entrained with periodic visual stimulation. In the present study, the assumption of alpha entrainment permeates some analytical decisions - where it would be possible to separate stimulus-driven from intrinsic rhythms more strongly than is currently the case, potentially yielding deeper insights into the oscillopathy of USN - and, ultimately, the interpretation of the results. Another potential issue to consider here is the analysis of gamma rhythms in EEG data, absent a control of miniature eye movements, a known problem (YuvalGreenberg et al., 2008, https://doi.org/10.1016/j.neuron.2008.03.027) that may be exacerbated here, given that USN patients could show different auxiliary gaze behaviour.

      We thank Reviewer #1 for the careful and constructive evaluation of our study, and for recognizing the strengths of the patient cohort and our combined empirical and computational approach. We also appreciate the reviewer’s concern that our original wording could be read as assuming that flicker stimulation necessarily entrains intrinsic alpha rhythms. In the revised manuscript, we have clarified that our interpretation is based on frequency-specific stimulus-locked responses and model-based inference, rather than on an a priori assumption of entrainment. We have also revised the relevant parts of the Introduction and Discussion to distinguish more clearly between stimulus-locked responses and intrinsic oscillatory dynamics. In addition, we have expanded our discussion of the potential influence of miniature eye movements on gamma-band activity and PAC. These points are addressed in detail in our responses to the Recommendations for the Authors below.

      Reviewer #2 (Public review):

      This study investigates how altered neural oscillations may contribute to unilateral spatial neglect (USN) following right-hemisphere stroke. By combining steady-state visual evoked potentials (SSVEPs), phase-amplitude coupling (PAC), transfer entropy (TE), and computational modeling, the authors aim to show that USN arises from disrupted hemispheric synchronization dynamics rather than simply from lesion extent. The integration of empirical EEG data with a mechanistic model is a major strength and offers a valuable new perspective on how frequency-specific neural dynamics relate to clinical symptoms.

      The work has several notable strengths. The combination of experimental and modeling approaches is innovative and powerful, and the findings provide a coherent mechanistic framework linking abnormal neural entrainment to attentional deficits. The study also provides concrete evidence to support the potential for frequency specific neuromodulatory interventions, which could have translational relevance.

      At the same time, there are areas where the evidence could be clarified or contextualized further. The manuscript would benefit from more detailed characterization of lesions, since differences in lesion topography (white vs. gray matter, occipital vs. parietal areas) could greatly improve our understanding of the physiopathology causing unilateral spatial neglect and the altered neural oscillations reported. Methodological choices, such as focusing analyses on occipital electrodes rather than parietal sites, and the potential influence of volume conduction in transfer entropy analyses, also need clearer justification/elaboration. In addition, while the authors report several neural metrics, it is not always clear why SSVEP power was chosen as the primary correlate of clinical severity over other measures. More broadly, the manuscript would be strengthened by clearer definitions of dependent variables and reporting of software and toolboxes used.

      Overall, the study makes a significant contribution by demonstrating that USN can be conceptualized as a disorder of disrupted oscillatory dynamics. With some clarifications and expansions, the paper will provide readers with a clearer understanding of both the strengths and the limitations of the evidence, and it will stand as a valuable reference for future work on oscillatory mechanisms in stroke and attention.

      We thank Reviewer #2 for the positive assessment of our integrated empirical and computational approach, and for highlighting the potential contribution of our findings to understanding oscillatory mechanisms in USN.

      In response to the reviewer’s comments, we have added new supplementary figures showing lesion overlap maps and lesion-volume analyses (Supplementary Figure 1), additional analyses related to electrode selection and SSVEP topography (Supplementary Figure 2), and clinical-correlation analyses of hemispheric imbalance measures (Supplementary Figure 3). We have also revised the manuscript to better contextualize lesion topography and lesion extent, clarify the rationale for the occipital-electrode and clinical-correlation analyses, and provide additional methodological details. These issues are addressed in detail in our point-by-point responses to the Recommendations for the Authors below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I found your manuscript well-written and hence easy to follow. Allow me to provide some recommendations to tackle the "weaknesses":

      (1) I have mentioned that the entrainment assumption is not warranted based on the current evidence, but I also think that the analysis and interpretation is critically constrained by this. The way the analysis is carried out conflates stimulus-driven and intrinsic brain rhythms, especially in the alpha band. In other words, spectral representations of the EEG data will likely be dominated by natural alpha rhythms, whereas the stimulus-driven signals could be accentuated by a different analysis, time-locked to the stimulation. This is explained in greater detail in Keitel et al. (2019, https://doi.org/10.1523/JNEUROSCI.1633-18.2019). Looking into alpha and stimulus-driven responses separately, without the assumption of entrainment, may actually allow a more complete picture of the impact of USN because the stimulus driven SSVEPs are taken to indicate different cortical processes than alpha (see e.g., Duecker et al., 2021, https://doi.org/10.1523/JNEUROSCI.3134-20.2021, though for gamma). Importantly, this all does not exclude the possibility that alpha rhythms were indeed entrained here, but given the current situation, this should be an outcome of the study rather than an a-priori assumption. I suggest re-framing the manuscript this way.

      We agree that the current wording could be read as presuming alpha entrainment. We have revised the framing and terminology to reflect that entrainment is inferred from the 9-Hz–specific stimulation effect and the absence of a corresponding hemispheric bias at rest, with further support from the resonance mechanism demonstrated by our computational model. We also point readers to the Discussion “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, where we explain why the present 9-Hz–specific findings are interpreted in terms of alpha-range resonance/phase alignment and how this is expressed in EEG. We revised the Introduction SSER description (p.4) to:

      “SSERs are elicited by rhythmic sensory stimulation and provide a measure of stimulus-locked neural responses. When the stimulation frequency is close to the system’s intrinsic resonance frequency, these responses may include an entrainment component, reflecting the alignment of endogenous oscillations to external input (Pikovsky, Rosenblum, and Kurths 2003; Okazaki et al. 2021)”

      We have also revised the Introduction hypothesis statement (p.5) to avoid implying an a priori entrainment assumption, replacing it with:

      “We hypothesized that frequency-specific stimulus-locked synchronization dynamics in response to rhythmic stimulation would be selectively disrupted in one hemisphere, resulting in an interhemispheric imbalance in USN.”

      Finally, to ensure consistent framing in the Discussion, “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, we have revised the opening sentence (p.20) as follows:

      “We observed a hemispheric bias in stimulus-locked responses to flickering stimuli in USN patients in the alpha range, which corresponds to the natural frequency of the visual system (Rosanova et al. 2009; Okazaki et al. 2021).”

      (2) Transfer Entropy is used as a measure of directed connectivity and applied to narrow-band filtered EEG signals. Methodological issues have been pointed out with regard to that (Daube et al., 2022, https://doi.org/10.48550/arXiv.2201.02461). Has this been considered?

      We thank the reviewer for raising this important methodological concern. As Daube et al. (2022) noted, Transfer Entropy (TE) can be overestimated when applied to narrow-band signals with strong autocorrelation. However, in our analysis, TE was computed not from the narrow-band 9-Hz waveform itself but from its instantaneous amplitude (amplitude envelope), which fluctuates nonperiodically on a slower timescale, thereby reducing the risk of spurious causality driven by sinusoidal autocorrelation. We clarified this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “After applying an 8.5–9.5 Hz FIR bandpass filter to the EEG responses to 9-Hz flickering stimuli, we extracted the instantaneous amplitude (amplitude envelope) from the analytic signal using the Hilbert transform. Because this amplitude envelope fluctuates nonperiodically at a slower timescale than the carrier 9-Hz oscillation, it provides a broadband measure of signal dynamics while minimizing the strong autocorrelation inherent in narrow-band periodic signals that can spuriously inflate TE estimates (Daube C et al., 2022).”

      We have also clarified the relationship between TE and volume conduction in the same section:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than an instantaneous coupling.”

      Finally, we have added an explicit interpretation emphasizing the direction-specific nature of the effect in the Discussion, “Biased information transfer in USN patients” (p.23):

      “This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube, Gross, and Ince 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (3) If a closer control of miniature eye movements is not possible, I suggest removing any gamma analysis, or at least prominently mentioning the caveat that gamma activity may be contaminated by eye movement artifacts.

      We agree that the contribution of miniature eye movements cannot be completely excluded. However, we consider it unlikely that the hemispheric asymmetry in alpha– gamma PAC reported in this study mainly arises from eye-movement artifacts, for the following reasons. First, SP (saccadic spike potential)-related PAC would require saccade timing to be tightly phase-locked to the 9-Hz cycle. However, SPs are time-locked to saccade onset and typically cluster around 200–300 ms after stimulus onset (Yuval-Greenberg et al., 2008; Keren et al., 2010). Thus, it is unlikely that they would be consistently phase-locked to the 9-Hz cycle (≈111 ms), and even modest temporal jitter would markedly blur PAC. Second, because SPs are brief spike-like transients with broadband high-frequency components, periodic SP contamination would be expected to yield a broadband gamma profile, rather than the relatively narrow band (35–45 Hz) observed here. In addition, we directly compared gamma-band power (35–45 Hz) at O1 and O2 during 9-Hz stimulation using the same gamma range as in the PAC analysis and found no significant hemispheric differences in any group (Author response image 1), arguing against a systematic unilateral increase in gamma power driven by asymmetric saccade behavior. Taken together, these considerations make it difficult to attribute the observed alpha–gamma PAC asymmetry primarily to SPs arising from eye movements. Nonetheless, residual eye-movement artifacts cannot be completely ruled out, and we now explicitly state this limitation in the Discussion, “Limitations” (p. 25) by adding:

      “Fourth, the hemispheric bias in alpha–gamma PAC should be interpreted in light of potential contamination from miniature saccades (Yuval-Greenberg, Tomer, Keren, Nelken, & Deouell, 2008). However, several observations make it unlikely that such artifacts are the primary source of the effect. For saccadic spike potentials to account for the PAC under 9-Hz stimulation, they would need to occur in a highly periodic and phase-locked manner relative to the 9-Hz cycle. This scenario is unlikely given the stimulus-locked dynamics of miniature saccades. (i.e., post-stimulus inhibition followed by a rebound around 200–300 ms) (Yuval-Greenberg & Deouell, 2009). Moreover, SP-related contamination would be expected to yield a broadband gamma profile (~20–90 Hz) (Keren, Yuval-Greenberg, & Deouell, 2010; Yuval-Greenberg & Deouell, 2009), rather than the relatively narrow band (35–45 Hz) observed here. In addition, we found that gamma-band power in the 35–45 Hz range did not show any hemispheric difference between O1 and O2 during 9-Hz stimulation (data not shown).”

      Author response image 1.

      Hemispheric differences in gamma-band (35–45 Hz) power during 9-Hz flicker stimulation. Left (O1) − Right (O2) gamma power did not differ from zero in any group (non-USN: p = 0.45; USN: p = 0.34; healthy: p = 0.35). Error bars represent 2 standard errors of the mean.

      (4) Please provide more methodological detail on the resting state recordings. When, how, and under which circumstances were these recorded?

      Thank you for pointing this out. We clarified when the resting-state interval was taken in the Methods, “Steady-state visual evoked potential (SSVEP)” (p.7) by adding:

      “For the resting-state interval, power was estimated using the same procedure from the ‘off’ interval immediately preceding the 3-Hz flicker block (see Fig. 1)”

      (5) Provide power spectra of the EEG data for illustration - ideally for resting state and stimulation conditions. These should allow the reader to visually evaluate the effects of different stimulation frequencies, as well as differences between participant groups.

      Figure 2 already presents spectra normalized to the resting-state baseline. To make this explicit for readers, we clarified this point in the Figure 2 caption (p.11) by adding:

      “Spectra are normalized to the baseline from the resting-state interval.”

      Reviewer #2 (Recommendations for the authors):

      (1) The authors indicate L.608 "the precise extent and topography of brain lesions could not be fully homogenized across patients", but the manuscript would highly benefit from any additional detail that could be obtained from characterization of the lesion sites. At minimum, it would be important to indicate for USN and non-USN groups whether lesions predominantly affected white matter or gray matter, and whether occipital versus parietal cortices were involved (e.g., using MRI atlas templates). If this coarse information can be obtained, the authors could test whether any of these anatomical details can distinguish are different between USN and nonUSN patients. This information could help the reader evaluate whether the reported neural asymmetries might be driven by lesion topography rather than purely by oscillatory dynamics.

      We understand this comment as raising the important concern that lesion location and lesion volume may differ between the USN and non-USN groups and could contribute to the observed neural asymmetry. We agree that the presence of USN and the alteration of oscillatory dynamics should be interpreted in relation to which regions and networks are damaged, and to what extent. In response to the reviewer’s suggestion, we generated lesion overlap maps based on the available structural images and additionally quantified lesion volume for each patient (replaced Supplementary Figure 1). This analysis confirmed that lesion volume was significantly larger in the USN group than in the non-USN group. Thus, lesion volume is an important anatomical factor to consider when interpreting group differences in USN and neural responses.

      At the same time, the present results suggest that lesion volume and coarse lesion topography alone are not sufficient to explain the 9 Hz-specific imbalance in interhemispheric synchrony. Interhemispheric synchrony depends on distributed network functions involving multiple cortical and subcortical regions and their connecting pathways, and similar functional imbalances may arise from different patterns of anatomical damage. Moreover, although the newly added lesion maps confirmed more extensive lesions in the USN group, lesions in both groups predominantly involved the right MCA territory, and lesion extent and location varied across patients. Importantly, the hemispheric asymmetry in neural responses emerged selectively in the 9 Hz condition, whereas SSVEP responses at other frequencies were largely balanced between hemispheres. If lesion volume or coarse lesion topography alone were sufficient to explain the effect, one might expect a more uniform reduction across frequencies or a simpler pattern corresponding to lesion extent.

      Accordingly, we do not treat lesion topography and oscillatory dynamics as competing explanations. Instead, we regard them as hierarchically related: anatomical damage alters network components, and this in turn gives rise to a frequency-specific imbalance in synchronization capacity. To clarify this point, we replaced the previous Supplementary Figure 1 with a new figure showing lesion overlap maps and lesion volume information, and revised the Discussion section “Distinct neural responses in non-USN and USN patients” (p. 24) as follows.

      “To further characterize the anatomical background of these group differences, we generated lesion overlap maps and quantified lesion volume in the USN and non-USN groups (Supplementary Figure 1). Both groups predominantly showed lesions involving the right MCA territory, but lesion volume was significantly larger in the USN group than in the non-USN group (USN: 72,419 ± 73,778 mm<sup>3</sup>; non-USN: 11,739 ± 23,907 mm<sup>3</sup>; Mann–Whitney U = 361.0, p = 1.41 × 10<sup>-5</sup>). These anatomical differences indicate that lesion extent is an important factor associated with USN. However, they do not by themselves fully explain the frequency-specific neural effect observed here. Importantly, this frequency specificity coincides with the intrinsic alpha frequency of the visual system. This correspondence suggests that the present finding may not simply arise from lesion location or lesion volume alone, but may instead reflect a more complex mechanism involving selective functional disruption of oscillatory networks. From this perspective, lesion topography and oscillatory dynamics should be regarded not as competing explanations, but as different levels at which the same pathological condition can be understood. The key question is which network components are affected and how their dysfunction gives rise to the frequency-selective hemispheric imbalance in synchronization capacity at 9 Hz. Thus, even if USN and non-USN patients differ in lesion extent and aspects of lesion topography, this does not undermine the present interpretation, but rather highlights the need to examine how anatomical damage relates to frequency-specific network dysfunction.”

      We have also referred to our computational account and clarified its implication for the functional mechanism in the same subsection (p. 24):

      “Our computational model illustrates a plausible mechanism for such a process. When two coupled oscillators sharing the same intrinsic alpha frequency are connected via asymmetric interhemispheric couplings, the model selectively produces an imbalance in synchrony at the resonant frequency, whereas responses at non-resonant stimulation frequencies remain relatively balanced between hemispheres. This model result suggests that post-lesion asymmetry in interhemispheric coupling may bias alpha-band information processing (e.g. phase-dependent sampling/synchrony) between hemispheres, and may consequently manifest as systematic biases in perceptual and attentional allocation.”

      We have also revised the Limitations (p. 25) to acknowledge that the present lesion analyses characterize the distribution and extent of lesions across groups, but do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony.

      “Second, although we added lesion overlap maps and quantified lesion volume, these analyses primarily characterize where and how extensively lesions were distributed across the two patient groups. They do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony. Future studies with larger cohorts will be necessary to combine detailed lesion-symptom mapping and assessments of white-matter disconnection with neural synchrony analyses to determine which anatomical network disruptions lead to frequency-specific alterations in oscillatory dynamics after stroke.”

      (2) The rationale for focusing on O1 and O2 electrodes should be clarified. Given that neglect is classically associated with parietal dysfunction, one would expect analyses of parietal electrodes to be informative. Could the authors justify their choice, and possibly report whether similar effects were (or were not) observed at parietal sites?

      Our primary SSVEP analyses focused on O1 and O2 because flicker stimulation is designed to drive the visual system, and the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015). We also note that hemispheric differences at central and frontal electrodes are already shown in Fig. 4 and described in the Results, demonstrating no significant left–right differences at these sites across any stimulation frequency conditions, including 9 Hz. In addition, we added a supplementary figure showing the scalp topography of the mean fundamental-frequency SSVEP power averaged across all stimulation conditions, which displays the expected occipital maximum and thus a typical SSVEP spatial profile (Supplementary Fig. 2, shown below). We added the rationale and the reference in the Methods, “Steady-state visual evoked potential (SSVEP)” (p. 7) section by adding:

      “SSVEP power at the fundamental (stimulated) frequency was then extracted for subsequent analyses. Because the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015), we evaluated occipital electrodes (O1/O2) as primary sites. The scalp topography of the mean fundamental-frequency SSVEP power averaged across stimulation conditions is shown in Supplementary Fig. 2.”

      (3) For the mutual information and transfer entropy analyses, the potential influence of volume conduction should be acknowledged. Numerous studies mitigate this issue by applying source reconstruction or connectivity metrics that are insensitive to zerolag correlations. Even if the present study did not use such approaches, the authors should discuss the extent to which volume conduction might confound their results, and ideally provide some justification for why their findings remain valid.

      Regarding directed connectivity, our primary analysis uses Transfer Entropy (TE), which evaluates time-lagged prediction and is therefore, by definition, not directly sensitive to zero-lag common-source correlations. Consistently, our key finding is direction-specific (feedforward only), which is not readily explained by symmetric, zero-lag relationships typical of volume conduction. We stated this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than instantaneous coupling.”

      We have also clarified the direction-specific logic in the Discussion, “Biased information transfer in USN patients” (p.22):

      “This hemispheric difference appeared only in the Feedforward (visual-to-frontal) direction, and no such difference was observed in the Feedback (frontal-to-visual) direction. This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube C, Gross J and Ince RAA, 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (4) In line 579, the authors write 'patients with USN may have larger lesions than those without USN.' It would be useful to clarify whether this claim is supported by the present data (i.e., lesion size comparisons between the two patient groups) or whether it reflects prior literature. If based on the current sample, please provide the corresponding statistical evidence.

      This issue has been addressed as part of our response to comment (1), with corresponding revisions made to the Discussion (p. 24), the Limitations (p. 25), and Supplementary Figure 1.

      (5) The dependent variable for SSVEP analyses is not clearly described, which makes it difficult to understand the results (e.g., L.214, L285-297). It seems that for all stimulation conditions, one value was extracted (and then compared between O1 and O2 electrodes), but is that the amplitude at the stimulated frequency (e.g., for a visual stimulation at f Hz, comparing the amplitude of the power spectrum at f Hz between O1 and O2)?

      Thank you for pointing this out. We clarified that the dependent variable for the SSVEP analyses was the power at the fundamental (stimulated) frequency f for each condition. We added this explicitly to the Methods, “Statistical analysis” (p.9):

      “The SSVEP power at the fundamental (stimulated) frequency f was analyzed for each condition as the dependent variable...”

      (6) Most analyses are described relatively precisely mathematically, but no toolbox or software is mentioned. Were they implemented using custom scripts? If toolboxes/software was used, the authors should indicate which ones and their versions, and add references. For instance, I might be wrong, but I guess the linear mixed model analyses were not performed using custom code. MATLAB is mentioned for the simulations, but no version is indicated. This is particularly important for reproducibility since the authors did not provide direct access to their scripts.

      The EEG preprocessing, PAC, and cluster-based permutation tests were conducted using the FieldTrip toolbox integrated with custom MATLAB scripts. Linear mixed-effects models were performed in SPSS. We added this information to the Methods, “Preprocessing” (p.7):

      “All EEG analyses were implemented using custom scripts in MATLAB (MathWorks, Natick, MA, USA) and the FieldTrip toolbox (Oostenveld R et al., 2011).”

      We have also specified the tools used for the linear mixed-effects models and the cluster-based permutation test in the Methods, “Statistical analysis” (p.9):

      “…using a linear mixed model in SPSS… a cluster-based permutation test (Maris E and Oostenveld R, 2007) using FieldTrip.”

      (7) The rationale for correlating BIT scores specifically with SSVEP power (rather than with the hemispheric imbalance measure, or with MI/TE results) is not clearly articulated. The results section highlights several neural metrics (SSVEP imbalance, PAC, TE), so it would strengthen the manuscript if the authors explained why SSVEP power was prioritized for correlation analyses. Is there a theoretical or empirical reason for this choice?

      While our main group-level results demonstrate a frequency-selective hemispheric imbalance, Fig. 5 suggests that the lesioned and intact hemispheres are not necessarily simple mirror images of each other and can show distinct response profiles. Therefore, to clarify which hemisphere’s response changes are associated with symptom severity, we assessed correlations with BIT using hemisphere-specific SSVEP power. At the same time, it is also informative to directly test the extent to which symptom severity covaries with hemispheric imbalance metrics (i.e., left–right differences in SSVEP, PAC, and TE). Accordingly, in the revised manuscript we additionally analyzed correlations between BIT and hemispheric imbalance measures of SSVEP, PAC, and TE, and report these results in the Supplementary Information (Supplementary Fig. 3). We added this clarification to the Results, “Correlation between SSVEP power and USN severity” (p.17):

      “In addition, correlations between BIT and the hemispheric imbalance of SSVEP power are provided in the Supplementary Information (Supplementary Fig. 3A). We also report, as additional exploratory analyses, correlations between BIT and hemispheric imbalance measures derived from PAC and TE (Supplementary Fig. 3B and 3C). None of these correlations reached statistical significance after correction, although the feedforward TE imbalance to the left frontal region showed a marginal uncorrected association with BIT score (r = -0.51, uncorrected p = 0.05).”

      We have also added the following statement to the Discussion, “Biased information transfer in USN patients” (p.23).

      “This interpretation is also consistent with the trend-level correlation shown in Supplementary Figure 3C. Although the correlation did not survive correction for multiple comparisons, patients with a stronger feedforward TE bias toward the left frontal region tended to show higher BIT scores. This observation raises the possibility that asymmetric information transfer from the right visual cortex to the intact left frontal cortex may contribute to compensatory network reorganization associated with milder neglect symptoms.”

      (8) The discussion could be enriched by considering whether the observed abnormal alpha-band entrainment relates to the well-known alpha-lateralization phenomenon in spatial attention. In healthy individuals, covert attentional orienting is typically accompanied by lateralized modulations of alpha power (increased ipsilateral, decreased contralateral). Could the asymmetric alpha entrainment reported here be interpreted as a pathological exaggeration of this mechanism, thereby linking the electrophysiological findings more directly to attentional orientation deficits in neglect?

      We thank the reviewer for raising this valuable point. To integrate our results with the alpha-lateralization literature, we added a new paragraph to the Discussion, “Local hemispheric bias of alpha-band entrainment and attentional dysfunction in USN patients” (p.22) subsection

      “Clinically, USN presents as neglect of the left visual field and is generally thought to reflect a relative dominance of rightward orienting. Because alpha power is often interpreted as reflecting functional inhibition (Worden et al. 2000; Kelly et al. 2006; Foxe and Snyder 2011), classic alpha-power lateralization associated with rightward orienting would typically predict increased alpha power over the right hemisphere and decreased alpha power over the left. However, in our data, right-hemisphere alpha power during stimulation was comparable to that of controls, whereas the intact left hemisphere exhibited stronger stimulus-locked synchronization (SSVEP) and enhanced alpha–gamma coupling. Thus, the present findings do not appear to reflect a simple amplification of tonic spatial alpha-power lateralization. Rather, they suggest a dynamic, temporally selective form of alpha-mediated inhibition that organizes the timing of local excitability. From this perspective, alpha-power lateralization and stimulus-locked entrainment may reflect related but distinct aspects of alpha-based inhibitory control, with the former regulating spatial gating and the latter regulating the timing of sensory processing (Jensen and Mazaheri 2010).

      Moreover, such a timing-based bias may arise even in the absence of explicit flicker. Visual input is continuously sampled under the influence of intrinsic alpha activity, and perceptual sensitivity fluctuates with the phase of ongoing oscillations (Romei et al. 2008; Iemi et al. 2017; VanRullen 2016). As a consequence, processing is relatively facilitated when incoming events coincide with high-excitability phases and relatively suppressed otherwise (Busch, Dubois, and VanRullen 2009; Mathewson et al. 2010). Thus, a hemispheric bias in alpha-band synchronization capacity could bias the timing with which sensory input is sampled across the two hemispheres, even without explicit rhythmic stimulation. This view is also consistent with the possibility that the bias is less apparent during eyes-closed rest with minimal visual input, yet becomes behaviorally expressed in everyday settings where continuous input is sampled in a phase-dependent manner (Landau and Fries 2012; VanRullen 2016). Overall, a hemispheric bias in synchronization capacity within the alpha range likely disrupts the temporally coordinated sampling of visual input required for balanced spatial attention, contributing to the attentional deficits observed in USN.”

      (9) All reported t-tests should include degrees of freedom (df). This is essential for transparency and allows readers to assess the robustness of the statistical results.

      We added the degrees of freedom for the t-tests reported in the Results, “Hemispheric imbalance of TE” (p.15):

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This is an extensively revised version of a previously submitted manuscript that, as detailed in their 20-page response to the first reviews, satisfactorily addresses the reviewers' comments. In particular, the revised manuscript makes it much clearer how this work fits into and advances the field. The added experiments strengthen the rigor of the manuscript as well. Overall, this paper is ready to go.

      We thank the reviewer for the positive evaluation and for recognizing the improvements made in our revised manuscript.

      Reviewer #2 (Public review):

      The revised manuscript has been substantially improved. The authors have addressed many of my previous concerns through the addition of new data, analyses, and discussion. The characterization of epithelial folding in the ascidian Ciona provides valuable insight into a comparatively less explored morphogenetic system, and the imaging and quantitative analyses are overall compelling. That said, a few important points remain to be addressed.

      We thank the reviewer for the positive assessment and for acknowledging the improvements made in our revised manuscript. We are also grateful for the reviewer’s continued constructive suggestions.

      One remaining issue concerns the mechanistic novelty of the actomyosin redistribution described in this study. The authors emphasize that the key novelty lies in the stepwise translocation of actomyosin from the lateral membrane to the apical domain during the initial stage (apical constriction), followed by redistribution from the apical domain back to the lateral domain during the accelerated stage (invagination). I agree that the dynamic redistribution itself is potentially interesting and may represent an underexplored aspect of epithelial morphogenesis. However, as I discussed in my previous review comments, from a mechanics perspective, the role of apical actomyosin in driving apical constriction and of lateral actomyosin in contributing to tissue folding/invagination have already been demonstrated in multiple systems, although to varying extents depending on the model. Therefore, while the current study convincingly documents a distinct spatiotemporal sequence of actomyosin localization in Ciona atrial siphon tube formation, it could be clarified further to what extent this work advances new mechanical principles underlying epithelial folding, as opposed to revealing a variation in the deployment of previously described force-generating modules.

      Importantly, I think the manuscript has the potential to provide deeper conceptual insight if the authors more explicitly consider the significance of the "redistribution" process itself. Redistribution does not only involve the appearance of actomyosin at a new membrane domain; it also necessarily involves its disappearance from the previous domain. The latter aspect has, in my view, been much less explored in the literature. For example: Is the removal of lateral actomyosin during the early phase important for efficient apical constriction? Conversely, is the reduction of apical actomyosin during the later accelerated phase important for proper invagination mechanics? These questions are particularly interesting because they address whether redistribution between domains serves an active mechanical regulatory role, rather than focusing on the role of force-generating actomyosin at a given location.

      I acknowledge that addressing these questions experimentally could be technically challenging. One potentially powerful way to address this would be through the revised computational model. For example, the authors could test whether tissue folding is altered when actomyosin is allowed to accumulate at a new domain without being concomitantly depleted from the original domain. Such analyses could help distinguish whether redistribution itself has functional mechanical importance, rather than merely reflecting sequential recruitment to different cellular regions. In my opinion, incorporating this aspect would substantially strengthen the conceptual and mechanistic novelty of the study.

      We thank the reviewer for raising this important point. We agree that the mechanical significance of actomyosin redistribution, beyond the individual roles of apical and lateral contractility, deserves further clarification. In our study, quantitative analysis of F-actin dynamics revealed a bidirectional reorganization of the actomyosin network during siphon morphogenesis: the increase of F-actin intensity in one domain was accompanied by its reduction in the other domains during both the initial and accelerated invagination stages. These observations suggest that actomyosin redistribution may represent an active mechanical regulatory process rather than merely sequential recruitment to different cellular domains. Following the reviewer’s suggestion, we used the computational model developed in this study to examine its functional significance.

      Specifically, we modified the temporal dynamics of actomyosin activity by maintaining apical contractility during the accelerated stage or by prematurely enhancing lateral contractility during the initial stage in simulations. We found that sustained apical tension during the accelerated stage primarily induced stronger central cell elongation (Figure 6—figure supplement 1A-C), whereas elevating lateral actomyosin activity during the initial stage (14–16 hpf) suppressed central cell elongation and reduced the inward movement of surrounding cells toward the central region (Fig. 6—figure supplement 1D-F), resulting in earlier bending deformation with a flatter invaginating morphology (Fig. 6—figure supplement 1F). Although both perturbations eventually converged to similar final shapes, likely because they reached similar final actomyosin distributions, their distinct morphogenetic trajectories demonstrate the mechanical importance of the temporal sequence of actomyosin redistribution. These results further clarify the significance of the previously identified apico-basal tension imbalance and lateral contraction by revealing how their sequential activation coordinates tissue deformation. The initial dominance of apical contractility, together with limited lateral contraction, promotes cell elongation and convergence of the active region, whereas the subsequent shift toward lateral contractility facilitates cell shortening and deep tissue invagination.

      Thus, our results highlight that the bidirectional redistribution of actomyosin is not merely a consequence of morphogenesis, but contributes to the dynamic regulation of epithelial folding. We have incorporated these new simulation results and analyses into the revised manuscript as Figure 6—figure supplement 1.

      My other concern relates to the new optogenetic data presented in Figure 4-figure supplement 2. In the "Dark" samples, active myosin does not appear to be clearly enriched along the membrane, but instead seems relatively diffuse within the cytoplasm. This appears distinct from the images shown in Figure 2, where active myosin exhibits clear membrane enrichment. Could the authors provide top-view images for the samples shown in Figure 4-figure supplement 2? This would help clarify whether active myosin is indeed enriched along the apical membrane at 16 hpf and along the lateral membrane at 17 hpf in the "Dark" condition.

      We thank the reviewer for this careful observation. We agree that the p-MLC signal in the "Dark" samples of the original Figure 4—figure supplement 2 appeared less membrane-enriched than in Figure 2. This was due to a technical adjustment: because the optogenetic system occupied the 568 nm channel, we have to switch the p-MLC signal from Alexa Fluor 568 anti-rabbit IgG to Alexa Fluor 647 anti-rabbit IgG, which yielded relatively weaker membrane signal under our imaging conditions. To address this, we have replaced the cross-sectional images with better-representative examples and added top-view images (Figure 4—figure supplement 2). These new panels clearly showed that active myosin was enriched at apical junctions (16 hpf) and lateral membranes (17 hpf) in the Dark condition, consistent with that in Figure 2. Since the Dark and Light groups were processed identically, the relative comparison remains valid.

      In addition, the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4-figure supplement 2 appears noticeably different from that shown in Figure 4. Specifically, the apical side of the tissue in Figure 4 appears substantially more relaxed than in Figure 4-figure supplement 2. Based on the authors' interpretation of the optogenetic experiments, apical active myosin is not strongly affected by the treatment described in Figure 4. If so, one would expect apical constriction to remain largely intact. However, the more relaxed apical domain shown in Figure 4 seems to suggest that apical constriction may in fact be perturbed by the optogenetic manipulation. This apparent discrepancy complicates the interpretation of the experiment and seems somewhat inconsistent with the authors' main conclusion from this figure.

      We thank the reviewer for this very careful and insightful observation. We agree that the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4—figure supplement 2 appears less relaxed than that in Figure 4B.

      We acknowledge that optogenetic inhibition of myosin activity might indeed have a partial effect on activity of apical myosin, which can lead to apical relaxation after the initial constriction, as shown in the representative embryo in Figure 4B. However, because Ciona embryos are not always perfectly synchronized at 16 hpf (the time point when light illumination was initiated), individual embryos exhibit slight variations in the extent of apical constriction that has already been achieved during the initial stage (13.5–16.0 hpf). For embryos that had completed a relatively stronger apical constriction by 16 hpf, the apical domain can maintain its constricted morphology even after light exposure (Figure 4—figure supplement 2). Embryos with relatively weaker apical constriction at the time of light onset are more prone to exhibit apical relaxation upon optogenetic manipulation, as illustrated in Figure 4B. This developmental heterogeneity is the primary reason for the morphological variability observed between individual embryos in the optogenetic groups. Importantly, despite this morphological variability, the quantitative comparison of apical p-MLC intensity between the Light and Dark groups in Figure 4—figure supplement 2B showed no statistically significant difference (t-test, ns), which is consistent with the fact that during normal development, apical myosin activity naturally declines after 16 hpf (Figure 2B). In contrast, lateral p-MLC intensity was significantly reduced in the Light group compared to the Dark control (Figure 4—figure supplement 2B). This reduction in lateral contractility is the key factor responsible for the blockade of invagination progression.

      We have revised the statements accordingly. Hopefully, these clarifications have adequately addressed the reviewer's concern.

      Reviewer #3 (Public review):

      Concerns raised in the initial submission were addressed in the revised manuscript.

      We thank the reviewer for the encouraging feedback and for acknowledging our revisions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further revisions suggested.

      We thank the reviewer again for the positive assessment.

      Reviewer #2 (Recommendations for the authors):

      Here are several additional comments and suggestions in addition to the concerns described in the public review.

      Line 86 - 88: the authors state "However, this transition is from apical to basolateral (Sherrard et al., 2010), rather than a bidirectional redistribution between apical and lateral domains." This statement feels somewhat out of context because the concept of "bidirectional redistribution" has not yet been introduced. It may fit more naturally in the Discussion section, after the relevant observations and interpretations have been fully presented.

      We thank the reviewer for this suggestion. We have revised the sentence in the Introduction to avoid introducing the concept of "bidirectional redistribution" prematurely.

      Line 92 - 95: the authors raised the questions of "Whether a bidirectional redistribution of actomyosin between apical and lateral domains operates as a core mechanism for sequential invagination, and whether lateral contractility is essential for the accelerated phase, remain unclear." These questions also feel somewhat out of context, for the same reason mentioned above.

      We thank the reviewer for this suggestion. We have revised the text to avoid prematurely introducing the concept of "bidirectional redistribution" in the Introduction.

      Line 162 - 163: The authors state that "This redistribution pattern was consistent with that of F-actin in the corresponding phases." This conclusion should be revised, as the reported increase in apical F-actin and reduction in lateral F-actin during the initial stage do not appear to reach statistical significance, which is different from that of active myosin.

      We thank the reviewer for this careful observation. We agree that the F-actin changes during the initial stage did not reach statistical significance, unlike the active myosin data. We have revised the sentence to state that the myosin redistribution showed a similar trend to F-actin, while acknowledging the lack of statistical significance for F-actin.

      Line 349 - 350: "and the invagination speed (represented by slopes of curves in Figure 5A) gradually slows down at later stages." It seems that Figure 5A should be Figure 5B.

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      Reviewer #3 (Recommendations for the authors):

      We appreciate the efforts made by the authors to address the questions and comments raised in the initial submission. The only remaining concerns are regarding grammar and spelling and a couple of minor errors.

      (1) Lines 346-347, "These trends are consistent with the experimental mutant data (Figure 5A, B)."

      This was confusing - was this meant to be Figure 3A, B?

      We sincerely apologize for this confusion. We have corrected it in the revised manuscript (now Figure 3B).

      (2) Lines 349-350, "... the invagination speed (represented by the slopes of curves in Figure 5A) gradually slows down at later stages..."

      Similar to above, was this meant to be Figure 5C?

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      (3) There is a typo in Line 336 ("dimmish") and some minor grammatical concerns in the introduction.

      We have corrected the typo "dimmish" to "diminish". We have also carefully proofread the entire manuscript and fixed any remaining grammatical issues.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study presents a novel analysis of MRI resources for 16 avian species, spanning major (though not all) clades and ecological niches. This is a significant step towards large-scale datasets on internal parcellation and long-range connectivity, central to evolutionary studies for understanding the evolution of the bird brain.

      Strengths:

      The integration of high-resolution T2-weighted and diffusion-weighted MRI with histological validation (Nissl and Luxol Fast Blue staining) provides a strong, cross-validated framework for studying avian brain anatomy. Data on long-range connectivity are particularly useful for understanding how relationships between brain components evolved. The approach is also scalable, allowing for more detailed evolutionary analyses compared to what is currently possible.

      Weaknesses:

      The sampling supports evidence of modular evolution in the bird brain, but it is limited for broad evolutionary claims, as the effects of sizes and phylogenies can be hard to disentangle without enough species per clade.

      We appreciate the reviewer highlighting this limitation and agree that it warrants explicit consideration. Although the 16 species included in the current study encompass diverse avian lineages and ecological characteristics, our sampling was not designed to rigorously disentangle the effects of phylogeny and brain size within individual clades.

      To improve the taxonomic coverage of our dataset, we plan to expand the MRI dataset in the revised manuscript by including additional specimens that are currently available to us. Specifically, we will acquire and analyze MRI data from the slaty-backed gull (Larus schistisagus), black-tailed gull (Larus crassirostris), budgerigar (Melopsittacus undulatus), and crow (Corvus sp.). In addition, we plan to incorporate publicly available MRI data from the ostrich (Struthio camelus). These additions will broaden the phylogenetic and anatomical coverage of the comparative dataset.

      Nevertheless, we fully acknowledge the reviewer’s point that, even with these additional species, the number of species sampled within individual clades will remain insufficient to rigorously separate phylogenetic effects from size-related effects or to support broad generalizations across all avian lineages. We will therefore explicitly state this limitation in the revised manuscript and temper evolutionary claims that extend beyond what can be supported by the present taxonomic sampling.

      Accordingly, we will more clearly position the primary contribution of this study as the establishment of an MRI-based comparative resource and analytical framework applicable across diverse avian species, rather than as a comprehensive phylogenetic test of avian brain evolution. Within this scope, we will present the observed variation among brain compartments as evidence consistent with mosaic/modular diversification, while carefully limiting the broader evolutionary interpretation of these patterns.

      Tractography-based claims should be treated cautiously without sensitivity analyses. This is particularly important when comparing brains with different sizes and tissue properties.

      The reviewer raises an important point regarding the interpretation of tractography-based comparisons. We agree that particular caution is required when comparing datasets across species that differ in brain size, tissue properties, and imaging characteristics.

      In the revised manuscript, we will perform additional sensitivity analyses to evaluate the robustness of the major tractography-derived patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction, rather than relying solely on the current threshold of 0.1, and assess whether the major reconstructed trajectory patterns and cross-species differences are robust to this parameter.

      We will also re-analyze the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the underlying diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret these comparisons within the limitations of the currently available datasets, without implying that either reconstruction approach provides a definitive representation of the underlying fiber architecture.

      In addition, we will provide a more detailed description of the relevant diffusion MRI acquisition parameters and spatial and angular resolution of the datasets so that potential technical differences among species can be more clearly evaluated. We will also discuss how such differences may affect cross-species tractography comparisons.

      Finally, we will revise the interpretation of the tractography results throughout the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid treating reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the potential for false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      Existing literature is not acknowledged sufficiently. This makes some claims of novelty misleading, and prevents readers from understanding the current state of knowledge in this research area.

      We fully agree with this assessment. Although substantial comparative neuroanatomical work has established important principles of avian brain evolution, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a result, some statements regarding the novelty of the present study are overstated.

      In the revised manuscript, we will substantially expand our discussion and citation of the relevant literature, including the extensive comparative studies of individual avian brain regions and brain subdivisions highlighted by the reviewers. We will revise the Introduction and Discussion accordingly to more accurately describe the current state of knowledge in comparative avian neuroanatomy and to clearly distinguish the contributions of previous studies from those of the present work.

      We will also carefully revise statements regarding the novelty of our study. Rather than implying that our study provides the first comparative framework for examining internal avian brain organization, we will more appropriately emphasize its contribution as an MRI-based comparative resource that enables standardized visualization and analysis of internal brain anatomy and connectivity across diverse avian species.

      Reviewer #2 (Public review):

      This manuscript presents a comparative MRI dataset from 16 avian species and uses MRI and tractography to examine variation in brain organization across birds. The authors argue that their analyses support mosaic brain evolution and provide a framework for comparative neuroanatomy. Although the dataset represents a useful resource, particularly given the inclusion of understudied species like penguins, toucans, and hornbills, I have substantial concerns regarding the novelty of the study, the anatomical interpretation of the results, and the validity of the tractography analyses. In its current form, I do not believe the manuscript provides sufficient new biological insight to support many of its conclusions.

      Major Concerns

      (1) The authors repeatedly state that comparative neuroanatomical studies in birds have largely been unable to examine internal brain organization or "internal parcellation". This claim is inaccurate and reflects limited engagement with a substantial body of literature. For decades, comparative studies have examined variation in the size of major avian brain subdivisions as well as specific sensory, motor, and associative nuclei. For example, the extensive work of Andrew Iwaniuk and colleagues has documented variation in numerous brain regions across birds and related these differences to ecology, behavior, and sensory specialization (e.g., Gutierrez-Ibanez et al., 2009; Iwaniuk et al., 2006, 2008, 2010; Corfield et al., 2015). Other authors have also made important contributions in this area (e.g., Boire and Baron, 1994; Burish et al., 2004; Moore and DeVoogd, 2011, 2017). Importantly, previous work has already examined variation in major subdivisions of the avian brain using relatively standardized datasets (e.g., Iwaniuk et al., 2004; Iwaniuk and Hurd, 2005), including datasets that contain more species and greater taxonomic diversity than the current study. In other words, these studies have already provided detailed analyses of internal brain organization across broad taxonomic samples.

      The manuscript should therefore be reframed as providing a new MRI-based resource rather than introducing the first comparative framework for studying internal avian brain organization. The current framing significantly overstates the novelty of the work.

      We appreciate the reviewer drawing attention to this issue. Although a substantial body of comparative neuroanatomical work has already examined internal brain organization in birds, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a consequence, some statements regarding the novelty of our study are overstated.

      In the revised manuscript, we will expand our discussion and citation of previous work, including the studies highlighted by the reviewer on interspecific variation in major avian brain subdivisions, sensory and motor nuclei, and other anatomically defined brain regions. We will revise the Introduction and Discussion to more accurately represent the existing body of comparative avian neuroanatomy and to clarify how the present study complements and extends these established approaches.

      Most importantly, we will reframe the manuscript so that its primary contribution is presented as the establishment of an MRI-based comparative resource for visualizing and quantitatively analyzing internal brain anatomy across diverse avian species. We will remove or revise statements implying that the present study provides the first comparative framework for examining internal avian brain organization. Instead, we will emphasize the complementary advantages of the MRI-based approach, particularly its ability to provide non-destructive three-dimensional visualization of internal brain structures in intact specimens and to enable comparison of multiple brain compartments within a common analytical framework across diverse avian species.

      We believe that this revised framing will more accurately position the contribution of our study within the existing literature and clarify the specific value of the dataset.

      (2) A second significant concern is the lack of anatomical specificity in the tractography analyses. The authors repeatedly refer to regions such as "anterior cortex," "dorsal cortex," and "temporal cortex." These terms are not standard anatomical designations in avian neuroanatomy and provide little information about the actual structures being analyzed.

      For example, the "temporal cortex" could potentially include portions of the nidopallium (including the caudolateral nidopallium, NCL), mesopallium, and arcopallium. Similarly, the "anterior cortex" could correspond to the somatosensory or visual Wulst, the anterior nidopallium, or several other structures. The designation "dorsal cortex" is similarly difficult to interpret. Because these seed regions may encompass multiple functionally distinct systems, it is impossible to evaluate the biological significance of the reported connectivity patterns.

      I strongly encourage the authors to define their seed regions using accepted avian neuroanatomical terminology and to provide detailed anatomical maps. More informative analyses would focus on well-defined structures with known connectivity, such as the Wulst, arcopallium, entopallium, or NCL. As currently presented, the tractography results are too coarse to support meaningful biological conclusions.

      This is an important concern, and we agree that the anatomical nomenclature and definition of the regions used in the tractography analyses require substantial improvement.

      In the revised manuscript, we will carefully re-evaluate the anatomical description of each region with reference to established avian neuroanatomical terminology, anatomical atlases, and available histological information. Where the analyzed regions can be reliably assigned to established anatomical structures, we will replace broad mammalian-style positional terminology such as “anterior cortex,” “dorsal cortex,” and “temporal cortex” with more appropriate avian neuroanatomical terminology. For example, we will re-examine whether the region currently referred to as the “optic lobe” can be more precisely defined as the optic tectum, while the cerebellum can be retained as an anatomically well-defined region.

      At the same time, we recognize an important limitation of the present tractography analysis. The spatial resolution and analytical framework of the present comparative datasets do not necessarily permit reliable assignment of all analyzed regions or reconstructed trajectory patterns to fine pallial subdivisions such as the entopallium, arcopallium, or NCL across species. We therefore do not intend to assign such specific anatomical identities where they cannot be supported with sufficient confidence.

      Instead, for regions that cannot be unambiguously assigned to a single established anatomical subdivision, we will define their location and extent using reproducible anatomical landmarks and clearly indicate the level of anatomical resolution supported by the data. We will also provide revised anatomical maps showing the locations of the analyzed regions and their relationship to major avian brain subdivisions. This will allow readers to evaluate more clearly which anatomical structures may contribute to the reconstructed trajectory patterns.

      Accordingly, we will revise the biological interpretation of the tractography results to match this anatomical resolution. Rather than attributing the reconstructed patterns to specific fine-scale pallial structures or functional systems when these cannot be reliably distinguished, we will interpret them more conservatively as broad patterns of fibre organisation among anatomically defined brain regions. These revisions will improve the anatomical transparency of the analysis while avoiding anatomical or functional interpretations that exceed the resolution of the present datasets.

      (3) I am not an MRI specialist, but I have concerns regarding the interpretation of the tractography results. Bird brains are small, and diffusion MRI tractography is already known to be challenging even in substantially larger brains. The manuscript provides limited information regarding image resolution, diffusion sampling, and the expected accuracy of tract reconstruction in these specimens. More importantly, there is little validation of the tractography results. Diffusion tractography is prone to both false positives and false negatives, and reconstructed pathways cannot be assumed to represent true anatomical connections.

      The authors should provide evidence that their tractography pipeline can accurately recover known pathways. For example, they could compare reconstructed tracts with well-established anatomical pathways such as the anterior commissure or major visual pathways, which would substantially strengthen confidence in the results. Without such validation, it is difficult to determine whether the observed species differences reflect biological variation or methodological artifacts.

      We agree with the reviewer that the tractography results require cautious interpretation and that further evaluation of the robustness and anatomical plausibility of the tractography pipeline would strengthen the study.

      In the revised manuscript, we will provide a more detailed description of the diffusion MRI acquisition parameters, spatial and angular resolution, diffusion reconstruction, and tractography procedures so that the methodological limitations of the analyses can be more clearly assessed.

      Recent studies using DTI under imaging conditions comparable to those employed in the present study have demonstrated successful reconstruction of fiber trajectories in the mouse brain, which is smaller than the avian brains examined here (Janz et al., eLife, 2017). We therefore consider that the size of the avian brain itself does not preclude DTI-based fiber reconstruction and that such analyses are technically feasible at the brain sizes examined in the present study.

      Importantly, we will perform additional sensitivity analyses to evaluate the robustness of the reconstructed trajectory patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction. We will also re-analyse the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret the comparison between reconstruction approaches within the limitations of the currently available datasets, without implying that either approach provides a definitive reconstruction of the underlying fiber architecture.

      We will also evaluate the robustness of the k-means clustering used to summarize tractography patterns. Because the current choice of k = 10 was not based on an independently established biological criterion, we will examine alternative values of k and assess whether the major trajectory patterns and cross-species differences are robust to the choice of cluster number.

      In addition, we will examine anatomically well-characterized pathways, including major commissural and visual pathways, and assess whether the reconstructed trajectories are consistent with known avian neuroanatomy. We will use these comparisons as an assessment of anatomical plausibility rather than as definitive validation of tractography accuracy.

      Finally, we will revise the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid interpreting reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the possibility of false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      (4) I also have some methodological concerns regarding the comparisons of anterior commissure (AC) size and cerebellar foliation. First, the authors measure the AC in a coronal section. I would recommend measuring the AC area in a midsagittal section instead. Furthermore, the authors use the cross-sectional area of the same coronal section as the scaling variable. This seems problematic because the area of any given section will depend on the angle of sectioning and other technical factors. If the objective is to compare the relative size of the AC, then total brain volume or telencephalon volume would be more appropriate scaling variables.

      With respect to cerebellar foliation, the authors developed their own metric. I would encourage them to use methods already established in the literature, such as the foliation index described by Iwaniuk et al. (2006). Their approach may yield similar results, but using the foliation index would facilitate direct comparisons with existing datasets and would allow incorporation of additional published data (e.g., Cunha et al., 2021, which includes foliation index measurements for 54 bird species). The authors should also be aware that the foliation index scales with body size. Consequently, the high degree of foliation observed in penguins may not necessarily indicate cerebellar expansion or increased demands for sensorimotor integration associated with their specialized locomotion. I therefore believe that the conclusions regarding variation in AC size and cerebellar foliation should be re-evaluated after more appropriate analyses are performed.

      These methodological suggestions are very helpful. We agree that both the anterior commissure analysis and the assessment of cerebellar foliation should be re-evaluated to enable more anatomically and quantitatively appropriate comparisons across species.

      Anterior commissure:

      We agree that the current analysis, in which AC area was measured from the coronal section showing the largest cross-sectional profile and normalized to whole-brain area in the same section, may be influenced by differences in brain geometry and section orientation among species. In the revised manuscript, we will re-evaluate the method used to quantify AC size, including measurements from sagittal or midsagittal views where anatomically appropriate. We will also examine normalization against volumetric measures, such as total brain or telencephalic volume, rather than relying solely on a single-section whole-brain area. The corresponding results and interpretations will be revised accordingly.

      Cerebellar foliation:

      We also agree that our current branch-counting approach should be considered in relation to established quantitative measures of avian cerebellar foliation. In the revised manuscript, we will evaluate whether the cerebellar foliation index described by Iwaniuk et al. (2006), or a comparable standardized measure applicable to our midsagittal MRI datasets, can be used to re-analyze cerebellar foliation. This will also allow us to place our observations more directly in the context of the larger comparative datasets reported previously, including that of Cunha et al. (2021).

      Importantly, we recognize that cerebellar foliation is strongly influenced by allometric and phylogenetic factors. We will therefore re-evaluate the interpretation of interspecific differences in cerebellar foliation in light of brain/body size relationships and the existing comparative literature. In particular, we will revise the interpretation of the pronounced cerebellar foliation observed in penguins and other species and avoid attributing these differences directly to locomotor or sensorimotor specialization unless supported by the revised analyses.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Naina Gour and colleagues provide a detailed observational study in which they demonstrate that MRGPRX4, a human G-protein coupled receptor (GPCR), is expressed exclusively in human melanomas and, when expressed in mouse melanocytes, drives the development of melanomas in mice. These findings provide evidence that MRGPRX4 has the properties of an oncogene, at least in certain cellular environments.

      Strengths:

      A strength of this work is the nice historical note in which overexpression of MAS1, a GPCR, led to classic studies of transformed fibroblasts in culture and tumors in nude mice. Cloning of MAS1 led to the identification of the MRGPR family of receptors, now known to be key players in neuroimmune and neurosensory phenomena. Here, the story comes full circle with a member of the MRGPR family being linked to a tumor, specifically melanoma. Perhaps the story is not entirely surprising given that the neural crest serves as a precursor for both nerves and melanocytes. But it is nice to see.

      Additional strengths include the vast array of tools and techniques employed, from public databases to engineered mice, to establish firmly that MRGPRX4 is expressed in melanomas, although not in every malignant cell.

      We thank the reviewer for the positive assessment of our work and for noting the arc back to the original MAS1 studies, which we agree makes for a satisfying full-circle story. We would add one nuance: while the shared neural crest origin of melanocytes and nociceptors offers a plausible explanation for why an MRGPR family member could be co-opted in melanoma, we were nonetheless surprised that this property is specific to MRGPRX4. When we tested the other MRGPRX family members, some of which are also known to be expressed in sensory neurons, using the same genetic strategy (MRGPRX1, MRGPRX2, MRGPRX3; Supplementary Fig. 2B), none induced melanoma, despite arising from the same family of receptors. This suggests the oncogenic property is not simply a generic consequence of neural crest ancestry, but reflects a biology specific to MRGPRX4 itself, which we find is one of the more intriguing open questions this work raises.

      Weaknesses:

      Given the power of the strengths of the data and story, the following comment is only sort of a weakness, as the topic is addressed while being saved for future studies. Specifically, what leads to the expression of MRGPRX4? The authors posit that it is an epigenetic phenomenon, look briefly at methylation, and rather than going down the proverbial rabbit hole of what comes first, have reasonably decided to punt.

      We agree with the reviewer that understanding the precise mechanisms through which MRGPRX4 gets activated during oncogenesis is important, and we had noted this as a limitation of current findings. As we mentioned in the discussion, we hypothesize that one or more environmental modulators - UV exposure, inflammatory cues, or an aging skin microenvironment could initiate the epigenetic reprogramming underlying this expression. This will require dedicated studies and is suited for future work.

      Another concern is that given what comes across as the initial observation of MRGPRX4 being expressed in melanoma, what do all of the additional studies add?

      The expression data in human melanoma are correlative and serve only as the entry point. The subsequent studies establish MRGPRX4 as a causal, druggable driver of melanoma: it is sufficient to induce 100% penetrant metastatic melanoma in vivo, required for melanoma proliferation and invasion, signals through ligand-independent basal PI3K-AKT/MAPK activity, and is pharmacologically targetable.

      For the non-cognoscenti, and to make the manuscript more accessible, the abbreviation NC/EMT, which is also inverted to EMT/NC, should be spelled out periodically as neural crest/epithelial-mesenchymal transition.

      This is corrected in the revised version.

      Please explain how this study came about. Was it a result of someone deciding to look at expression in the GTEx project and compare it to a tumor database?

      This was a serendipitous discovery. As MRGPRs can modulate immune cell function, we were driving expression of various human MRGPRX genes in immune lineages in mice. We noticed that mice overexpressing MRGPRX4 developed spontaneous melanoma-like growths on the ear and tail. Upon investigation, we found that while the Cre driver we had used is appropriate for immune cells, it is also expressed at low levels in melanocytes. We hypothesized that our initial "immune-knock-in" strategy likely resulted in low-level knock-in in mouse melanocytes, thereby allowing MRGPRX4 to directly drive melanocytic transformation. This led us to explore MRGPRX4 expression in human malignancies, where we found MRGPRX4 was highly expressed in melanoma. Next, to directly test our hypothesis that MRGPRX4 transforms melanocytes, we overexpressed MRGPRX4 using Tyrosinase-CreER, the standard Cre driver for melanocytes. This led to a fully penetrant spontaneous melanoma phenotype (Fig. 2).

      A comment could be made to explain that while NSG and normal mice were used, the former are immunocompromised, and drawing conclusions without specifying these differences is a weakness.

      This is a fair point, and we agree the distinction deserves clarification. The two mouse systems in this study serve different purposes and are not interchangeable. The transgenic TyrCreER+; MRGPRX4LSL+/- model, in which melanoma arises spontaneously from endogenous mouse melanocytes, was studied in immunocompetent mice, allowing us to fully characterize the tumour, including the tumour immune microenvironment (Fig. 2L–O), in an intact immune setting. In contrast, the human A2058 melanoma cell studies (proliferation and metastatic seeding; Fig. 4B–J) required an immunodeficient host, since human cells would otherwise be rejected by a competent murine immune system. For these experiments, we used NSG mice, the standard immunodeficient strain used in human xenograft studies. We have updated the text to reflect this.

      In Figure 1A, the p-value of -145 begs for a little explanation. I don't recall seeing such a p-value.

      The small p-value reflects two features of the comparison: (1) both GTEx normal skin and TCGA-SKCM contain large sample sizes (hundreds of samples per group), which gives the Mann-Whitney U test (also known as the Wilcoxon rank-sum test) statistical power, and (2) MRGPRX4 expression shows minimal overlap between the two groups; it's essentially undetectable in normal skin but broadly expressed across melanoma samples. Together, these yield such a p-value. This is expected for rank-based tests applied to large, cleanly separated datasets.

      Have you considered treating the murine melanomas with murine via PD-L1? I appreciate that this comment is somewhat superfluous given the inhibition of MRGPRX4 with compound 31-2, but given the human therapeutics combined with the fact that you have done 'everything else', I wonder what might happen.

      We agree with the reviewer that testing checkpoint blockade in the MRGPRX4-driven model is a relevant direction. Our finding that MRGPRX4-driven tumours develop a PD-L1hi, neutrophil-rich microenvironment (Fig. 2M–O) makes anti-PD-L1 treatment a natural next experiment, and we would predict it to be informative both on its own and in combination with an MRGPRX4 inhibitor, given that the two target distinct compartments- the tumor-immune microenvironment versus tumour-intrinsic proliferative/invasive signaling. We consider it an important direction for future work.

      Given the basal ligand-independent signaling, might engineering variants of MRGPRX4 that do not signal be of value?

      This is a great point. A variant that disrupts basal (ligand-independent) signaling specifically would be informative. Identifying and validating such a basal-activity-disrupting variant, including confirming normal receptor expression and trafficking, is an important direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study presents a fundamental new finding - the identification of a sensory-neuron itch receptor, MRGPRX4, as an unexpected melanoma oncogene through a mechanism of lineage-inappropriate expression rather than mutation. The evidence supporting the core observation (tumor-specific upregulation, restriction to invasive transcriptional states, and sufficiency to drive fully penetrant metastatic melanoma in vivo) is compelling, drawing on convergent human genomic datasets and a well-controlled genetic mouse model. However, several of the mechanistic and translational claims - particularly regarding causal drivers of invasion, the immunosuppressive tumor microenvironment, and in vivo pharmacological efficacy - remain incomplete, relying on correlative evidence.

      Strengths

      The authors propose that MRGPRX4, normally restricted to a subset of peripheral sensory neurons, is aberrantly re-expressed in melanoma rather than through mutational mechanisms, and that this re-expression is sufficient to drive tumorigenesis through basal, ligand-independent GPCR signaling. This is a genuinely novel model for oncogenesis, and the manuscript deploys an impressive range of approaches - bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, phosphoproteomics, and pharmacology - to support it.

      Strengths:

      The claim that MRGPRX4 is selectively upregulated in melanoma and confined to neural-crest-like/invasive transcriptional states is well supported, with consistent results across multiple independent human scRNA-seq datasets. The claim that ectopic MRGPRX4 is sufficient to drive melanoma is convincingly demonstrated by the fully penetrant, metastatic phenotype in the Tyr-CreER;MRGPRX4-LSL model, with appropriate specificity controls showing that MRGPRX1, MRGPRX2, and MRGPRX3 do not phenocopy this effect.

      The claim that MRGPRX4 signals through basal, ligand-independent activity is reasonably well supported by bile-acid quantification showing endogenous ligand concentrations well below the EC50 required for activation.

      We are thankful to the reviewer for the positive feedback.

      Weaknesses:

      The claim that MRGPRX4 remodels the tumor microenvironment toward an immunosuppressive state rests on flow cytometric frequency data (altered neutrophil/eosinophil ratios, increased PD-L1+ myeloid populations) but lacks any functional immune assay to demonstrate that this remodeling actually impairs anti-tumor immune responses.

      We acknowledge the reviewer’s suggestion that functional immune assays will demonstrate that MRGPRX4-driven tumor remodeling impairs checkpoint-driven anti-tumour immune response. We consider this an important direction for future work.

      The claim that the two MRGPRX4-enriched tumor subpopulations (ECM-rich and NC-like/invasive) underlie the observed invasive and metastatic phenotype is not directly tested; the authors appropriately acknowledge this as an open question, but it is worth noting explicitly that this leaves the mechanistic link between the identified cell states and the functional phenotype (proliferation, invasion, metastasis shown in Figure 4) unresolved.

      We agree with the reviewer. To test the direct role of these subpopulations in driving metastasis in our model, we need to ablate them specifically. This is possible with an intersectional genetics strategy, wherein we could use a state-specific Dre or Flp driver (e.g., a Prrx1-DreER or Prrx1-FlpO knock-in) crossed to a dual-recombinase-dependent effector allele (e.g., a Frt- or Rox-gated diphtheria toxin receptor), and then crossed to our TyrCreER-MRGPRX4-LSL model. This quadruple-transgenic strategy would allow ablation restricted specifically to MRGPRX4-driven tumour cells occupying the target mesenchymal state, for example. These genetic lines can be created but require generating or sourcing new dual-recombinase-dependent alleles and a substantially longer breeding and validation timeline (approximately 18-24 months). Together, these represent possible, if long-term, strategies for directly testing the causal contribution of these MRGPRX4-enriched subpopulations to melanoma invasion and metastasis.

      Finally, the comparison with BRAF- and NRAS-driven GEMMs (Figure 3K-L) establishes overlap in transcriptional cell states but does not report whether these canonical models themselves upregulate endogenous Mrgprx4. This omission leaves unclear whether MRGPRX4 acts as a convergent node downstream of canonical oncogenic signaling, or represents an independent, parallel route to a similar phenotypic endpoint - a distinction that matters considerably for how broadly the finding should be interpreted.

      MRGPRX4 is a primate-specific receptor with no mouse ortholog in melanocytes; thus, we cannot assess endogenous expression of MRGPRX4 in BRAF/NRAS-driven GEMMs. Nevertheless, the following observations argue against MRGPRX4 functioning solely downstream of canonical oncogenes. First, TyrCreER<sup>+</sup>; MRGPRX4LSL mice develop fully penetrant melanoma without engineered BRAF or NRAS activation or tumour-suppressor loss. Second, MRGPRX4 loss in BRAF V600E mutant A2058 cells reduces pERK1/2, pAKT, and pS6K, showing that MRGPRX4 sustains these signaling outputs even in the presence of activated BRAF. Together, these findings support a model in which MRGPRX4 could provide a distinct oncogenic input.

      Overall assessment:

      The manuscript's central, most novel claim - that lineage-inappropriate expression of a sensory GPCR is sufficient to drive melanoma - is compellingly supported. The secondary mechanistic and translational claims built around this finding are convincing and consistent with the broader literature but are currently supported by correlative rather than causal or functional evidence.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of our manuscript. We appreciate the reviewers’ insightful comments and suggestions, which will help strengthen the mechanistic rigor of our work. Below, we outline the key revisions we plan to undertake in the revised version.

      Response to Reviewer 1

      (1) Dynamic regulation of PA lactylation during infection.

      We plan to examine PA lactylation levels at multiple time points post-infection to assess whether PA lactylation changes dynamically during the viral replication cycle.

      (2) Epistasis experiments linking K605/K609 to lactate- or enzyme-dependent phenotypes.

      We acknowledge that multiple viral proteins undergo lactylation and that ATAT1/SIRT1 may mediate lactylation of multiple viral proteins. We will perform viral replication assays in the context of K605/K609 mutant viruses under lactate supplementation or ATAT1/SIRT1 manipulation conditions. These experiments will allow us to assess whether the effects of lactate or ATAT1/SIRT1 manipulation on viral replication are dependent, at least in part, on PA K605/K609.

      (3) Conservation across viral strains and cell lines.

      We will further examine key phenotypes in additional influenza A virus subtypes (e.g., H1N1 swine influenza and H9N2 avian influenza strains) and in additional cell lines to assess the extent to which the observed mechanism is conserved beyond the PR8 laboratory-adapted strain.

      (4) Direct biochemical evidence for ATAT1 and SIRT1 activity on PA.

      We plan to perform in vitro lactylation and de-lactylation assays using purified recombinant PA, ATAT1, and SIRT1 proteins to investigate whether ATAT1 and SIRT1 can directly modulate PA lactylation, respectively.

      (5) Disentangling lactylation from charge/structural effects.

      We acknowledge the reviewer’s point that K609R shows reduced lactylation without obvious changes in polymerase activity. We will include K-to-Q substitution mutants (e.g., K605Q/K609Q) in functional assays to further assess the functional consequences of these substitutions. Although K-to-Q substitutions do not strictly mimic lysine lactylation, these mutants may help distinguish effects related to lysine charge/chemical properties from those specifically attributable to lactylation. We will also temper our conclusions and acknowledge that charge and/or structural effects may contribute independently to the observed phenotypes.

      (6) Enzymatic activity dependence of ATAT1 and SIRT1.

      We will further investigate the enzymatic activity-dependent versus -independent contributions of ATAT1 and SIRT1 using catalytically inactive mutants, together with the epistasis experiments described above. We will also revise the text to clarify the interpretive limitations of these experiments.

      (7) Improved loss-of-function approaches.

      We will complement the existing siRNA experiments with CRISPR/Cas9 knockout cell lines for ATAT1 and SIRT1 and, where feasible, repeat key assays in ATAT1- and SIRT1-knockout cells.

      In the revised manuscript, we will also add a detailed methodological explanation in the figure legend and Methods section to clarify how the luciferase complementation system distinguishes asymmetric from symmetric polymerase dimers, show individual data points overlaid on bar graphs with error bars, and correct spelling errors throughout the manuscript.

      Response to Reviewer 2

      (1) Direct in vitro lactylation/de-lactylation assays.

      As noted above, we will perform in vitro modification assays with purified proteins to investigate whether ATAT1 and SIRT1 directly modulate PA lactylation and de-lactylation, respectively.

      (2) Functional assays with lysine-to-glutamine substitution mutants.

      Although we recognize that K-to-Q substitutions do not strictly mimic lysine lactylation, we will generate K-to-Q substitution mutants (e.g., K605Q/K609Q) and evaluate their effects on polymerase activity and viral replication to further assess the functional relevance of these sites.

      (3) Complementation experiments linking ATAT1/SIRT1 phenotypes to PA K605/K609.

      As noted above, we plan to perform viral replication assays in the context of K605/K609 mutant viruses under ATAT1/SIRT1 manipulation conditions to assess whether PA K605/K609 contributes to the effects associated with ATAT1/SIRT1 manipulation.

      (4) IP-MS comparison of PA WT versus mutant host interaction profiles.

      Our study focuses on the mechanism by which PA lactylation modulates viral polymerase activity and replication. A comprehensive host interactome analysis via IP-MS represents a broader systematic investigation beyond the scope of this focused work. We will discuss this as an important future research direction in the revised manuscript.

      (5) Exploration of host antiviral immunity downstream of PA lactylation.

      This work focuses on the direct effects of PA lactylation on viral polymerase activity and replication. As the PA mutants exhibit altered replication capacity, differences in IFN/ISG expression would be largely secondary and difficult to disentangle from the direct effects of altered viral replication. We therefore consider this question beyond the scope of the current study and will add it as a future research direction in the Discussion section.

      In the revised manuscript, we will also examine PA lactylation levels under increasing lactate concentrations to assess their relationship with the dose-dependent changes in viral titers. We will revise the text to clarify the interpretive limitations and, where feasible, perform endogenous co-immunoprecipitation experiments to further assess the interactions between PA and ATAT1/SIRT1 under physiological expression conditions.

      We believe these revisions will strengthen the mechanistic evidence and help address the core concerns raised by both reviewers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a detailed ultrastructural analysis of the larval pharyngeal sensory organs, including the dorsal pharyngeal sensilla, dorsal pharyngeal organ, ventral pharyngeal sensilla, and posterior pharyngeal sensilla. Using electron microscopy and 3D reconstruction, Richter et al., present a comprehensive mapping and classification of pharyngeal sensory structures, defining the morphological type of pharyngeal sensilla based on ultrastructure and generating a neuron-to-sensillum map. These findings significantly advance our understanding of internal larval sensory systems and establish a robust framework for future functional studies in coordination with external sensory systems.

      Strengths:

      The application of high-resolution electron microscopy and 3D imaging analysis successfully overcomes technical challenges associated with visualizing deep internal structures. This enables an unprecedented level of anatomical detail of the larval pharyngeal sensory system. Thus, the study complements and completes existing maps of larval sensory circuits, contributing a comprehensive neuroanatomical characterization of larval sensory input pathways. These insights will inform future studies on larval behavior, sensory processing, and may also have applied relevance for insect control strategies.

      Weaknesses:

      While the manuscript is concise, clearly written, and methodologically rigorous, it primarily addresses a specialized readership with expertise in insect neuroanatomy.

      We thank the reviewer for the positive assessment of our study and for the helpful suggestions. In response, we have clarified the visual presentation of the pharyngeal sense organs in Figure 1, expanded the discussion of adult pharyngeal sensory systems, briefly broadened the comparison to other insect species, checked and corrected the scale bars, and added further methodological detail where appropriate.

      Reviewer #2 (Public review):

      Summary:

      This manuscript documents the structure of the pharyngeal nervous system of the Drosophila larva. The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools, they have achieved their goal. The paper is written clearly and illustrated beautifully with 3D models and annotated sections. The data will significantly enrich the field of Drosophila neurobiology.

      Strengths:

      Given the dataset, the findings presented are solid and will be an important work of reference for the future.

      Weaknesses:

      Previous work, including EM, on the pharyngeal sensory organ is not sufficiently referenced and used for comparison with the data presented in this study.

      We are grateful for the reviewer’s thoughtful comments and for the suggestion to strengthen the historical and comparative context of the work. We have revised the introduction to better acknowledge and discuss the relevant previous EM-based literature on adult and larval internal gustatory sensilla, clarified the organization of the shared pore structure in T1–T3, highlighted the DPO multidendritic neurons more explicitly, and added a comparison that emphasizes the added value of the complete serial EM dataset.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) For improved clarity, highlight the pharyngeal sense organs in Figure 1B. Consider using the color schemes to differentiate between peripheral and internal sensory organs.

      We thank the reviewer for this helpful suggestion. We have revised Figure 1 to more clearly separate the pharyngeal sense organs from the external sense organs in the head region. This revision improves visual clarity and accessibility for readers.

      (2) In reference to lines 80-84, expand the discussion to address how future studies could explore the conserved morphological and functional characterization of the adult pharyngeal sensory system.

      We appreciate this suggestion and have expanded the discussion accordingly. We now briefly address how future work could compare the larval and adult pharyngeal sensory systems to examine conserved morphological and functional features.

      (3) To broaden the manuscript's appeal and emphasize its relevance beyond Drosophila, briefly discuss similarities, differences, or conserved roles of pharyngeal sensory systems in other insect species.

      Thank you for this valuable recommendation. We have added a paragraph placing the Drosophila pharyngeal sensory system in a broader insect context, including similarities, differences and potential conservation across species.

      (4) Recheck the scale bars in all figures, including the supplemental material.

      We thank the reviewer for pointing this out. We carefully rechecked all scale bars across the main and supplemental figures and corrected the missing ones.

      (5) Consider including additional details on image processing or provide appropriate citations for further reading.

      We appreciate this suggestion. We have expanded the methods section to include additional information on technical details and provide the relevant reference for further reading.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 57ff: The previous literature describes internal gustatory sensilla in considerable detail.

      (a) Adult: These sensilla form three complexes, the labral sensory organ, and the ventral and dorsal cibarial sensory organ (Nayak & Singh, 1983, 1985; Singh, 1997; Stocker & Schorderet, 1981; Kendroud et al., 2017). The work by Nayak and Sing includes TEM and presents detailed EM-based schematics. This should be referenced and discussed.

      (b) Larva: Gendre et al. 2004, describes the internal gustatory organs and relates them to their adult counterparts:

      - Dorsal pharyngeal sense organ (DPS) and dorsal pharyngeal organ DPO) are the forerunners of adult labral and ventral cibarial sensory organs

      - Posterior pharyngeal sensory organ (PPS) is the forerunner of the adult dorsal cibarial sensory organ

      - Ventral pharyngeal sensory organ (VPS), derived from the labial segment, undergoes apoptosis during metamorphosis

      This work, connecting larva and adult (and containing detailed diagrams comparing adult and larval pharyngeal sensilla) should be presented in the introduction.

      We thank the reviewer for this important comment. We have revised the introduction to better cite and discuss previous EM-based studies of internal gustatory sensilla in both adult and larval stages, and we now place our findings more explicitly in the context of this prior work.

      (2) Line 180: the relationship between the ending of T1-T3 in one shared pore, and the individually wrapped sensilla should be explained; maybe a simple diagram would help. I did not understand how it works. Normally, in a gustatory sensillum, you have one or more sensory neurons, surrounded by thecogen, trichogen, and tormogen cells. The trichogen generates the shaft with the pore at its tip. Now here, in T1-T3, you have three sets of thecogen/trichogen/tormogen. Do all three trichogen cells somehow participate in the shaft with the common pore? Or only a single one, and the other two generate no shaft? It is possible this cannot be resolved, but the authors should address the problem and suggest a possible scenario.

      We appreciate the reviewer’s concern and agree that this point required clarification. We have revised the relevant text to better explain the organization of T1-T3 and their shared pore and the organization of the support cells.

      (3) Line 205: the DPO multidendritic neurons with dendrites into the hemolymph should be shown; in Figure S4G, I could see only cell bodies. These MD neurons in the gustatory system are, I believe, a true novelty and should be emphasized more if the material allows (text figure!)

      Thank you for highlighting this point. We have revised the results and supplementary material to show these neurons more clearly and to emphasize their novelty and potential relevance to the pharyngeal sensory system.

      (4) A somewhat detailed comparison between the ultrastructure of the DPS as extracted from the serial EM stack of this study, and the conclusions of Nayak and Singh 1983 as depicted in their diagram Figure 7a would be productive. The idea being: what additional details can (only) a complete EM stack provide, compared to conventional EM.

      We appreciate this suggestion. Rather than directly comparing larval and adult structures in detail, we now emphasize what the complete serial EM dataset adds beyond conventional single-section EM, namely a more comprehensive and complete reconstruction of the sensory organs and associated cell types (multidendritic neurons, papilla sensilla, and chordotonal organs that were not described before, organization of support cells)

      (5) To round off the work and connect it to the previously published analysis of gustatory terminal arborizations and connectivity in the brain (Miroschnikow et al.,2018), it would be helpful to add an analysis of the distribution of axons from the different sensilla in the nerves. Miroschnikow analyzes the central terminations of the same sense for which the peripheral structure is described here, only that in their L1 connectome, the periphery was cut off. Do the findings of the current study match their predictions, as to the number of sensory neurons, etc? It should be possible to follow, even at the lower resolution of the dataset presented here, to follow axons of sensory neurons through the nerves to the neuropil entry, and thereby make the connection. I consider this to be of great importance for the field, for authors who want to use the data of this study, and the Miroschnikow et al analysis, for their own studies.

      We thank the reviewer for this thoughtful and constructive suggestion. We fully agree that linking the peripheral sensory anatomy described in this study to the central projections analyzed by Miroschnikow et al. would be highly valuable and of broad interest. However, a systematic analysis of axon distributions from the different sensilla through the nerves to their neuropil entry points is beyond the scope of the present work. Owing especially to the dataset’s resolution and inherent limitations, tracing the connections from sensory organs through the nerves to their projections in the brain is technically highly challenging and extremely time-consuming, since much of the process would need to be performed manually. We therefore do not include a detailed comparison with the predictions from Miroschnikow et al. in this manuscript. Nevertheless, we appreciate that such an analysis would be an important next step for the field and a useful resource for future studies.

      We are grateful for the reviewers’ thoughtful feedback, which has helped us improve the manuscript substantially. We hope that the revised version addresses the concerns raised and better conveys the significance of our work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study uses a tripartite transdiagnostic computational framework to distinguish depression-specific, anxiety-specific, and shared psychopathology dimensions, in their relationships to mood variability and mood reactivity to reward prediction errors across multiple large non-clinical cohorts and a clinical sample. The evidence is convincing overall because the study combines large samples, a well-characterized gambling task and in-depth computational and psychometric analyses, and it replicates the depression-specific association with blunted reward prediction error-sensitivity in a clinical sample. However, the anxiety-specific effects are less consistently supported across individual datasets, may be underpowered in the clinical cohort because of comorbidity, and some aspects of the factor-analytic, risk-attitude, and mediation analyses would benefit from clearer explanation. These findings advance a mechanistic account of how distinct symptom dimensions differentially shape reward-based mood updating and variability, providing a principled framework for future transdiagnostic modeling.

      We thank the editors and reviewers for this important assessment.

      Regarding inconsistent results for anxiety-related effects in healthy datasets. Although the anxiety-specific factor showed associations in the expected direction across healthy datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity in non-clinical participants (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants. We have clarified this point below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants.”

      We also admit that the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      Regarding be underpowered sample size in the clinical cohort. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity. We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β_RPE; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      We have also revised the manuscript accordingly to make the factor-analytic, risk-attitude, and mediation analyses clear.

      Reviewer #1 (Public review):

      Summary:

      This is a very interesting paper. The research question is intriguing, allowing the authors to address commonly observed comorbidities between depression and anxiety and their dissociable and opposite relationship to mood fluctuations and sensitivity to reward prediction errors. The computational analyses are very in-depth, including many state-of-the-art checks and validations. Another strength is the inclusion of several large or very large samples, including a patient sample in addition to the general population sample.

      I have the following questions:

      (1) Factor analysis

      I found the hierarchical organization of the factors interesting. While this is a very common procedure in, for example, the field of intelligence (producing sub-scores and a general g factor), it is not yet very commonly used in the field of computational psychiatry (though it has been validated before for anxiety/depression, so it is used here with good reason). I was also impressed by the methodological depth. In particular, it was of note how thoroughly done it was (for example, repeating the EFA on the second half of the data set). I have one question though: is the sample size too small for the exploratory analyses, given the number of items? Given the stability across the half-split, I imagine it is not. Perhaps the authors could spell out how many items, what would be the recommended standard for a subject-to-item ratio, and comment on this. A very technical point, the authors should specify how they extracted the factor scores from the other data sets (is it using the Thurstone or Bartlett method)? From experience (though not doing a hierarchical factor analysis), Bartlett can be somewhat better compared to the default (Thurstone) - better as in the resulting factors more closely recapitulating the factor correlations in the original sample (and independence of responses of other participants in a sample for computing a person's factor score). Could you also comment on similarities or divergences in this hierarchical factor analysis approach from another one recently used transdiagnostically in Wise et al. (2026, Translational Psychiatry)?

      We thank the Reviewer for the positive evaluation of our hierarchical factor-analytic approach and for recognizing the methodological depth of our analyses. We are particularly grateful for the Reviewer’s constructive suggestions regarding the participant-to-item ratio, factor score extraction, and the relation between our approach and recent transdiagnostic hierarchical factor-analytic work.

      First, hierarchical organization. As the Reviewer noted, hierarchical and bifactor representations have a well-established tradition in intelligence research, where they are used to model the g factor alongside domain-specific abilities (e.g., Reise, 2012; Rodriguez et al., 2016). Crucially, the use of such hierarchical structures in the present study was motivated primarily by theory and evidence from anxiety and depression research, rather than by analogy to intelligence research alone. This tradition can be traced back to the tripartite model of anxiety and depression (Clark & Watson, 1991), which distinguished a broad shared component of general distress or negative affect from more specific anxiety- and depression-related components. Subsequent psychometric work has further supported bifactor and hierarchical representations of anxiety and depression symptoms, including models that separate a general internalizing/distress factor from symptom-specific dimensions (e.g., Simms et al., 2008). More recently, similar hierarchical symptom structures have also been adopted in computational psychiatry to relate shared and specific affective symptom dimensions to task-derived computational parameters (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Thus, the bifactor structure used here provides a theoretically motivated way to capture both the variance shared by anxiety and depression and the symptom-specific variance relevant to our computational analyses. We have clarified this point in the revised manuscript as follows:

      Pages 3-4:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor. Although bifactor and hierarchical models have long been used in psychometrics, e.g., intelligence research[35,36], their application to anxiety and depression is grounded in the tripartite model and subsequent psychometric work distinguishing general internalizing/distress from symptom-specific dimensions. This framework has recently been extended to computational psychiatry, where shared and specific affective symptom dimensions have been linked to task-derived computational parameters. For example, Gagne et al. (2022) used bifactor analysis to show that depression was associated with weaker prior beliefs, whereas anxiety was associated with a stronger negative bias in belief updatin31.”

      Second, the participant-to-item ratio. We agree that the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for exploratory factor analysis, including the often-cited recommendation of five participants per item (Costello & Osborne, 2005). However, as the Reviewer noted, the split-half analysis showed a stable factor structure, and the independent CFA in the other split-half sample further supported the robustness of the solution. Importantly, participant-to-item ratios are only one criterion for evaluating factor recovery. De Winter, Dodou, and Wieringa (2009) demonstrated that reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items. These conditions were largely met in our data. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “In addition, the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for EFA, including the often-cited recommendation of five participants per item[8]. However, the split-half EFA yielded a stable factor structure, and the independent CFA further supported the robustness of this solution. Moreover, reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items9. These conditions were largely met in our data.”

      Next, factor score extraction. We followed prior work using bifactor modeling in computational psychiatry (Gagne et al., 2020) and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it produces orthogonal scores and therefore avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. In addition, as suggested by the Reviewer, we extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of estimating the unique associations of shared and symptom-specific components with task-derived parameters, because it separates the general distress/internalizing factor from the statistically separable residual anxiety- and depression-specific components. For this reason, we retained the Anderson–Rubin factor scores in the main analyses. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “For factor score extraction, we followed prior work using bifactor modeling in computational psychiatry[10] and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. As a robustness check, we also extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of isolating the unique contributions of shared and symptom-specific variance, because it separates the general distress/internalizing factor from residual anxiety- and depression-specific components. Thus, we retained the Anderson–Rubin factor scores in the main analyses.”

      Finally, the similarities and differences between our hierarchical factor-analytic approach and the recent transdiagnostic hierarchical factor-analytic approach of Wise et al. (2026). Both approaches fit EFA models with different numbers of factors and use cross-level correlations to characterize hierarchical symptom structure. However, Wise et al. (2026) applied this framework to a broader symptom battery covering transdiagnostic and neurodevelopmental dimensions, identifying a hierarchy that included a general psychopathology factor and more specific dimensions such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal. In contrast, our study focused more narrowly on anxiety and depression dimensions, with the goal of deriving symptom factors that could be linked to task-derived computational parameters. Accordingly, whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance remains unknown. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026), to better determine whether the links among symptom dimensions, RPE-related mood sensitivity, and mood variability are disorder-specific or transdiagnostic. We have discussed this point in the revised manuscript as follows: Page 19:

      “Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      (2) Linking factors to task parameters

      As I understand it, the authors relate the orthogonalized depression/anxiety to task parameters (sensitivity to RPEs on mood and mood variations) using correlations. In order to have a better understanding of how this relates to other commonly used approaches, I would pose two questions:

      (i) What are the correlations when the full (non-orthogonalized) factor scores for depression and anxiety are used? Are the signs the same?

      (ii) What are the results when, instead of the independent correlations, the authors perform b_RPE ~ anxiety + depression (again using the non-orthogonalized factors)? I'm assuming all of these analyses should give the same results if the authors' hypothesis of opposing effects of anxiety and depression holds true.

      We thank the Reviewer for these helpful comments. Our original analyses used orthogonalized depression- and anxiety-specific factor scores because this approach is aligned with our analytic aim of separating shared and symptom-specific variance within the tripartite/bifactor framework of anxiety and depression (Clark & Watson, 1991), and has been used in prior work (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Orthogonalization allows us to statistically separate the shared distress component from the symptom-specific components of anxiety and depression, which was central to our hypothesis regarding their opposing associations with task-derived parameters. As expected, the orthogonalized anxiety- and depression-specific factor scores were uncorrelated in the combined dataset (r < 0.001, p = 1.000; n = 1,026). By contrast, the full non-orthogonalized depression and anxiety scores retained substantial shared variance and were highly correlated (r = 0.638, p < 0.001; n = 1,026), making their separate associations less straightforward to interpret.

      Nevertheless, we agree that analyses using the full non-orthogonalized depression and anxiety scores provide an important comparison with more commonly used non-orthogonal symptom-score approaches. We therefore conducted the analyses suggested by the Reviewer. When the full depression and anxiety scores were entered separately into linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas the simultaneous regression model reveals their opposing symptom-specific associations.

      Together, these supplementary analyses support our original interpretation that RPE-related mood sensitivity is associated with the separable anxiety- and depression-specific components in opposite directions. We have revised the manuscript as follows:

      Supplementary Pages 3-4:

      “We further analyzed non-orthogonalized full depression and anxiety scores to assess the robustness of our results. When full depression and anxiety scores were entered in separate linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas simultaneous regression reveals their opposing symptom-specific associations. These results support our interpretation that RPE-related mood sensitivity is linked to the separable anxiety- and depression-specific components.”

      Minor comments:

      (1) The authors should write down when the data were collected for each study. This is because AI capabilities have massively increased since ~2020 in quite specific steps (with the public release of new AI models), meaning that AI is likely to have been able to do tasks and questionnaires without detection if data were collected recently.

      We thank the Reviewer for this important comment. We have now added the data collection periods for each dataset in Table 1. The laboratory and clinical dataset were collected in a controlled laboratory setting rather than through online testing. As shown in Table 1, all online experiments were conducted before November 2022, prior to the public release of ChatGPT and its broad entry into public awareness. Therefore, our data were unlikely to have been substantially affected by AI-assisted responding. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (2) The authors should include a statement in the methods section that checks for AI were done. If none yet, could you do any? Recent papers (Westwood, PNAS 2025; van der Stigchel PNAS, 2026) point to the risk since at least the release of o4-mini (used in the cited paper to create very human-like behaviour).

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (3) It would have been good to collect questionnaires of other, thought to be unrelated psychiatric traits, like compulsivity or schizophrenia symptoms, to check the specificity of the results, also under the assumption that higher scores on either of these skewed questionnaires can pick up individual differences in 'bad questionnaire completion'. The authors should comment on the absence of other questionnaires in the discussion in the limitations section.

      We thank the Reviewer for this helpful comment. We agree that the absence of broader psychiatric trait measures limits our ability to evaluate the specificity of the observed associations. Our symptom assessment focused specifically on anxiety and depression because the study was motivated by hypotheses about their potentially opposing links with RPE sensitivity and mood variability. Although previous research has shown intact mood sensitivity to RPEs in individuals with suicidal thoughts and behaviors (Wang et al., 2026), we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-style variance.

      Regarding the concern that higher scores on symptom questionnaires with skewed score distributions may partly capture individual differences in poor-quality questionnaire responding, we note that we implemented strict data-quality procedures for both questionnaire and task data. Four attention-check items were embedded throughout the questionnaire battery, requiring participants to select a prespecified response, for example, “Please select the second option for this item.” Similarly, four attention-check trials were embedded throughout the gambling task. For example, participants were asked to choose between a certain gain of 20 points and a gamble with possible outcomes of 35 and 55 points, for which the dominant response was to choose the gamble option. Participants who failed any of these attention checks were excluded. In addition, our behavioral and mood data reproduced key patterns reported in previous studies (Rutledge et al., 2014 & 2015) using momentary mood ratings during gambling tasks, including higher mood following gains than following losses and systematic mood drift over time (all ps < 0.001). These procedures and validation checks reduce the likelihood that the present findings were driven by poor questionnaire or task completion. We have clarified this point in the revised manuscript as follows:

      Page 19:

      “Second, our symptom assessment focused specifically on anxiety and depression. This choice was motivated by our primary hypotheses, but it limits our ability to evaluate the specificity of the observed associations. Recent work has shown that individuals with suicidal thoughts and behaviors exhibit reduced mood sensitivity to certain rewards (CR), but not to RPEs[49], suggesting that the current RPE-related effects are not driven by suicide-related processes. However, because we did not assess other psychiatric dimensions, such as compulsivity or schizophrenia-spectrum symptoms, we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      Page 20:

      “Participants were excluded if 1) they failed any of the attentional checks (4 items); 2) they made the same choices for all items; 3) they responded with extreme inconsistency in two similar questionnaires (difference in z-scores out of ±2).”

      Page 21:

      “There were four items for attentional checks, which required the participants to make a specific choice and were embedded in the entire measurements, e.g., ‘please select the second option for this item’.”

      Page 22:

      “We also set 4 trials embedded in the entire task for attentional checks. For example, participants were asked to make a choice between a certain gain 20 and a gamble 35/55, where the correct response for this trial was the gamble choice.”

      Page 5:

      “Choice data (e.g., gambling rates) and mood data (e.g., initial mood, mean mood, and mood variation) showed patterns similar to those reported in previous studies measuring momentary mood during gambling tasks (Figure S2 & S3)[10,45]. We also replicated established effects on momentary mood: mood was higher following gains than following losses, and mood drifted over time (all ps < 0.001; Figure S4).”

      (4) The authors could include a more explicit sentence in the abstract stating that the anxiety result did not hold up in the clinical population.

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #2 (Public review):

      Summary:

      Despite their common co-occurrence, depression and anxiety are known to alter mood fluctuations in opposite ways. Here, the authors aimed at distinguishing depression-specific from anxiety-specific from psychopathology-general effects of reward processing on mood fluctuations, focusing on reward prediction errors (RPEs), which are known to be linked to mood fluctuations. This mechanistic study aims at uncovering the process through which these psychopathologies are associated with mood modulations. The authors were able to appropriately test their hypothesis and obtained results corroborating their conclusions.

      This work provides a convincing demonstration of the relevance of computational psychiatry (Huys et al, 2016) and the use of decision neuroscience to shed light on the interplay of anxiety, depression, and mood.

      Strengths:

      The authors used a tripartite model to distinguish depression vs anxiety, as well as a computational model distinguishing reward expectation (EV in the model) from outcome processing through RPE, which are two sequential cognitive processes.

      The manuscript adequately addresses the concerns one would have regarding risk-attitudes and regarding referring to trending statistical results.

      Weaknesses:

      The sample size of the clinical sample (N=116) may not be sufficient to detect anxiety-specific effects due to the high rate of comorbid anxious depression. It would be beneficial to include the number of MDD vs GAD vs anxious depression diagnoses in the clinical population, as this would likely shine light on the power limitations.

      We thank the Reviewer for this helpful comment. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (see Table S8 for diagnosis, illness duration, and medication status). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity.

      We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β<sub>RPE</sub>; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #3 (Public review):

      Summary:

      In this submission, Wang and colleagues jointly examine the association between depression and anxiety symptoms and individuals' affective reactivity to reward prediction errors in Ruttledge et al.'s gambling paradigm. Taking a bifactor approach to anxiety and depression in several non-clinical (and one clinical sample), the authors find that anxiety-specific symptoms relate to over-reactivity of mood to reward prediction errors (RPEs) as well as heightened mood variability, while depression-specific symptoms relate to blunted mood sensitivity to RPEs. These depression- but not anxiety-specific relationships replicated in patient samples.

      Strengths:

      I was impressed that the data-driven, transdiagnostic approach employed by the authors uncovered specific relationships between anxiety and depression-specific factors and RPE reactivity in a well characterized task and computational model, especially in a non-clinical sample. This sheds new light on how these affective processes may be perturbed-and importantly, in different ways-by anxiety and depression symptoms. Likewise, the replication of the depression-specific finding (RPE hypo-reactivity) in a clinical sample was nice to see.

      Weaknesses:

      (1) While the anxiety- and depression-specific factors had differential effects on mood variability (Figure 2A-D) and RPE reactivity (Figure 2E-G) in all samples, such that the correlations between the two factors and these mood parameters were significantly different, the anxiety factor was not consistently (significantly) associated with either mood-related parameter across samples. However, the authors resolve anxiety-specific predictive effects when they collapse across datasets. While it is intuitive that achieving a larger effective sample size would afford the power necessary to detect such individual differences, this struck me as a major caveat for this set of results.

      We thank the Reviewer for this important comment. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.

      However, we have clarified in the revised manuscript that the anxiety-related effects were less robust than the depression-related effects and require further replication in larger samples.

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (2) The authors observe associations between the 'common factor' of depression and anxiety and risk-attitude tendencies, presumably the alpha (exponent) parameter in a prospect theory-type subjective value model. But where is this analysis explained? (i.e. how was this model formulated and how were risk attitude parameters estimated?) And what is the interpretation of this finding - is there precedent for looking at risk attitudes in this task? And why would these predictive effects only be observed in relation to the common, but not unique, factors of anxiety and depression?

      We apologize for the unclear statement. We have added a description of the computational modeling of choice behavior. Please see our revisions below:

      Page 13:

      “Choice parameters were estimated using an established approach–avoidance prospect theory model[10,45,49], which included loss aversion, domain-specific risk attitude parameters in the gain and loss domains, and value-independent Pavlovian approach and avoidance parameters (see Supplementary Note 7 for details of the computational choice models). In this model, risk attitude was quantified by the exponent parameter α in a prospect-theory-inspired subjective value function. Lower α values reflect greater risk aversion, whereas values closer to or above 1 reflect more linear or risk-seeking valuation.”

      Supplementary Pages 8-9:

      Note 7: Computational model of gambling choice

      To quantify how different events impacted participants’ momentary moods during the gambling In line with previous studies[14,15], our choice model space included expected value model (cM1), prospect theory model (cM2)[16], and approach-avoidance prospect theory model (cM3)[14]. For cM2 (Equations 6-9), there were 3 parameters, including risk aversion (α, range: [0.3, 1.3]), loss aversion (λ: [0.5, 5]), and inverse temperature (μ: [0, 10]).

      Where V<sub>gain</sub> and V<sub>loss</sub> are the objective gain and loss from a gamble, respectively. Please note thatV<sub>gain</sub> is 0 in loss trials and V<sub>loss</sub> is 0 in gain trials. V<sub>certain</sub> is the objective value for the certain option. U<sub>gamble</sub> and U<sub>certain</sub> denote the subjective utilities of the gamble and the certain option, respectively. Choice probability for gamble (P<sub>gamble</sub>) is determined by the softmax rule. Building on cM2, cM3 decomposes the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components (Equations 6-8 & 10-12). That is, choice probability for P<sub>gamble</sub> in cM3 is jointly determined by the softmax rule and approach/avoidance parameters (β<sub>gain</sub>: [-1, 1], β<sub>loss</sub>: [-1, 1]). Approach/avoidance parameters are not applied in mixed trials. Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking[17,18]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.

      Model comparison using BIC revealed that the winning model for each dataset was the approach-avoidance prospect theory model (cM3; mean R<sup>2</sup> = 0.51 for the laboratory dataset, 0.49 for the online dataset1, 0.54 for the online dataset 2, and 0.40 for the clinical dataset; Table S9).

      Please also see our interpretation of this finding below:

      Page 18:

      “With respect to decision-making, prior literature using risky decision-making tasks without feedback has linked pathological anxiety to greater risk aversion[58]. In line with this, our results from a risky decision-making task with feedback suggest that the common factor, rather than anxiety-specific variance per se, is more consistently associated with risk aversion. This suggests that heightened gain-domain risk aversion may be a transdiagnostic feature of internalizing psychopathology, rather than being uniquely attributable to anxiety.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you very much for giving me the opportunity to review this very interesting paper.

      Recommendations:

      (1) Add more specific ethics information than "study was approved by ethics committee of Beijing normal university".

      We thank the Reviewer for this important comment. We have added approval number. Please see our revision below:

      Page 20:

      “The study was approved by the Ethics Committee of Beijing Normal University (approve number: ICBIR_A_0016_028). Written or electronic informed consent was obtained from all participants before participation.”

      (2) Add information on how participants were recruited. I think the websites listed only hosted the experiment/questionnaires?

      We thank the Reviewer for pointing this out. We have revised the relevant text as follows:

      Page 20:

      “A total of 2634 participants via online platforms (questionnaires from https://www.wjx.cn and tasks from https://www.naodao.com) took part in five experiments, including a psychometric experiment, a laboratory experiment, two online replication experiments. Participants were recruited through participant pools and study advertisement. For online experiments, interested participants accessed the study through an online link and completed the questionnaires and task remotely. For the laboratory experiment, participants completed the study in a controlled laboratory setting.”

      (3) Typo in Figure 1A, grey panel - psychometric.

      We apologize for the typo. We have corrected typographical errors throughout the manuscript.

      (4) In the Discussion, there is a section on r-to-z transformations, and I was not quite sure what in the Results this links to.

      We thank the Reviewer for pointing out the unclear statement. Please see our revision below:

      Page 18:

      “First, although anxiety- and depression-related associations differed consistently, the anxiety-specific associations themselves were less robust across datasets.”

      **Reviewer #2 (Recommendations for the authors):&&

      The Results sections 2 (depression) and 3 (anxiety) could be improved by reducing the back and forth between factors throughout the results. It may be useful to split them into 3 sections: depression only, anxiety only, and depression vs anxiety.

      We thank the Reviewer for this helpful suggestion. As suggested, we have reorganized this part into three sections: depression, anxiety, and depression versus anxiety. Please see our revision below:

      Page 12:

      “Differential associations of depression and anxiety with mood fluctuations. To directly test whether depression- and anxiety-specific factors differed in their associations with mood dynamics, we compared the corresponding correlations. These comparisons showed that depression-specific associations were significantly more negative than anxiety-specific associations for both mood variation (laboratory dataset: Z = -1.84, p = 0.033; online dataset 1: Z = -5.36, p < 0.001; online dataset 2: Z = -3.42, p < 0.001) and β_RPE (laboratory dataset: Z = -1.77, p = 0.038; online dataset 1: Z = -3.67, p < 0.001; online dataset 2: Z = -4.00, p < 0.001; Figures 2A–C and 2E–G). These results support distinct associations of depression- and anxiety-specific factors with RPE-related mood dynamics.”

      In the discussion, the authors could have mentioned the brain areas most likely to be involved in these processes, both cognitive and psychopathological, as previous studies (such as Cecchi et al, 2022) have aimed at identifying regions involved in RPE processing while modulating mood in health. A short section on this would be useful to the neuropsychiatric community.

      We thank the Reviewer for this helpful suggestion. We agree that the Discussion would benefit from a more explicit consideration of the neural systems that may support RPE-related mood updating and their relevance to psychopathology. We have revised the Discussion accordingly, as shown below:

      Page 16:

      “Although the present study did not include neuroimaging, the observed computational dissociation may map onto partially distinct neural systems involved in reward learning, mood updating, and affective psychopathology. RPE processing has been consistently linked to striatal–midbrain dopaminergic reward-learning circuits[8,50]. The integration of these reward-learning signals into subjective mood and value-based decision-making may further involve the ventral medial prefrontal cortex and orbitofrontal cortex[44]. In addition, the anterior insula may be particularly relevant for integrating feedback-related signals with affective and interoceptive states[8,44], potentially linking RPE processing to anxiety- and depression-related mood dynamics. Consistent with this view, Cecchi et al. (2022)[51] used intracranial EEG to show that feedback-related neural activity tracks mood fluctuations and risky choice. Future neuroimaging studies should test whether depression-related reductions and anxiety-related increases in RPE-related mood sensitivity are associated with altered interactions among striatal, prefrontal, and insular circuits.”

      Reviewer #3 (Recommendations for the authors):

      (1) The authors need to present a clearer definition of the terms "bifactor analysis" and "tripartite model" in the Introduction. What does tripartite mean in this context? What are the assumptions of such bifactor analyses (e.g. as used in Gagne et al. and the present work) and how, in broad strokes, are they carried out? These are important constructs to clarify for readers outside the computational psychiatry niche.

      We thank the Reviewer for this helpful suggestion. We have revised the Introduction accordingly, as shown below:

      Page 3:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor.”

      (2) Previous examinations of depression and RPE reactivity in this task paradigm, as the authors note (e.g. Rutledge et al., 2017), observed that individuals diagnosed with depression showed an intact association between RPEs and mood. In other words, there was no previously observed relationship between depression and affective reactivity to RPEs in this task context. Here, the authors find that the "unique" depression factor identified by the authors (in a non-clinical sample) is associated with blunted RPE sensitivity - this is worth commenting on specifically.

      We thank the Reviewer for this helpful suggestion. We have discussed this point in the Discussion. Please also see it below:

      Pages 15-16:

      “Our computational model not only replicates the important role of RPEs in mood dynamics but also highlights the divergent mediating roles of RPE-related mood sensitivity in the associations of depression and anxiety with mood fluctuations. The opposite associations of depression and anxiety with mood sensitivity to RPEs complement previous findings of apparently intact RPE-related mood sensitivity in depression[12,25,37]. These findings further underscore the necessity of decomposing shared and specific components of depression and anxiety in studies of mood dynamics, which can enhance our understanding of their distinct associations with emotion processing and cognitive flexibility. This point is consistent with bifactor-based work showing that shared and specific symptom dimensions can have different computational correlates. For example, Gagne et al. (2020) showed that bifactor-derived symptom dimensions differentially relate to maladaptation to environmental volatility[32], complementing previous findings that trait anxiety is associated with inflexible adjustment to volatility[32].”

      (3) There is a note (line 247) about the interpretation of the correlations in Figure 2, which attempts to explain away the inconsistent relationships between the anxiety-specific factor and mood variability as well as RPE reactivity observed in Figure 2. I can't say I understand the authors' point here about "signs of positive correlations", so I would say the authors need to clarify their logic here. More to the point, the authors only resolve anxiety-specific predictive effects when they collapse across these datasets. As discussed above (see 'weaknesses'), this is a serious limitation in my view and needs to be discussed as such in the paper.

      We apologize for the unclear statement. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). We have revised it to make it clear. Please see our revisions below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β_RPE were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect, we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001).”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (4) I expected to see that the authors would also investigate relationships between anxiety/depression related factors and the decay (gamma) parameter in the 'Happiness equation', which is presumably estimated from the data here. While I don't have a strong intuition about directions of (or presence of) predictive relationships here, doesn't it stand to reason that different aspects of psychopathology examined here might map onto how long- (versus short-) lasting the effects of, say, RPEs are, upon mood?

      We thank the Reviewer for this important comment. In the healthy datasets, we fitted a linear mixed-effects model predicting the decay parameter (γ) from the three bifactor scores, with dataset included as a random intercept. None of the factors showed a significant association with γ (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we fitted a linear model predicting γ from the three bifactor scores and again found no significant associations (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959). We have clarified this point in the revised manuscript as follows:

      Supplementary Page 5:

      “In the healthy datasets, we conducted a linear mixed-effect model against decay parameter (gamma) with all three factors, with dataset as a random factor. Results did not show significant effect (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we conducted a linear model against decay parameter (gamma) with all three factors and found no significant effect (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959).”

      (5) The rationale for and interpretation of the mediation model, which presumably aims to explain the relationships between anxiety- and depression-specific factors, RPE reactivity, and mood variability was barely explained by the authors. At present, I'm not sure what the added value of this analysis is. The authors should either remove or explain/motivate the mediation more clearly.

      We thank the Reviewer for this helpful comment. The rationale for the mediation analysis is that mood variability in the task is not only a descriptive behavioral outcome, but may also arise from the degree to which momentary mood is updated by RPEs. Therefore, if depression is associated with reduced RPE-related mood sensitivity and anxiety with increased RPE-related mood sensitivity, these alterations should statistically account for their opposite associations with mood variability. The mediation model directly tested this possibility by examining whether RPE-related mood sensitivity accounted for the association between symptom-specific factors and mood variability. We have clarified this rationale in the revised manuscript as follows:

      Page 10:

      “Given the strong correlation between β<sub>RPE</sub> and mood variation (rs > 0.67, ps < 0.001), we further conducted a mediation analysis to examine whether individual differences in RPE-related mood sensitivity statistically accounted for the association between depression loading and mood variation. This analysis was motivated by the hypothesis that depression-related dampening of mood variability may arise, at least in part, from reduced mood sensitivity to RPEs.”

      (6) This submission would benefit from extensive English language copy editing. There are many passages in the paper (in fact, too many to list here) that suffer from either grammatical errors or clarity issues.

      We apologize for these mistakes. We have corrected the typographical errors throughout the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      The study clearly shows the role of Sox17 in Sertoli cells as being important to SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with the authors' prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      At the same time, the available data do not yet fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

      Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulates the SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

      Reviewer #3 (Public review):

      Summary:

      These studies are based on previously published work that showed that deletion of expression of the Sox17 gene in the testis essentially deleted the formation of the Sertoli valve in the Rete testis. The authors extended this work by constructing a vector that resulted in increased Sox17 expression by Sertoli cells and enhanced formation of the Sertoli valve in both wild type and Sox17 knockout mice. The work provides strong evidence supporting the requirement for Sox17 expression to allow formation of the Sertoli valve.

      Strengths:

      The general approach was to express Sox17 from a Tg mouse that expressed Sox17 from Sertoli cells. This Tg mouse was bred into both the WT and the Sox17 KO mouse. The Sertoli valve was enhanced in both the WT/Tg mouse and KO/Tg mouse, showing that ectopic Sox17 could compensate in the Sox17 Ko and act in a concentration-dependent manner in the WT mouse. The results are strong and support the conclusions from the authors. The results were as expected from the original paper describing the KO of Sox 17. These results strengthen these conclusions and provide ideas for additional conclusions. These studies were technically challenging, and the authors provided a very solid manuscript.

      Weaknesses:

      The authors refer several times to high or low expression, but it all appears to be based on immunohistochemistry, and there is no real quantification using PCR, for example. The process used for cell quantification lacks a rationale for why certain numbers were assigned.

      We sincerely thank the reviewers for their careful evaluation of our manuscript and for their constructive and encouraging comments. We are grateful for the recognition of the significance of the Sertoli valve as an understudied transition region between the rete testis and seminiferous tubules, as well as for the positive assessment of our genetic approach and the evidence that ectopic SOX17 expression can modulate SV formation. We have carefully considered all points raised in the assessment and have revised the manuscript accordingly. The major revisions include:

      (1) Clarification of the scope and limitations of the study (Reviewers #1 and #2):

      In response to the comments that the developmental assembly of the Sertoli valve and the precise consequences of its functional disruption remain incompletely understood, we clarified the scope and limitations of the present study at the end of 7th paragraph in the Discussion. Although our findings support a model in which SOX17 regulates SV formation through paracrine signaling, the downstream effectors and precise molecular mechanisms remain to be identified. We therefore revised the Discussion to avoid overinterpretation of the molecular mechanisms and to emphasize that comprehensive mechanistic analyses, including transcriptomic analyses using the Tg mouse model, represent an important direction for future research. We also added histological analyses of the earliest detectable lesions at 4 weeks of age and low-magnification images of adult Sox17 cKO testes (new Figure S1), revealing selective sloughing of round spermatids despite preserved Sertoli cell architecture and subsequent mosaic spermatogenic defects among individual seminiferous tubules. These observations provide additional insights into the altered luminal microenvironment and suggest that spermatogenic defects may progress in a tubule-by-tubule manner.

      (2) Clarification of quantitative analysis and methodology (Reviewer #3):

      In response to concerns regarding the basis and methodology of cell quantification, we revised the Methods to provide detailed information on tissue preparation, fixation, orientation of the rete testis–Sertoli valve region, and the criteria used for quantitative analysis of SV-associated Sertoli cells (new Figure S4). We clarified that Sertoli cells were counted within the SV region extending approximately 100 μm from the RT boundary, including Sertoli cells protruding into the RT lumen, based on previously established criteria (Aiyama et al., 2015).

      (3) Clarification of the limitations of expression-level assessment (Reviewer #3):

      In response to concerns regarding the quantitative assessment of SOX17 and other SV-associated molecules, we clarified the technical limitations of selectively isolating the very small SV region and obtaining sufficient material for quantitative molecular analyses such as qPCR at the end of 7th paragraph in the Discussion. We therefore clarified that expression of SV-associated molecules in the present study was primarily evaluated using histological and immunohistochemical approaches and added the relevant text to acknowledge these limitations.

      We also made additional revisions to clarify each mouse Tg line, phenotypic descriptions, standardize gene nomenclature, improve methodological descriptions, and refine the relevant Discussion where appropriate.

      We sincerely appreciate the reviewers’ thoughtful and constructive comments. Their feedback has helped us clarify the scope of our conclusions, strengthen the methodological descriptions, and improve the overall presentation of the study. All changes have been incorporated into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (i) Although the current paper is not responsible for interpreting the 2022 findings, both datasets show reduced spermatid production accompanied by multinucleated giant germ-cell syncytia. This phenotype has been attributed to backflow of tubular fluid and consequent microenvironmental perturbation. While this is a reasonable hypothesis, it is not entirely consistent with earlier experimental observations. Complete ligation of the efferent ductules reliably produces giant cells, whereas estrogen-receptor knockout, which also causes massive luminal fluid accumulation, does not. In addition, ligation of the testicular artery itself can induce giant-cell formation. Although this may have already been answered in the papers, can you be sure that a direct or indirect effect on the vasculature can be excluded in the Sox17-cKO model?

      We thank the reviewer for this important comment. In our models, SOX17 expression was manipulated specifically in the Sertoli cell lineage, either by SF1-Cre-mediated Sox17 deletion or by ectopic SOX17 expression under the hAMH-promoter. SOX17-expressing vascular endothelial cells were not targeted in either model, making a direct effect of Sox17 manipulation on the testicular vasculature unlikely. Moreover, the partial rescue of the Sox17 cKO phenotype by hAMH-Sox17 supports the interpretation that the phenotype primarily results from altered SOX17 function in Sertoli cells and RT epithelia.

      However, indirect effects on the vascular or interstitial environment by aberrant luminal flow cannot be completely excluded, particularly with the substantial accumulation of sloughed round spermatids (giant cells) within the rete testis. Addressing the potential for an initial luminal flow defect, we newly added histological images of 4-week-old testes (Figure S1), where selective post-meiotic germ cell sloughing occurs despite preserved Sertoli cell process architecture, suggesting an altered adluminal microenvironment that impairs Sertoli–spermatid adhesion. Furthermore, low-magnification images of adult mature Sox17 cKO testes (Figure S1B) display a mosaic pattern of spermatogenic defects across individual tubules. While 3D reconstruction was not conducted, this structural pattern supports the view that spermatogenic failure progresses on a tubule-by-tubule basis, potentially linked to the structural integrity of individual Sertoli valves.

      (ii) A related and important unresolved issue is the total number of Sertoli cells per testis in cKO males. The number of Sertoli cells per tubule cross-section is reported to be equivalent to controls; however, the substantial reduction in testis weight implies a corresponding reduction in tubule length. Under these conditions, maintenance of a normal per-cross-section count would still be compatible with an overall decrease in total Sertoli-cell number. Although it is generally accepted that murine Sertoli cells exit the cell cycle around postnatal day 15, continued growth of the testis may still occur in the Sertoli valve region, where Sertoli cells retain proliferative capacity. Your discussion of possible heterogeneity in the embryonic origin of Sertoli cells near the rete testis is therefore particularly intriguing and commendable. Should this hypothesis be substantiated, it would raise the possibility that Sertoli cells derived from the valve region, especially those that migrate into the seminiferous tubules, are intrinsically less competent to support full spermatogenesis than those of classic gonadal-ridge origin.

      To help readers appreciate the overall severity and topographic distribution of the spermatogenic defect (particularly in tubule segments distant from the rete), inclusion of a low-magnification photomicrograph of a well-fixed (Bouin's) testicular cross-section would be very useful.

      We thank the reviewer for this important comment. We agree that the maintenance of Sertoli cell numbers per seminiferous tubule cross-section does not necessarily indicate preservation of the total Sertoli cell number per testis, particularly given the substantial reduction in testis size and potential reduction in overall seminiferous tubule length. Although total Sertoli cell numbers can theoretically be estimated using stereological approaches, such analyses are technically demanding and beyond the scope of the present study.

      We also appreciate the reviewer’s insightful suggestion regarding potential heterogeneity among Sertoli cell populations. Sertoli cells associated with the Sertoli valve region may have distinct developmental origins or functional properties compared with classical gonadal ridge-derived Sertoli cells, which could potentially influence their capacity to support complete spermatogenesis. Although this hypothesis was not directly tested in this study, we have expanded the Discussion to highlight the developmental and functional heterogeneity of Sertoli cell populations associated with the Sertoli valve as an important topic for future investigation.

      In addition, as requested, we have added a low-magnification image of well-preserved testicular cross-sections in Supplementary Figure S1B to better illustrate the overall severity and topographic distribution of spermatogenic defects throughout the testis.

      Specific Comments:

      (1) Figure 3A and associated fertility/histology data. The results state that epididymal spermatozoa were detected in only 2 of 7 cKO;Tg males at 8 weeks of age, yet Materials and Methods indicate that spermatogenesis was evaluated in only 5 males. a) Were the remaining two males also examined histologically? b) It would be interesting to determine if the severity of pathological changes was the same in regions more distant from the rete testis, or possibly different tubules. See: Nakata H, Wakayama T, Sonomura T, Honma S, Hatta T and Iseki S (2015). "Three-dimensional structure of seminiferous tubules in the adult mouse." J Anat 227(5): 686-694. c) In addition, mating trials were performed with four independent cKO;Tg males, two of which sired offspring. It is unclear whether the testes of these four mating males were included among the five (or seven) animals evaluated for histology, and whether the two fertile males correspond exactly to the two individuals that retained epididymal sperm. Please clarify these relationships explicitly so that readers can correctly interpret the link between histological findings and fertility.

      We thank the reviewer for this important comment. We apologize that the relationship among the groups of animals used for histological analysis, epididymal sperm detection, and fertility assessment was not sufficiently clear in the original manuscript. Because this study focused specifically on the anatomically minute RT–SV region, our sampling strategy had to prioritize the maximal utilization of this limited tissue. In this study, the RT–SV region, the remaining testicular tissue, and the epididymis were processed separately as three tissue blocks for each animal (Figure S4) and were independently evaluated for distinct analysis sets. Briefly, the proximal quarter containing the rete testis and Sertoli valve region was used for SV analysis, whereas the remaining three-quarters of the testis were used for evaluation of spermatogenesis, and the epididymis was analyzed separately for the presence of spermatozoa. Therefore, due to these technical requirements, tissue allocation, and independent analytical evaluation, the numbers of animals used for RT–SV analysis, testicular histology, epididymal sperm detection, and fertility testing were not identical.

      For quantitative histological analyses, we also used virgin males to minimize potential variation associated with mating experience and to allow comparison with age-matched littermate controls. Therefore, these animals were not used for fertility testing. Fertility assessment was performed using an independent cohort of cKO; Tg males that were subjected to long-term mating trials with wild-type females. Thus, fertility outcomes and histological findings were not designed to be directly matched at the individual level.

      In response to the reviewer’s suggestion, we have clarified the selection of experimental animals and the relationship among fertility assessment and histological analyses in the Materials and Methods and added a schematic illustration of the sampling strategy in Figure S4. We also corrected the citation for Nakata H et al., 2015 in the revised manuscript.

      (2) Page 8, line 301 (Sertoli-cell quantification). The description of the counting method-"counted in each ... (~100 μm from the edge of the RT; Fig. 4C)"-is ambiguous.

      (a) Does this mean that cells were counted beginning at the rete boundary and extending radially outward for approximately 100 μm, or is a circumferential sampling area intended? (b) Figure 4C shows a large standard deviation, indicating substantial variability with the current approach. An alternative strategy (for example, counting Sertoli cells within standardized areas or per tubule specifically within the valve region) might reduce variability and improve reproducibility. Regardless of the method ultimately chosen, a more precise, step-by-step description of the quantification protocol is required so that it can be reliably replicated by other laboratories.

      We thank the reviewer for pointing out that the description of the Sertoli cell quantification method was not sufficiently clear. The Sertoli cell quantification was performed using the same criteria as previously described (Aiyama et al., 2015; Uchida et al., 2022), in which SOX9-positive Sertoli cell nuclei within the SV-associated region were counted.

      In the revised manuscript, we have clarified that the SV region was operationally defined as comprising (i) the terminal 100 μm segment of the seminiferous tubule immediately adjacent to the rete testis (RT) and (ii) the protruded SV extending into the RT lumen. Based on the distribution of spermatogonial stem cells, the ~100 μm region extending from the RT boundary along the seminiferous tubule toward the ST side was defined as the SV region (Aiyama et al., 2015). Only sagittal sections showing a continuous RT–SV–ST axis and sectioning the SV approximately through its mid-sagittal plane were included for quantitative analysis.

      Furthermore, to improve reproducibility, we have added a more detailed description of the tissue preparation and quantification procedures in the Materials and Methods and provided a schematic illustration of the quantification strategy in the new Figure S4.

      Reviewer #2 (Recommendations for the authors):

      (1) Phenotypic differences between transgenic lines: the phenotypic differences between the tg26 and tg27 lines are intriguing and warrant further clarification. While tg27 mice exhibit infertility and defective spermatogenesis, tg26 animals remain fertile with SV expansion. Could the authors elaborate on the underlying causes of these differences? In particular, is infertility in tg27 mice due to excessive SOX17 expression impairing Sertoli cell function? A comparison of Sox17 expression levels between tg26 and tg27 lines would be informative. In addition, it would be useful to assess whether acetylated tubulin (Ac-Tub) expression is present in the Sertoli cells of the tg27 mouse testis.

      We thank the reviewer for this highly constructive and insightful comment. We clarified in the revised manuscript that the analysis of the Tg27 mouse was performed using the F0 founder male and added an explanation that only the Tg26 line could be established as its heterogenous SOX17 expression in Sertoli cells did not impair overall fertility. We agree that the phenotypic differences between the Tg26 line and the Tg27 mouse provide important clues regarding the dosage-dependent effects of SOX17 in Sertoli cells. Unfortunately, we were unable to establish a stable, multi-generational transgenic line from this Tg27 founder (F0) male. Consequently, we could not perform detailed molecular or immunohistochemical analyses on this line beyond the initial histological evaluation of the F0 generation presented in Figure 1. For this reason, we cannot provide a quantitative comparison of Sox17 expression levels or evaluate acetylated tubulin (Ac-Tub) expression in Tg27 Sertoli cells.

      To address the reviewer's concern without overstepping the available data, we removed direct quantitative comparisons of Sox17 expression levels between the two lines from the text. Instead, we added a clear description of their contrasting cellular expression patterns - specifically, the mosaic, heterogeneous SOX17 expression in Tg26 Sertoli cells versus the ectopic, uniform SOX17 expression in the infertile #27 F0 male - in the 'Animals' section of Materials and Methods. This mosaic pattern in Tg26 testes suggests the presence of Sertoli cells with low or undetectable SOX17 levels, which may be associated with sustaining overall fertility.

      (2) Mechanism of SOX17 action: although SOX17 is a transcription factor, the author's studies indicate it regulates SV formation via paracrine and/or autocrine signaling. The underlying mechanisms remain unclear. Which downstream factors mediate this effect? The observed upregulation of RSPO1 and WNT4 is suggestive, but more direct evidence would strengthen this conclusion. For example, does SV expansion in tg26 mice depend on the activation of RSPO1/WNT signaling? Additional molecular analyses, such as bulk RNA-seq comparing control and transgenic testes, could help identify pathways regulated by SOX17 and clarify its mode of action.

      We thank the reviewer for this important and insightful suggestion. At present, comprehensive analyses, including scRNA-seq of Sox17 cKO and littermate control testes, have not identified definitive downstream targets of SOX17 (Uchida et al., 2022). As the reviewer rightly points out, the Tg26 mouse model generated in this study represents a valuable tool for investigating SOX17-dependent molecular pathways. To this end, we are currently conducting transcriptomic analyses of Tg26 seminiferous tubules to identify genes altered in SOX17+ Sertoli cells. However, determining whether these candidate genes represent direct transcriptional targets of SOX17 and whether they function specifically in the rete testis-associated region during Sertoli valve formation will require extensive functional and expression studies. Therefore, we feel it would be premature to draw definitive conclusions regarding the underlying molecular mechanisms, including the precise involvement of the RSPO1/WNT signaling pathway, in the present manuscript. Accordingly, rather than overinterpreting the available data, we have revised the Discussion to clarify this limitation (at the end of 7th paragraph in the Discussion). Furthermore, incorporating initial insights from our ongoing Tg26 transcriptomic analyses, we have added a brief discussion, supported by relevant literature, on the possibility that SOX17 may regulate Sertoli valve formation by modulating cell adhesion and extracellular matrix (ECM) organization and altering the responsiveness of SOX17-positive Sertoli cells to morphogenetic signals originating from the rete testis (new 5th paragraph in Discussion).

      (3) Minor comment: Gene nomenclature should be standardized: e.g. line 245, Sox17 and hAMH should be italicized.

      We thank the reviewer for pointing this out. All gene names have been italicized throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      No suggestions except to quantify some of the changes in concentration of agents by PCR rather than eyeball levels with immunocytochemistry. Verify the cell quantification procedure used.

      We thank the reviewer for this comment. The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 sites per mouse testis. As a result, selective isolation of the SV region to collect sufficient material for molecular analyses, such as quantitative PCR, remains technically challenging. We have therefore added this limitation to the Discussion.

      Regarding the cell quantification procedure, we have clarified the methodology in the revised Materials and Methods and added a schematic illustration in Figure S4. Specifically, the Sertoli cell number in the SV region was quantified by counting SOX9-positive Sertoli cell nuclei within a standardized SV-associated region in RT–SV–ST sagittal sections.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In this manuscript authors examined the effect of rif1 knockout on replication timing and transcription in early embryos of zebrafish. Contrary to the expectation, genome-wide replication timing domains did not significantly change upon Rif1 knockout, although the replication timing became less dynamic in the mutant, meaning the entire genomes are replicated toward the mid S. In contrast, transcriptional profiles change by rif1 mutation throughout the embryo stage. These effects were more predominantly observed after gastrulation at the early stages of zebrafish development.

      The results presented in this manuscript provide new information on the effects of rif1 mutation on early zebrafish development, although the underlying mechanism has not been explored. The information is useful for researchers in the field of early development, with specific focus on replication and transcription regulation.

      The genome wide analyses of replication timing has been conducted and analyzed properly. The transcriptional analyses are conducted by RNA-seq and SLAM-seq (determining the nascent mRNA), and the results convincingly show the overall transcriptional patterns at different developmental stages.

      This work shows that Rif1 regulates replication timing and transcription in zebrafish embryos, while the extents of the effects vary during the developmental process. Although the data convincingly illustrate the whole picture of Rif1 KO on replication and transcription during zebrafish development, the mechanistic insight is missing. Especially, how Rif1 may or may not coordinately regulate replication and transcription during the zebrafish development has not been addressed.

      We thank the reviewer for recognizing the value of combining genome-wide replication-timing, RNA-seq, and SLAM-seq analyses across zebrafish development. We agree that the original study did not establish a molecular mechanism linking Rif1-dependent transcriptional and replication-timing effects. To address whether these effects are locally coordinated, we added a gene-centred analysis comparing replication-timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5--figure supplement 2). Differentially expressed genes did not show a clear enrichment in early- or late-replicating regions, either at Dome or at pre-MBT. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes. We also expanded the Discussion to explain the limitations of the current study and the need for future measurements of origin use, fork progression, chromatin state, and cell-type-specific effects. The new discussion of Nakatani et al. (2025) further places our findings in the context of evidence that Rif1-dependent replication-timing changes can be uncoupled from transcriptional changes.

      Reviewer #2 (Public Review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report on the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. The manuscript is well written, though there are aspects of the presentation that could be improved for a broader scientific audience. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The work is descriptive and feels like two relatively unconnected studies, transcription and replication plus a small bit of development, and the difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. There isn't a lot of new insight into the mechanism of how Rif1 affects either replication timing or gene expression. As such, the overall study is an useful set of findings and detailed data for future work, but it doesn't make a big step forward in understanding the role of Rif1 or the biological processes it affects.

      Weaknesses worth addressing include the following:

      (1) Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no explanation for the loss of females from the data gathered here.

      (2) The approach to distinguish nascent zygotically expressed mRNAs from maternal mRNAs is a strength. Are the differentially expressed genes related at all to regions of the genome whose replication timing is most affected? Are any of them related to the sex determination or developmental phenotypes?

      We thank the reviewer for recognizing the rigor of the analyses and the value of studying replication timing in a developing vertebrate. We revised the manuscript extensively to make the experimental logic, zebrafish developmental context, replication-timing analyses, and figure legends more accessible to a broad audience. We also quantified the gastrulation phenotype, showing an approximately one-hour delay in completion of epiboly in maternal-zygotic rif1 mutants rather than a persistent developmental arrest.

      We agree that the mechanism underlying the sex-ratio phenotype remains unresolved. The transcriptomic experiments were performed in whole embryos at stages much earlier than zebrafish sex determination and therefore cannot resolve changes in primordial germ cells or supporting gonadal somatic cells. We have avoided making a mechanistic connection between the early embryonic transcriptional changes and the adult sex-ratio phenotype and identify this as an important area for future study. To address the relationship between transcription and replication timing, we added Figure 5--figure supplement 2. Genes with increased or decreased transcript abundance at Dome were not preferentially associated with early- or late-replicating regions. Together with the distinct developmental timing of the transcriptional and replication-timing phenotypes, this supports the interpretation that Rif1 has separable roles in the two processes rather than a single local mechanism that directly couples them.

      Reviewer #3 (Public Review):

      Using the zebrafish model system, this manuscript assessed the roles of Rif1 protein in replication timing control and transcription during early development, and successfully demonstrated the differential impact of Rif1 protein in replication timing control and transcription. Moreover, the comprehensive assessments of the impacts of mutating Rif1 on animal development (including animal survival and sexual development) were assessed. Although there are works that examined Rif1's implications in replication timing and transcription separately, this work is unique in assessing all these points at once.

      The strength of this manuscript is the genomic analyses of replication timing and transcription being combined in a single model system. Consequently, this manuscript clearly demonstrates the differential impact of Rif1 in these processes during zebrafish development.

      The weakness of this manuscript is, as the authors comment in the Discussion, analyses of replication timing and transcription were performed using bulk embryos. There is a possibility that tissue-specific changes could have been masked. Tissue-specific or single-cell analysis in the future will fill the gap in the knowledge.

      Some of the findings presented in this manuscript are consistent with previous findings using different models such as Drosophila and mice, whereas other findings do not necessarily agree. I hope further studies will reveal more clearly what is common in these systems, and what is different.

      Also, the suggestion that the Rif1 protein may be implicated in a function similar to Fanconi-Anemia genes/proteins is very intriguing.

      Overall, the data presented in this manuscript sufficiently justify the authors' claims. Moreover, this manuscript provides interesting insights into Rif1's function, as well as how development could be controlled.

      We thank the reviewer for highlighting the strength of analyzing replication timing, transcription, and developmental phenotypes in the same vertebrate model. We agree that bulk-embryo measurements may mask tissue- or cell-type-specific effects. We now emphasize this limitation and the need for future tissue-specific or single-cell studies, particularly in the cell populations relevant to sex determination. We also expanded the cross-species context by discussing the recent mouse-embryo study by Nakatani et al. (2025), which supports a conserved role for RIF1 in consolidation of the replication-timing program while also indicating that replication-timing and transcriptional effects can be uncoupled. We agree that defining which Rif1 functions are conserved across zebrafish, mouse, Drosophila, and other systems, including possible relationships to Fanconi-anaemia pathways, will be an important direction for future work.

      Reviewing Editor:

      While the paper was under revision, a relevant paper from the Torres-Padilla lab was published (Nakatani et al., Developmental Cell, 2025). It complements these studies and cites the previous version of this manuscript. I suggest adding a reference in the Discussion to support the conclusions.

      We thank the Reviewing Editor for bringing the recent study by Nakatani et al. to our attention. We have added a standalone paragraph near the end of the Discussion explaining how this work complements our findings, and we have added the complete reference to the bibliography. The new Discussion text reads:

      “A recent study in mouse embryos independently identified RIF1 as a regulator of the developmental consolidation of the RT program. RIF1 depletion produced a less-defined, developmentally immature RT program, while RIF1-dependent RT changes were not correlated with transcriptional changes (Nakatani et al., 2025). Together with our findings in zebrafish, these results support a conserved role for RIF1 in sharpening replication timing during vertebrate development and indicate that its effects on replication timing can be uncoupled from changes in gene expression.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The results presented in this manuscript provide new information on the effects of rif1 mutation on replication and transcription during early zebrafish development, although the underlying mechanism has not been explored. I suggest authors consider conducting the following experiments.

      (1) Does replication timing domains have any role in Rif1-mediated regulation of transcription? It is not clear from the data presented whether transcriptionally affected genes are in the early replicating domains or late replicating domains (that appear after the shield stage). This should be examined.

      We thank the reviewer for this helpful suggestion. To address whether transcriptional effects in rif1 mutants are associated with replication timing, we assigned each gene the nearest smoothed replication timing value and compared replication timing distributions for genes whose transcript levels increased at Dome, decreased at Dome, or were not significantly changed. This analysis is now shown in Figure 5—figure supplement 2. Genes with increased or decreased transcript abundance at Dome did not show a clear enrichment for either early- or late-replicating regions relative to genes with no significant transcript change. This was also true when replication timing was examined at pre-MBT, the stage preceding the major transcriptional changes detected at Dome. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes observed in rif1 mutant embryos. We have revised the Results to describe this analysis and added Figure 5—figure supplement 2.

      (2) It is of interest whether the Rif1-mediated regulation of transcription and replication are mediated by a common mechanism, e.g. through alteration of chromatin structures. Close look at the data in Figure 3D indicates that some genome segments convert replication timing or undergo significant changes of replication timing. It would be informative to know whether these segments (Rif1-regulated replication domains) are associated with the genes whose expression change upon rif1 knockout.

      We thank the reviewer for this insightful suggestion. We agree that an association between Rif1-dependent replication timing changes and Rif1-dependent transcriptional changes would be informative, and we considered this analysis. We attempted to identify Rif1-regulated replication timing domains using the same approach that we previously used to define developmentally regulated timing domains. However, the effect of Rif1 loss differed qualitatively from the developmental timing switches described in our prior work. Rather than producing a limited set of discrete timing-domain transitions, Rif1 loss caused a broad reduction in the dispersion of replication timing values across the genome, consistent with a general flattening of the timing profile. Under these conditions, an unbiased domain-calling approach preferentially identifies genomic regions with the most extreme early or late timing values in wild-type embryos, because these regions show the largest shift toward the mean in rif1 mutants. Thus, the resulting “Rif1-regulated replication domains” largely reflect the strongest wild-type timing domains rather than a discrete set of Rif1-specific regulatory intervals. For this reason, we do not think that assigning differentially expressed genes to such domains would provide a meaningful test of whether Rif1 regulates transcription and replication timing through a common local mechanism. Instead, we have now added a gene-centred analysis comparing replication timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5—figure supplement 2), which directly addresses whether transcriptionally affected genes are associated with early- or late-replicating regions.

      (3) Replication is analyzed only by timing analysis. Authors need to analyze frequency of origin firing and replication fork rate by DNA fiber analyses to see whether they are affected by rif1 knockout at various stages of development.

      We agree that measuring origin firing frequency and replication fork rate would provide valuable additional information about how Rif1 loss affects the replication program. However, performing DNA fibre analyses across multiple zebrafish developmental stages and genotypes would require substantial optimization and experimental expansion beyond the scope of the current revision. The current study was designed to measure genome-wide replication timing and transcript abundance across developmental stages, rather than single-molecule replication dynamics. We therefore have not added DNA fibre experiments. Instead, we have revised the Discussion to acknowledge this limitation and to clarify that replication timing reflects the combined effects of origin usage, fork progression, fork directionality, and fork stability. We added the following text to the Discussion:

      “A further limitation of this study is that we concentrated on replication timing without directly measuring other features of the replication program that contribute to this timing. These features include origin usage, replication fork spacing, fork directionality, fork progression, and fork stability. A more comprehensive understanding of how Rif1 loss affects these parameters will be important for defining the relationship between Rif1-dependent changes in replication timing and transcription.”

      Figure 4B, D and F: I did not see the blue lines which represent preMBT in the panels shown.

      We thank the reviewer for identifying this error. The pre-MBT data were not intended to be shown in Figures 4B, 4D, and 4F. We have corrected the figure legend by removing the reference to the blue pre-MBT line.

      Line 270: Figure 4G should be Figure 6G.

      We thank the reviewer for identifying this error. We have corrected the figure reference from Figure 4G to Figure 6G.

      No description of Figure 6E and 6F in the main text.

      We thank the reviewer for noting this omission. We have added text to the Results describing Figures 6E and 6F. The revised text explains that Dome Up-DEGs are normally upregulated from pre-MBT to Shield stages but show earlier upregulation in rif1 mutant embryos, whereas Dome Down-DEGs normally decrease between Dome and Shield stages but show earlier reduction in mutant embryos.

      Reviewer #2 (Recommendations For The Authors):

      (1) This study is an extension of the lab’s previous work which established the wild-type genome-wide replication timing pattern during zebrafish development. The experimental details and analysis are described in the methods, but the general strategy is sometimes treated very cursorily. A non-expert can only understand parts of it by going back to the Seifert study.

      We thank the reviewer for pointing this out. We agree that the replication-timing strategy should be understandable without requiring readers to consult our previous study. We have revised the manuscript to explain the general logic of the assay more clearly. Specifically, we now state that replication timing was inferred from copy-number differences between S-phase and G1-phase genomic DNA: genomic regions that replicate early in S phase are enriched in S-phase DNA relative to G1 DNA, whereas later-replicating regions are less enriched. We also clarified that pre-MBT, dome, and shield embryos were treated as S-phase samples because most cells are in S phase at these stages, whereas nuclei from bud and 24 hpf embryos were sorted by DNA content to isolate G1 and S-phase fractions. These additions make the experimental design and interpretation of the replication-timing profiles clearer in the main text and Methods.

      Figure 2 is meant to document developmental delay in early embryos, but the differences between the single wt and mutant examples in 2D are poorly described and labeled. Most readers will be unfamiliar with the specifics of zebrafish development. There is also no quantification of this developmental phenotype, and that quantification should be included along with better labeling and description of 2D.

      We thank the reviewer for pointing this out. We agree that the developmental delay shown in Figure 2D required clearer explanation and quantification for readers who are less familiar with zebrafish gastrulation. We have revised the Results to explain that epiboly is the process by which the blastoderm and yolk syncytial layer move toward the vegetal pole to envelop the yolk cell, and that zebrafish gastrulation stages are commonly described by the percentage of yolk coverage. We also added quantification of this phenotype. At 10 hpf, most wild-type embryos had completed epiboly, whereas most rif1 mutant embryos had not: 18 of 24 wild-type embryos, but only 2 of 24 mutant embryos, had reached 100% yolk coverage. By 11 hpf, all wild-type and mutant embryos had completed epiboly. These revisions clarify that rif1 mutant embryos show an approximately 1-hour delay in epiboly completion rather than a persistent arrest in gastrulation.

      (3) The presentation could be greatly improved with additional information about the experimental approach and display. As written, the text and figure legends assume readers are intimately familiar with replication timing experiments, zebrafish development, and differential gene expression analysis. Most of the figure legends are not sufficient to understand the figures themselves, and the necessary information is also not always in the results. An example is Figure 3 which is not well described (other than the PCA plots); the term “lag” which is the x-axis in 3C is not defined.

      We thank the reviewer for this helpful comment. We agree that several aspects of the replication-timing analysis required clearer explanation for readers who are less familiar with replication-timing experiments. We have revised the Results to explain the logic of the replication-timing assay more clearly and have added a more detailed description of the autocorrelation analysis in Figure 3C. Specifically, we now explain that autocorrelation measures how similar replication-timing values are across increasing genomic distances along the same chromosome, providing a quantitative readout of the peak-and-valley structure of the timing profile. We also clarified that increasing autocorrelation across hundreds of kilobases reflects the progressive establishment of broader replication-timing domains during development. In addition, we changed the x-axis label in Figure 3C from “lag” to “Genomic distance (Mb).” Together, these changes should make the experimental approach and display easier to understand without requiring readers to consult our previous replication-timing study.

      Figure 4 is generally poorly described and labelled (4B, D, and F graph legends indicate preMBT in the data, but there are no blue lines on the graphs), and Figures 6 and 7 are quite busy.

      We thank the reviewer for pointing this out. We agree that the Figure 4 legend incorrectly described the data shown in panels B, D, and F. The pre-MBT data were not intended to be plotted in these panels, and we have removed the corresponding reference from the figure legend. We recognize that Figures 6 and 7 contain several analyses, but we have retained the current organization because the panels in each figure address a connected set of questions. Figure 6 summarizes how Rif1 loss affects abundance of developmentally regulated transcripts, whereas Figure 7 extends this analysis by directly measuring nascent transcription using SLAM-seq.

      Reviewer #3 (Recommendations For The Authors):

      I do not think any additional experiments are required to justify the authors’ claims. Well done! However, for readers’ benefit, I propose the following changes or adding more explanations:

      (1) Page 2, line 86: I guess “single copy” means “single copy per haploid”. Better to clarify this point.

      We thank the reviewer for this helpful clarification. The reviewer is correct that “single copy” refers to a single copy per haploid genome. We have revised the text to state that the zebrafish genome has a single copy of the rif1 gene per haploid genome.

      (2) Related to the data presented in Figure 2C, do you have an explanation for why sex determination is affected in the heterozygotes, despite the change in Rif1 expression being subtle (Figure 1C)?

      We thank the reviewer for raising this point. We agree that the reduction in whole-embryo rif1 mRNA levels in heterozygotes appears modest relative to the sex-ratio phenotype. At present, we can only speculate about the basis for this difference. One possibility is that whole-embryo mRNA measurements do not accurately reflect Rif1 abundance in the specific cell populations that influence zebrafish sex determination, such as primordial germ cells or their supporting somatic cells. We have therefore avoided making a strong mechanistic conclusion from the heterozygous phenotype.

      (3) Related to the data presented in Figure 2D, did you observe a delay in heterozygotes?

      We thank the reviewer for this question. We have not quantitatively analyzed epiboly progression in heterozygous embryos. However, we did not observe an obvious developmental delay in heterozygotes during early development. The delay shown in Figure 2D was observed in maternal-zygotic rif1 homozygous mutants.

      (4) Figure 3D: it is not easy to distinguish WT and mutant lines, particularly for the Bud stage. Please consider changing the colour schemes or other aspects. For example, making colour lines thinner may help.

      We thank the reviewer for this helpful suggestion. We agree that the wild-type and mutant profiles in Figure 3D, particularly at the bud stage, were difficult to distinguish in the original version. We have revised Figure 3D by reducing the line width of the colored profiles, which improves the contrast between the wild-type and mutant traces.

      (5) Figure 3E: Could you avoid overlapping of WT and mutant plots?

      We thank the reviewer for this suggestion. We considered separating the wild-type and mutant density plots in Figure 3E, but we have retained the overlaid format because the purpose of this panel is to directly compare the distributions of replication timing values between genotypes at each developmental stage. Overlaying the plots makes the reduced dispersion of timing values in the rif1 mutants easier to visualize relative to the corresponding wild-type distribution.

      (6) Figure 4C and 4E: the point legends (WT and mutant) do not match the points used in the graph.

      We thank the reviewer for noting this potential source of confusion. In Figures 4C and 4E, point shape indicates genotype, with open squares representing wild-type samples and open circles representing rif1 mutant samples. Point color indicates developmental stage. We used separate visual encodings for genotype and stage to avoid a large legend containing every genotype-stage combination. To make this clearer, we have revised the figure legend to state explicitly that point shape denotes genotype and point color denotes developmental stage.

      (7) Figure 4D: Very difficult to recognise 24 hr mutant line. Please improve the way there are shown.

      We thank the reviewer for this helpful suggestion. We agree that the 24 hpf mutant profile in Figure 4D was difficult to distinguish in the original version. We have revised the figure by changing the appearance of the mutant lines to make them more visible while preserving the stage color scheme.

      (8) Related to data presented in Figure 4B. Is it possible to show a statistical evaluation of all (or a reasonably large number of samples from) DARs?

      We thank the reviewer for this suggestion. Figure 4A already provides a genome-wide analysis of the DAR set shown by example in Figure 4B. Specifically, Figure 4A plots the change in replication timing from shield to 24 hpf for all 2,498 putative enhancer-associated DARs in both wild-type and rif1 mutant embryos. The strong correlation between wild-type and mutant values indicates that DAR-associated timing changes are largely preserved in rif1 mutants. Because all DARs used for this analysis are included in the scatterplot, we did not add a separate statistical analysis of selected examples from Figure 4B.

      (9) Page 8, line 220: It is unclear what “all” means. Is it all the available replication timing values genome-wide? Please clarify.

      We thank the reviewer for noting this ambiguity. In this sentence, “all” refers to all genome-wide replication timing values calculated from the genomic windows used in our replication timing analysis. We have revised the text to make this clearer.

      (10) Figures 6C and 6D: Colour labels are too dark and it is almost impossible to read texts inside. Please reconsider the colour scheme.

      We thank the reviewer for pointing this out. We agree that the labels in Figures 6C and 6D were difficult to read because of insufficient contrast. We have changed the text colour inside the colored boxes to white to improve legibility.

      (11) Related to overall transcription studies: Is there any sign that Rif1 mutation affects the transcription of genes involved in sex determination?

      We thank the reviewer for raising this interesting question. We have not specifically analyzed whether genes involved in sex determination are differentially expressed in the early embryonic transcriptome data. Because zebrafish sex determination occurs substantially later than the embryonic stages analyzed here, and likely depends on specific cell populations such as primordial germ cells and supporting gonadal somatic cells, we do not think the current whole-embryo RNA-seq data can directly resolve this question. We therefore avoid drawing a mechanistic connection between the early transcriptional changes and the adult sex-ratio phenotype. Determining whether Rif1 mutation affects transcription in the cell populations that regulate zebrafish sex determination will be an important direction for future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have addressed all the concerns and recommendations by the reviewers, in particular, the requested control experiments using alternative super-resolution microscopy approaches and analysis of the data using Voronoi tessellation in addition to DBSCAN, as requested by reviewer 3. We also provide additional data on tensin3 as suggested by reviewer 2. Finally, to provide a first insight on the role of mechanical forces in the distribution of integrin nanoclusters inside FAs as recommended by reviewer 1, we have performed experiments at different cell seeding times where it is known that FA maturation over time requires mechanical forces.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In recent years, it has become increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant in order to obtain a better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αvβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αvβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αvβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      We thank the reviewer for their feedback and have addressed their recommendations in our updated manuscript and accompanying reply (see recommendations to the authors). In summary, we have now performed experiments at different seeding times as FA maturation requires mechanical forces, and enquired whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs. The results are now shown as new Fig. 2 and discussed in pages 9 and 10. In addition, in order to get a first insight into the biological implications of our findings we performed dual-colour super-resolution experiments of tensin-3 and α<sub>5</sub>β<sub>1</sub> in FAs, as tensin-3 has been implicated in fibronectin fibrillogenesis. The results are now shown in Fig. S8 and we discuss their potential implications in pages 22 and 23 of the revised manuscript (see more details in the reply to the recommendation to the authors).

      Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters, and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation, preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with the help of conformation-specific antibodies, and this enabled us to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins.

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at an early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin, and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible that antibodies containing two binding sites for antigen influence the nanoscale organization (and also activation) of the receptors. Control experiments to study the possible contribution of antibodies to the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with a small ligand and detected using a monovalent binder.

      We understand the concern of the reviewer regarding the use of antibodies for imaging. Nevertheless, we would like to clarify that antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation.

      Nevertheless, and although it is highly unlikely to happen in fixed cells, there could be two potential sources of antibody (Ab) labelling artefacts. As the reviewer noted, a primary Ab containing two binding sites could bind to two adjacent proteins (within ~10 nm from each other), potentially underestimating the stoichiometry of the nanoclusters, i.e., number of receptors or proteins per nanocluster. However, in our manuscript we never attempted to provide an estimation of the nanocluster stoichiometry, as it is highly challenging (and prone to artefacts) to provide quantification of the number of proteins using super-resolution-based single-molecule localisation methods which rely on the stochastic blinking of individual fluorophores.

      A second source for potential artefacts comes from the use of the secondary Ab, which (albeit unlikely) could bind to two different primary Abs. To exclude this potential artefact, we performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. These new data are included now in Fig. S4. As can be observed, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab (for more details, please see the reply to the recommendations for authors section).

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with a certain adapter protein, such as tensin.

      We fully agree with the reviewer and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). We provide more details of our answer in the section of “recommendation to the authors”. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      Reviewer #3 (Public review):

      Summary:

      In their study, the authors reveal using dual-color super-resolution STORM microscopy modality and immunolabeling in fixed adherent cells, that β1 and β3 integrins as well as adaptors (paxillin, talin and vinculin) are all organized in nanoclusters of similar size (50nm) and molecular density (20 copy number) inside FAs but also outside. Using activityspecific immunolabeling of β1 and β3 integrins, they revealed that active integrin subpopulations were both clustered but in distinct exclusive nano-aggregates in agreement with Spiess et al. (2018). Once more, the "active" integrin nanoclusters displayed similar properties in terms of size and molecular density, suggesting that molecular organization in nanoclusters is an intrinsic property of integrins in plasma membrane multimerizing independently of their location (inside or outside FAs), their level of activation, or their connection to the cytoskeleton. Then the authors followed up by analyzing at the mesoscale how these "universal" nanoclustered adhesive units are distributed spatially. Inspecting the surface density of nanoclusters revealed that the density of integrin nanoclusters in FAs was 5x larger, compared to integrin nanoclusters outside adhesions. Interestingly, whereas the density of total integrin nanoclusters was 2-4x larger than adaptor nanoclusters, the density of "active" integrin nanoclusters stoichiometrically matches that of talin and vinculin nanoclusters, and was slightly outnumbered by paxillin nanoclusters. These findings suggest that inside FAs, among the total number of integrin nanoclusters, the subset of "active" integrin nanoclusters could be engaged with "adaptor" nanoclusters on a 1:1 ratio. Using analysis of the nearest neighbor distance (NND) between distinct integrin clusters and each of the adaptors, the authors report that they found negligible spatial colocalization of integrins with these adaptor proteins and that spatial segregation is essentially determined by the density of nanoclusters within the FAs. As authors reported that α5β1 and αvβ3 do not intermix at the nanoscale, the authors finally highlighted how α5β1 and αvβ3 distinct nanoclusters are differently organized and segregated inside FAs. Adapting the NND analysis in order to inspect how far the nanoclusters are from the edges of FAs they are located in, authors revealed that α5β1 but not αvβ3 integrin nanoclusters are enriched on FA edges and that similar FA edge-enriched distribution for "active" α5β1 and adaptor protein nanoclusters was found for talin and paxillin but not vinculin. The latter results suggest that FA edges could constitute multiprotein hubs for enhanced colocalization and activation for α5β1 integrin nanoclusters and adaptors such as talin and paxillin. Unfortunately NND analysis could not confirm this enhanced colocalization hypothesis.

      General Assessment:

      While the study presents some valuable findings, it reads currently as a compilation of intriguing but preliminary observations derived primarily from a single methodology (dual-color STORM and DBSCAN clustering analysis). As the initial findings often lack confirmation through additional data analysis (such as the NND analysis the authors used), there's a critical necessity to bolster the methodological approach. This should involve replicating the main findings using alternative single-molecule super-resolution techniques (such as quantitative DNA-PAINT) or employing different clustering analytical tools (such as voronoi-tessellation). Furthermore, the manuscript feels incomplete, focusing solely on describing molecular organization without offering substantial insights into how these observations correlate with the regulation, activation, and functionality of integrins at the cellular level.

      We appreciate the comment of the reviewer and have taken their recommendation to heart in order to validate our methodology. In summary, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      The manuscript presents extensive datasets and utilizes methodologies in which the investigators demonstrate expertise. Nevertheless, there's uncertainty regarding the novelty and broad appeal of the findings. For instance, the observation of integrin nanoclustering has been previously reported in several publications (e.g., Changede et al., Dev Cell 2015; Spiess et al., JCB 2018; Fujiwara et al., JCB 2023). Similarly, the accumulation of specific proteins at the periphery of FAs has been documented elsewhere (e.g., Sun et al., NCB 2016; Stubb et al., NatComm 2019; Nunes-Vicente TCB 2023), as well as the differential dynamic organization of α5β1 and αvβ3 integrins inside FAs (e.g., Rossier et al., NCB 2012). Beyond the universal organization of adhesive proteins, there's a need to identify novel insights that significantly advance the field. One potential avenue could involve pinpointing the molecular determinant controlling the FA edge enrichment of active α5β1 integrins and talin nanoclusters. For instance, could there be an interplay between α5β1 and αvβ3 integrin nanoclusters visible on one's organisation when suppressing the other using deletion (KO) or depletion (SiRNA)? Also, could KANK, which also exhibits enrichment and regulates talin activity (e.g., Sun et al., NCB 2016), play a role in this process? Identifying the molecular players that regulate even partially the mesoscale organization of nanoclusters of proteins would really benefit the breadth of this manuscript.

      We could not agree more with the reviewer and in fact, we are currently investigating the mechanisms that control the enrichment of α<sub>5</sub>β<sub>1</sub> and adaptors at the edges of FAs. However, considering the amount of work needed to determine the spatiotemporal organization of other molecular players using super-resolution imaging constitutes a major tour de force.

      To get a first insight into the process of active α<sub>5</sub>β<sub>1</sub> enrichment at the FA edges, we hypothesised that mechanical forces exerted by the actomyosin machinery could influence the lateral distribution of both integrin subsets (α<sub>5</sub>β<sub>1</sub> and α<sub>v</sub>β<sub>3</sub>) inside FAs. Since FA maturation and strengthening over time requires mechanical forces, we performed experiments at different cell seeding times (90 min, 3 hours and 24 hours) and used STORM imaging to follow the evolution of integrin nanoclustering in time as well as their spatial distributions inside FAs. Interestingly, while nanoclustering of both integrin sub-sets inside FAs is not influenced by seeding times, their lateral distribution was markedly different, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at earlier seeding times, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appeared as rather random at earlier seeding times and progressively organized reaching a well-defined lateral spacing at 24 hours of spreading time. These initial data strongly suggest that mechanical forces might play a role in the distinct lateral distribution of both subsets of integrin nanoclusters over time. We have now included these data as new Fig. 2 of the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also discuss potential avenues for further research along the directions suggested by the reviewer.

      In addition, since it has been recently shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022) which are enriched in β<sub>1</sub> integrins, we performed dual-colour super-resolution STED microscopy of β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, our initial data show co-enrichment of both tensin-3 and active β<sub>1</sub> nanoclusters at the FA periphery, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original manuscript) or to tensin. Our current working hypothesis is that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to facilitate the translocation of α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, most probably in a talin-tensin-dependent manner. We have now included these data as Fig. S8 and accompanying discussion in pages 22 and 23 of the revised manuscript.

      Echoing the previous concern, the manuscript described a novel and rather surprising finding related to molecular clustering of adhesion proteins. Indeed, the fact that nanoclusters exhibit uniform size and molecular density regardless of the protein type, location, or activation level is indeed surprising and raises many questions about the methodology used to assess molecular clustering. I feel that the description and characterization of integrin nanoclusters appear incomplete and need to be expanded by comparing different analytical strategies for protein clustering. Furthermore, a lack of the manuscript in its actual form concerns the quantification of integrin numbers inside the observed nanoclusters. I agree that the path from optical microscopy to protein stoichiometry quantification is hard and full of drawbacks. But the authors do not fully address these issues that are extremely important when discussing protein nanoclustering. This quantitative aspect should be discussed.

      We appreciate the comment of the reviewer as indeed, the existence of “universal” nanoclusters is intriguing. Recently, together with Prof. S. Mayor we have written a short review in Curr. Opin. Cell Biol 2024 proposing that nanoclustering constitutes a molecular-scale organisation principle that governs cellular information flow at the plasma membrane. Our proposal is supported by an extensive number of recent papers showing that most cell membrane receptors and downstream signalling components are organized as pre-assembled nanoclusters. We posit that these nanoclusters serve as modular units whose concatenation in a specific spatiotemporal sequence leads to distinct signalling outputs. Thus, the existence of universal nanoclusters of integrin receptors and adaptors is indeed intriguing but not surprising to us.

      In any case, the concern of the reviewer is well-taken, and as mentioned above, we have used a different algorithm to detect and quantify nanoclustering, obtaining similar values using either Voronoi tessellation or DBSCAN approaches. These data are now included as Fig. S5 in the manuscript.

      Regarding the quantification of integrin numbers inside the observed nanoclusters, we agree with the reviewer that determining protein stoichiometry using single-molecule localization microscopy or STED remains a major technical challenge and is highly prone to artefacts. For this reason, we refrain from making claims about absolute protein numbers per nanocluster. Our relative comparison of nanoclustering among the different proteins investigated is thus exclusively based on the number of single-molecule localisations contained in each nanocluster which is a fair approach since we always use the same reporter fluorophore and maintain similar excitation conditions throughout our experiments. We have now included a few lines on page 9 regarding quantification of the absolute protein numbers inside the nanoclusters and further discuss in the revised manuscript the limitations of single-molecule localisation methods towards the stoichiometry determination of the nanoclusters (see page 20 of the revised manuscript).

      First, it is crucial for the authors to carefully examine and discuss in their manuscript whether there are any potential biases or limitations in the experimental techniques (dual-color STORM) or data analysis methods employed (DBSCAN). Second, the authors did not in the current manuscript, but should provide control samples to demonstrate the sensitivity and dynamic range of their experimental strategy.

      As already mentioned, we have validated the STORM data using both DNA-PAINT and STED and, validated our data analysis obtained with DBSCAN using the Voronoi tessellation algorithm. See Figs. S3, S4 and S5. In terms of sensitivity and dynamic range of our methodology: our set-up has single-molecule detection sensitivity which is demonstrated by the fact that we observe and detect discrete blinking events, a property of single-molecule fluorescence emission and key ingredient to super-resolution single-molecule localisation microscopy. The dynamic range (if we understand correctly the question of the reviewer) is given by the number of frames used to accumulate single-molecule localisations. In our case, we stop acquisition after we deplete most of the single-molecule spots in the imaging view, which typically occurred after 70,000 frames acquisition, as correctly mentioned in the material & methods section.

      In STORM images displayed in Figure S1, the authors highlighted localization clusters detected by DBSCAN as a signature for integrin nanoclusters. But the authors do not discuss the localization spots that were not detected by DBSCAN. Could they be individual integrins? And if so, they should also be considered as useful information? This brings me to another related technical question about how DBSCAN handles the case where fluorescent molecules are blinking. This is important as multiple emissions by a single fluorophore could be detected as a nanocluster of several molecules where it would be an artefact due to the photophysics of the fluorophore. Could the authors comment on these points?

      As mentioned in the original manuscript, between 20-30% of the localizations were not assigned to nanoclusters (Fig. S1H, I) since we imposed a minimum of ten localizations within the radius defined by DBSCAN to be considered as a true nanocluster. This essentially means that regions with less than 10 localizations were not considered in our nanoclustering analysis. However, we cannot be certain as to whether these lower number of localizations correspond to individual integrins, stochastic blinking of the fluorophore or small aggregates containing only a couple of integrins, for the same reasons that we cannot provide quantification of the absolute number of proteins included in each nanocluster: stoichiometry determination by means of single-molecule super-resolution methods is highly prone to artefacts.

      Regarding the concern of how DBSCAN handles fluorophore blinking, the reviewer is completely right as the photophysics of the fluorophore can influence the analysis of the data and the identification of true nanoclusters. To decouple the photophysics of the fluorophore we first assess the number of blinking events within the DBSCAN radius, i.e., number of localizations corresponding to individual antibodies sparsely distributed on the glass surface. In our case, the median values for the two activator-reporter pairs corresponded to 5 localizations for Alexa 405-Alexa647-conjugated Abs and 3 localizations for Cy3-Alexa 647-conjugated Abs (see Fig. 1E). Yet, despite these median values, the number of localizations per individual Ab naturally shows a distribution. Thus, to avoid any overestimation in the degree of nanoclustering, we impose an additional constrain to our analysis and consider true nanoclusters only those ones containing at least 10 localizations. We have now significantly extended the explanation in the main text (see page 6) as well as materials & methods so that it becomes clearer to the reader.

      Also, using isolated and stochastically physisorbed fluorophores (Ab coupled with activator /reporter pairs used in this study) on glass helped define the signature in STORM of a single isolated molecule. To obtain the signature of clustered fluorophores, the authors could use anti-donkey antibodies to cross-link those STORM-specifically labeled Ab as a means to artificially obtain clustered fluorophores. Ultimately, to avoid the bias effect of the glass surfaces on the photophysics of fluorophores and be in the same imaging conditions as for the described nanoclusters, the authors should use model systems composed of multimers of GFP vs. single GFP, immunolabeled with a GFP-binding monoclonal antibody. This will permit evaluation of the cluster signature obtained with DBSCAN analysis of STORM data for single vs. multimers of known stoichiometry. This would constitute an undisputable molecular stoichiometry ruler.

      We appreciate the suggestions of the reviewer. Regarding the potential bias effect of the glass surface on the photophysics of the fluorophores we would like to clarify that the “calibration” for the number of blinking events per individual Ab on glass were performed on the same sample containing the cells that we image, so that we maintain exactly the same experimental and imaging conditions avoiding any potential artefacts. To our understanding this approach is more accurate than performing the calibration on glass substrates and then moving to samples containing the cells. This information is now contained in page 6 of the revised manuscript and in the materials and method section. Once the number of blinking events from individual Abs on glass within the DBSCAN radius are determined, one can then determine the number of localizations within the same DBSCAN radius on other parts of the sample. More localizations within the same DBSCAN radius basically means more molecules, and thus nanoclusters. This approach has been extensively used by other experts in the field as we properly acknowledge in our manuscript (Pageon et al, Mol. Cell. Biol 2016; Spiess et al. J. Cell Biol 2022).

      Using anti-donkey antibodies to cross-link those STORM-specifically labelled Ab in order to artificially obtain clustered fluorophores, as suggested by the reviewer, is indeed a sound approach to retrieve signatures of clustering. Nevertheless, we have preferred not to use this approach because those artificially induced clusters would have very little resemblance to the real nanoclusters and would only allow us to validate the performance of DBSCAN for cluster recognition. As mentioned above, DBSCAN is a well-established algorithm and used by many different experts in the field and thus can be trusted by the community. Instead, and following the recommendation of the reviewer, we now provide results using an alternative cluster analysis algorithm (Voronoi tessellation) reaching similar conclusions regarding the existence of integrin and adaptor nanoclustering inside FAs.

      Finally, the suggestion of using monomeric vs multimeric GFPs to determine the stoichiometry of the nanoclusters is highly appreciated. Indeed, we have used this approach in the past to identify nanoclustering of the chemokine receptor CXCR4 in living T cells (Mol. Cell 2018 and PNAS 2022). However, these experiments are best performed at sub-labelling conditions, which inherently underestimate the degree of nanoclustering. Combining GFPs with PALM to enable super-resolution is another approach but also subject to artefacts regarding the photo-conversion efficiency of GFPs as we reported earlier (Nature Methods 2017) and leading to underestimation of nanocluster stoichiometry.

      In summary, providing nanocluster stoichiometry from single-molecule localisation images remains a major technical challenge and is highly sensitive to methodological assumptions. We have therefore focused here on providing robust evidence for the existence of integrin and adaptor nanoclustering, using three different superresolution approaches and two independent analytical methods for cluster determination.

      Due to the surprising finding of the nanoclusters' "universality", it is imperative for the authors to validate the findings through complementary methodologies and analytical tools. This should involve replication of results using alternative super-resolution techniques (quantitative DNA-PAINT) and exploring different clustering algorithms (VoronoïTesselation) to ensure the robustness and reliability of the observations.

      As already mentioned, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This work already, as is, provides significant and novel information on IAC.

      The unexpectedly low spatial colocalisation of integrins with adaptor proteins might indeed be caused by the potentially quite long extension of talin upon force exposure and the ample zones of activity of IAC proteins, imaging the involved proteins in scale, as can be seen in Barnett and Goult (2022, doi: 10.3389/fncel.2022.1014629)? In super-resolution microscopy, this spatial separation might, in fact, become apparent. It would be interesting to see whether lowering the actomyosin contraction by different concentrations of blebbistatin lowers this separation. In general, it would also be interesting to understand whether lowering the forces can disrupt the nano-organisation and the strong separation of the two analysed integrin subtypes. It is true that nascent adhesion formation is force-independent, but maybe the forces play a role in establishing the particular integrin subtype nano-organisation within more mature IAC. I am also aware that a lot of work has already gone into the conclusion of this project.

      We thank the reviewer for these thoughtful comments and suggestions. Our most recent preliminary data (not yet included in this manuscript) indeed indicate that the physical separation of integrin nanoclusters (and adaptors) inside focal adhesions (FAs) is force-dependent. We are currently reproducing these experiments using lipid bilayers of varying viscosities and controlled ligand density to explore how ligand mobility (i.e., equivalent to force exerted from the extracellular side) controls the degree of IAC nanoclustering and their spatial segregation in FAs. This approach is more amenable to super-resolution microscopy as lipid bilayers are quite thin and optically transparent, yet the experiments are still challenging, time-consuming, and thus ongoing.

      To obtain a first hint as to whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs as the reviewer suggests, we have performed experiments at different seeding times (90 min, 3 hours and 24 hours). Our results show that even at earlier times (90 min), when a lower number of mature FAs are established, nanoclustering of integrins and main adaptors are similar to 24 hours. In contrast, and as suggested by the reviewer, the spatial distribution of the different subsets of integrin nanoclusters inside FAs is markedly different as a function of seeding time, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at 90 min, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appears rather random at earlier seeding times and progressively organizes reaching a well-defined lateral spacing at 24 hours of spreading time. As FA strengthening over time requires mechanical forces, and α<sub>v</sub>β<sub>3</sub> is preferentially involved in FA strengthening (Roca-Cusachs et al PNAS 2009), these data strongly suggest that forces play a differential role in the lateral distribution of both integrin nanoclusters over time. We have now included these data as a new Fig. 2 in the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also mention in the discussion additional experiments, as suggested by the reviewer, to further substantiate this hypothesis.

      Considering the high quality of the work and the new insight about the inner organisation of IAC, maybe the summary Figure 5 should be elaborated a bit, taking into account e.g. different lengths of extended talin proteins and also the various positions of vinculins on talin proteins (depending on opened cryptic binding sites), as well as the possibility that various actin filaments might be associated with single talins. What I mean is, the authors impressively demonstrate the complexity of IAC nano-organisation, which should be paid more tribute in the concluding figure. The quality of the figure should be adapted to the quality of the work.

      We have adapted Figure 5 (now Figure 6) as suggested by the reviewer.

      I would be curious to hear a bit more about the further speculations of the authors in the discussion, e.g., about why the integrin subunits are organised in this way. Why might the α<sub>5</sub>β<sub>1</sub> be preferentially located in the periphery? What is the potential physiological relevance of this organisation? Is this organisation different in other cell types (have the authors looked at other cells)? Is the organisation lost in pathophysiological situations, such as cancer?

      Although we do not know yet what drives the preferential location of α<sub>5</sub>β<sub>1</sub> nanoclusters to the FA periphery, it is known that Kank2 also exhibits enrichment at the FA periphery, regulates talin activity and it is involved in the formation of α<sub>5</sub>β<sub>1</sub>-enriched fibrillar adhesions (Sun et al, Nature Cell Biol 2016). Thus, it is highly probable that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery is a necessary step for their translocation from mature FAs to fibrillar adhesions to then assemble fibronectin into the fibrillar networks as found and needed in connective tissues. Consistent with this idea, we have observed similar α<sub>5</sub>β<sub>1</sub> distribution on other fibroblast cell lines (MEFS), which are the primary cells that produce fibrillar adhesions. Thus, α<sub>5</sub>β<sub>1</sub> nanocluster distribution inside FAs might be physiologically important for the process of fibronectin fibrillogenesis.

      Since it has been documented that tensin is important for fibronectin fibrillogenesis (Pankov et al J Cell Biol 2000) and more recently, it has been shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022), we thought to investigate the spatial distribution of tensin-3 and its relationship with α<sub>5</sub>β<sub>1</sub> inside FAs by means of dual colour super-resolution STED microscopy. Interestingly, our initial data on HFF cells seeded for 24 hours show both enrichment of tensin-3 and α<sub>5</sub>β<sub>1</sub> nanoclusters at the edges of mature FAs, supporting our working hypothesis that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to translocate α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, probably in a talin-tensin-dependent manner. While these initial data are quite exciting, many more experiments that include simultaneous super-resolution mapping of α<sub>5</sub>β<sub>1</sub>, talin and tensin in mature FAs are required to fully validate our hypothesis. Yet, because of their relevance we consider it appropriate to include these data as Fig. S8 and discussing their potential implications in pages 22 and 23 of the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform control experiments to confirm that the nanocluster size/composition is not affected by the antibodies used.

      As explained in the response to the public reviews, antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation. In addition, we have performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. As can be observed in new Fig S4, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab. Finally, we would like to highlight that our results on the nanoclustering of integrins in terms of their size and number of localizations is consistent with previous results obtained by other groups around the world using similar labelling protocols as us (Spies et al, J. Cell Biol 2022), or relying on halo-tag strategies, as suggested by the reviewer (see Fujiwara et al, J. Cell Biol. 2023). The consistency of these results amongst different groups gives us further confidence that the nanocluster size/composition are not affected by the antibodies used.

      (2) Extend the study by inspecting a set of integrin adapter proteins for their association with inactive integrins, focusing on adapters associated with the maintenance of the inactive state. Possible candidates would be tensin and filamin, for example.

      We thank the reviewer for the suggestion and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). These results might be surprising at first, since tensin competes with talin for the same binding site to the cytoplasmic β-tail of integrins, and thus believed to act as integrin inactivator, as the reviewer indicates. Nevertheless, recent data has shown that tensin is capable to activate integrins (in particular if β<sub>1</sub> is phosphorylated) by interacting with the actin cytoskeleton, providing mechanical coupling for integrin activation (Georgiadou & Ivaska, Trends Cell Biol. 2017). We have now included these new data as Fig. S8 in the revised manuscript. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      (3) While the methods are described in sufficient detail, it is important to ask if the findings are based on sufficient data. Table S5 provides detailed information about the number of samples studied, and it appears that only small numbers of samples were investigated for certain protein pairs. This should be discussed, and perhaps more data should be obtained to strengthen the data.

      We have now performed additional experiments using DNA-PAINT as alternative super-resolution imaging technique (as also requested by reviewer 3) which adds additional data to the whole manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 M<sup>pro</sup> enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the M<sup>pro</sup> monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how M<sup>pro</sup> inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      The major concern is the absence of mutagenesis data to support the proposed inhibitory mechanisms, particularly regarding the role of the inhibitor binding site.

      We thank the reviewer for the recognition and comments. We agree that mutagenesis is critical for validating the proposed role of C300. We therefore generated the C300S and C300F mutants and characterized their oligomeric states and proteolytic activities. C300S was designed to remove the reactive thiol group while minimally affecting M<sup>pro</sup> structure and dimerization. C300F was introduced to mimic the steric perturbation associated with C300 modification and assess its impact on M<sup>pro</sup> dimerization. Native PAGE showed that WT and C300S M<sup>pro</sup> predominantly formed dimers, whereas C300F was mainly monomeric. Consistently, C300S retained approximately 70% of WT activity, whereas C300F retained only approximately 10%. Because C300F itself strongly disrupted dimerization, C300S was used as the principal mutant to evaluate the specific contribution of the C300 thiol to ebselen action. Native MS showed that ebselen could still bind to both monomeric and dimeric C300S M<sup>pro</sup> but did not markedly shift its monomer-dimer equilibrium toward the monomeric state. In parallel, ebselen reduced WT activity to approximately 53% of the untreated control, whereas C300S retained approximately 78% activity at the same 1:3 M<sup>pro</sup>-to-ebselen molar ratio. These results provide experimental support for the contribution of C300 to ebselen-induced dimer destabilization and functional inhibition, while the residual binding and inhibition observed for C300S suggest the involvement of additional C300-independent interactions. The corresponding revisions have been made to Methods (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16).

      Reviewer #2 (Public review):

      Summary:

      This is a mechanistic study that provides new insights into the inhibition of SARS-CoV-2 M<sup>pro</sup>.

      Strengths

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      We thank the reviewer for the positive comments and recognition of our study.

      Weaknesses:

      The primary weaknesses relate to linking the biophysical observations more directly to functional enzymatic outcomes and providing more quantitative rigor in some analyses. While the study is overall strong, addressing its weaknesses and limitations would elevate the impact and translational relevance of the current manuscript.

      We thank the reviewer for these comments, which have helped to iM<sup>pro</sup>ve the quality and impact of our manuscript.

      (1) Correlation with Functional Activity:

      The most significant gap is the lack of direct enzymatic activity assays under the exact conditions used for MS and HDX. While EC50 values are listed from literature, demonstrating how the observed dimer stabilization (by peptidomimetics) or dimer disruption (by ebselen) directly correlates with inhibition of proteolytic activity in the same experimental setup would solidify the functional relevance of the biophysical observations. For instance, does the fraction of monomer measured by native MS quantitatively predict the loss of activity? Also, the single inhibitor concentration used in each MS experiment needs to be specified in the main text and legends. A discussion on whether the inhibitor concentrations required to observe these dimerization effects (in native MS) or structural dynamics (in HDX-MS) align with EC50 values would be helpful for contextualizing the findings.

      We thank the reviewer for these important points. To link the biophysical observations more directly to function, we compared the oligomeric states and proteolytic activities of WT, C300S, and C300F M<sup>pro</sup>. C300F was predominantly monomeric and retained only approximately 10% of WT activity, whereas C300S remained predominantly dimeric and retained approximately 70% activity. We further evaluated ebselen inhibition using a matched 1:3 M<sup>pro</sup>-to-ebselen molar ratio. Ebselen reduced WT activity to approximately 53% of its untreated control but reduced C300S activity only to approximately 78%, demonstrating that removal of the C300 thiol significantly attenuated the functional effect of ebselen. These data support a relationship between C300-dependent dimer destabilization and reduced proteolytic activity. The Methods (Lines 627–648), and Results (Lines 397–453) have been revised accordingly, with Figures S14–S16 newly added, in the revised manuscript. We did not expect a linear relationship between the monomer fraction measured by native MS and enzymatic activity loss, because ebselen can modify multiple cysteine residues, and individual modification events may have distinct effects on M<sup>pro</sup> dimerization and catalytic function. The concentrations and molar ratios used in the native MS, HDX-MS, and activity assays have now been stated in the figure legends. The ebselen concentrations used for native MS and HDX-MS were optimized for biophysical characterization and comparison, and therefore, these concentrations might not be directly related to their IC<sub>50</sub> or EC<sub>50</sub> values. In these experiments, ebselen was applied at a 3-fold molar excess relative to M<sup>pro</sup>, consistent with the enzymatic assay. The observed dimer disruption and conformational changes were consistent with functional inhibition, supporting their mechanistic relevance.

      (2) For the two Cys residues found to be targeted by ebselen, what are their respective modification stoichiometry related to the ebselen concentration? Especially for the covalent binding site C300, which is proposed in this study to represent a novel allosteric inhibition mechanism of ebselen, more direct experimental evidence is needed to support this major hypothesis. Does mutation or modification of C300 affect the M<sup>pro</sup> dimerization/monomer equilibrium and alter the enzymatic activity? If ebselen acts as a covalent inhibitor linked to multiple Cys, why is its activity only in the μM range?

      We thank the reviewer for the insightful comments. Our LC-MS/MS data identified C44 and C300 as ebselen-modified residues, but they do not permit reliable site-resolved occupancy measurements because modified and unmodified peptides can differ in digestion efficiency and MS response. We have therefore clarified that these data provide qualitative site identification rather than absolute modification stoichiometry. To obtain direct functional evidence for C300, we generated C300S and C300F mutants. C300S preserved dimer formation and substantial activity, whereas C300F was mainly monomeric and showed severe activity loss. Importantly, although ebselen-bound C300S species were still detected by native MS, ebselen did not markedly redistribute C300S toward the monomeric state, and its inhibition was reduced from approximately 47% for WT to approximately 22% for C300S. These results indicate that C300 is an important contributor to ebselen-induced dimer disruption, while residual binding and inhibition indicate additional reactive sites. Corresponding revisions have been made to the (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16). The moderate micromolar potency of ebselen is consistent with its heterogeneous, multi-site covalent reactivity: modification occupancy and functional consequence are site-dependent, and not every adduct produces complete inhibition.

      (3) For the allosteric inhibitor pelitinib with low-μM activity, no significant differences in deuterium uptake of M<sup>pro</sup> were observed. In terms of the binding affinity, what is the difference between pelitinib and ebselen? Some explanations could be provided about the different HDX-MS results between the two non-peptidomimetic inhibitors with similar activities.

      We agree with the reviewer that the absence of significant HDX changes for pelitinib requires clarification. Different from ebselen that forms covalent bond with multiple cysteine residues of M<sup>pro</sup>, which could lead to sustained conformational changes that are more readily detected by HDX-MS, pelitinib non-covalently binds M<sup>pro</sup> and might not induce significant perturbations in backbone dynamics that are detectable at the peptide level by HDX-MS. These points have been integrated into the revised manuscript (Lines 333-337).

      (4) Native MS Quantification: 

      The analysis of monomer-dimer ratios from native MS spectra appears qualitative or semi-quantitative. A more rigorous and quantified analysis of the percentage of dimer/monomer species under each condition, with statistical replicates, would strengthen the equilibrium shift claims. For native MS analysis of each inhibitor, the representative spectrum can be shown in the main figure together with quantified dimer/monomer fractions from replicates to show significance by statistical tests.

      We thank the reviewer for the suggestion. We have performed a quantitative analysis of the monomer-dimer equilibrium based on triplicate native MS measurements for each condition. Representative spectra, quantified monomer/dimer ratios, and statistical analyses have been added to Figures 1 and S3. The quantitative results have also been described in the Results section (Lines 158–161, 165-168, 172-174, 177-179, 199-200).

      (5) Changes of HDX rates in certain regions seem very subtle. For example, as it states 'residues 296-304 in the C-terminal region of M<sup>pro</sup> were more flexible upon ebselen binding (Figure 4c)', the difference is barely observable. The percentage of HDX rate changes between two conditions (with p values) can be specified in the text for each fragment discussed, and any change below 5% or 10% is negligible.

      We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296–306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The study lacks validation through inhibitor binding site mutagenesis assays, especially peptidomimetic inhibitor PF-07321332 and ebselen, which would strengthen the mechanistic conclusions.

      We appreciate this suggestion. For PF-07321332, the inhibitor forms a covalent interaction with the catalytic residue C145 and inhibits M<sup>pro</sup> activity through a distinct mechanism. Previous studies have shown that mutation of C145, such as C145A, completely abolishes M<sup>pro</sup> catalytic activity (Bhandari, D, et al. Communications Biology 2025, 8, 1061), making it difficult to directly evaluate the contribution of this residue to inhibitor-induced inhibition using enzymatic assays alone. This limitation and the relevant literature have now been discussed in the revised manuscript (Lines 272–277). Therefore, we focused on C300-dependent regulation of ebselen, which represents a distinct inhibitory mechanism involving modulation of M<sup>pro</sup> structural dynamics and dimer stability.

      To validate the role of C300 in ebselen-mediated regulation of M<sup>pro</sup>, we generated C300S and C300F mutants and performed additional biochemical and structural characterization. The enzymatic assay showed that the C300F mutation significantly affected M<sup>pro</sup> activity, and the inhibitory effect of ebselen on C300S M<sup>pro</sup> was markedly reduced compared with WT M<sup>pro</sup>. Furthermore, native MS analysis demonstrated that ebselen could still bind to C300S M<sup>pro</sup> but failed to induce a significant shift in the monomer-dimer equilibrium observed for WT M<sup>pro</sup>. These results indicate that C300 is not the only site involved in ebselen binding but is critical for mediating ebselen-induced structural perturbation and dimer destabilization. The manuscript has been revised accordingly for the (Lines 627–648), and Results (Lines 397–453), with new figures (Figures S14–S16) included, further supporting the functional contribution of C300 in ebselen-mediated M<sup>pro</sup> regulation.

      (2) MR6-31-2 is an ebselen derivative and exhibits a lower EC50 (1.78 μM) compared to ebselen (4.67 μM). It would be helpful to discuss why their activities differ, probably based on the assay conditions or binding behavior.

      We agree with the reviewer that the difference in antiviral activity between MR6-31-2 and ebselen requires further clarification. The lower EC<sub>50</sub> of MR6-31-2 may result from iM<sup>pro</sup>ved cellular properties, including compound stability, permeability, intracellular exposure, and potentially altered interactions with M<sup>pro</sup> and/or iM<sup>pro</sup>ved cellular properties. Although MR6-31-2 shares the ebselen scaffold, the modified chemical structure may affect its binding behavior and biological activity. However, EC<sub>50</sub> values obtained from cellular assays cannot directly reflect the biochemical inhibition potency against purified M<sup>pro</sup>. These points have been integrated into the revised Introduction (Lines 98–101).

      (3) In Figures 2, S1, S2, S4, S6, and S11, adding the drug name under each panel would make the data much clearer for readers.

      The corresponding drug names have been added to panels to iM<sup>pro</sup>ve figure clarity.

      Minor points:

      (1) Line 62-63 refers to the "long linker loop," while Figure 1a labels it as the "long loop linker." Please keep this consistent.

      The terminology has been unified as “long loop linker” throughout the manuscript.

      (2) Table 1 should be cited at line 80, and PDB code 7BAK should be included in Table 1.

      PDB code 7BAK has been included in Table 1, and Table 1 has been cited in the context, as suggested.

      (3) Figure 1a should include the corresponding PDB code in the figure legend.

      The corresponding PDB code has been added to the Figure 1a legend, as suggested.

      (4) It would be helpful to indicate in Figure 1a that the upper structure represents the dimer and the lower structure represents the monomer.

      The upper and lower structures in Figure 1a have been indicated as dimeric and monomeric M<sup>pro</sup>, respectively, as suggested.

      (5) In the Figure S1 legend, it should mention that some inhibitor structures (like ebselen and MR6-31-2) are not fully resolved. Also, the Se atom in ebselen should be shown in Figure S1f (PDB: 7BAK).

      The Figure S1 legend has been revised to indicate that some inhibitor structures, including ebselen and MR6-31-2, are partially unresolved, and the selenium atom of ebselen has also been shown in Figure S1f, as suggested.

      (6) Pelitinib is an allosteric, non-covalently binding inhibitor. However, in Figure S3, the native MS profile shows dimer species (13+ to 15+) compared with unbound M<sup>pro</sup> (14+ to 17+). Please clarify this difference.

      We thank the reviewer for raising this good point. Protein charge-state distributions can be influenced by solution-phase conformation, conformational flexibility, solvent properties, and electrospray droplet charging (Susa AC, et al. J Am Soc Mass Spectrom 2017, 28, 332-340). The observed shift in charge state distribution in native MS might suggest that the addition of pelitinib caused changes in the protein conformation, solvent property and electrospray droplet charging. The relevant literature and discussion have been added in the revised manuscript (Lines 200–204).

      (7) Line 172: "S1are" should be corrected to "S1 are."

      Corrected.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Comments on revised version:

      I thank the authors for their extensive revisions and detailed responses to the reviewers' comments. The manuscript has been substantially improved, and most of the major concerns raised in the initial review have been adequately addressed. In particular, the additional analyses of PKD2L1 channel activity, the incorporation of physiologically relevant pH conditions, the clarification of ASIC involvement, and the expanded Discussion have significantly strengthened the study.

      Major scientific concerns largely addressed:

      Quantification of PKD2L1 channel activity

      The authors appropriately addressed my previous concerns regarding the use of Po as the sole measure of channel activity. The inclusion of additional parameters such as apparent Po, open time, nmax, holding current, and membrane charge provides a more robust assessment of PKD2L1 activity and substantially strengthens the conclusions.

      Physiological relevance of pH modulation

      The inclusion of experiments at pH 6.5 and the additional analyses of holding current and resting membrane potential are valuable additions. These experiments considerably improve the physiological relevance of the study.

      ASIC contribution

      The additional pharmacological experiments using ASIC blockers are helpful and support the conclusion that the photolysis-evoked response in the apical process is predominantly mediated by PKD2L1 channels.

      Functional implications

      The expanded Discussion regarding Ca2+-dependent signaling, neurosecretion, and the potential physiological roles of CSFcNs considerably improves the manuscript.

      Remaining concerns:

      Continued overstatement regarding "exclusive" localization and function:

      Although the authors softened some statements in the revised manuscript, the term "exclusive" remains in several key locations, including the title.

      For example:

      "PKD2L1 channels segregated to the apical compartment are the exclusive dual-mode pH sensor..."

      The data clearly demonstrate strong enrichment of functional PKD2L1 channels in the apical process. However, the available evidence does not fully justify the term "exclusive," particularly because:

      - PKD2L1 immunoreactivity is still detectable outside the apical process.

      - ASIC-mediated responses are present in CSFcNs.

      - The authors themselves use more appropriate terminology such as "predominantly located" in the Discussion.

      Therefore, I recommend replacing "exclusive" with more conservative terminology such as:

      - predominant

      - predominantly localized

      - enriche

      - functionally segregated

      throughout the manuscript, including the title, Abstract, Introduction, Results, and Discussion.

      We agree with the reviewer that the world “exclusive” is misleading and should be replaced. Following the reviewer’s suggestions, we have deleted the word “exclusive from the title, which now reads: “PKD2L1 channels segregated to the apical compartment are the functional dual-mode pH sensors in cerebrospinal fluid-contacting neurons.”

      In addition, the word “exclusive” has been changed with more conservative terminology in other parts of the text: lines 80, 420, 466 and 551.

      Use of the term "tonic current"

      The manuscript continues to use the term "PKD2L1 tonic current."

      While the dibucaine-sensitive holding current is clearly present, the precise mechanism generating this current remains uncertain. Indeed, the authors themselves acknowledge in the Discussion that:

      - an alternative conducting state may exist, or

      - unresolved brief channel openings may account for the current.

      Therefore, the data support the existence of a sustained PKD2L1-associated current, but do not yet definitively establish a distinct tonic gating mode of the channel.

      I therefore recommend replacing:

      "tonic current" with a more neutral expression such as:

      - sustained current

      - PKD2L1-associated holding current

      - sustained PKD2L1-mediated current throughout the manuscript.

      Continued use of "off-current" and "off-response":

      The revised manuscript has improved considerably in this regard. However, the terms "off-current" and "off-response" still remain in portions of the text and figure legends.

      Because the manuscript itself demonstrates that the response reflects recovery from transient acidification rather than a separate OFF signaling mechanism, these terms remain potentially misleading.

      I recommend replacing them with terminology such as:

      - photolysis-evoked PKD2L1 current

      - recovery current

      - proton-removal-induced current

      throughout the manuscript, including figure legends.

      We apologize, as the word “tonic” and the terminology “off-current” should have completely disappeared after the first round of revisions. We have now replaced those all along the text. “Tonic” has been replaced by “sustained”.

      “Off-current” or “off-response” have been replaced by appropriate terms in lines: 339, 342, 544, 545, 546, 549, 552, 555, 556, 560, 576, 580, 585, 588, 807 and 929. We have nevertheless conserved the term “off-current” in line 552 as we are referring to terminology used by other authors.

      Minor editorial corrections

      Figure 1Bd Please change: "po" to "Po" for consistency with standard channel physiology nomenclature.

      Figure 1Ca Please add units (mV) to the voltage labels shown on the left side of the traces.

      Figure 3E Please change: "Norm po" to "Norm Po".

      Figure 4Fb Please replace: "sec" with "s" to conform with SI unit conventions.

      Done.

      The authors have addressed the majority of my previous concerns and the manuscript has been substantially improved. The remaining issues are primarily related to terminology and overinterpretation rather than experimental deficiencies.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

      Weaknesses:

      Not weaknesses, there are only minor requests to complete the beautiful study.

      (1) The apical extension's response to removal of acidification is nicely illustrated in Figure 4C,G. There's something puzzling there: while the response to Glutamate is immediate, the channel responses to H+ is extremely delayed by 100ms - 2s, and even sometimes came in bursts separated by few hundreds of ms. H+ diffuse even faster than glutamate. Why is that?

      I don't quite understand how the response is so delayed & how to explain the recurring bursts of channel opening in the figure panel ?

      The kinetic of the response to proton uncaging is analyzed in Figure 4E, where the charge of the current traces is plotted against time. What this analysis shows is that the response lasts a few hundred ms (τ 250 ms) and then the PKD2L1 activity increase subsides to baseline. The peak of the response is at 100 ms (Figure 4G), but the increase in activity happens as soon as the uncaging pulse ends (Figure 4D, G and H). This behavior has already been shown in expression systems, where the channel activity is blocked by protons and the blockage is released when the acid is withdrawn. In an intact cell as the CSFcNs studied here, the exact kinetics of the recovery response are probably more complex (and variable) than in expression systems. Indeed, it is known that the recovery of this current depends, for example, on pH and extracellular calcium. Also, PKD2L1 are inhibited by intracellular calcium (de Caen et al, eLife 2016) but are themselves permeable to Ca<sup>++</sup> ions. The interaction of these effects could give rise to the “bursts” that are observed in some cases. However, this is merely speculative at this point.

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      Following the reviewer’s suggestion, we have added a trace in Figure 4C (upper blue trace) showing the spontaneous activity of the cell, prior to uncaging, as it is already shown for another example in Figure 4D.

      - Could the authors use a fluorescent pH sensor to monitor pH in the extracellular space and in the cell ?

      This is an important point that was already addressed by the reviewing editors in the previous round of revisions. Indeed, we have attempted to perform pH calibrations in the setup using the pHsensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2 and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, the calibration under the conditions of a real experiment is not possible because the photolysis in the slice occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml). In these conditions, the 405 nm uncaging pulse bleaches the dye in the photolysis spot and any useful information is lost. In addition, our imaging system is not fast enough to follow the pH change. As discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 791), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond range.

      - Could the authors investigate whether in the apical extension, PKD2L1 channels are mainly at the outer membrane in the apical extension OR whether many channels are located in inner membranes ?

      PKD2L1 channels are probably subject to a high rate of turnover, and they are certainly localized in the plasma membrane of the apical process as well as in the inner membranes. Although this is a very interesting point, we believe it is out of the scope of this work.

      (2) Suppl Fig 4 is very cool and should be moved to main figure. The coupling of Soma and AP is very tight, yet there is a clear difference in targeting of channels that respond to cues in the CSF. In the context of an intact spinal cord, we can wonder how and when the contribution from ASIC in the some would be relevant to physiology. Can the authors think of experiments with an intact central canal to test the sensitivity and condition of recruitment of pH sensing in the soma (ASIC) versus the apical extension (PKD2L1)?

      We have followed the suggestion of the reviewer and have made Supplementary Figure 4 a main figure.

      The fact that the normal interphase between the spinal cord parenchyma and the cc is lost is already acknowledged in the discussion, lines 486 to 489. As the reviewer suggests, PKD2L1 and ASIC channels seem both to be important in the response of CSFcN to pH changes. However, both channels are activated in very different physiological contexts, as is discussed in the section “The involvement of ASIC channels”. Keeping the central canal intact in order to be as close as possible to physiological conditions, as suggested by the reviewer, would be ideal. However, as CSFcNs are in the middle of the cord, it would require the use of optical techniques that allow to penetrate deep into the tissue (e.g., 2-photon microscopy) that unfortunately are not available in our labs.

      (3) The Reissner fiber is missing after slicing the spinal cord. From our observations in fish, the fiber being under tension triggers lots of activity in CSF-cNs (Bellegarda et al Elife 2023) that also relies on PKD2L1 (Bohm et al NC 2016; Sternberg et al NC 2019). Could the authors discuss the contribution of the Reissner fiber to the PKD2L1 mediated modulation of CSFcN excitability ? Could the authors conceive a way to slice along the anteroposterior axis (sagitally) the spinal cord to keep the Reissner fiber in the central canal when recording CSF-cN apical extension ?

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      As discussed in the previous point, the in vitro slice preparation has technical limitations that are mainly related to the alterations of the normal structure of the tissue. Although keeping the Reissner fiber intact in a sagittal slice seems possible, accessing the CSFcNs with electrophysiological methods would still be a challenge.

      We have now added a sentence in the Discussion, lines 561 to 564, where we discuss that CSFcN excitability is modulated by the Reissner fiber and that it remains to be explored whether in rodents the gating of PKD2L1 channels is modulated by the Reissner fibre, as has been shown in zebrafish.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study examines the context-dependent modulation of auditory cortical neurons in response to expected sensory input, either self-generated sounds or expected perturbations of self-generated sounds. Specifically, using songbirds, the authors ask whether social context (the presence of a female conspecific) affects 1) the response of auditory cortical neurons to the bird's own song when he is singing; and 2) the response of neurons to perturbations of auditory feedback that the bird has been trained to expect.

      Strengths:

      First, the authors report that across the population, the responses of the neurons does not differ when a male bird sings alone or if he sings to a female. A fraction of auditory cortical neurons, however, do show significant differences in the firing rate, precision, and/or degree of burst firing when males sing alone vs. when they sing to females. This finding is broadly consistent with the literature showing that sensory neurons (visual, auditory, somatosensory, etc.) can be rapidly reconfigured into different "information processing modes" depending on behavioral state (e.g., quiescence vs. vigilance).

      For the perturbation experiments, the authors trained birds to expect distorted auditory feedback during a particular syllable. They found that some neurons showed greater responses during perturbation when a female was present (compared to when males were alone) while other neurons had smaller responses during perturbation when a female was present. In addition, the response of a small number of auditory cortical neurons were not affected by behavioral state. These results contrast with their prior report that the responses of midbrain dopaminergic neurons that project to the basal ganglia are "uniformly reduced" in the presence of a female, raising a question of how an evaluation signal is transformed in the circuit from the primary sensory region to the midbrain.

      Weaknesses:

      While the experiments and analysis are solid, the finding that social context can alter responses of auditory cortical neurons in a multitude of ways (increase, decrease or no change) raises several questions that can be examined with additional analysis. For example, do context-dependent differences in auditory responses derive from context-dependent differences in the songs? Are context-dependent differences present in all classes of neurons and throughout the auditory system?

      The observed heterogeneity in the firing properties of auditory cortical neurons, both in response to self-generated sounds and during perturbations of auditory feedback, raises the question of which neurons are sensitive to social context (which likely can be addressed by the authors in a revision). The authors should provide additional details about the recordings:

      (a) What are the locations of the recording sites? Prior work has shown that there is an organized map of spectrotemporal features of sounds in the auditory cortex of songbirds; spectral tuning widths change along the medial-lateral axis and temporal tuning widths differ between the input and output layers of Field L. Were the recordings primarily in Field L2 (thalamo-recipient region), L1 or L3? Were some recordings lateral to Field L in secondary auditory regions? Were the neurons that showed context-dependent changes in firing properties localized or distributed throughout Field L (i.e., were the context-dependent differences in neural responses truly brain-wide)? At a minimum, the authors should include a schematic showing the different regions of Field L and a summary of the location of the recording sites. Images of the processed tissue with electrolytic lesions would also be helpful.

      We agree that the anatomical targeting and limits of localization should be made explicit. In the original manuscript, we referred broadly to recordings from "Field L" and described targeting coordinates in the Methods. In the revised manuscript, we have softened the anatomical claim from "Field L" to "auditory pallium" where appropriate, while explicitly stating that electrodes were aimed at Field L. We also added anatomical caveats and a new supplemental figure.

      The revised title and abstract now reflect this more conservative anatomical framing. For example, the abstract now states: "Here we recorded neural activity from the auditory pallium in zebra finches practicing singing alone and directing courtship songs to females." In the Introduction, we now explicitly state both the intended target and the limitation: "We targeted our recording electrodes to Field L, a primary auditory pallial area that projects into multiple higher auditory areas that, in turn, project to VTA."

      We then added the caveat: "Field L is composed of multiple subdivisions and surrounds the interfacial nucleus and because the implanted wire bundles spread in a small radius of up to ~0.5 mm, our recordings likely included large territories of the auditory pallium (Figure S1)."

      We also added mechanistic/anatomical context for why these recordings may reflect activity shaped by broader auditory forebrain circuitry: "Although Field L is classically described as a primary auditory thalamorecipient region, its activity may also be shaped by contextual signals related to the courtship context, potentially via recurrent interactions with higher-order auditory forebrain regions such as the caudal mesopallium (CM) and caudomedial nidopallium (NCM) (Bauer et al., 2008; Figure S1C)."

      Changes made in revision: We added Figure S1, which includes anatomical subdivisions, an example histological slice showing the area where the cannula was implanted, and auditory pathway connectivity. We also revised the wording throughout the manuscript from "Field L neurons" to more conservative phrasing such as "auditory pallium neurons" or "pallial auditory neurons" when appropriate. We did not claim layer-specific localization, because the revised manuscript explicitly states that we cannot make such claims.

      (b) Was the context-dependent modulation limited to a particular class of neurons (distinguished by spike waveform shape, spontaneous firing rate, or other feature)?

      We agree that identifying whether context-dependent modulation is associated with specific neuronal classes is important. In the revision, we added analyses examining relationships between spike width and various firing characteristics. We also looked for potential relationships between mean rate and DAF response, IMCC and DAF response, and found no clear trend. Overall, we did not observe any clear relationship between DAF response, DAF response modulation, and metrics like spike width or mean firing rate.

      The revised manuscript states: "Action potential width of individual neurons has previously been used to classify putative interneurons or putative principal cells in the zebra finch auditory pallium (Calabrese and Woolley, 2015)."

      We then describe the new analysis: "We tested if DAF-response scores, mean firing rates during singing, burst fraction, IMCC, and the change in all of these between undirected and directed singing was correlated with spike half width (spike half-width measured as peak-to-trough time; Figure S5)."

      The revised result is: "DAF response in either condition, the change in DAF response across conditions, and the change in firing rate, burst fraction, and IMCC were not significantly correlated with spike width (Figure S5A-B)."

      We also report that some general firing properties did correlate with spike width: "Consistent with previous literature, firing rate was significantly correlated with spike width (Pearson's correlation, p=0.003; Figure S5C). Interestingly, burst fraction (p=8.3x10-4) and IMCC (p=5.4x10-5) were also significantly correlated with spike width (Figure S5C)."

      Changes made in revision: We added Figure S5 and associated text analyzing whether context-dependent changes in DAF response, firing rate, burst fraction, and IMCC were correlated with spike half-width. These analyses did not support the conclusion that context-dependent DAF modulation was restricted to a waveform-defined neuronal class.

      (a) Prior work has shown that songs of zebra finches differ slightly when males sing alone compared to when they sing to females: songs are faster; pitch is less variable; and the number of introductory elements is greater when males sing to females. Do some of the observed social context-dependent differences in the responses of auditory neurons reflect differences in the songs in the two conditions? Did the authors of this study also find premotor activity in Field L, and if so, did it differ between the two social contexts? Might differences in Field L responses reflect motor/song differences?

      The revised manuscript now addresses the issues of context-dependent changes in song in several ways. First, we explain why motif-aligned comparisons are meaningful: "The acoustic structure of undirected and directed motifs is highly similar in adult finches, enabling singing-related neural activity to be precisely aligned and compared across contexts."

      Second, we added analysis and discussion of premotor-related activity. Changes made in revision: New results paragraph and new Fig S4. "A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context-dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."

      (b) For the perturbation experiments, this raises a question of whether perturbation amplitude is different when a male is alone and when a female is present. It would be useful to know if (and how much) perturbation amplitude varied depending on the location inside the cage as well as whether the sound pressure level of the underlying song was higher (e.g., Lombard effect).

      We previously calibrated the perturbation amplitude in Roeser et al., 2023, in an identical recording setup. Two speakers deliver the feedback on either side of the bird's home cage. We acknowledge the possibility that the position and orientation of the bird can affect the way the sound hits either of the bird's eardrums and thus potentially affect a neural response. However, neural activations following the absence of distortion playbacks were a major feature of the dataset and were context-dependent in some cases.

      The Methods state: "DAF was implemented with a custom LabVIEW acquisition program that analyzed song syllables in real-time and delivered syllable-targeted feedback." and "DAF (50 ms broadband noise bandpass filtered at 1.5-8 kHz to match frequency range of zebra finch song) was played over speakers in the recording chamber on top of a specific target syllable randomly on 50% of motif renditions."

      The revised manuscript also makes clear that experiments occurred in the bird's home cage: "Experiments were carried out in the male's home cage, which was inside a sound isolation chamber."

      Importantly, the revised Results show that not all DAF-related responses were simple activations to additional sound. Some neurons were activated by the absence of distortion: "Unexpectedly, some neurons were not activated by the song distortion but rather by the lack of target syllable distortion." and "These activations following undistorted renditions could also depend on the courtship context."

      Changes made in revision: We clarified the DAF stimulus and recording setup in Methods and added Figure 3 showing neurons activated following undistorted renditions. These data argue that context-dependent responses are not simply explained by DAF sound amplitude, although we do not claim that position-dependent acoustic variation was fully eliminated.

      Finally, it would be helpful if the authors could include a model and/or more discussion of how the uniform attenuation in midbrain dopaminergic neurons may arise given the heterogeneous responses in Field L.

      The revised manuscript provides evidence for context-dependent retuning upstream of VTA, but does not offer a direct mechanistic explanation for the uniform attenuation seen in dopaminergic neurons. The revised Discussion states: "Because the main goal of this study was to test if courtship-associated reduction in DAF signaling, recently observed in VTA DA neurons (Roeser et al., 2023), resulted from a local process in VTA or reflected a retuning of auditory responsiveness, we explicitly tested for changes in DAF responsiveness between alone and female-directed singing."

      It then explicitly contrasts auditory pallium and VTA: "Surprisingly, we discovered that Field L neurons could retune at the transition from lone to courtship singing in diverse ways, consistent with a more widespread process in the brain that does not fully explain the uniform DAF-signal attenuation observed in VTA."

      Changes made in revision: We expanded the Discussion to explicitly state that auditory pallium retuning is heterogeneous and therefore does not fully explain the uniform attenuation observed in VTA. We do not present a formal circuit model, but we now more clearly frame the result as evidence for broader sensory retuning that is likely transformed downstream.

      Reviewer #2 (Public Review):

      Summary:

      In the manuscript, Jones and Goldberg study auditory cortex in male zebra finches. They explore song-related responses in two different contexts, when the male is either alone or in the presence of a female. They find a heterogeneity of responses, in line with auditory cortical neurons computing the social modulation of responses found in VTA.

      Weaknesses:

      Stability of responses has not been studied: some neurons seem to have responses that slowly drift in time, which could lead to observed differences between alone and with-female conditions. Also, possible motor confounds and sound-of-audience confounds should be addressed. The language is often imprecise.

      Stability and Reversal: It is a bit unfortunate that stability of effects seemingly has not been studied by reversing experimental conditions. The work would be much stronger if authors could show that audience-dependent tuning is robust in individual cells. Did they record from some neurons during reversal back to the alone condition?

      We agree that recording stability is essential. A reversal experiment was not feasible for this dataset, as it is difficult to confirm whether song motifs produced immediately following female presence represent undirected singing or are directed to an unseen but recently present female. Instead, the revised manuscript adds a strict unit-stability criterion based on waveform similarity across conditions.

      The revised Results state: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." The exact criterion is: "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      Changes made in revision: The strict waveform-stability inclusion criterion and new Figure S2 directly showcase unit stability across the time course of the experiments.

      Motor responses: Does DAF playback change song? If so, especially if it applies only in one of the two conditions (audience/no audience), then the observed response differences could be motor-related rather than auditory responses.

      We agree that motor confounds must be minimized. We previously found that DAF did not affect the acoustics of the subsequent syllable (Gadagkar et al., 2016). The revised manuscript clarifies that DAF and undistorted trials were randomly interleaved and analyzed by comparing matched renditions within conditions. Importantly, we only analyzed motif-aligned activity, ensuring that all syllables within the song motif are the same.

      Changes made in revision: We clarified the DAF analysis framework and added a more conservative permutation-based analysis comparing distorted and undistorted trials within each context, then comparing those DAF-response vectors across contexts. We do not claim that all possible motor consequences of DAF are eliminated, but the analysis directly tests neural responses to randomly interleaved distorted versus undistorted renditions.

      Similarly, motif-aligned spiking activity was time warped to the median duration of undirected or directed motifs. Could the shorter motifs during directed song lead to alignment differences that would account for the different error responses in alone/with-female conditions?

      We agree this is an important technical point. The time-warping we conducted, standard in the field, compensates for the tempo differences between directed and undirected song. Importantly, our main analysis of change in error response no longer uses a 100 ms response window, but rather includes all windows in the motif.

      Changes made in revision: We clarified that the revised DAF response analysis uses motif-aligned, time-warped spike trains. Importantly, the revised analysis moves away from relying on a single scalar response window and uses bin-wise permutation tests with family-wise error correction.

      Audience versus sound of audience: Is it truly the audience that causes the difference in error responses or is it the sounds the audience makes?

      We agree that the sensory cues defining "audience" cannot be fully separated in this experiment. The reviewer raises an important point that female zebra finches occasionally call at the male. We have excluded all song motifs from analyses that include an overlapping female call.

      The revised Methods now explicitly state that motifs overlapping with female calls were excluded: "Any motifs that had overlapping time with a female call in directed motifs was excluded from analysis."

      We also revised the Discussion to treat the mechanism by which auditory pallium receives information about the female as an open question: "An open question is how auditory pallium receives information about whether a female is present, and how this information influences neural activity."

      Changes made in revision: We excluded motifs overlapping with female calls and added discussion explicitly acknowledging that how female presence is represented in auditory pallium remains unresolved. We do not claim to distinguish visual, auditory, social, or motivational components of the female-present condition.

      Reviewer #3 (Public Review):

      Summary:

      In this study, Jones et al. examine how neural activity in a primary auditory area (field L) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). Prior work has demonstrated that the presence of an audience attenuates the responses of dopaminergic neurons to distortions of auditory feedback (DAF). Here the authors report that even in a region that is primarily considered sensory, responses to DAF are also modulated by the audience, although in a heterogeneous manner. However, to be fully persuasive, additional analyses will be required to address how much of the apparent modulation by audience may be explained by other factors such as changes in recorded neurons or their properties over time.

      (1) A central concern relates to whether the main reported effects associated with differences in singing directed versus undirected song reflect only those changes in conditions, versus contributions from changes in unit isolation or response properties over time.

      We completely agree that unit stability is critically important in this study. To address this concern, we now quantify stability and apply strict inclusion criteria adopted from a study that assessed unit stability over days (Dickey et al., 2009). Additionally, we now include average waveform overlays for all example units across conditions as supplemental Figure S2.

      Changes made in revision: We added: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." and "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      (2) A second concern has to do with the categorical definition of 'error neurons'. The authors define a subset of neurons as error responsive only if their responses to DAF exceed a specific threshold (2.5 standard deviations). The problem is that for some neurons categorically defined as being responsive to DAF in only one condition, there is almost certainly not a significant difference in the actual responses to DAF between conditions.

      We overhauled our analyses characterizing DAF responses. Rather than relying only on a 2.5 z-score threshold, we now use a more conservative permutation-based approach that directly tests DAF responsiveness and context-dependent changes in DAF responsiveness.

      The revised Results state: "Statistical tests defining auditory neurons as DAF-responsive or not in a binary fashion may not be suitable if the underlying population of DAF-related responses exist on a continuum from responsive to non-responsive."

      The updated result is: "This more conservative approach identified 48/147 neurons as DAF-responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (3a) Some discussion of what is already known about the auditory tuning of Field L, and the extent to which responses associated with distortion of feedback may reflect the frequency tuning of Field L neurons versus something that might be construed as more specifically as detecting an error in perceived feedback.

      We agree that DAF-related changes in firing do not necessarily imply that neurons are explicitly detecting an "error" between predicted and actual feedback. Field L neurons can have spectrotemporal receptive fields and frequency tuning such that a broadband DAF stimulus could drive excitation or inhibition simply because the stimulus overlaps with excitatory or inhibitory regions of a neuron's receptive field. We therefore revised the manuscript to use more cautious language and to describe these responses as DAF-related or feedback-related signals rather than categorically as "error responses".

      Changes made in revision: The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We also added a sentence to the Discussion: "However, it is important to note that DAF-related changes in firing in auditory neurons do not necessarily imply that neurons compute sensory prediction errors. DAF-related responses could arise from ordinary auditory tuning to the broadband distortion stimulus."

      (3b) It would also be useful to discuss further previous work on differences in auditory tuning or responses between conditions when subjects are vocalizing, versus when vocalizations are played back, and to what extent efference copy signals might contribute to the processing of feedback distortions.

      We agree these are important points. Our experimental design did not include sufficient passive bird-own-song (BOS) playback trials to permit quantitative comparisons with vocalizing conditions, and we therefore cannot draw firm conclusions about the contribution of efference copy signals to the DAF responses described here. We did observe robust motif onset-associated neural activations, including some activity preceding motif onset, which were present across both social contexts (see new Figure S4). These observations are consistent with prior reports of premotor-related signals in Field L (Keller and Hahnloser, 2009), but whether such signals contribute differentially to DAF processing across contexts remains an open question that we now acknowledge in the Discussion.

      (3c) To what extent did the current study control for any vocalizations or other sounds produced by females during the directed singing, and could this have contributed to differences in Field L activity between conditions?

      Please see response R2.4 above, in which we describe the exclusion of all song motifs that overlapped in time with a female call. This exclusion criterion was applied throughout all analyses of directed singing.

      Figure 1D: In the directed condition there are no spikes at all following the first handful of motif renditions. Were the directed and undirected recordings interleaved here?

      Undirected and directed trials were not interleaved. The raster plots are presented in chronological order; however, for each behavioral condition, rows are sorted with the earliest renditions at the bottom and the most recent at the top. We have clarified this in the figure legend.

      A minor issue: the raw example trace with male alone does not seem to have a corresponding set of points in the raster plot. For panel E, I also cannot find rasters that correspond to the example recordings shown at top.

      In the original version, we randomly downsampled the condition with more trials to equalize trial counts across conditions in the example rasters, while performing all analyses on the full set of recorded trials. As a result, the example spike shown in the raw trace was drawn from one of the downsampled trials not displayed in the raster.

      Changes made in revision: For greater transparency, we now include all trials from both conditions for each example neuron in Figure 1.

      Figure 2A also shows a neuron that looks like it has non-stationarity; for the alone condition without altered feedback, the main peak has no spikes for the bottom half of the rasters.

      In the original version, example neurons were selected to illustrate the DAF-response scoring method, which in some cases highlighted neurons with less stable response profiles. In the revised manuscript, we have replaced this example with neurons that exhibit more robust and stable DAF-related responses, and we now provide a broader set of example neurons illustrating both increases and decreases in DAF responsiveness across conditions.

      Other figures show firing rate distributions that appear to be very non-Gaussian, with some motifs during which there is a lot of activity, and others in which there is little activity. Please consider applying non-parametric tests as appropriate.

      We agree. In general, some neurons exhibited non-uniform firing rate distributions across trials. All of our main analyses are now conducted using non-parametric permutation tests, which do not assume a Gaussian distribution of trial-by-trial firing rates.

      Approaches to addressing the non-stationarity issue could include more specifically indicating examples in which recordings from the alone condition and directed condition are interleaved and exhibit reversible changes in the pattern of responses.

      Unfortunately, nearly all of our undirected and directed recording periods were not interleaved, as the experimental design required a block of undirected singing followed by directed singing with female presence. We find it informative, however, that DAF-response modulation was observed in both directions, with some neurons losing DAF responsiveness during directed song and others gaining it, a pattern that is difficult to attribute to a simple unidirectional drift in recording quality. We now provide additional examples illustrating both directions of modulation in Figures 2 and 3.

      The methods and/or raster plots should include some further explanation of the time periods over which recordings were made in the alone versus directed conditions, and the extent to which they are interleaved or not.

      We have clarified this in the revised Methods. In brief, recording began when the home cage lights came on each day, with the male left to sing alone until at least 40 undirected song motifs were collected. A female was then introduced in approximately 10-minute intervals until at least 40 directed song motifs were collected. The total recording duration on a given day ranged from 0.56 to 10.27 hours, reflecting variability across birds in the time required to elicit sufficient singing in each context. We have added this information to both the Methods and relevant figure legends.

      It would be most helpful to assess the stability of waveforms and unit isolation across time.

      We now apply strict inclusion criteria based on waveform stability, as described in R3.1 above. SNR was quantified as Vpp/(2*sigma_noise), where Vpp was the peak-to-peak amplitude of each filtered spike waveform and sigma_noise was estimated from the median absolute deviation of the filtered voltage trace. This combines the peak-to-peak normalization used by Nordhausen et al. (1996) with the robust noise estimator described by Rey et al. (2015). Waveform overlays for all included example units are provided in Figure S2.

      It would be reassuring to see that significant differences between conditions are equally or more prevalent under the conditions of greatest unit isolation and recording stability.

      The average SNR of neurons ultimately included in the analysis was 9.47 +/- 3.57, with a minimum of 4.69. Neurons that exhibited significant DAF-response modulation did not have a significantly different SNR than neurons that did not exhibit significant modulation (Wilcoxon rank-sum test, p=0.38). The mean SNR for significantly modulated neurons was 8.70, compared to 9.5 for non-modulated neurons, indicating that the detection of context-dependent modulation was not systematically biased toward neurons with lower recording quality.

      One other way that the authors might be able to address the main concern would be to look at the stability of firing patterns within conditions.

      We agree that stability of firing patterns within conditions is an important consideration, and this concern directly motivated the adoption of the permutation-based analysis described above. In this framework, the observed DAF-response difference between conditions is compared to a null distribution generated by shuffling condition labels across trials. This approach inherently accounts for within-condition trial-by-trial variability and does not assume stationarity of firing rates.

      It would be helpful to have additional explanations of the criteria used for counting spikes, and assessing stability of recordings.

      Spike waveforms were visually inspected for consistency using our custom MATLAB GUI on a 12-second file basis. Interspike interval violations below 1 ms were explicitly checked as an indicator of multi-unit contamination. Detection thresholds were manually set, and each recording file included in the analysis was independently inspected. We have added a more explicit description of these procedures to the Methods section.

      For the specific examples shown in figures, it would be useful to indicate by small tick marks or otherwise which spikes were counted as single units.

      We appreciate this suggestion. In the revised figures, we have improved the clarity of the example raw voltage traces by annotating the detection threshold and, where multiple units were present on a channel, indicating the waveform amplitude range corresponding to the isolated single unit. We believe this provides sufficient transparency regarding spike identity without requiring tick marks on every individual spike, which would substantially reduce legibility of the example traces.

      What were the criteria for determining multi-unit versus single-unit activity?

      In the context of this manuscript, "multi-unit activity" refers to channels on which no single neuron could be reliably distinguished from others based on waveform shape and amplitude. Units ultimately included in the study were those for which a single, consistent waveform cluster could be identified and isolated in the custom GUI. In cases where a second distinguishable unit was present on the same channel, it was manually excluded from the sorted single-unit record. We have clarified this distinction in the Methods.

      Categorical scores: This definition results in cases where responses of 2.45 vs 2.55 are described as 'retuned', even if these responses are not significantly different. Retuning would be more persuasively demonstrated if the authors could provide a test of whether or not the responses for individual neurons differ significantly between conditions.

      We completely agree, and thank the reviewer for motivating us to develop a more rigorous statistical approach. Our revised analysis uses a non-parametric permutation test that explicitly tests for significantly different DAF responses between undirected and directed singing conditions, with correction for multiple comparisons. This replaces the previous threshold-based categorical classification and directly addresses the concern that neurons near the threshold boundary were being treated as categorically different.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Minor comments:

      (1) Please include a schematic of the brain, including the different subregions of Field L and the connections between auditory regions and the midbrain.

      Done. Figure S1 has been added, including a schematic of Field L subdivisions and auditory pathway connectivity.

      (2) The authors should include some additional information about the recordings, such as the proportion of Field L neurons that exhibited singing-related changes in firing rate. It would be helpful to include some examples of spontaneous activity when the bird is quiescent in Figs. 1-2, especially for cells that do not show firing locked to song.

      We appreciate this suggestion. Given the scope of the current revision and the primary focus on DAF-response modulation, we have elected not to add spontaneous activity examples to Figures 1-2 at this time. We agree this would be a valuable addition in future work and have noted it as a limitation in the Discussion.

      (3) Methods, p. 10: Surgery and awake-behaving electrophysiology: "The of the cannula" - this is the only mention of a cannula. Do the authors mean the ends of the probes?

      Cannula placement and wire bundle extension from the end of the cannula has been clarified in the Methods.

      (4) Bottom of p. 10: Fix reference for biorxiv paper: "ref andreas paper"

      Fixed.

      (5) Methods, p. 12: Redundant sentences regarding significant error response criteria.

      Fixed. The redundant sentences have been removed.

      Reviewer #2 (Recommendations For The Authors):

      (1) The abstract is too vaguely formulated. Authors should try to quantify the statements already in the abstract.

      We have reworded the abstract to align with the revision's more conservative claims regarding social context modulation of auditory feedback, and have added specific quantitative statements where possible.

      (2) Authors repeatedly refer to 'perceived song errors' without performing experiments or reporting on behavioral readouts of how birds perceive the jamming sounds. The wording should be changed to something more neutral, e.g. 'DAF responses'.

      We revised the manuscript throughout to use more neutral language centred on "DAF-related" or "feedback-related" responses rather than "perceived errors" or "mistakes". The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We similarly revised the abstract and all relevant passages in the Results and Discussion.

      (3) Authors write that 33 neurons were DAF responsive in both conditions. How should we interpret this overlap relative to independence and identity assumptions?

      We agree that the original presentation made the interpretation of overlap across conditions unclear. The observed overlap is greater than expected under a strict independence assumption but smaller than expected if responsiveness were identical across conditions, consistent with partial but incomplete sharing of DAF responsiveness across social contexts. In the revised manuscript, however, we have moved away from this binary classification framework because DAF responsiveness appears to vary continuously across neurons. The permutation-based analysis now directly tests for changes in DAF responsiveness across contexts without requiring categorical assignment.

      (4) Only 10 neurons were not affected by courtship state or only 10 error responsive neurons were not affected? I suggest authors do a multivariate analysis or use a mixed effect model and summarize the result as a table.

      We agree that the categorical accounting of neurons across conditions was difficult to follow in the original manuscript. In the revised manuscript, we clarified the distinction between neurons responsive to DAF within a condition and neurons exhibiting significant modulation of DAF responsiveness across conditions. We now explicitly report: "This analysis identified 71/147 neurons as DAF responsive in at least one behavioral condition, whereas 76/147 were not responsive in either condition." and "This more conservative approach identified 48/147 neurons as DAF responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (5) It would help if authors could define 'z-scored difference'. Better known is d prime, is this the same?

      For each neuron, the z-scored DAF response was computed as the z-scored firing rate difference between distorted and undistorted trials. Importantly, our revised main analysis avoids any normalization such as z-scoring, and instead uses a permutation-based approach applied directly to spike counts.

      (6) Is the 'retuning' assessment a bit conservative? Neurons could also retune by showing error scores greater than 2.5 in both conditions but a shifted response time.

      We agree that neurons could retune by shifting the latency of DAF responses. Although potential latency shifts are beyond the scope of the current study, we did observe suggestive evidence of possible latency changes in some example neurons across conditions. We have noted this as an interesting direction for future analysis.

      (7) Could the stability of DAF response across trials be described? E.g. as the ratio between intra versus inter condition variability?

      We agree that stability of DAF responses across trials is an important concern. In addition to imposing strict waveform stability requirements, our permutation-based statistical test explicitly accounts for trial-by-trial variability by constructing null distributions from within-condition trial shuffles. We have also replaced the previously shown unstable example neuron with neurons that exhibit more consistent DAF-related responses across trials, and provide additional examples in Figures 2 and 3.

      Minor:

      (8) 'significant increase in burst fraction': specify effect size of t test in results section.

      We now specify in the main text: "A small but significant increase in burst fraction was observed (paired t-test, p=9.3x10-6, n=138 neurons, mean +/- SEM: 0.11 +/- 0.006 vs 0.15 +/- 0.007, Figure 1J)."

      (9) The IMCC parameter should be specified in the main text.

      The Gaussian smoothing parameter (20 ms) has now been specified in the main text.

      (10) Fig. 2: indicate the windows within which error scores are computed.

      This is no longer applicable, as the revised permutation-based analysis does not rely on scoring error responses within a fixed window.

      (11) In Fig. 2A, the neuron has an error score of -2.54 (significant), but the red and blue curves look almost the same.

      We agree that the previous error score quantification did not always capture firing rate differences in an intuitive way. This example neuron has been replaced in the revised manuscript, and the new analysis avoids scalar error scores in favor of the permutation-based approach.

      Reviewer #3 (Recommendations For The Authors):

      Minor points:

      (1) "(ref andreas paper)." Add reference here?

      Fixed.

      (2) Hessler and Doupe 1999 is a good reference for premotor signal re-tuning during courtship.

      We agree. The reference has been included in the revised manuscript.

      (3) Page 5: "discharge depended on courtship state, using" - should this be "depending"?

      The original wording was intentional: "we tested how discharge depended on courtship state." We have verified this reads correctly in context and made no change.

      (4) Page 9: "consistent with a brainwide process" - what is meant here?

      We have revised this wording. The revised manuscript replaces "brainwide process" with clearer language describing a distributed modulation of auditory responsiveness that is not confined to a single nucleus.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study provides mechanistic evidence that tea-adapted two-spotted spider mite overcomes green tea catechin defenses via the horizontally transferred dioxygenase TkDOG15, supporting a two-step adaptation model, combining enzyme refinement and inducible upregulation. The evidence is convincing because multi-omics signals converge with functional validation (RNAi knockdown and recombinant enzyme assays) and well-controlled behavioral/toxicity assays to link TkDOG15 activity and expression to survival and feeding on tea.

      We thank the editors and reviewers for this positive assessment of the importance of our study and the strength of the evidence. We would like to point out one factual correction. The assessment describes the tea-adapted mite as the "two-spotted spider mite" (TSSM, Tetranychus urticae), but the species adapted to tea in this study is the Kanzawa spider mite (KSM, Tetranychus kanzawai). We would suggest revising "tea-adapted two-spotted spider mite" to "tea-adapted spider mite" or "tea-adapted Kanzawa spider mite" accordingly.

      Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling crossvalidation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Weaknesses:

      Although TkDOG15 is assumed to "detoxify" catechins by ring cleavage, the study doesn't identify or characterize the breakdown metabolic products. If the metabolites are indeed non-toxic compared to the parent catechins, that would strengthen the detoxification hypothesis. Also, the transcriptomic and proteomic analyses identified other potential detoxification enzymes, such as CCEs, UGTs, and ABC (Supplementary Tables 3-1 & 3-2), which were also upregulated. The manuscript focuses almost exclusively on TkDOG15, potentially overlooking a multigenic adaptation mechanism, where these other enzymes might play synergistic roles, although it was mentioned in the discussion section.

      Reviewer #1 (Recommendations for the authors):

      There is no need for additional experiments, but I suggest revising the discussion section to mention the weaknesses pointed out above.

      We thank the reviewer for the positive assessment and helpful suggestions. We have revised the Discussion (L276-283) to address both points as limitations. First, we now note that we did not characterize the products of TkDOG15-mediated catechin cleavage, and that confirming their reduced toxicity relative to the parent catechins would further support its detoxification role. We note this as a direction for future work. Second, we note that DOG15 in KSM on tea was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 and warranting functional validation.

      Minor corrections below:

      (1) Figure 1a: For better readability, I recommend adding "KSM" and "TSSM" to the two pictures, respectively.

      Done.

      (2) L165: tetur20g01790 refers to a TSSM gene, while TkDOG15 refers to a TSM protein. Revise it accordingly. (Same for L442 and L481).

      The reviewer is correct that tetur20g01790 is the TSSM gene ID. As the KSM genome is not yet available, we identified the TkDOG15 gene, the KSM ortholog of tetur20g01790, by de novo assembly of our RNA-seq reads. We have revised L169 and L455 accordingly.

      (3) L264: Supplemental Table 3-2.

      Done (L271).

      Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Weaknesses:

      While transcriptome/proteome analyses reported changes in the expression of other detoxification-related enzymes, including CCEs, UGTs, ABC transporters, DOG1, DOG4, and DOG7, it is regrettable that the contribution of each enzyme, including its interaction with TkDOG15 and the functional analysis of each enzyme within the overall catechin detoxification system, was not investigated.

      We thank the reviewer for the encouraging assessment and this comment. We agree that the contributions of the other detoxification-related enzymes, including their interaction with DOG15, remain to be investigated. As this point overlaps with a comment from Reviewer 1, we have revised the Discussion (L276-283) to note that DOG15 was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15, and that the functional analysis of their individual and combined contributions warrants future work.

      Reviewer #2 (Recommendations for the authors):

      The manuscript titled "Adaptation of an Herbivorous Arthropod to Green Tea Plants by Overcoming Catechin Defenses" presents a well-designed, mechanistically insightful study that advances our understanding of herbivore adaptation to plant chemical defenses. The work is scientifically sound and of potential interest to a broad readership in chemical ecology and evolutionary biology.

      However, before the manuscript can be considered for acceptance, the authors must adequately address the comments outlined below regarding clarity, presentation, and interpretation across the manuscript.

      We thank the reviewer for the positive evaluation of our study. We have carefully addressed each of the specific comments below regarding clarity, presentation, and interpretation, and we believe these revisions have substantially improved the manuscript.

      Specific comments on each section:

      (1) Abstract

      (a) The authors are encouraged to add a concise concluding sentence summarizing the broader significance of the study and indicating potential future research directions or limitations, which would strengthen the impact of the abstract.

      We have added a concluding sentence to the Abstract summarizing the broader significance of the study and indicating future directions (L38-40).

      (b) The authors may consider adding representative quantitative results to the abstract, as this would enhance clarity and increase the impact and interpretability of the study for readers.

      We have added representative quantitative results to the Abstract. Specifically, we now state that the mRNA and protein levels of DOG15 in tea-adapted T. kanzawai are up to 31.6 and 12.1 times higher, respectively, than in T. urticae fed on tea plants (L30-32). For consistency, we now refer to the gene as "DOG15" throughout the Abstract (L29, L30, and L36).

      (2) Introduction

      (a) While the paragraph is informative, it reads more like a summary of the main results than a statement of study objectives. The authors are encouraged to reframe this section to explicitly define the study's aims and hypotheses.

      We have reframed the final paragraph of the Introduction to explicitly state the study's aims and hypotheses rather than to summarize the results (L73-81).

      (b) The authors should avoid excessive citation of multiple references for a single thematic statement when one key reference is sufficient. Where appropriate, inclusion of more recent literature is encouraged.

      We have reduced multiple citations for single statements to the most representative references: Cabrera et al. (2006) for the health benefits of catechins (L46) and Grbić et al. (2011) and Dermauw et al. (2013) for the DOG gene count (L66-67).

      (3) Materials and Methods

      (a) The Materials and Methods section is comprehensive and technically sound; however, its length and density reduce overall clarity. The authors are encouraged to streamline descriptions of standard or well-established protocols and rely on appropriate citations where possible.

      We agree that clarity can be improved by removing redundancy. The Materials and Methods are intentionally detailed to allow independent replication of our protocols, so we have retained this detail and instead removed the overlapping methodological descriptions from the figure captions, where the same information was repeated (see our response to comment 6a).

      (b) Greater consistency is needed in reporting biological and technical replicates across different experiments (e.g., performance assays, transcriptomics, proteomics, and enzymatic activity assays) to enhance reproducibility.

      We have standardized the reporting of replicates across all experiments to the format "x independent experimental runs (n = y per run)." Throughout the manuscript, "independent experimental runs" denotes biological replicates, with technical replicates specified separately where applicable (three technical replicates for qRT-PCR).

      (c) The authors should provide brief justification for key methodological parameters, such as catechin concentrations, exclusion criteria in behavioral assays, and thresholds used for defining DEGs and DEPs, to improve transparency and interoperability.

      We have added brief justifications for the three parameters. 1) The catechin concentration range was chosen to encompass the individual catechin levels measured in fresh tea leaves (L340-341). 2) In the behavioral assays, inactive mites were excluded because their movement was insufficient to determine chemo-orientation behavior, and escaped mites were excluded because they did not complete the assay (L365-367). 3) The thresholds for DEGs and DEPs follow criteria commonly applied in mite transcriptomic studies (Vidal-Quist et al., 2025, newly added to the references) (L414-416) and are consistent with our previous spider mite proteomic analysis (Arai et al., 2025) (L444-445).

      (4) Results

      (a) While significant differences in survival and fecundity are reported, briefly indicating the magnitude of these differences (e.g., percentage or fold change) would improve clarity and strengthen the presentation (Lines 91-96).

      We have added the magnitude of the differences (L94-97). The revised text now states that after 10 days, almost 90% of tea-adapted KSM survived, compared with about 5% of non-adapted KSM and 33% of TSSM, and that tea-adapted KSM laid up to about 2 eggs/surviving female daily, whereas the other two populations laid almost no eggs.

      (b) The final sentences include interpretative and concluding statements regarding catechins as key metabolites and mite adaptation. These statements would be more appropriate for the Discussion section rather than the Results (Lines 127-130). Follow the same for the rest of the Results section also.

      Following the reviewer's suggestion, we have removed the interpretive and concluding statements from the end of the Results section, so that it now reports only the observations (L129-130). The interpretation regarding the multiple modes of action of catechins and the insensitivity of tea-adapted KSM is already presented in the Discussion (L206-213 and Conclusions), so we did not duplicate it there. We also reviewed the remaining Results subsections and confirmed that they report the experimental observations and their direct conclusions without broader interpretation.

      (c) The comparison among catechin classes is clear; however, briefly listing the mean concentrations of each catechin (as shown in Figure 2a) in the text would improve readability without duplicating the figure (Lines 135-139).

      We have added the approximate mean concentration of each catechin to the text (L136-137).

      (d) Please clarify in the Results whether the same exposure concentration and duration were applied for all catechins and mite species, or explicitly direct readers to the Methods section (Lines 141-142).

      We have clarified in the Results section that all four catechins were tested at the same concentration series (0, 10, 10<sup>2</sup>, 10<sup>3</sup>, 10<sup>4</sup>, and 10<sup>5</sup> ppm) and the same exposure duration (24 h) for both mite populations (L143).

      (e) The phrase "lower sensitivity" should be explicitly linked to LC<sub>50</sub> estimates to ensure that the basis of comparison is immediately clear to readers (Lines 143-144).

      Following the reviewer's suggestion, we have linked the sensitivity comparison to the LC<sub>50</sub> values (L143-147). The comparison is now stated relative to TSSM based on the LC<sub>50</sub> estimates, and for ECg and EC we note that the LC<sub>50</sub> of tea-adapted KSM exceeded the highest concentration tested.

      (f) This section clearly identifies TkDOG15 as a key gene underlying tea adaptation in KSM; however, the authors are encouraged to briefly clarify the criteria used to define "highly enriched" mRNAs and proteins (e.g., fold-change and statistical thresholds) in the Results text or by explicitly directing readers to the Methods. This would improve transparency and facilitate interpretation of the multi-omics comparisons (Lines 147-173).

      We have added the criteria used to define the enriched mRNAs and proteins (log2 fold change ≥ 1 with adjusted p-value < 0.05 for mRNA and p-value < 0.05 for protein) and referred readers to the Materials and Methods (L159-160).

      (g) The enzymatic comparison between TkDOG15 and TuDOG15 is well presented; however, the authors are encouraged to briefly discuss whether the two amino acid substitutions (Q127A and T203A) were individually or jointly responsible for the increased catalytic efficiency, or to acknowledge this as a limitation and potential direction for future functional studies (Lines 176-190).

      We have added a brief discussion of whether the two substitutions (Q127A and T203A) act individually or jointly (L254-257). We note that T203A is adjacent to the active-site residue Y202 and may contribute more directly to catalytic efficiency, and we acknowledge that dissecting their individual contributions by site-directed mutagenesis is a direction for future work.

      (5) Discussion

      (a) The authors appropriately acknowledge that the molecular basis of chemosensory insensitivity and the contribution of additional detoxification enzymes remain unresolved. To further improve clarity, these statements could be explicitly framed as hypotheses or future research directions to clearly distinguish them from experimentally supported mechanisms (Lines 205-208; 266-270).

      We have reframed the statements on chemosensory insensitivity (L209-213) and the contribution of additional detoxification enzymes (L272-274) as hypotheses and future directions, distinguishing them from the experimentally supported mechanisms.

      (b) While DOG15 is convincingly identified as a key contributor to tea adaptation, a brief clarification of its relative importance compared with other upregulated detoxification enzymes would strengthen interpretative balance, even if the roles of these enzymes remain unresolved (Lines 259-265).

      DOG15 was the only enzyme upregulated at both the mRNA and protein levels (Figure 3d,e), and the only enzyme functionally validated in this study, by RNAi silencing (Figure 3f) and recombinant enzyme assays (Figure 4c). We have established that DOG15 contributes to tea adaptation, but because the other upregulated enzymes were not functionally tested, their relative contributions cannot be determined at this stage. As we note in the Discussion, tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 (L277-283). We therefore did not add further text, to avoid duplication.

      (c) The discussion linking host plant adaptation to reproductive isolation and ecological speciation is interesting and well contextualized; however, these evolutionary implications should be slightly tempered or explicitly framed as potential long-term outcomes beyond the immediate scope of the present study (Lines 271-281).

      We have tempered the evolutionary implications (L292-294). The revised sentence now frames the link to reproductive isolation and ecological speciation as a potential outcome over longer evolutionary timescales rather than a direct finding of the present study.

      (6) Figure captions

      (a) The figure captions (Figures 1-4) are exceptionally detailed and, in several places, repeat methodological information already described in the Materials and Methods. The authors are encouraged to shorten the captions by retaining only information necessary to interpret the figures, while referring readers to the Methods for experimental details.

      We have shortened the figure captions (Figures 1-4) by removing methodological details that are described in the Materials and Methods, retaining only the information needed to interpret each figure. Where appropriate, readers are now referred to the Materials and Methods or to Supplemental Figure 1-2 for the full experimental procedures.

      (b) Several captions contain long, multi-sentence descriptions that may hinder readability. The authors may consider simplifying the wording, grouping related panels more concisely, and removing procedural details (e.g., extraction conditions, exposure durations, and instrument settings) to improve clarity and visual accessibility.

      As described in our response to comment 6a, we have simplified the figure captions by removing procedural details such as extraction conditions, exposure durations, and instrument settings, and by grouping related panels more concisely. These details are retained in the Materials and Methods.

      (c) In Figure 1, the panel labels (a-h) do not appear in a clear sequential order. For consistency with the other figures and to improve readability, the authors should ensure that panel lettering is arranged in a logical, sequential order throughout the manuscript.

      We appreciate the reviewer's attention to panel ordering. In the current layout, the panel lettering follows the order in which the panels are first cited in the text. Arranging the panels in a strict left-to-right, top-to-bottom sequence would require reducing the size of several panels, including the HPLC chromatogram in panel (e) and the survival and fecundity time courses in panels (c) and (d), which would compromise their readability. We have therefore retained the current arrangement, in which related panels are grouped together and the larger panels are kept at a legible size. We hope the reviewer finds this acceptable.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary: 

      The manuscript by Yang, Wang, and Cléry presents a lightweight pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. Overall, the authors offer a framework for identity tracking that prioritizes real-time inference. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification. 

      Strengths: 

      (1) The proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. 

      (2) Evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. 

      Weaknesses: 

      (1) The pipeline's reliance on controlled animal isolation and small visual markers raises questions about the approach's generalizability to unconstrained multi-animal environments. The provided confusion matrices (Figures 6-8) indicate that the most common misclassifications are background-related, possibly suggesting that detection specificity is the primary source of error. All things considered, these findings raise concerns about performance in its use in socially dynamic and visually complex environments. 

      Thank you for the comment. The background column of the confusion matrix can be explained by several occasions: a) the model detects an object where there is no object, b) there is more than one prediction label for the same object, or c) an object appeared in the image but not manually labeled, however the program was able to detect that object. The value of the background column does not necessarily mean that the detection is incorrect, as the precision score for the detection labels are good. We have rephrased the relevant sections for clarification to include the sources of the increased value in background columns in confusion matrices, as follows:

      “The background class of the confusion matrix showed frequently predictions as marmoset faces and collar beads for the training (Figure 6A) and validation set (Figure 6B). However, it does not necessarily indicate incorrect predictions or misclassifications. Instead, these values were mostly explained by multiple detections of the same object class. For instance, additional marmoset faces were predicted when multiple animals were present within a single video frame. The long collar structure or motion blur of the marmosets could also cause multiple detections of beads that belong to the same collar. This also corresponded to the high precision and recall scores observed across prediction classes (Figure 5D), suggesting that the increased background false positives were mainly related to the object-count discrepancies, instead of poor detection performance.”

      Prediction misclassification is one source of the background false positive. The misclassification could not be avoided in automatic prediction algorithm, but we included the manual filtering and majority-voting during our real-time classification to reduce this effect. Multiple detection of the same class may also be considered as the background, since only one object may be labelled in the ground truth, such as multiple collar beads or automatic face extraction. In addition, blurry objects were not labeled manually during training but can be detected during prediction, which also resulted in background false positive. It was clarified in the main text as follows:

      “The normalized confusion matrices showed high accuracy and consistency of most marmoset faces and collars detection in training (Figure 8A) and validation (Figure 8B) tests, with some exceptions. Particularly, the background was frequently identified as the collar of Young2 marmoset. This elevated background score was likely contributed by the multi-color design of the Young2 marmoset collar, making it more difficult to distinguish compared to collars with a single bead color. In this occasion, if one bead is occluded, blurred, or outside the field of view, the other visible collar bead could affect the prediction and lead to an incorrect identification from the ground truth.”

      (2) The manuscript claims performance comparable to that of human experimenters but provides no explicit evidence to support these claims. While it is plausible that human experimenters may be less accurate in facial recognition tasks involving closely related marmosets, the authors don't provide evidence. Moreover, while that might be the case, the color-coded beads provide a salient identity cue for the model, which complicates the interpretation of this comparison grounded in facial recognition. 

      Thank you for pointing out this concern. The aim of the facial recognition tool is to collect data from marmosets without having experimenters to check the identity continuously. The program is not aimed at outperforming the experimenters’ role but avoid having constant human intervention that can disrupt a more ecological in cage data collection. It is also essential for having more flexibility to collect data in case a specific experimenter is not here and thus to not disrupt the project. Human experimenters have extensive experience closely working with marmosets, having the unique collar beads associated to each marmosets allows human experimenters to hardly make mistakes identifying marmosets and to do it quickly. We collected identification accuracy of human experimenters by presenting 10 clips of the five marmosets involved in the manuscript (2 clips per marmosets), with 2 random clips repeated twice. The results were plotted by each experimenter. The identification accuracy of the experimenters correlates with the time spent with the animals, as the animal health technicians (responsible for daily health check and husbandry) achieved 95.83% average accuracy in identifying the marmosets. We clarified those points in the main text as follows:

      “Its automated pipeline substantially reduces the time and work required for traditional manual identity labeling, while maintaining an expert-level human performance and reproducibility across experimenters (95.83% average accuracy for animal health technicians, responsible for daily health check and husbandry while lab experiments ranges between 25 to 80% of accuracy depending on the amount of time spent with each animal, Supplementary figure 1). The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      We filtered the prediction of the collars, and the identification result solely based on the faces for the 2 young marmosets was correct. The prediction results were plotted on Video 7, Video 7—video supplement 1, and Video 7—video supplement 2 and added to the Results section as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #2 (Public review):

      Summary: 

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting. 

      Strengths: 

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieve human level performance, both in terms of precision and recall of the trained models.

      Weaknesses: 

      The robustness of the system is not clear from the limited datasets presented. The use of color-coded beads undercuts the study's premise that the system achieves truly non-invasive tracking. Although the system achieves good performance in face detection, it does not perform as well for classification using faces alone (especially when the faces are similar, as in twin animals). Here, too, the color-coded beads play a key role in identity discrimination. The stated goals of the study and the actual results presented are therefore at odds.

      Thank you for the comment. First, we would like to clarify the role of the collar beads in our system. Compared to the faces, a unique identity marker, the collar beads were not used as the main identity classifier but rather as an external visual marker. The color-coded bead was not used solely for the purpose of marmoset video classification; it was also used as an additional source of identification for one marmoset. As the marmosets usually move very fast inside the cage, it is mainly used as a visual marker for experimenters to recognize them in a distance in a short time.

      The mislabelling is more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for the model training, as discussed in the paragraph #4 of the Discussion section. Collar beads are small and less frequently detected by the camera, since it could be occluded by the marmoset fur. In addition, it was invisible to the camera if the marmoset turned sideways or was far from the camera. Therefore, higher weight was assigned to the beads due to their small size and less frequent detections compared to face labels, such that it was only an element to confirm the identity, instead of the main classifier.

      The inclusion of the collar beads doesn’t affect the prediction results of the marmoset faces. The model achieved a good precision/recall score for the identity labeling in the manuscript. In the revision, with the majority-vote strategy, we filtered the detection of all collar beads and showed that the model was able to correctly identify the marmosets solely by their faces. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #3 (Public review):

      Summary: 

      In this manuscript, Yang et al introduce a new method for automatically identifying marmosets in their home cage using a supervised deep learning method that recognizes the face and colored beads on marmoset collars. The authors show a high precision rate of identifying marmosets to levels comparable to a human experimenter. The method overall seems robust at identifying marmosets at different life stages and different settings; however, given the current form, I'm struggling to see the generalizability and experimental utility of this method. 

      Strengths: 

      (1) The authors provide a near-perfect automatic identification of marmosets in their home cage. 

      (2) This method is robust across lightning, camera angles, etc., making it potentially useful for marmoset (and other NHP) identification outside the housing cage as well 

      Weaknesses: 

      (1) Despite the almost perfect precision, in its current form, I'm failing to see how this method can be useful to other labs. 

      Thank you for your comment. This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Future work will focus on extending the program application on identification from different housing conditions, in combination with various behavioral tasks such as in-cage touchscreen system or manual tasks, and in the wild that precludes handling or isolation of marmosets for collecting behavioral data. The program solely requires a camera, a computing device, and marmosets, as there are no hardware restrictions. In addition, we are currently collaborating with other labs on the marmoset identification from videos taken from other setups. The program achieved effective face extraction from the marmoset in the video, without the need for additional program modifications.

      (2) This is a nice methods manuscript, but the authors do not present results to show how their method can be used outside of identifying marmosets inside their home cages in a small field of view. 

      Thank you for your feedback. The method developed was applied in combination with other touchscreen behavioral tasks, aiming to extract data without human intervention. This approach was discussed in paragraph #6 in the Discussion section. While this manuscript focuses on the methods of close-view face identification when marmosets perform behavioral tasks, the identification and automatic face extraction program could also be applied to marmoset videos taken from a larger view, including phone cameras. Even though the marmosets are still housed in their home cage, the example videos presented the program’s application in a larger field of view. We have added examples of the videos/photos from a different experimental setup to respond to this comment in the Discussion section as follows:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10).”

      (3) Reading the manuscript is strenuous, given its repetitive nature. Consolidating and shortening the results, as well as adding some definitions to the results section, would be helpful. 

      Thank you for pointing this out and your suggestions. We have rephrased the Results section for simplification to facilitate the understanding of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The weight of the color-coded beads was increased to improve identification accuracy. From a brief look at the code provided on GitHub, the weight assigned to the beads seems substantial. This calls into question the need to use facial recognition in the identification strategy. As the code currently stands, facial identity appears to serve primarily as a fallback when bead detection fails to register. To strengthen the methodological justification, the paper would benefit from the authors providing a rationale for choosing this weighting scheme and, if available, supplemental figures showing performance across a range of different weights to demonstrate why that specific value was assigned in the algorithm.

      We have added the model prediction results without the collar beads showing that the facial recognition algorithm works even without the collar beads and that those collar beads are not the main classifier. The Methods section has been modified as follows:

      “For each detected bounding box, the scripts returned a corresponding label of marmoset face and collar bead color. We assigned the detected collar beads as the corresponding marmoset identity with a higher weight, which improved the detection confidence across frames.”

      The Discussion section has been modified as follows:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      The weight assigned to the beads in the GitHub page is the highest weight that we would suggest. The actual weight can be customized by the experimenters based on the actual experimental setup. For example, we used the weight of 2 in our real-time version of the marmoset face identification, while marmosets were presented with their corresponding tasks once identified. The GitHub page has been edited to clarify this point.

      (2) The overall utility of this approach, other than the real-time detection component, needs more clarification. It is currently unclear why this approach, and in which specific experimental or observational settings, is particularly advantageous compared to existing methods for assigning animal identity.

      In addition to the advantages mentioned in paragraph #1 of the Introduction and paragraphs #1-3 in the Discussion, we have added more details in the Discussion section:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (3) Although it appears that performance based on faces and color beads was evaluated separately, this was not clearly presented, leading to confusion about whether face detection performance also benefited from color beads on the animals.

      Prediction of the different labels in the same model is independent, so the prediction of color beads is not affecting the prediction results of marmoset faces. Correct identity classification could be achieved without depending on the color beads, as we have filtered out the color beads detection class. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Reviewer #2 (Recommendations for the authors):

      (1) I found the paper quite confusing as written. The term "model" is overused and highly conflated: there are the YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model. The flowchart in Figure 2A is equally confusing. The mapping from the flowchart to the results is not straightforward, and I needed several passes to grasp it. I would recommend that the authors simplify the terminology and the mapping of the methods to the results.

      Thank you for pointing this out! The YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model were indeed separately trained object detection models. They all have different weights/parameters but share the same YOLOv8 architecture/backbone. We removed some of the “model” term in the manuscript and replaced them with “classifier/framework/pipeline” to avoid misunderstanding. This information has been clarified in the revised manuscript of the Methods section and Figure 2, which provides an overall clearer explanation of the workflow of the methodology of the program.

      (2) It is not clear how robust these results are, given the limited data sets analysed.

      We agree that only five marmosets were involved in this manuscript, this unfortunately limited the robustness of the prediction results. Indeed, the limited number of animals that can be used per study has been a main limitation in non-human primate research, as they are very valuable animal models. However, we included approximately 3400 images in the training dataset, which were collected across days. New videos and photos that were captured from different devices were also used in the testing to ensure that the program can be used on new marmosets, different housing cage, and from different recording devices as indicated here:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”

      (3) There are two paradoxes regarding the stated motives of the study:

      (a) If the objective was to truly use non-invasive methods for the identification of animals, then why use the color-coated beads?

      As mentioned previously, identity detection can be made without collar beads, still with correct prediction results as indicated here:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The color-coated beads are used for easier and quick marmoset identification during daily care, health check, for weekend staff, training or handling.

      (b) If the objective was to achieve high identification performance, and the color-coated beads are sufficient for this purpose, then why bother with faces at all?

      Collar beads are small compared to the face, and less visible due to fur occlusion and motion blur. Moreover, it is possible that some marmosets do not have collar beads due to their young age or when involved in other procedures such as imaging scans. The collars need to be checked and changed regularly in growing marmosets and it is not always convenient (some marmosets do not support the collar, some can have sensitive skin that would lead to abrasion) thus the need to develop a facial recognition system. Furthermore, marmosets who are from other labs or in the wild might not wear a collar with colored beads, thus face is the main classifier in this model to be more generally applicable. It is highlighted here:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      (4) I was puzzled by the face similarity results in Figure 9. It appears that the face similarity measures were stronger (higher cosine similarity, lower Euclidean distance) for the adult data set compared to the twin data set. If so, why was it more challenging for the system to handle the twin data set?

      Face similarity can only be compared within models (therefore within adults and within twins). As this is calculated from different models, the adult face similarity cannot be compared with twins’ face similarity. It has been clarified in the Methods section as follows:

      “Statistical tests were performed only within the face classifier of each marmoset family, as embedding spaces may vary in scaling, learned features, and baseline metrics making cross-model comparison of inter-individual face similarity unreliable (Bollegala, 2017).”

      And in the Results section as follows: “We performed the statistical tests only on the face classifier for the adult marmoset family, as the twin marmoset model only involved two individuals and thus not valid for within-model statistical analysis (Table 2, 3).”

      The twin dataset aims to represent a test for the program utility in new marmosets, especially for testing if the program can still distinguish between the marmosets with similar faces. Thus, the number of twin data collected is less than the adult dataset, as explained by Discussion paragraph #2 “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.” This explains the challenge the system faces when differentiating the twins, while increasing the training dataset is required to solve this issue.

      Reviewer #3 (Recommendations for the authors):

      Major issues:

      (1) My main issue is regarding the utility of this method in scientific experiments. This manuscript is a "methods paper" introducing a face recognition method to identify a single marmoset in their home cage in a very specific and confined field of view. This comprises a limitation on what experiments can be performed using this method. On the contrary, if (a) the authors can show that this method can be used for a bigger field of view, where the social structure/interactions can be studied for neuroethological, cognitive or social studies that will make this method significantly more robust; or (b) design an experiment that can be performed using the current method to show that this method in its current form is sufficient.

      (a) Our method worked in larger home cage (larger view) with videos taken inside the cage / outside the cage, with multiple marmosets moving around, while the camera and its fixation are also moving. A new figure (Figure 10) has been added to highlight this wide application:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”.

      (b) We are currently using this method to collect in-cage touchscreen data with multiple marmosets without the need to isolate such animals to acquire the data, avoiding social separation. The collection of data in nonhuman primates is still a long process, so we wanted to share the facial recognition system first, aligned with our commitment towards open science, to benefit the broader community (we have already been contacted by two labs since the publication of this preprint) while we keep collecting data for the scientific project. We have added the touchscreen application as example in the Discussion section as follows:

      “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      (2) The authors claim a longitudinal identification of marmosets, yet I think the data to fully support this are deficient. This might be a result of unclarity of this experiment. How was this experiment done? Was the training done on the 7 months and then applied to the 11 months? Are there more continuous data that track the precision of the identification in time? For example, how does the twin identification evolve in time?

      This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Ongoing work in the lab, the main research focus of which is the longitudinal assessment of cognitive functions, either during neurodevelopment or in preclinical ageing model, is benefiting from such algorithms to help identifying the animals to collect in cage behavioral data. As such, we have done some testing in one young cohort. The training of the young marmosets’ identification was done only on the 7-month data, and then we applied the identification program to the videos of the same marmosets when they were 11 months old and 16 months old (for the no-collar results) as indicated as follows in the Methods section:

      “Moreover, we evaluated the model performance and its generalization across developmental stages using new videos: 1) from the adult marmosets and 2) from the same young marmosets at 11 months, which were not involved during initial program training”.

      And in the Results section as follows: “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The identification program was shown to correctly identify the marmosets; however, we found that “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.”

      The mislabeling was more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for model training. The face images used in the identification model training were less compared to the adult model, which contributed to a less accurate prediction result. As the marmoset is still developing before adulthood, their face features will become more different as they age. By increasing the number of training images from different ages of the young marmosets, this could be solved as it is therefore possible to build efficient identification program for longitudinal study. Thus, instead of the current classifiers presented in this manuscript, we suggested that the method/tool could be beneficial for longitudinal studies, not restricting to the individual-based identification program mentioned in this manuscript, as “The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      (3) How does this method compare to other methods that were used in the past?

      The advantages were mentioned in paragraph #1 of the Introduction and paragraph #1-3 in the Discussion. Current approaches for marmosets are usually visible markers (ear dye, collar, etc.), RFID, or observation, of which manual works and human interventions are required. These methods usually need continuous adjustment due to tighter collar, dye fading, etc. This can affect marmoset behaviors, especially during their behavioral task performance, as mentioned as follows:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (4) Did the authors think of adding a continuity or a space constraint? For example, video 6 shows misidentification of the twins; in this specific case, adding a probabilistic continuity or space constraint that will limit identity switches might be useful. This can also be using a retroactive correction - for example, video 3.

      We would like to thank the reviewer for this suggestion. We agree that these approaches will be valuable improvements for future offline analysis.

      The probabilistic continuity constraint can indeed help decrease identity switches. In our current application, we have implemented a temporal smoothing through majority voting across a 30-frame (1 second) window, of which the program outputs the most frequent prediction of identity. With this strategy, we could reduce the occasional frame misprediction and maintain the real-time performance. Our animals are free-moving and may appear in any location within the camera field of view and housing cage. Therefore, position is not strongly associated with the identity of individual.

      We agree that retroactive correction could improve the detection consistency for offline analysis by correcting past detection by future prediction results. However, the current pipeline is incorporated with behavioral tasks, meaning that the prediction results aim to be transmitted with minimal time delay. As additional frame analysis and extra computational power may be needed for retroactive correction, the increased latency can be limited to the utility of real-time system and task control.

      Minor issues:

      (5) In its current form, I think the manuscript can be significantly shortened and the results/figures can be consolidated (confusion matrices with validation figures for example).

      We have followed the reviewer’s suggestion and shortened the Methods and Results sections.

      (6) The term "unseen" that the authors use in their results is confusing. Are the authors referring to monkeys that are hidden from their view, or "unseen" before by the model? The video indicates the latter, but I think the term can be changed to something less confusing, like novel, new, etc.

      We have changed the term “unseen” by “new” to avoid confusion.

      (7) Can the authors add information about the relationship between the number of manually labeled images and the identification precision?

      The relationship between number of manually labeled images and identification precision has been described in the Discussion section as: “Moreover, the performance of the system is strongly dependent on the amount and variability of the training data, with identity classification improving as more marmoset images are involved in the model training.”

      This means that more manually labeled images (i.e. larger training datasets) could improve the identification precision. However, model performance will plateau regardless of training dataset size, referred in the Results: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).”

      More manually labeled images could help improve the variability of model prediction, but too many of these images are also risky for overfitting. In this case, overfitted model might not be able to make valid predictions on new videos/images.

      (8) Can the authors expand a bit about the difference between YOLO Nano, small, and medium in the methods?

      Thank you for the suggestion. We have added a brief description of the pre-trained models in the methods: “These pre-trained models share the same object detection backbone but differ in number of parameters and computing power. Larger models, such as YOLOv8 medium, provide higher detection accuracy but require greater computational resources and longer inference time. In contrast, smaller models prioritize the computational efficiency.”

      (9) Clear and short definitions of what IoU, Recall, F1, and other terms represent should be added to the results section (not formulas, short sentences).

      The definitions and formula for the evaluation metrics were described in detail in the Methods section. To help readers while avoiding repetitions with the detailed methodology definitions, we have added a brief description of these terms at their first mention in the Results section: “Model performance included the precision (the proportion of correct positive predictions), recall (the proportion of corrected predicted ground-truth labels), and mAP@50–95 (the average detection accuracy across different IoU object localization thresholds; see Methods and Materials section for detailed definitions).”.

      (10) In the methods-"video collection" section, can the authors please include more information? Is this a motion-sensitive camera? Otherwise, what's the size of the data that is collected? This will help in reproducibility and system requirements. If this is not continuously collected data, discuss what can be done to make an online identification tool.

      We have included the camera as industrial color/RGB camera, which functions like any webcam and has no motion-sensitive functions. The size of data collected was described in detail in the Methods section that we slightly modified for clarity for: “Three adult marmosets from one family were recorded for 1 hour across each of the 5 recording days, with unrestricted voluntary access to the primate chair space. For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C). Two young marmosets were briefly isolated and recorded separately for testing and improving the automatic face extraction program. We recorded them at two developmental time points, 7 months old and 11 months old (an additional time point at 16 months old has been added for one marmoset to test the identification without collar). During video collection, a sliding panel and an in-cage box were positioned near the housing cage door to temporarily isolate individual marmosets from other family members. Individual isolation was kept brief (approximately 10 minutes) to prevent disturbance and potential stress due to family separation. “To capture sufficient variability in postures, individuals, and lighting conditions, clips were sampled throughout the adult marmoset videos (approximately 5 hours) (Figure 2B).”

      The size of the training dataset was also described in detail in the Methods section, referred as: “To minimize image computations and data storage, we created a dataset of 2498 annotated images from the three adult marmosets. All images were manually annotated to label marmoset faces, individual identities, and the collar bead colors (Figure 2C). The annotated images were used for training models of multi-marmoset face classification and the automatic identity extraction, which can automatically detect, localize, and identify marmoset faces (Figure 2A, Step 1 – 4). We created another dataset of two young marmosets at 7 months old (total images = 502) for testing the automatic facial and identity extraction (Figure 2A, Step 4 – 5). For both adult and young marmoset datasets, images were randomly divided into a training set and a validation set at a ratio of 8:2.”

      Our manuscript is not describing an online identification tool (i.e. the described program does not require connection to internet). Instead, once trained and the program is performing well with new marmoset videos, we could use the trained weights for real-time marmoset identification (no need to collect new training data) as we are doing it for our touchscreen data collection in cage. It is referred in the Discussion as follows: “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      This program can be used for online applications if you are using a camera that is connected 24/7. For now, we are only using it while using our behavioral testing chair due to limitations issues (safety recordings, removing all electrical apparatus during night, overheating of camera if use continuously).

      (11) Would increasing the size of the beads help with their identification? In the images included in Figure 2C, it's very difficult to see these beads.

      Yes, collar beads can be occluded by fur, blurred by motions, or outside the field of camera view. We have now included increasing the collar beads size in the Discussion, referred as: “An alternate experimental solution is to improve collar visibility, including using distinct color code across individuals within a family, increasing the size of the beads, or increasing collar beads number to reduce occlusion.” However, the size of the beads needs to be appropriate to avoid being inconvenient and disruptive to the animals to ensure their welfare.

      (12) It's difficult to understand the setup of the camera in regard to the housing cage. Can the marmosets go into the primate's chair at any point (from the videos, it seems so, but the dashed line in 1C might indicate otherwise), or is the primate chair there only to mount the camera? Consider redrawing 1C in a more clear way.

      The figure 1C is a simplified drawing of the photo 1A, we have edited the drawing of Figure 1C to highlight the free access when the chair is mounted to the cage as stated here: “For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C).”

      (13) Given the repetitiveness of the figures, an icon atop each figure specifying what's being tested will be helpful.

      We have added an icon for each figure 3-8.

      (14) I feel the supplemental figures for Figure 9 are more compelling than the main figure. Consider including some of the panels in the main figure.

      Thank you for the suggestion. We included heat maps of the adult family relationships in Figure 9, including the cosine similarities and Euclidean distances. The current legend of Figure 9 is changed as follows: “Figures 9. Across-model visualization of the face similarity between marmoset pairs. Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the 3 relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) and 1 (exactly same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly same) and 1 (very different) for marmoset faces.”

      (15) Related to this, I feel that the display of Figure 9 obscures the differences that the authors report.

      The display of Figure 9 has been improved following the above reviewer’s suggestions.

      (16) In Figure 9, given the large effect sizes but non-significant p-values, will adding more training points/epochs improve the differences.

      For the trained recognition programs, we have already implemented the “early stopping criteria” as shown in the Results section: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).” This means that the program has already plateau with its performance.

      (17) Line 30: either "within" or "in".

      This has been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how distinct honey bee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. However, some of the stronger mechanistic conclusions regarding direct regulation of octopamine signaling remain incomplete without more specific validation of receptor-level effects and direct quantification of octopamine levels or signaling activity.

      We revised some of text in the manuscript, since we agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence—including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. The revised text better reflects the data included in this paper.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honey bees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honey bees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study is the high prevalence of preexisting background DWV and SBV infections in the honey bee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      We thank Reviewer #1 for these comments.

      (1) The authors utilize Epinastine to block octopamine signaling, describing it as a highly specific OA receptor antagonist. However, pharmacological inhibitors often lack absolute specificity. Epinastine might bind to other octopamine receptor subtypes present in honey bee neural and flight muscle tissues, or it could potentially cross-react with tyramine and dopamine receptors. Without further genetic validation (e.g., RNA interference targeting specific receptors), it is difficult to definitively conclude that the altered flight performance is solely due to the blockade of the specific Oβ−2R pathway.

      We thank the reviewer for this thoughtful comment and agree that pharmacological approaches have inherent limitations with respect to receptor specificity. However, among the available octopamine receptor antagonists, epinastine is considered one of the most selective compounds for insect octopamine receptors. Roeder et al. (1998) reported that epinastine exhibits affinities for octopamine receptors that are at least four orders of magnitude greater than those for other insect biogenic amine receptors, including dopamine, tyramine, histamine, and serotonin receptors. We updated the text to include this information.

      Honey bees encode four β-adrenergic-like receptors (AmOARβ1- AmOARβ4) and one αadrenergic-like receptor (AmOARα1). Our transcriptomic analyses indicated that expression of AmOARβ2 was substantially higher than that of other octopamine receptor genes. Specifically, AmOARβ4 transcripts were not detected in our RNA-seq datasets, while AmOARβ1 and AmOARβ3 were expressed at very low levels in most samples (Supplementary Table S9; Figure S5). Although AmOARα1 transcripts were detected in some samples, expression levels were consistently lower than those of AmOARβ2. These observations support the interpretation that the physiological effects observed following epinastine treatment are primarily mediated through disruption of AmOARβ2 signaling. We updated the text to include this information.

      We agree that receptor-specific genetic approaches would provide valuable complementary evidence. RNAi-mediated knockdown of AmOARβ2 is an attractive future direction; however, RNAi efficacy in honey bees is variable and influenced by factors including transcript turnover rates. In addition, dsRNA treatments can induce sequence-independent antiviral effects that could confound interpretation in studies involving viral infection (Flenniken and Andino PONE 2013; Brutscher, Daughenbaugh, and Flenniken Sci Reports 2017). We have revised the manuscript to more explicitly acknowledge these limitations and to clarify the basis for our interpretation of the epinastine experiments.

      (2) As a natural neurotransmitter, insects have evolved highly efficient "cleanup" mechanisms. OA is rapidly cleared from the synaptic cleft via reuptake transporters and quickly inactivated by enzymes such as N-acetyltransferase (NAT) or Monoamine Oxidase (MAO). Consequently, an injection of OA produces only a transient "pulse" of activity. It is often a poor "tool" for inducing prolonged physiological effects compared to synthetic formamidines like Amitraz.

      We thank the reviewer for this important point regarding the pharmacokinetics of octopamine. We agree that octopamine is rapidly metabolised and cleared under physiological conditions and that exogenous administration is unlikely to precisely mimic endogenous signaling dynamics. Our goal was not to induce a prolonged pharmacological activation of octopamine signaling comparable to that produced by synthetic agonists such as amitraz, but rather to determine whether increasing octopaminergic signaling could mitigate the flight impairments associated with DWV infection. Octopamine was administered either by injection or through feeding (Lines 86-89), both of which resulted in significant improvements in flight performance in DWV-infected bees (Figure 2). The observation that two independent delivery methods produced similar outcomes supports the conclusion that enhanced octopaminergic signaling can partially rescue the DWV-associated flight phenotype. We have revised the manuscript to clarify this distinction and to acknowledge that exogenous octopamine administration likely produces transient elevations in signaling rather than sustained receptor activation.

      (3) The study relies heavily on transcriptomics and quantitative PCR to measure the mRNA expression of key synthesizing enzymes, namely tyrosine decarboxylase (tdc) and tyramine βhydroxylase (tβh), to infer the activation or suppression of the octopamine pathway. However, changes in enzyme synthesis at the RNA level are often insufficient to accurately reflect the true physiological levels of biogenic amines. To robustly prove the authors' hypothesis of a "feedback loop that regulates intracellular OA concentrations", direct quantification of actual octopamine and tyramine titers in the bees (e.g., using high-performance liquid chromatography or mass spectrometry) is necessary.

      We thank the reviewer for this comment and agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. Previous studies have successfully quantified OA in honey bees using HPLC-based approaches, including KayaZee et al. (2022, eLife), who measured OA in honey bee muscle tissue (both naturally occurring levels and levels post-treatment with 10 mM OA), and Cook et al. (2017, J. Exp. Bio) who quantified OA in pooled honey bee brain samples.

      Prior to submission, we inquired with our institutional mass spectrometry facility regarding the feasibility of measuring OA in individual honey bee samples. The expected concentrations of OA in our samples was below their limit of detection, so we did not pursue these analyses at that time. During the review process, we explored the possibility of analyzing a subset of samples at external facilities that may have the sensitivity required to quantify OA and tyramine in honey bee tissues. Since such analyses would require substantial resources, with estimated costs of approximately $5,000–10,000 for 12–15 samples that have been stored in the -80C since the study, rather than flash-frozen in liquid nitrogen as described by Zee et al. 2022. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence— including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. We thank the reviewer for this valuable suggestion. While these analyses are beyond the scope of this study, we will consider using this approach in future studies.

      Reviewer #2 (Public review):

      Summary:

      This highly original and well-designed study provides insight into how honeybee picorna-like viruses, Deformed wing virus (DWV) and Sacbrood virus (SBV), affect flight performance, and reveals the role of the octopamine (OA) pathway in virus-honeybee interactions. The authors used a flight mill to quantify the flight performance of bees with different levels of DWV and SBV. Bees were treated with OA and/or epinastine (EP) - an OA receptor antagonist; the study also quantified virus loads and expression of two key genes involved in OA biosynthesis.

      The results showed that reduced flight performance associated with high DWV levels could be alleviated by OA administration. In contrast, increased levels of SBV had the opposite effect, leading to enhanced flight performance. This suggests distinct physiological responses to DWV and SBV infections. Administration of EP had led to a reduction of flight performance in SBVinfected bees, indicating the involvement of the OA pathway.

      The authors also quantified levels of mRNAs of enzymes involved in OA synthesis, tyrosine decarboxylase (TDC) and tyramine beta-hydroxylase (TbH), and concluded that DWV induced expression of TbH, while SBV upregulated expression of TDC. Furthermore, the study identified upregulated and downregulated genes in response to SBV, DWV and DWV in combination with OA.

      Strengths:

      The study reported opposing effects of infections of related viruses, SBV and DWV, on honeybee flight performance, and identified the central role of the octopamine (OA) signaling pathway in the effect of viruses on honeybee flights.

      These findings were achieved by using a combination of approaches, including experimental measurement of flight distance, virus infections, and introduction of OA and EP. Experimental work with honeybees is technically challenging and requires specialized expertise, which makes the results produced in this study more valuable.

      DWV and SBV are among the most important honeybee pathogens affecting honeybee health and threatening the pollination service. Therefore, an understanding of the mechanisms underlying DWV and SBV pathogenesis has the potential to develop novel approaches to mitigate the negative impact of these viruses.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank Reviewer #2 for these comments

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions for the manuscript.

      (1) L. 45-46

      Please note that not only high virus levels have a negative impact on honeybees. Low levels of DWV, typical of covert infections, can have long-term deleterious effects on honeybee foraging and survival. Please include citation (e.g., Benaets et al, 2017, Proc Biol Sci (2017) 284 (1848): 20162149. https://doi.org/10.1098/rspb.2016.2149).

      We thank the reviewer for these comments and edited the text accordingly, and apologize for our inadvertent omission of Benaets et al 2017, which is cited in our previous publication.

      (2) L. 113

      Clarify what is meant by "high DWV levels"

      "i.e., 10^8 copies / 2 ug RNA" -> "i.e., above 10^8 copies / 2 ug RNA"

      We thank the reviewer for this comment and corrected this in the text.

      (3) L.115

      "..mock infected bees.." /Figure 2A.

      Did these bees have low levels of DWV, below 10^8 / 2 mg RNA? What was the level of DWV in these bees?

      Mock-infected bees had an average of 3x10<sup>5</sup> DWV copies and 2x10<sup>3</sup> SBV copies per 2 µg RNA (reported in Lines 107-109 in revised manuscript, a few lines before in original manuscript).

      (4) Figure 2 / Legends to Figure 2

      Note that in Figure 2 legends, the grey areas show 95% confidence intervals for regression lines.

      We thank the reviewer for this comment and added this in the text.

      (5) Figure 2 / Legends to Figure 2

      Consider including correlation coefficients (R) and p-values for each of the regression lines in Figures 2A-F. (These could be included in the Figure 2 legends).

      We thank the reviewer for this suggestion and agree that providing sufficient statistical information is important for data interpretation. Because the analyses presented in Figure 2 are based on linear mixed-effects models that incorporate both fixed and random effects, the statistical outputs are more complex than those associated with simple linear regressions. For figure clarity, we chose not to include all model statistics within the figure panels or legends and the key statistical results, including p-values and model fit metrics (R<sup>2</sup> values), are reported in the main text (Lines 121+). In addition, complete model outputs, including all relevant coefficients, correlation estimates, and associated statistics, are provided in Supplemental Data Sheet S4. To address the Reviewer’s comments, we revised the figure caption to improve clarity and include key p-values.

      We believe this approach better balances accessibility in the main figures with comprehensive reporting of the statistical analyses and thank the Reviewer for this useful suggestion.

      (7) L.357-377 - virus-specific responses

      A previous honeybee transcriptome analysis study, which showed different responses to DWV and SBV, could be cited (Ryabov E. 2016. PeerJ 4:e1591 https://doi.org/10.7717/peerj.1591).

      We thank the reviewer for this point and included this citation in line 338 of original manuscript (line 349 in revised, tracked-changes manuscript).

      (7) L. 412

      "bees were collected 24 hours prior to eclosion" -> e.g. "bees were collected at pupal stage 24 hours prior to eclosion"?

      Specify if dark-eyed pupae were collected to make sure eclosion in 24 hr.

      We thank the reviewer for making this point, and we revised the methods and results text to improve clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

      We appreciate this positive assessment and thank the reviewers for their time and constructive comments. We have made the following changes in the revised manuscript to address all points raised by reviewers.

      (1) Clarify the distinction between suppression efficiency and functional cost.

      (2) Add controls: smFRET experiments in the presence of monovalent 10E8.4 and iMab individually.

      (3) All of the smFRET population contour plots have been removed, as suggested.

      (4) Repeat neutralization experiments of tagged viruses (carrying nc-AA-incorporated, amber-suppressed Env), add and compare infectivity profiles between before and after click-chemistry labeling of tagged viruses.

      (5) Add a section (Complementary views from smFRET and structural studies) to the Discussion on how these approaches complement each other.

      (6) Further clarify three prefusion conformational states identified by smFRET, the relation with previously identified States 1, 2, 3, and asymmetry, the heterogeneity of Env presentations and virion morphology, and the focus of this study.

      Please find below our point-by-point responses to the public reviews and recommendations for the authors.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Thank you for the positive comments!

      In validating the functionality of the S401TAG/R542TAG Env, the authors performed infectivity assays and observed 20% infectivity as compared to wild-type (Figure S2A). However, the text equates this with "20% dual-amber suppression efficiency". This would benefit from some explanation. Why do the authors interpret infectivity as reporting on amber suppression efficiency, and not the functional cost of modifying Env, which is probably unavoidable? Or a combination of both? Is there data to suggest that 100% amber suppression would leave Env 100% functional? If so, this would be valuable to show. If not, the text should be clarified.

      We acknowledge this concern and have clarified the distinction between suppression efficiency and functional cost in this revised manuscript. The observed reduction in infectivity does not translate into functional loss; instead, it more reflects the efficiency of suppression (one of the critical limitations of applying genetic code expansion in mammalian cells). To support the preservation of Env functionality, we performed dose-response neutralization experiments of tag-free and 100% dual-ncAA-incorporated Env virions by two trimer-specific neutralizing antibodies, which exhibited similar dose-dependent neutralization sensitivity (Fig. 1D), providing stronger validation than infectivity assays. We also compared infectivity between labeled and unlabeled virions and observed no significant difference (Fig. S3B).

      We have previously discussed several limitations of amber suppression in mammalian cells when combined with smFRET viral systems (PMID: 38232732; PMID: 40716060) and, more recently, in our methodology chapters (PMID: 42349953; PMID: 42349954). In brief, orthogonal tRNA/aaRS pair–mediated amber suppression (reassigning/repurposing amber stop codons to non-canonical amino acids) of the introduced ambers in the target protein (Env in our case) must compete with the cellular translation system, particularly release factors that recognize amber codons and terminate translation. Readthrough of endogenous amber codons in virus-producing cells (in our case, HEK293T) can disrupt normal protein expression and virus production. Similarly, readthrough of pre-existing amber codons in HIV-1 ORFs other than the targeted ambers in Env can disrupt virus assembly, which we addressed by generating an amber-free provirus (PMID: 38232732). Introducing two amber codons into Env further reduces efficiency, as dual suppression requires two sequential successful suppression events within the same Env molecule.

      The authors state that the contour plots in Figure 2E reveal "dynamic sampling" of the observed FRET states. Strictly speaking, as presented, the contour plots (and FRET histograms) provide no information on dynamics per se. They indicate only the relative thermodynamic stabilities of the FRET states; transitions between states are a matter of interpretation. The TDPs, shown later in Figure 5A, nicely display the dynamics. More importantly, interpretation of the contour plots is challenging, as some seem to suggest an evolution toward lower FRET states. This is especially evident in Figures 2F and 3D, which suggest that the system evolves into a stable 0.1-FRET state (CO) after about 3 sec. Unless the authors want to conclude something from this, I would suggest that they consider removing the contour plots, since their interpretations are fully supported by the FRET histograms alone.

      We agree and have removed the contour plots, as they do not add meaningful information beyond what the histograms show.

      The data indicating that Env conformation is manipulated by 10E8.4/iMab is interesting. If I understand correctly, 10E8.4/iMab is an engineered antibody with one Fab targeting MPER and the second Fab targeting CD4. In the absence of CD4, could the difference between 10E8.4/iMab and the other MPER antibodies be due to 10E8.4/iMab being monovalent with respect to MPER binding?

      We appreciate this question. To address this, we have performed important controls: smFRET experiments in the presence of 10E8.4 and iMab individually in the absence of CD4. The results are shown in Fig. S9 in the revised manuscript, which indicates that 10E8.4 behaves similarly to other MPER-directed bNAbs we tested in this study, whereas iMab does not appear to affect the conformational populations of Env. The dual effect exerted by the bivalent 10E8.4/iMab is therefore very unexpected and thus interesting, as discussed in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      We appreciate these positive comments!

      Weaknesses:

      (1) The limited controls on how click chemistry affects Env (as labelled Env HIV virions were not evaluated).

      We agree. Our previous validation focused on ncAA-incorporated Env HIV-1 virions, but not the fluorescently labeled virions. To address this, we have added infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity.

      We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation.

      Nevertheless, we have previously demonstrated real-time tracking of single click-labeled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labeled Env retains functional competence.

      (2) Photobleaching of donor and acceptor molecules occurs right after 10sec exposure.

      We acknowledge this limitation and have included it in the revision.

      (3) Other limitations are well described in the corresponding section.

      We appreciate this comment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As a means of clarifying the mechanism of 10E8.4/iMab, the authors might consider performing separate smFRET experiments in the presence of the normal 10E8.4 antibody and the normal iMab antibody (a negative control). Alternatively, they could consider imaging in the presence of the DH511.2_K3 and VRC42 Fabs (as opposed to full-length Ig) to make a cleaner comparison, although this may be less informative given the high concentrations of antibodies used.

      We thank the reviewer for this excellent suggestion. To enable a direct comparison, we performed the most informative control by examining virus-associated Env in the presence of 10E8.4 alone and iMab alone. The corresponding smFRET results are presented in Fig. S9. We found that 10E8.4 behaves similarly to other MPER-directed antibodies, whereas iMab alone does not appear to have a notable effect on the conformational propensity of Env. For transparency and to facilitate future antibody design, we have also included the Fab region sequences of the antibodies in Table S2.

      Reviewer #2 (Recommendations for the authors):

      The article is well-written, the findings are of high interest for the community. The article should be shared once the points stated below are clarified and revised by the authors.

      We appreciate this comment and the points raised by the reviewer and have revised the manuscript accordingly.

      (1) In Figure 1C, the tomographic slices showing HIV-1 WT as compared to HIV-1 decorated with EnvBG505 S401ncAA R542nCAA are quite different morphologically. The micrographs chosen show a big particle with two capsids close to a smaller one without a capsid and at least in this plane bold (no Env incorporation) for the WT; whilst for the HIV-1 decorated with EnvBG505 S401ncAA R542nCAA no capsid is apparent in both particles, one (the right one) is very small and the right one does present a number of Envs but no apparent capsid is visible here. Please comment - perhaps it would be important to average the morphological traits of both and look at average diameter, average Env incorporation, morphology of the capsid, percentage of immature particles, capsid abnormalities (as the one shown in the upper micrograph).

      We thank the reviewer for this thoughtful comment.

      The tomograms in Fig. 1C were included to demonstrate the overall size and shape of the viral particles rather than to provide a quantitative structural comparison. HIV-1 viral particles are inherently heterogeneous, and the original slides were selected as representative examples. Following the reviewer's suggestion, we replaced the representative wild-type (Fig. 1C, top panel) and tagged virus (Fig. 1C, bottom panel) tomographic slides with those that better reflect the overall quality of each sample. To further address this concern, we refer the reviewer to the nanoparticle tracking analysis (NTA) shown in Fig. S3, which shows no significant difference in particle diameter between the wild-type and tagged viruses. In the revised manuscript, we now replace "morphology" with "shape" or "size," as these terms better reflect what our results can say.

      We agree that a quantitative analysis of capsid morphology, Env spike incorporation, and the proportion of immature particles would be informative. However, such analyses would require a substantially larger cryoET dataset, which is beyond the scope of the present study; nevertheless, it is certainly in our interest to pursue a cryoET-focused study of EnvCA interactions, with Env complexed with 10E8.4/iMab. Our primary objective is to study Env conformational dynamics by smFRET rather than viral morphogenesis or capsid maturation, whose relationship to Env dynamics remains largely unexplored. It is also worth noting that the optimal particle populations for smFRET and cryoET differ. smFRET measures the conformational dynamics of individual Env trimers and therefore selectively analyzes virions containing a single dually labeled Env trimer, whereas cryoET structural analyses typically benefit from particles with higher Env spike densities. Therefore, the particle populations favored for the two techniques are not identical.

      (2) In Figure 1D, there is a difference in neutralisation with PGT151 - how different are these two curves - how does the labelling affect neutralisation for bNAbs targeting gp41? Would it be possible to assess also the infectivity, fusion and neutralisation profiles of particles where the flurophores are included? This would be without diluting the Env for single particle analysis, but just to understand how harsh the organic reaction is and how it affects Env function (as all experiments and conclusions in the manuscript are based on labelled Env).

      Again, we sincerely appreciate these questions, which have helped us improve the manuscript. Neutralization assays for the tagged viruses were performed using the ncAA-incorporated, amber-suppressed viruses, whereas the engineered wild-type is amber-free. The differences between these two dose-response curves in the original Fig. 1D are small and within the experimental variation routinely observed under even identical conditions (same virus and same bNAb). We have repeated these experiments, and the new results are shown in the revised Fig. 1D. Although minor variations remain, the overall neutralization profiles and IC50 values are highly consistent.

      To assess whether the fluorophore labeling reaction affects Env functionality, as noted above, we have included infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity. We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation. Nevertheless, we have previously demonstrated real-time tracking of click-labelled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labelled Env retains functional competence.

      We believe that the unchanged infectivity of labeled viruses relative to their unlabeled counterparts, together with our previously observed real-time trajectories of click-labeled virions in live cells, provides strong evidence that our labeling strategy does not measurably impair Env function.

      (3) In Figure 2E and 2G, the authors employ a three Gaussian fit approach to recover the three populations (pre-triggered - pre-fusion closed - CD4 bound open). Can you please relate these with State 1, 2 and 3 from the original article (PMID: 25298114). Comment on the possibility that more than three populations could be fitted and what this could mean - pre-triggered and partially open (one gp120 asymmetrically open) could occur? Could this labelling approach account for this asymmetry?

      Thanks for this suggestion. In this study, we compared our results obtained using the gp120-gp41 structural axis with those obtained using the referenced gp120 V1-V4 structural axis to confidently assign the FRET-identified states to the previously reported three primary populations. The referenced axis is comparable to those used in the original article (PMID: 25298114) and later confirmed using the amber-click strategy (PMID: 38232732). We observe the same structural changes from these two distinct structural angles, as probed under ligand-free conditions (Fig. 2E and 2G) and CD4-triggered open conditions (Fig. 2F and 2H).

      The pre-triggered state corresponds to State 1; the pre-fusion closed state corresponds to the symmetric State 2 (which the SOSIP-based soluble Env primarily adopts; PMID: 30971821); and the CD4-bound open state corresponds to the fully open State 3. The assignment of the FRET states observed from the gp120 V1-V4 structural axis to States 1, 2, and 3 was originally reported in two studies (PMIDs: 27795397 and 29561264). In the asymmetric trimer configuration, the State 2 FRET signal originates from the free protomer, while the other one or two protomers bind CD4 and adopt the open conformation (PMID: 29561264). The asymmetric intermediate (PMID: 29561264) was identified using a heterotrimer experimental design consisting of a mixture of wild-type and CD4-binding-incompetent D368R protomers, which was not used in the present study. Therefore, our labeling approach cannot unambiguously resolve this asymmetry.

      Regarding the possibility of more than three populations, evidence from current and previous studies (PMIDs: 25298114, 27795397, 29561264, 38232732, 30971821) strongly supports the presence of three primary states of virus-associated Env, with additional substates that can be resolved under specific triggering conditions (PMIDs: 30974085, 41326374, 39640534). The assignment of such substates requires well-controlled experimental designs (PMIDs: 30974085, 41326374, 39640534).

      We have related PT, PC, and CO to States 1, 2, and 3, and added comments on multiple states and asymmetry in the revised manuscript.

      (4) When comparing smFRET with CryoET or structure, one can see that in smFRET there are always many potential conformations for big sub-populations of Env. Indeed, there is a trend, and the addition of bNAbs (Figure 4) clearly has an impact on increasing and stabilizing a particular state as defined by the authors (e.g. PT at 45% upon addition of 8ANC195, but also 32% PC and 23% CO). I assume that when analysing single particle CryoET or single virus CryoET, one needs to discard after template matching different scenarios that do not necessarily contribute to the highest resolution and this information is not always discussed. It would be interesting to address this in the discussion as the effect on Env dynamics of adding ligands (including CD4 and 17b) is not inducing in all Envs a drastic conformational change - this could be derived from the Ka of the ligands, but also from the intrinsic Env heterogeneity in both dynamics and architecture - I think that addressing these matters in the discussion could be of interest for the community. In this regard, the transition density plots are very helpful.

      We completely agree and appreciate this insightful suggestion. We have expanded the Discussion to better address the complementary insights provided by smFRET and structural approaches. In single-particle cryoEM, we do not observe the full spectrum of Env conformations for technical reasons, not because particles are intentionally discarded to obtain only the highest-resolution structures. One reason is that open Env conformations are much more sensitive to radiation damage than closed Env. Likewise, ligand-free closed Env is more sensitive to radiation damage than a bNAb-stabilized closed Env. Thus, the outcome of an SPA cryo-EM study depends strongly on the biological question being addressed and the conformational state that is preferentially preserved under the experimental conditions. Although one could hypothetically collect much larger datasets to recover lower-abundance conformations, this would be both cost-prohibitive and unlikely to faithfully represent the relative conformational populations due to differential, conformation-dependent radiation damage.

      We agree that the smFRET data highlight an important aspect of Env dynamics. Ligands, including bNAbs, CD4, and 17b, generally shift the conformational equilibrium toward particular states rather than driving all Env trimers into a single conformation. This likely reflects both differences in ligand binding properties and the intrinsic conformational heterogeneity of Env. We therefore believe that structural studies and smFRET provide complementary information. Structural methods resolve the molecular architecture of individual conformational states at atomic (by cryoEM) and near-atomic (by cryo-ET) levels, whereas smFRET quantifies their relative populations and dynamic interconversion. As the reviewer pointed out, the transition density plots are particularly valuable in illustrating these dynamic changes.

      We have incorporated these points into the revised Discussion (Subtitle: complementary views from smFRET and structural studies in general).

      (5) One of the very interesting findings of the paper is that the effect of bNabs (at least the ones tested) have an impact on the Env dynamics and how this shift can alter entry - therefore the structural view is perhaps less important - In spite of this, we still employ a structural jargon to refer to "Env open conformation stabilisation" for instance - even if the data shows that upon ligand exposure dynamics are still important but shifted. Please comment.

      This is a great point. Our smFRET data show that bNAb binding generally shifts the conformational distribution toward and stabilizes particular Env states by lowering their free energy and increasing their occupancy. We have clarified that ligand-induced stabilization reflects a redistribution of the conformational ensemble while preserving the intrinsic dynamic nature of Env.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on the excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons, and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting, and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability, given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarf the effects of the test compounds. The experiments may suffer from undersampling.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure (Figure S2) that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of neuron-to-neuron variability in the effect of inhibiting Nav1.7 channels even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. We also now note that there is a similar high degree of cell-to-cell variability in relative functional expression of Nav1.7 and Nav1.8 channels in capsaicin-sensitive mouse DRG neurons.

      (2) A potential confounder when using post-mortem human DRG neurons is heterogeneity of cell types. The methods clearly state that the cells selected for recording were of 'generally' small size, but specific criteria for what constitutes 'small' or other unstated selection criteria were not provided. A table of individual cell capacitance and input resistance values, along with information about individual donors (age, sex, ethnicity), is important to include. Additionally, some discussion of how DRG neuron heterogeneity impacts the findings. This relates to concern #1 about sample size determination and how cell heterogeneity factored into this calculation.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). As noted, we have also added a figure (Figure S2) that breaks out the effects on these parameters for each donor, illustrating that there is a high degree of neuron-to-neuron variability in the effects even within a single donor. We have added several sentences to the Discussion concerning the neuron-to-neuron variability in the effects of Nav1.7 inhibition, including the possibility that this may reflect heterogeneity of cell function.

      Reviewer #2 (Public review):

      Summary:

      The authors examine the functional role of Nav1.7 voltage-gated sodium channels in human sensory neuron electrogenesis using a Nav1.7 selective inhibitor and human dorsal root ganglion neurons obtained from organ donors. Patch-clamp electrophysiology is used at physiological temperature to measure the impact of Nav1.7 inhibition on sensory neurons' action potential firing. This is an important topic as Nav1.7 and Nav1.8 have been identified as therapeutic targets for the treatment of pain, but there has been mixed success with isoform-specific inhibitors in clinical trials. The data suggest that Nav1.7 and Nav1.8 have overlapping yet complementary functions in nociceptor neurons and that targeting both may be most effective for reducing nociception.

      Strengths:

      The data are of high quality. Action potential properties are measured at 37 degrees Celsius. Threshold is measured using brief pulses. The Nav1.7 inhibitor has been reported to be highly selective for Nav1.7 over Nav1.8 and moderately selective for Nav1.7 over Nav1.1 and Nav1.6. Data are collected using identical conditions and protocols to a previous study on the role of Nav1.8 in similar neurons.

      Weaknesses:

      The study relies on a single Nav1.7 inhibitor that has not been extensively characterized. One prior study indicates that the IC50 is around 140 nM, thus the 600 nM concentration used in this study could be predicted to reduce Nav1.7 currents by 80%. However, there is no voltage-clamp data in the current study to confirm this, and therefore, it is unclear if the batch of AM-2099 is as potent as reported in the paper that initially described its selectivity. The impact of Nav1.7 inhibition is compared to data from a previous study by this lab, and this is a minor concern. It would have been interesting to see if the combined inhibition of Nav1.7 and Nav1.8 completely blocked action potential generation in the human DRG neurons.

      We have done experiments to directly characterize the potency of the AM-2099 sample we used on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sub>50</sub> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      We have also added a new figure (Figure S3) showing the effects of a different Nav1.7 inhibitor, PF04856264. The effects of this inhibitor were qualitatively identical but quantitatively smaller than those of AM-2099. When we realized this, we did voltage clamp experiments on cloned Nav1.7 channels and discovered that the potency of PF-04856264 at 37°C was weaker than expected from the published IC<sub>50</sub>, which was determined at room temperature.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for on-going and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al. investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons and that Nav1.7 blockers may have limited efficacy as analgesics. This is surprising, given that patients with loss-of-function mutations in Nav1.7 suffer from congenital insensitivity to pain. While it may indeed be true that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia, the present study was limited to a single concentration of AM-2099. The manuscript would be significantly strengthened by a more careful and thorough pharmacological characterization of this compound, which has not been widely used or validated in native human DRG neurons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Weaknesses:

      Only a single concentration of AM-2099 was used for all experiments. This compound was reported to be selective for cloned human Nav1.7 channels in heterologous systems, but has not been validated in other studies after the original publication in 2016. Since the original study reported a substantial statedependent block of recombinant Nav1.7 channels, more detailed pharmacological characterization of AM-2099 is needed in human DRG neurons to fully support these claims. This study would be significantly strengthened by the inclusion of dose-response curves to assess how much of the sodium current is inhibited at this concentration, confirming selectivity in hDRG, and whether maximal inhibition of Nav1.7 still has limited efficacy in reducing the firing of native human sensory neurons.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2). These show that 600 nM AM-2099 produces about 85% inhibition of Nav1.7 channels at 37°C. We have added a paragraph to the Discussion explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      With regard to the broader point about reconciling the variable and sometimes relatively modest effects of Nav1.7 inhibition with the complete loss of pain sensation in humans with loss-of-function mutations, we have modified the Introduction and Discussion to eliminate any implication that the results in the manuscript suggest that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia. Our experiments are only on action potential firing in the cell body, and it is perfectly possible that inhibiting Nav1.7 channels in the axon could disrupt generation or propagation of action potentials, either in the main axon or in the fine axon terminals in the spinal cord. We have modified the Discussion to explicitly point this out, which would reconcile the loss of pain sensation in humans with loss-offunction mutations with the incomplete effects of Nav1.7 inhibitors on excitability of cell bodies.

      Recommendations for the authors:

      Reviewing Editor comments:

      In addition to the points noted in the eLife assessment summary above, the study has several important strengths, including use of human primary neurons and recordings performed under physiologically relevant conditions (at 37 {degree sign}C using brief current injections). However, reviewers also identified several key weaknesses that must be addressed to support the central conclusions. In particular, multiple reviewers raised concerns regarding the lack of voltage-clamp data evaluating the efficacy and specificity of AM-2099 inhibition of Nav1.7 currents. A single dose of 600 nM was used based on the report of Marx (2016) in recombinant systems (Marx, 2016). Since no other studies other than the single Amgen report exist on this compound, it is important to validate its effects directly in the human DRGs used here. Additional concerns include the lack of dose-response analysis, as well as the large variability in baseline properties, which complicates the interpretation of the results. To assist the revision of the study, we outline below the key issues that should be addressed.

      Recommendations for authors:

      (1) Add voltage-clamp experiments to directly measure Nav1.7 current inhibition by AM-2099 in hDRG neurons. Given the limited previous characterization of this compound, it is important to confirm that the concentration used here (600 nM) effectively blocks Nav1.7 currents in the native system used here.

      (2) Related to point 1 above, perform a dose-response of AM-2099 on hDRGs on Nav1.7 currents in human DRGs. Since this study, at least in part, is framed as a comparative analysis of Nav1.7 vs Nav1.8 channel subtypes in DRGs, it seems important to establish pharmacological equivalence to ensure that the comparisons are made at functionally comparable levels of channel block.

      We have added results from experiments to directly characterize the potency of AM-2099 on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sup>50</sup> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      (3) Reviewer 1 notes that there seem to be 10-20-fold differences in baseline firing properties, which would exceed the effects of the test compound. This raises concerns about undersampling. Additional analysis or experiments would strengthen the conclusions.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of cell-to-cell variability in the effects even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. Reviewer 1 made the excellent suggestion that because of the neuron-to-neuron variability in baseline properties, the effects of compounds could be better illustrated by displaying changes from baseline. Following this suggestion, we have added Tukey-style box plots displaying the data in this way. Together with the donor-to-donor breakout of data in the new Figure S2, these plots make it clear that the neuron-to-neuron variability reveals genuine differences in the channel make-up of each neuron and not experimental error.

      (4) Reviewer 2 notes an interesting experiment: does a combined block of Nav1.7 with the AM compound and Nav1.8 block action potential generation? If Nav1.7 controls threshold and Nav1.8 controls firing, then the combined inhibition should be highly effective in blocking nociceptive output, which could have therapeutic relevance.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for ongoing and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #1 (Recommendations for the authors):

      Concerns in addition to those in the Public Review:

      Major:

      (1) As per point 1 of the weaknesses in the Public Review, I'm concerned that the experiments suffer from undersampling. This requires a discussion of how the sample size was determined.

      We have added experiments from an additional 3 donors to the key results in Figures 3-5, more than doubling the number of neurons for these measurements.

      (3) The effect of compounds could be better displayed as a change from baseline in Figure 1C-E. Also, are the AP traces and phase plots shown in Figures 1AB and 2AB averages or representative?

      Thanks for this excellent suggestion. We have added box-plots that show changes from baseline for the various parameters. We have also clarified that the action potential traces and phase plots are from application of AM-2099 in a single representative neuron.

      Minor:

      (1) Provide source of VX-548 and report the purity of both compounds.

      We have provided the information for VX-548 and added the information on the purity of both compounds

      (2) Clinical failures of Nav1.7 blockers may not be solely due to pharmacodynamic limitations as implied by this study. Pharmacokinetic differences and toxicity (e.g., effects on the autonomic nervous system) may also have contributed.

      Thanks for raising this important point. We have added this point to the Introduction.

      Reviewer #2 (Recommendations for the authors):

      It is an interesting study, and the conclusions are reasonable. However, it would have been good to see validation of the potency of AM-2099 on native DRG sodium currents and/or recombinant human Nav1.7 channels expressed in a heterologous expression system.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2).

      Minor comments:

      (1) Page 3, middle paragraph - there is a "(" missing before Renganathan.

      Thanks, corrected.

      (2) Page 4: Is anything known about AM-2099 in terms of state-dependence? It seems like Marx 2016 is the only previously published study using it, so additional information on the inhibitor would be helpful.

      We have not characterized the state-dependence of AM-2099, but we characterized its potency in voltage clamp using holding voltages similar to the average resting potentials of the cells in current clamp conditions.

      (3) Page 6 discusses that there might be differences between human and rodent DRG neurons in terms of Nav1.7 and Nav1.8. It would be nice if this were directly tested with these same Nav1.7 and Nav1.8 inhibitors.

      We have recently done such a study on mouse DRG neurons which has just been published (J Physiol. 604:6104-6127, doi: 10.1113/JP290574).

      (4) Figure 2A, right panel: I could not figure out the difference between the red and green traces. Perhaps this could be explained in the figure legend?

      Thank you for pointing out that this was confusing. These two traces showed two different subthreshold responses, one of which was slightly regenerative without generating a full-blown spike. We have simplified the figure by now showing only a single subthreshold response.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusion that Nav1.7 inhibition has limited efficacy for inhibiting the firing of human DRG neurons is not fully supported by the data. This may be true, but it cannot be concluded without a more thorough pharmacological characterization of this compound. Dose-response curves and experimental confirmation of Nav1.7 selectivity (maybe just total Nav current, TTX-sensitive and TTX-resistant components) are needed.

      We agree and have now added two new figures with this data.

      (2) How was the 600 nM concentration chosen? Given that AM-2099 was reported to exhibit state dependent block, how much of the Na current is inhibited by this concentration at the initial voltage used in current clamp experiments (~-80 mV)?

      We have added a paragraph to the Discussion recognizing the limitation that 600 nM AM-2099 produces ~85% rather than complete inhibition of Nav1.7 current and explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      (3) It appears that the effects of AM-2099 on the refractory period are bimodally distributed, where neurons that recovered more slowly at baseline were preferentially affected by AM-2099 (Figure 4). Do these reflect different neuronal populations (e.g. smaller or larger diameter DRG) or different resting voltages in these experiments?

      We agree that there seem to be two groups based on initial refractory period. Examining the parameters for the cells, there is no clear correlation between the effects of AM-2099 on the refractory period with resting potential or cell diameter. At this time, it is not obvious what determines the differences in refractory period. We speculate that neuron-to-neuron differences in the potassium conductances that generate the after hyperpolarization may be different in these cells but it will take further work to explore this.

      (4) How much of the sodium current is mediated by Nav1.7 in hDRG neurons? How does inhibition of both Nav1.7 and Nav1.8 affect hDRG excitability?

      The new Figure 2 shows data quantifying the AM-2099-sensitive current in the DRG neurons. With regard to combined Nav1.7 and Nav1.8 inhibition, we are currently doing experiments examining inhibition of excitability by combined Nav1.7 and Nav1.8 inhibition, which we agree is the logical next step in exploring how the two components of current control excitability. These are still in progress. Because designing and interpreting these experiments is facilitated by the current experiments with Nav1.7 inhibition alone, we believe that reporting the current results now will serve the scientific community better than waiting to obtain and interpret a body of data on dual inhibition in a sufficient number of donors, which we obtain only sporadically.

      (5) Donor information and soma diameters should be included. Capsaicin sensitivity testing was mentioned in the methods, but I was unable to find any inclusion of these data in the results. These may be useful to potentially infer effects in different cell types.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). We have now clarified that data were confined to cells verified to be capsaicin-sensitive and that ~95% of all cells tested were capsaicin-sensitive.

      (6) Please check statistical tests and reporting. Several graphs do not appear to have paired responses (e.g. Figure 1E, Figure 4B). As a result, two-tailed Wilcoxon tests would not be appropriate. Also, check reported p-values (e.g. p=.0002), which are identical for multiple panels in the Results section.

      We have clarified that the symbols of action potential width in control without a corresponding value after AM-2099 represent neurons in which the action potential in AM-2099 had a peak < 0 mV. These cells were not included in the data set of paired parameters used for the Wilcoxon test. We have also checked and verified all statistical tests.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Freas and Wystrach present a computational and experimental study of ant navigation. The main innovation of the computational model is the insertion of an oscillatory element between the steering signal and the motor control that results in a trajectory whose heading oscillates around a goal direction. Additionally, the model imposes periodic cessations of forward movement and inversely couples rotational speed to forward velocity. As a result the model periodically makes larger reorientations reminiscent of those seen in behaving ants.

      The behavioral data consists of two experimental sets: experienced Melophorus bagoti foragers, recorded in 2010 and inexperienced M. bagoti foragers, recorded in 2023-2024 at the same site. The behavioral data is qualitatively compared to the model in Figures 3 through 6. In figures 3-5, all ant sets are grouped together while in Figure 6 they are separated. In Figure 6, the authors should do a careful job of making sure the reader is aware that comparisons are being made between behavioral data sets captured more than a decade apart and of justifying the validity of a quantitative comparison between these sets.

      We now make explicit in the methods and figure that the two datasets were recorded at different times: experienced foragers from 2010 (Deeti et al., 2023) and inexperienced foragers from 2023. Their comparison is used to test a qualitative difference predicted by the model.

      The manuscript also describes Myrmecia ants and makes comparisons between modeled Myrmecia ants and supplemental videos of these ants (Videos 3,4). These videos are not described in the methods. While the captions describe these as ants "homing in an unfamiliar environment," the videos show tethered ants walking on a ball. Without more information and absent any analysis, it is difficult for me to understand how these videos support granular points in the text about coupling between rotation and forward velocities.

      We have added a description of Videos 3 and 4 to their respective captions and now state explicitly that these clips were recorded from ants on a tethered trackball apparatus and that they provided as qualitative examples of behaviour discussed in the text.

      Strengths:

      The manuscript's main thesis, that an oscillatory element interspersed between the control signal and the motor unit can reproduce aspects of ant navigation, appears supportable.

      Weaknesses:

      Qualitative agreement between aspects of a model and aspects of a behavioral measurement do not prove the correctness of a model. In the section (802), "An ancestral design? Striking parallels with crawling Drosophila larvae," the authors argue that behavioral data in larvae support their model, despite the larva's lack of a (known) central complex. C. elegans navigation can also be segmented into longer runs and shorter exploratory behaviors (Chen 2025), comparable to the runs and scans described here. C elegans definitively does not have a central complex. In general, multiple internal mechanisms are capable of producing the same macroscopic behavioral outcome. This fact limits the ability of behavioral data to confirm the details of a particular model; it does not imply that observation of similar behaviors in multiple species shows that a particular model is correct or generalizable.

      Here the ability of the behavioral data to confirm or constrain the model is further limited by the qualitative nature of the comparisons. Some of the comparisons are trivial (e.g. Figure 5E-F: any first order process will produce a Poisson distribution, and in the model a Poisson process was explicitly coded in with parameters chosen (1070) to match the behavioral data). Finally, the number of adjustable parameters (13) is comparable to the number of comparisons made; it is unclear that the model could not be adjusted to fit any set of behavioral measurements.

      Our model is a minimal neuro-mechanical model. It is not a mathematical model where each parameter can be optimised to a final output.

      From our 13 parameters, 10 were either taken directly from independent prior studies. 5 concern the oscillator, and have been arbitrarily chosen and simply need to produce regular oscillation (as explained in supplemental material). 4 are the necessary motor gain and noise, which scale neural activations values into movement units, note that this conversion is backed up by previous evidence in drosophila and present in previous ant models). 1 parameter specifies the width of the bump of activity in the CX, and is roughly matched to neural imaging data in flies. None of these parameters have been introduced or adjusted to back up our claim.

      Only 3 parameters have been added to produce scannings. Two of them were tuned to match local scanning-specific data (the probabilistic trigger to stop (p_stop) and the threshold for triggering a saccade (θ_CPG) enabling us to tune fixation duration). This level of parametrisation enables the model to reproduce realistic scans, but does not influence the qualitative predictions of this article. For instance, we agree that the probabilistic trigger producing the Poisson distribution of scan duration (Figure 5’s E) is used to parametrise the model to scanning data, which does not constitute an emerging prediction of the model. We do not include it as evidence (see ~490). Finally, the CX_output_gain, is a new parameter we invoke to implement our hypothesis that CX steering modulates the oscillator. From these three added parameters emerge the large array of behavioural signatures and predictions. These are emerging consequences of the model's architecture rather than curve-fits.

      While the introduction is improved, there is still room to eliminate confusion as to what aspects of the model reflect hypothesized rather than measured neural circuits. For instance, if there is data showing LAL oscillations in insects, the authors should cite it and call it out clearly.

      Alternatively they should say that the oscillator is hypothesized based on measured bistability. They should also clarify whether they are discussing neural oscillations or motor oscillations and whether these oscillations are measured, modeled, or hypothesized.

      As one example: Lines 283-284 "This oscillator [referring to the model's intrinsic oscillator described in the previous paragraph], which is widespread in insects (Cheng, 2024; Kanzaki, 2005; Kanzaki and Mishima, 1996), resides in the lateral accessory lobes (LAL)" reads as though it is known that a neural oscillator occupies the LAL. Cheng 2024 is a brief review of behavioral oscillation. Kanzaki et al. 2005 describes numerical modeling and simulation with a physical robot. Kanzaki and Mishima, 1996 demonstrates bistability (flip-flopping) in moth descending neurons. None of these show neural oscillations and none of them describe the LAL. The authors should review the paper and be scrupulously careful that the claims made in the text are supported in the cited references. These difficulties were pointed out in a previous round of review; hopefully they can be fully corrected this time.

      Kevin S. Chen, Jonathan W. Pillow*, Andrew M. Leifer*, "State-switching navigation strategies in C. elegans are beneficial for chemotaxis," arXiv:2508.00191 31 July 2025.

      We have softened the text to be more explicit about what is modelled versus what is neurally shown (labelled each as behavioural, electrophysiological, or modelled)

      Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

      We thank Reviewer 2 for this assessment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The paper would benefit if the authors would take a more traditional/formal approach to the presentation and interpretation of results. They should avoid drawing conclusions or presenting interpretations in the figure captions (e.g. caption to figure 6) and avoid unnecessary modifiers (e.g. just use "supports" instead of "strongly supports"). This might help correct the tendency of the manuscript to overstate or over-interpret the correspondence between the model and the data.

      Figure captions are now revised to remove interpretive discussion. We also removed unnecessary modifiers throughout the manuscript.

      We also added clarifying text to the “An ancestral design?” section to make clear that we are not implying homologous neural implementations across taxa.

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed my comments fully and I only have a few minor, mostly editorial points:

      line 142: it appears that the references should refer to goal encoding, but both references are head direction papers (one review, one research paper). The only paper showing goal encoding in the CX is Mussels-Pires et al 2024. This should be fixed to ensure the citations are not misleading.

      Changed citations.

      line 165: maybe remove 'intrinsic' to not suggest that this reflects what it known from Biology? It is clear that in the model it is an intrinsic oscillator, but it should not leave the impression that this is an established fact for the LAL

      Removed when not discussing the model.

      line 168: remove either 'a diversity' or 'key qualitative'

      Changed to reproduce multiple key qualitative…. (~Line 170)

      section: 'Neural substrate of insect navigation':

      - '... compares to output steering commands.' grammar is misleading as 'output' might be read as an adjective rather than a verb, replace with 'generate'?

      Changed.

      - as above, the only paper showing goal directions is Mussels-Pires et al 2024. Any other paper either assumes this in models (such as Stone et al and the Honkanen review) or deal with head direction encoding. Pfeiffer and Homberg, 2014 is a general review. Please ensure that references are used more accurately.

      We revised the text to clarify the specific evidence provided by each cited reference.

      - The goal direction in the CX can be updated by various pathways....' After this, behavioral and modeling papers are cited, which is misleading. None of these papers deal with the CX or the neural representation of goals. Same with the rest of the sentence, referring to MB output and PI, which is only shown in modeling, not data.

      We now explicitly distinguish between behavioural evidence and modelling evidence.

      line 210: use CX, not central complex

      Changed.

      Figure 2: I suppose all data shown are modeling data? This should be more explicit in the figure caption (a bit unclear what 'using the neural circuit model.' in the caption heading means. Maybe rephrase to: 'Schematic of neural circuit model and its outputs across navigation relevant brain regions.' (or something like that, putting model first, not brain regions)

      Changed.

      Figure 7: Axis labels in the graphs are still much too small to be read on a printed version (ensure at least 5pt font size in the actual figure on the printed page)

      Enlarged axis labels.

      Line 654: mirror, not mirrors

      Changed.

      line 816: insert 'fly' before larva, as otherwise one might assume this refers to ant larva

      Added (~line 822).

      line 819: The CX does (as we currently know) not exist in fly larvae. At least not as a brain structure, or a set of homologous neurons. There might be equivalent circuits for action selection, but they have not yet been convincingly described. I suggest to rephrase to: 'Although no present as neuropils in fly larvae, the CX and LAL....'

      Changed as suggested (~Line 830).

    1. Author response:

      We thank the reviewers and editor for their thoughtful and constructive comments. Our goal was to connect theoretical work on habituation with empirical findings on intracellular habituation in Stentor coeruleus. We developed the phase-portrait analysis to provide a more formal way to examine habituation dynamics and to address the concern that apparent potentiation might simply reflect incomplete recovery. We are glad that the reviewers found this approach useful, and we hope to build on it in future work through mechanistic modeling.

      We will submit a revised version with the following changes:

      (1) We will include a supplementary figure that provides visual intuition for the habituation curves and phase portraits.

      (2) We agree that the anomalous behavior of the 2 min ISI / 1 hr ITI condition is noteworthy. This behavior arises from a subtle difference in the fitted shape of the trial 2 habituation curve: its Hill coefficient is less than 1, so the curve has no inflection point and its initial slope has nonzero magnitude. As a result, the corresponding phase portrait begins away from the x-axis, unlike the other conditions, whose Hill coefficients are greater than 1 and whose phase portraits are U-shaped. We will discuss this explicitly in the revision.

      (3) We will reconsider the layout to make the background and results easier to follow. In particular, we will consolidate the repeated material while preserving the context needed to interpret the theoretical consequences of the empirical findings.

      (4) We will revise the conclusions and discussion to leave claims about the decay of potentiation more open-ended.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Weaknesses:

      (1) The derivation of the main error term misses some important steps, which complicates peer review at this stage. In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The derivation of the main error term misses some important steps, which complicates peer review at this stage

      We thank the reviewers for this careful observation. This concern is associated with the error term. Thus, we first clarified the noise assumption explicitly. We assume that ξ(t) is i.i.d. over time with zero mean and an arbitrary (not necessarily diagonal) positive definite covariance matrix .

      In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The cited sources do contain the derivation for a noise term with full covariance. We have updated the citation that directly supports Eq. (S.2): Proposition 11.1 of Hamilton (1994, TimeSeries Analysis), which establishes the asymptotic distribution The proof, given in Appendix 11.A of Hamilton (1994), proceeds via a CLT for martingale difference sequences. See Theoretical Details of the Supplementary Materials.

      (2) The practical recommendation at the end of the paper also requires clearer guidance on how the design perturbations are constructed, and how many times and for how long the system is stimulated in each iteration of the experiment.

      Thank you for this helpful suggestion. We agree that the practical implementation of the experimental design should be explained more clearly. We have addressed this concern in two ways. First, we have revised the manuscript to explicitly describe the parameter design procedure. Second, we have revised the manuscript to clearly provide a reference to the detailed experimental condition table in the supplementary material. See Results - Main Manuscript.

      (3) Finally, there is no analysis of model mis-specification. In particular, the true dynamics are unlikely to be linear; the noise is unlikely to be either Gaussian or uncorrelated across time; and the B matrix is unlikely to be known perfectly. We’re not suggesting that the authors consider a more complex model, but it’s important to know how sensitive their method is to model mismatch. If nothing can be done analytically, then simulations would at least provide some kind of guide.

      We thank the reviewer for raising this important point regarding model mis-specification. We agree that it is important to run simulations to assess the impact of these model mismatches, therefore we conducted preliminary simulations to assess the sensitivity when two primary assumptions are violated: linear state dynamics and a perfectly known input matrix B. We added these preliminary results in the Supplementary Material, and revised the main manuscript to include a mention of these results. In summary, the simulations showed:

      The model estimation error increases with the strength of the nonlinearity; however, perturbation can increase the information and lead to accurate estimation of the hidden mode even under nonlinearity.

      The matrix A can be estimated roughly even if the assumed input matrix B differs from the true matrix in some cases.

      The matrices A and B can be estimated jointly without bias, provided the stimulation pattern excites the full state space.

      In such joint estimation, preferentially exciting the hidden modes directly leads to more accurate estimation of A than stimulating other modes. See Background and Results - Main Manuscript, Experimental Conditions and Results - Supplementary Material.

      Recommendations for the authors:

      (1) Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them. That’s important, because they are going to be used as control signals, so we need to know how accurately u(t) can be specified, and what its range is.

      Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them.

      We appreciate the reviewer’s helpful comment. We agree that it is important to describe these stimulation methods with appropriate references. We have added a new section titled “Neural Stimulation as Control Inputs” to the background, which connects our theoretical framework to practical experimental settings.

      We need to know how accurately u(t) can be specified, and what its range is.

      The accuracy and range of control inputs vary substantially depending on the specific stimulation technique and experimental setup, and a thorough discussion would require a dedicated review beyond the scope of this manuscript. Instead, we have added a sentence acknowledging the gap between practical experimental implementations and the theoretical formulation, and cited relevant references for readers interested in further details.

      Background - Main Manuscript

      “Neural Stimulation as Control Inputs

      This section describes how commonly used neural stimulation techniques can be related to input signals in control theory. Their adjustable parameters vary depending on how the stimulation inputs are modulated.

      Three non-invasive electrical stimulation methods illustrate how stimulation paradigms map onto basic control inputs. Transcranial magnetic stimulation (TMS) induces brief and transient perturbations via electromagnetic pulses [19], which are naturally represented as a sequence of impulse-like inputs, where the timing and intensity of each pulse are the primary controllable parameters. Transcranial direct current stimulation (tDCS) primarily modulates neural activity through approximately constant inputs [26], which can be viewed as a step-like signal whose main controllable parameter is the amplitude of the applied current. Transcranial alternating current stimulation (tACS) delivers oscillatory inputs [7], corresponding to sinusoidal signals characterized by amplitude, frequency, and phase. In control theory, impulse, step, and sinusoidal inputs are the basic components used to characterize system responses and dynamics [22, 21].

      The control input framework extends beyond non-invasive techniques to invasive and optogenetic stimulation. Invasive electrical stimulation, including intracranial microstimulation and deep brain stimulation (DBS), enables direct delivery of electrical inputs to neural tissue [17], providing flexible control over amplitude and timing through pulse trains or temporally structured waveforms. Optogenetic stimulation allows genetically targeted activation or inhibition of specific neurons using light [5], providing fine-grained control over multiple input dimensions, including amplitude (light intensity), temporal pattern, and cell-type specificity. In particular, recent developments enable stimulation at the level of individual neurons with high temporal precision [27, 16], allowing flexible construction of spatiotemporal input patterns.

      These stimulation examples demonstrate that the theoretical framework developed in this paper connects to practical experimental settings. While a substantial gap remains between idealized control inputs in theory and experimentally realizable stimulation, the core principles established in the following sections provide a foundation that naturally extends to these practical stimulation paradigms.”

      (2) Is the solid curve in the last panel of Figure 4b the prediction? If not, are the data points consistent with the prediction in Equation 8? This should be clear.

      Thank you for this question. The solid curve shows the empirical eigenvalues of the state covariance matrix, not the eigenvalues of matrix A. As shown in Equation 9, the estimation error is proportional to the inverse of the sum of the eigenvalues of the state covariance matrix. We have clarified these points by revising the main text and the figure captions. See Results - Main Manuscript.

      (3) Page 10, "Nodes 7 and 8 have only outgoing edges". It looks like node 7 has an incoming edge from node 8 (Fig. 6b). Or are we misinterpreting something?

      Thank you for pointing out this inconsistency. You are correct. Node 7 did have an incoming edge from Node 8 and contradicted the statement in the text. Moreover, we now think that two hub nodes (Node 7 and 8) are not necessary for demonstrating our primary theory. To resolve these issues, we have revised the simulation so that the network now contains a single hub node with only outgoing edges. Please refer to Fig. 6.

      (4) Figure 6f, g: why isn't the estimation error proportional to tr[Sig_x^{-1}]?

      We thank the reviewer for this observation. In the original manuscript, the estimation error in Figure 6f, g was plotted on a logarithmic scale, which obscured the proportional relationship with . The underlying values are indeed proportional, consistent with our theoretical prediction.

      In the revised manuscript, we have substantially reorganized Figure 6 to make the theoretical reasoning more transparent by using 1/µ instead of following Eq. 9. The sum across the column in panel (g) is proportional to the estimation error shown in panel (h).

      (5) Page 11, "The simulation was conducted with a single node receiving an impulse-shaped perturbation input." What’s an "impulse-shaped perturbation input"? A delta function? Please make this clear.

      Thank you for this clarifying question. We clarified the explanation as follows.

      Results - Main Manuscript

      “Each node was individually perturbed by an impulse input. Here, an impulse input is defined as a Kronecker delta at t = 0 with fixed amplitude α = 10, with no external input applied at any subsequent time step.”

      (6) Page 11, "Fig. 6d shows that the system possesses modes with small absolute eigenvalues." According to Figure 6d, all the absolute eigenvalues are between 9.9 and 9.98. So this statement appears not to be correct. Are we missing something?

      Thank you for pointing this out. You are correct. The previous statement was inconsistent with the figure. In the revised simulation, we have redesigned the network so that it clearly contains modes with distinct damping characteristics: heavily damped modes with absolute eigenvalues below 0.8 and lightly damped modes with absolute eigenvalues close to 1.0 (see revised Fig. 6c). This makes the relationship between mode damping and perturbation effectiveness much more transparent.

      Results - Main Manuscript

      “Figs. 6c shows the damping rates |λA| of the eigenvalues of A: the mode formed by Nodes 1 and 2 (15Hz) is heavily damped, while those formed by Nodes 3–6 are moderately damped. Node 7 serves as a hub with only outgoing edges.”

      (7) Page 11, "Crucially, the eigenvectors corresponding to these rapidly decaying modes (e.g., evec 7, evec 8) have their largest components concentrated at Nodes 7 and 8." If "nodes" are the same as "eigenvalue index", then the components on node 8 are zero (Figure 6e, bottom). In any case, it should be clear what you mean.

      Thank you for this comment. We agree that the previous description regarding eigenvectors was unclear and inconsistent with the figure. We now think that explaining with eigenvectors is not necessary for demonstrating our primary theory. We have therefore replaced the plots of eigenvectors with the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly and rigorously explained by Equation (9). These reciprocals clearly show that the modes associated with Nodes 1 and 2 are heavily damped and applying perturbations to these nodes contribute to the estimation error and the perturbations to hub node (Node 7) broadly increases all eigenvalues of Σ<sub>X</sub>, thereby reducing the estimation error across all modes. This change ensures that all simulation results are grounded in the theory presented in the paper. Please refer to the revised Fig. 6d–g for details.

      (8) It’s not clear to us what’s plotted in Figure 6e. The real part of the eigenvectors? Which would explain why some of the eigenvectors are the same (e.g., 1 and 2). But that does not seem like a good idea, since the eigenvectors can be rotated by an arbitrary complex phase. Also, eigenvectors 7 and 8 seem totally opposite, and there’s no weight at all on index 6 and index 8. Could there be a mistake in the figure? In addition, Figure 6e is explained and interpreted in two different paragraphs, discussing the same observation. It would be easier to understand if they were moved to the same paragraph. Also, in the second paragraph where plot 6e is referenced (page 11, line 19), what does ’concentrated’ mean?

      Thank you for this comment. As described in our response to Comment (7), we have removed the eigenvector-based analysis from the revised simulation to prevent confusing readers. The revised results focus on quantities that are directly explained by Equation (9). Please refer to the revised Fig. 6 for the updated results.

      (9) Page 11, "This simulation demonstrates that, given a tentative connectivity matrix, an effective perturbation input (such as TMS or tDCS) can be designed by targeting the node with the highest weighted out-degree." "Demonstrates" seems strong. There is only one simulation, and that wasn’t totally convincing: a perturbation applied to node 7 did almost the same as a perturbation applied to nodes 1-6, even though it had a much higher outdegree than those nodes. It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error. In particular, which properties of a hub node make it effective as a stimulation target? Why is it just its total output weights and not also its number of edges/centrality? How is the high output weight of a node related to its alignment with other eigenvectors, and is this always the case or just in this example?

      Demonstrates seems strong.

      Thank you for this important comment. We agree that "demonstrates" was too strong and have replaced it with "illustrates."

      It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error.

      We agree with this comment. We have revised the simulation and accompanying text to connect the results directly to the main theory based on . As responded to Comments (7) and (8), we have removed the eigenvector-based analysis and instead plotted the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly related to the estimation error via Equation (9). This change allows us to provide a clear theoretical explanation for why perturbing certain nodes minimizes the prediction error.

      Is this always the case or just in this example?

      We have also explicitly stated the limitations: whether a hub node or a specific subnetwork node is more effective depends on factors such as the outgoing edge weights from the hub and the individual damping rates of each mode. The revised text emphasizes that this simulation presents one example of perturbation location design, and the optimal strategy must be evaluated case by case using the theoretical framework of Equation (9).

      Results - Main Manuscript

      “As this simulation represents only one example of location design, its limitations and the corresponding countermeasures should be stated. In actual experiments, the most effective perturbation location depends on factors such as hub-node connectivity and modal damping rates, and B itself may not always be known a priori. In such cases, approaches such as iterative optimization of (described in a later section) and joint estimation of A and B (Supplementary Material B.3) provide systematic alternatives. Nevertheless, the results presented here provide an intuitive guideline: stimulation directed at hub nodes or at nodes driving heavily damped modes effectively excites the full set of dynamical modes and minimizes estimation error.”

      (10) What are the physical units for the time scales and stimulation amplitudes? For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic. By relating it to the eigenspectrum of A, which is supposed to be physiologically realistic, one can estimate the length of the stimulation window and support the above statement. Similarly, the impulse amplitude on page 13 is alpha=10<sup>20</sup> (a.u.). Also, Table 1 in the Supplementary has an extremely wide range of stimulation amplitudes. Is 10<sup>20</sup> a.u. a feasible amplitude in practice? And finally, please tell us which nodes the input was applied to.

      Thank you for your incisive comments. We have addressed each comment as follows.

      What are the physical units for the time scales and stimulation amplitudes?

      They don’t have physical units. This study is a theoretical investigation that focuses on the relative differences between passive observation and perturbation-based approaches, rather than providing precise predictions for specific experimental settings. The time scales and stimulation amplitudes are therefore expressed in arbitrary units.

      For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic.

      We agree that the original expression “unrealistic” was not appropriate given the arbitrary units. We have revised the text to clarify that the time scales are in arbitrary units and that the main point is about the relative difference in required data length between passive observation and perturbation-based approaches, rather than making an absolute claim about feasibility.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Is 10<sup>20</sup> a.u. a feasible amplitude in practice?

      Although we have already stated that the stimulation amplitudes are in arbitrary units, we agree that the original value of 10<sup>20</sup> was excessively large and could be misleading. We have revised the impulse amplitude from 10<sup>20</sup> to 10<sup>2</sup>, as 10<sup>20</sup> is physically unrealistic—it would imply a stimulus intensity many orders of magnitude beyond any conceivable experimental setting. The revised value of 10<sup>2</sup> is more plausible: for reference, TMS stimulation voltages exceed typical EEG amplitudes by roughly 4–6 orders of magnitude. The classification accuracy decreased slightly; however, our main conclusion regarding the efficiency of the perturbation-based approach under limited data remains unchanged. See Table B.1.

      Finally, please tell us which nodes the input was applied to.

      The stimulus location was optimized to minimize the . We have clarified this procedure in the main text and added a visual indication of the selected stimulation site (red circles) in Fig.7.

      Results - Main Manuscript

      “The neural signals were simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.

      (11) Page 13: "The controlled transition test was run with T = 1." Previously, T referred to the number of time steps. Is that the case here? If so, that seems hard to justify. If not, please tell us what T is (and, ideally, use a different symbol).

      Is that the case here?

      No. In this context, T does not denote the number of time steps.

      If not, please tell us what T is (and, ideally, use a different symbol).

      Thank you for your helpful comment. We have standardized the notation throughout the manuscript. In this paper, T consistently denotes the number of time steps (i.e., data length). In the sections “Neural State Classification” and “Neural State Transitions,” we had mistakenly used T to refer to time length. To resolve this inconsistency, we have added a separate column labeled “Data Length (T)” to Table B.1 for clarification and replaced the previous usage of T with “data length” where appropriate.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Results - Main Manuscript

      “The controlled transition test was run with T = 50. The control objective was to set nodes3 and 4 to 25 while keeping all other nodes at 0 without any movement. See Table B.1.”

      (12) Page 14: "where the matrix A (Fig. 10a) is designed to have 16 modes." What do you mean by has "16 nodes"?

      Thank you for pointing this out. By “16 modes,” we refer to 16 dynamical eigenmodes (i.e., 16 eigenvalue pairs). In the real-valued state-space representation used in the simulations, each complex conjugate pair corresponds to a 2-dimensional real block, resulting in a 32-dimensional system (32 nodes). We revised the wording to clearly distinguish between the number of dynamical modes and the dimensionality (number of nodes) of the state vector, to avoid confusion.

      Results - Main Manuscript

      “Iterative refinement of both the perturbation design and the estimation process progressively improves the accuracy of A. The time-series data is collected from 32 points, where the matrix A (Fig. 10a) is designed to have 16 oscillatory modes (i.e., 16 complex-conjugate eigenvalue pairs, yielding 32 eigenvalues in total).”

      (13) In Figure 10b, the y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps. This will make it clear that the estimation error drops by about 33%. It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      Thank you for this insightful comment. We agree with the reviewer that quantifying the expected improvement would be valuable for experimental practice. However, this simulation is a theoretical demonstration. Its primary purpose was to show that an iterative active perturbation approach can progressively converge to an optimal perturbation design even without prior knowledge of the true system, rather than to quantify a universally expected improvement rate.

      The y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps.

      We believe that the y-axis should start at the accuracy with optimal perturbation because the primary purpose of this simulation was to demonstrate that an iterative active perturbation approach can progressively converge. If we had started the y-axis at zero, the message would be visually obscured.

      Potential users of this method would want to know how much improvement they’re likely to see.

      We acknowledge that this is one example of the application of our method, and the magnitude of improvement depends on various factors such as network structure, noise level, and stimulation design. We should not mislead the readers. We have therefore maintained the y-axis starting point and added a clarifying statement in the revised manuscript to indicate that the simulation is case-specific rather than universal.

      Results - Main Manuscript

      “These results should be interpreted as a case-specific illustration rather than a universal gain, as the magnitude of improvement depends on factors such as network structure, recording duration, noise level, and stimulation design. This simulation demonstrates that our perturbation design framework enables the step-by-step refinement of system identification even without prior knowledge of the system.”

      (14) Please provide a derivation for the factorised covariance in Equation S.2. This is the equation that underpins the main result of the paper - the error in dynamical system estimation. Currently, it appears to be taken from [Hamilton, J. D. Time Series Analysis], yet we were unable to find this result in the book. Most of the derivations in Chapters 8.1 and 8.2 assume diagonal noise covariance, and even isotropic noise (cov = sigma*2 * I), which simplifies the particular case of the derivations. Could the authors provide a reference or the derivation for the case with full covariance?

      As described in our response to the Weakness above, the relevant result is Proposition 11.1 of Hamilton (1994), not Chapters 8.1–8.2. Proposition 11.1 states the asymptotic distribution of the vectorised OLS estimator in a VAR model with i.i.d. innovations whose covariance matrix Ω is an arbitrary positive definite matrix. The factorisation in our Eq. (S.2) corresponds directly to the revised manuscript we explicitly cite “Hamilton (1994), Proposition 11.1” at Eq. (S.1) and in Hamilton’s Proposition 11.1, with and . In state the i.i.d. assumption on ξ(t) in the preceding paragraph, so that the connection to this result is unambiguous. See Theoretical Details Supplementary.

      (15) Strongly perturbing/Exciting fast-decaying modes to give them more ’runway’ and increase observed variability makes intuitive sense for a normal system. But what would happen in the case of non-normal dynamics, where stimulating one dimension only transiently amplifies it, but then excites other dimensions? Non-normality breaks the alignment between PC components and dynamic modes (see Kumar, Ankit, Loren M. Frank, and Kristofer E. Bouchard. "Identifying feedforward and feedback controllable subspaces of neural population dynamics." arXiv preprint arXiv:2408.05875 (2024)), so eigenvectors of A and Σ<sub>X</sub> won’t align for non-normal dynamics. This alignment appears to be a hidden assumption of this paper. Should the normal dynamics then be stated as an assumption/limitation of the framework? Is the example in Figure 6a highly non-normal? Does considering out-degree provide an empirical approach, an alternative to ’enlargement’, to dealing with non-normality?

      We thank the reviewer for this insightful comment regarding non-normal dynamics.

      Should normal dynamics be stated as an assumption/limitation of the framework?

      No. Our framework does not assume normal dynamics. For example, Figure 5 and the revised Figure 6 illustrate non-normal cases. In Figure 5, the row and column norms of A differ substantially (row norms ≈ [1.96, 4.05, 1.33, 4.70]; column norms ≈ [6.14, 1.94, 1.39, 0.87]). In Figure 6, Node 7 is a hub with only outgoing edges, so row 7 of A is zero while column 7 has nonzero entries. In addition, the Nodes 5–6 and Nodes 3–4 are strictly one-way, leaving the corresponding off-diagonal block upper-triangular. Both features break the symmetry required for normality.

      We additionally computed a commutator-based non-normality index which equals 0 for any normal matrix and approaches for the canonical maximally non-normal example (the 2×2 nilpotent Jordan block). The non-normality index is 1.34 for Figure 5 and 0.41 for Figure 6. These values confirm that both are clearly non-normal.

      Is the example in Figure 6a highly non-normal?

      Yes. As stated above, the network structure of Figure 6a clearly indicates non-normality.

      Does considering out-degree provide an empirical approach, an alternative to ‘enlargement’, to dealing with non-normality?

      No. In the previous manuscript, the results of out-degree and eigenvector structure were provided as supplementary intuition. However, the central contribution of our framework lies in maximizing the minimum eigenvalue µ of the observed state covariance (Equation 9), and for systems where non-normality is strong and subnetwork structure is less modular, the iterative optimization of provides a principled, assumption-free method. We clearly mentioned this point in the revised manuscript as follows. See Results in the main manuscript.

      (16) In the iterative experiment at the very end of the paper (Figure 10), what was the strategy for designing ‘u<sub>design</sub>´? How many stimulations were applied in an iteration? Are you stimulating along eigenvectors? Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      Thank you for this important question. We have clarified the design rule and stimulation protocol in the revised manuscript, and address each sub-question below.

      How many stimulations were applied in an iteration?

      One stimulation session was applied in an iteration.

      Are you stimulating along eigenvectors?

      The optimal stimulation is designed as the target node of the perturbation is determined through numerical optimization that minimizes

      Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      The optimal stimulation is designed as a composite-frequency sinusoidal input encompassing all modes of the estimated Â, and the target node of the perturbation is determined through numerical optimization that minimizes . See Results - Main Manuscript.

      (17) While it is clear that the proposed active method performs better than passive observation, some results lack a comparison with stimulating random directions/nodes with a comparable control energy (Figure 7e & Figure 10b).

      Thank you for your helpful comment. Although the optimally designed stimulation outperforms random stimulation, the previous stimulation settings were not configured to explicitly demonstrate this difference. Therefore, we modified the network structure, recording length, and stimulation intensity so that both the main messages and the difference from random stimulation can be shown simultaneously. Accordingly, we have added a comparison with random stimulation in both Figure 7e and Figure 10b.

      In Figure 7e, we included a random stimulation condition where the target node is selected randomly, and the results show that the optimized stimulation outperforms random stimulation.

      These additions strengthen the evidence for the effectiveness of our proposed method compared to non-optimized approaches.

      (18) Is it reasonable to assume full observability of the system? It would be interesting to consider biases arising from the partial observability of the system, in the spirit of Figure 9, which looked at partial controllability.

      We thank the reviewer for this suggestion. We agree that partial observability is an important consideration, however it is out of scope for the current work. Thus, we have added a future direction in the Discussion addressing partial observability. We note that Takens’ embedding theorem and Hankel DMD enable recovery of a system’s eigenvalues from partial observations, and since our framework relies on the eigenvalue structure of A, the proposed perturbation design remains applicable under partial observability.

      Discussion - Main Manuscript

      “Two directions warrant further investigation: extending the framework to partial observability, and validating it through stimulation experiments. In experimental neuroscience, recordings are often limited to a subset of neural populations, resulting in partial observability. A growing body of work has leveraged delay-embedding techniques, represented by Takens’ embedding theorem [25], to reconstruct hidden dynamics from partial observations [2, 3, 23, 11]. Applying such techniques enables the estimation of the full connectivity matrix, thereby extending our framework to settings with partial observability. The second direction concerns experimental validation. Validating a theoretical framework through experimental design is an essential in bridging the gap between theory and practice.”

      Recommendations for improving the writing and presentation.

      (1) The word ’state’ is overloaded in Figure 7. When talking about neural state classification, the ’state’ refers to a regime guided by a distinct dynamics A (should it be A<sub>i</sub>? Figure 7A bottom). However, each dynamical system also has a ’state’. A different word should be used in Figure 7A and the corresponding text.

      We agree that the terminology is potentially confusing. To avoid the confusion, we replaced the term “state” with “task condition” and “neural signal” in the main text, and revised Fig. 7.

      Results - Main Manuscript

      “We designed a neural network with clearly distinct task conditions and considered a simulation setting in which these conditions are classified using signals of a fixed duration. These distinct task conditions consist of five types, each defined by a unique linear dynamical system characterized by differing eigenvalue spectra and connectivity topologies of matrix A (Fig. 7a). These task conditions are intended to mimic different cognitive or behavioral contexts. For example, in a typical motor task experiment, such conditions could correspond to motor execution or imagery involving the left or right hand, or resting state [1, 24]. The neural signal was simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.”

      (2) It would be helpful to provide dimensions of matrices around Equation S.2, since vectorization makes dimensions hard to track.

      We thank the reviewer for this helpful suggestion. We agree that explicitly stating the matrix dimensions improves readability, particularly around the Kronecker product where vectorization can obscure the size of the resulting covariance matrix. We have revised the text as follows (the equation S.2 is 3 now). See Theoretical Details of the Supplementary Materials.

      (3) The background sections of the paper would benefit from referring to similar active perturbation methods: Wagenmaker, Andrew, et al. "Active learning of neural population dynamics using two-photon holographic optogenetics." Advances in Neural Information Processing Systems 37 (2024): 31659-31687. Minai, Yuki, et al. "MiSO: Optimizing brain stimulation to create neural activity states." Advances in Neural Information Processing Systems 37 (2024): 24126-24149.

      We thank the reviewer for these helpful suggestions. We have incorporated Wagenmaker et al. (2024) and Minai et al. (2024) into the Introduction. Their works focus on developing algorithmic approaches to active stimulation design for specific experimental platforms, while our work aims to establish a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We cited these works and have clarified this distinction in the revised manuscript as follows.

      Introduction - Main Manuscript

      “In this paper, we propose a framework for designing the optimal perturbation input through control theory in neuroscience. We interpret neural dynamics as a control system [8, 6, 14, 12, 20, 24], and treat external perturbations as control inputs to design properties of neural stimulation (Fig. 1d). If the optimal perturbation input can be systematically designed, it becomes possible to steer the neural system toward states that are maximally informative (Fig. 1e), thereby enhancing the accuracy of the inferred connectivity (Fig. 1f). While recent studies have begun to develop algorithmic approaches to active stimulation design for specific experimental platforms [18, 28], a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We first describe how to formulate neural dynamics as a control system and how to estimate the model parameters from observed data. Building upon this formulation, we derive a theoretical basis that enables us to design the optimal perturbation inputs for the neural system identification. We demonstrate the validity and utility of this theoretical basis by exploring its implications for optimizing parameters of common neurostimulation techniques and by applying it to practical examples, including neural state classification [4, 9, 1, 24] and control of neural states [8, 13, 14, 12]. In these demonstrations, we define concrete problems and apply the theory to validate its practical utility.”

      Minor corrections to the text and figures.

      (1) In Figure 3, it would be helpful to point out that the plots are in the subspace spanned by the first three principal components.

      Thank you for pointing out. We revised the caption of the Fig.3 as follows. See Results - Main Manuscript.

      (2) Figure 4 caption: "eigenvectors" –> "eigenvectors of Σ<sub>X</sub>", just to make it crystal clear (since A also has eigenvectors).

      Thank you for your suggestion. We have revised the caption of Figure 4. See Results - Main Manuscript.

      (3) Page 5, Equation (8): Capital xi should be introduced in the main text of the paper as the covariance matrix for the noise term xi. Now it can only be understood after reading the Supplementary. Or maybe call the covariance matrix Σ<sub>ξ</sub>rather than Σ<sub>ξ</sub>? That will probably make it clearer.

      Thank you for this suggestion. We have made both changes. First, we introduced the noise covariance matrix explicitly in the main text immediately after Eq. (1), defining. Second, we replaced the notation Σ<sub>ξ</sub> with Σ<sub>ξ</sub>(lowercase subscript matching the noise variable throughout the main text and supplementary.

      (4) Page 7, Equation (14): The description of the equation states "The covariance matrices of x(t) for impulse inputs can be written as:... ". We assume this is supposed to be the "state vector," not "covariance matrices".

      Thank you for catching this. We have corrected the wording to “state vector” instead of “covariance matrices". See Results - Main Manuscript

      (5) Page 10, after introducing Figures 6a-b, potentially a sentence is missing (’...’ in the first line of the last paragraph).

      Thank you for pointing this out. We have removed the placeholder along with updating the stimulation settings for Figure 6.

      (6) Page 10, Figure 6 (e): A clearer labelling would be helpful, e.g., a title for the legend (e.g. node index) and a more informative title (e.g. Eigenvector alignment with nodes), a caption (what is the take-home message?), and y-axis labels (what is the ’value’?).

      Thank you for this helpful suggestion. We have revised the Figure 6 taking your suggestion into account. The new figure includes a clearer title, axis labels, and an informative caption that highlights the key take-home message. See Results - Main Manuscript

      (7) Page 11, paragraph 1: The sentence "...(as discussed in Section)" is missing a section reference.

      Thank you for pointing this out. We have replaced the section reference placeholder with the equation reference to Equation (9), which is the relevant theoretical result.

      (8) Page 13, caption of Figure 8c: compputed -> computed.

      We have corrected this typo. We also reviewed the manuscript for any similar typographical errors and corrected them.

      (9) Page 14 Figure 9: The y-axis label for "Controlled State Process" plots is missing.

      Thank you for catching this. The y-axis was hidden by other elements in the figure. We have revised the figure layout to ensure that the y-axis label is visible.

      (10) Page 16: The sentence "As described in Section, a preliminary ... " is missing a section reference

      Thank you for pointing this out. We have corrected the missing reference. The sentence now reads “as shown in Fig. 10” rather than the incomplete “as described in Section.”

      (11) Page 16, Fig. 10c: It looks like the errors have inconsistent color ranges. A shared colorbar would help.

      Thank you for your suggestion. We have revised Figure 10c to use a shared colorbar across all subplots.

      (12) Page 17, Paragraph preceding eq. 16: "the eigenvectors of X" --> "the eigenvectors of Sigma_X".

      Thank you for catching this. We have revised the text. See Methods - Main Manuscript.

      (13) S.1: This is a GLS, not an OLS estimator, if this assumes full noise covariance.

      As stated in our response to the Weakness above, we assume the noise term ξ(t) is i.i.d. with covariance matrix Σ<sub>ξ</sub>, and the estimator we analyze is the OLS estimator. See Theoretical Details - Supplementary Material.

      (14) S.2: Unclear that⊗is the Kronecker product (not outer), as it is not defined.

      We thank the reviewer for pointing out this ambiguity. In the revised Supplementary Material, we have explicitly explained the Kronecker product with a reference to Hamilton [10, Appendix A.4, p. 732]. See Theoretical Details - Supplementary Material.

      (15) S.51: There is an accidental comma between alpha and A after "xdiff(t) ="

      Thank you for pointing this out. We have removed the accidental comma. See Theoretical Details - Supplementary Material.

      (16) Section B.2: It would be useful to have the simulation details for that section (as is given for the other section in B.1).

      We thank the reviewer for this helpful suggestion. We have added a dedicated “Simulation details” paragraph to Appendix B.5 (the section containing Fig. B.4) so that the setup is now described with the same level of specificity as the other appendix sections. See Experimental Conditions and Results - Supplementary Material.

      References

      (1) Irma N Angulo-Sherman, Marisol Rodríguez-Ugarte, Nadia Sciacca, Eduardo Iáñez, and José M Azorín. Effect of tDCS stimulation of motor cortex and cerebellum on EEG classification of motor imagery and sensorimotor band power. J. Neuroeng. Rehabil., 14(1):31, April 2017.

      (2) Hassan Arbabi and I Mezić. Computation of transient koopman spectrum using hankeldynamic mode decompoisition. APS, page G1.009, November 2017.

      (3) Steven L Brunton, Bingni W Brunton, Joshua L Proctor, Eurika Kaiser, and J Nathan Kutz. Chaos as an intermittently forced linear system. Nat. Commun., 8(1):19, May 2017.

      (4) Adenauer G Casali, Olivia Gosseries, Mario Rosanova, Mélanie Boly, Simone Sarasso, Karina R Casali, Silvia Casarotto, Marie-Aurélie Bruno, Steven Laureys, Giulio Tononi, and Marcello Massimini. A theoretically based index of consciousness independent of sensory processing and behavior. Sci. Transl. Med., 5(198), August 2013.

      (5) Karl Deisseroth. Optogenetics. Nat. Methods, 8(1):26–29, January 2011.

      (6) Shikuang Deng, Jingwei Li, B T Thomas Yeo, and Shi Gu. Control theory illustrates the energy efficiency in the dynamic reconfiguration of functional connectivity. Commun. Biol., 5(1):295, April 2022.

      (7) Shrey Grover, Renata Fayzullina, Breanna M Bullard, Victoria Levina, and Robert M G Reinhart. A meta-analysis suggests that tACS improves cognition in healthy, aging, and psychiatric populations. Sci. Transl. Med., 15(697):eabo2044, May 2023.

      (8) Shi Gu, Fabio Pasqualetti, Matthew Cieslak, Qawi K Telesford, Alfred B Yu, Ari E Kahn, John D Medaglia, Jean M Vettel, Michael B Miller, Scott T Grafton, and Danielle S Bassett. Controllability of structural brain networks. Nat. Commun., 6:8414, October 2015.

      (9) Mark Hallett, Riccardo Di Iorio, Paolo Maria Rossini, Jung E Park, Robert Chen, Pablo Celnik, Antonio P Strafella, Hideyuki Matsumoto, and Yoshikazu Ugawa. Contribution of transcranial magnetic stimulation to assessment of brain connectivity and networks. Clin. Neurophysiol., 128(11):2125–2139, November 2017.

      (10) James Douglas Hamilton. Time Series Analysis. Princeton University Press, Princeton, 1994.

      (11) Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. InputDSA: Demixing then comparing recurrent and externally driven dynamics. arXiv [q-bio.NC], November 2025.

      (12) Shunsuke Kamiya, Genji Kawakita, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Optimal control costs of brain state transitions in linear stochastic systems. J. Neurosci., 43(2):270–281, January 2023.

      (13) Teresa M Karrer, Jason Z Kim, Jennifer Stiso, Ari E Kahn, Fabio Pasqualetti, Ute Habel, and Danielle S Bassett. A practical guide to methodological considerations in the controllability of structural brain networks. J. Neural Eng., 17(2):026031, April 2020.

      (14) Genji Kawakita, Shunsuke Kamiya, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Quantifying brain state transition cost via schrödinger bridge. Netw. Neurosci., 6(1):118– 134, February 2022.

      (15) Hassan K Khalil. Nonlinear systems. Prentice-Hall, Upper Saddle River, NJ, 2002.

      (16) Paul K LaFosse, Zhishang Zhou, Jonathan F O’Rawe, Nina G Friedman, Victoria M Scott, Yanting Deng, and Mark H Histed. Single-cell optogenetics reveals attenuationby-suppression in visual cortical neurons. bioRxivorg, page 2023.09.13.557650, May 2024.

      (17) Andres M Lozano, Nir Lipsman, Hagai Bergman, Peter Brown, Stephan Chabardes, Jin Woo Chang, Keith Matthews, Cameron C McIntyre, Thomas E Schlaepfer, Michael Schulder, Yasin Temel, Jens Volkmann, and Joachim K Krauss. Deep brain stimulation: current challenges and future directions. Nat. Rev. Neurol., 15(3):148–160, March 2019.

      (18) Yuki Minai, Matthew Smith, Joana Soldado-Magraner, and Byron Yu. MiSO: Optimizing brain stimulation to create neural activity states. In A Globerson, L Mackey, D Belgrave, A Fan, U Paquet, J Tomczak, and C Zhang, editors, Advances in Neural Information Processing Systems 37, volume 37, pages 24126–24149, San Diego, California, USA, 2024. Neural Information Processing Systems Foundation, Inc. (NeurIPS).

      (19) Davide Momi, Zheng Wang, and John D Griffiths. TMS-evoked responses are driven by recurrent large-scale network dynamics. Elife, 12(e83232), April 2023.

      (20) Ali Moradi Amani, Amirhessam Tahmassebi, Andreas Stadlbauer, Uwe Meyer-Baese, Vincent Noblet, Frederic Blanc, Hagen Malberg, and Anke Meyer-Baese. Controllability of functional and structural brain networks. Complexity, 2024(1), January 2024.

      (21) Norman S Nise. Control Systems Engineering. John Wiley & Sons, 8 edition, 2020.

      (22) Katsuhiko Ogata. Modern Control Engineering. Prentice Hall, 2010.

      (23) Mitchell Ostrow, Adam Eisen, and Ila Fiete. Delay embedding theory of neural sequence models. arXiv [cs.LG], June 2024.

      (24) Yumi Shikauchi, Mitsuaki Takemi, Leo Tomasevic, Jun Kitazono, Hartwig R Siebner, and Masafumi Oizumi. Quantifying state-dependent control properties of brain dynamics from perturbation responses. J. Neurosci., page e0364252025, December 2025.

      (25) Floris Takens. Detecting strange attractors in turbulence. In David Rand and Lai-SangYoung, editors, Dynamical Systems and Turbulence, Warwick 1980, volume 898 of Lecture Notes in Mathematics, pages 366–381. Springer, Berlin, Heidelberg, 1981.

      (26) Liam C Tapsell, Matheus D Pinto, Ann-Maree Vallence, Casey Whife, Maria Luciana Perez Armendariz, Shaswat Senger, Jack Andringa-Bate, Dana Hince, and Myles C Murphy. What are the optimal transcranial direct current stimulation parameters and design elements to modulate corticospinal excitability? a systematic review and longitudinal meta-analysis. Neurol. Res. Pract., 7(1):86, November 2025.

      (27) Lei Tong, Shanshan Han, Yao Xue, Minggang Chen, Fuyi Chen, Wei Ke, Yousheng Shu, Ning Ding, Joerg Bewersdorf, Z Jimmy Zhou, Peng Yuan, and Jaime Grutzendler. Single cell in vivo optogenetic stimulation by two-photon excitation fluorescence transfer. iScience, 26(10):107857, October 2023.

      (28) Andrew Wagenmaker, Lu Mi, Marton Rozsa, Matthew S Bull, Karel Svoboda, Kayvon Daie, Matthew D Golub, and Kevin Jamieson. Active learning of neural population dynamics using two-photon holographic optogenetics. Adv. Neural Inf. Process. Syst., 37:31659–31687, 2024.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Summary of Revisions Performed:

      We have clarified the qPCR methodology in the methods section and stated the housekeeping gene GAPDH to address potential misunderstandings.

      We have assessed hearing in the generated HA-tagged mouse lines and included an adequately powered ABR measurements analysis in the revised manuscript as a supplemental figure.

      We have included powered DPOAE experiments in both ATP8B1 and TMEM30B KO mice to strengthen the findings of the ABRs.

      We have clarified the presentation of the z-stack in Figure 1F.

      We have elaborated on the analysis for Figure 7B to strengthen comprehension by readers.

      We have revised the statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      While we appreciate the suggestion to examine TMEM30B localization on the ATP8B1 KO background, this is not feasible within a reasonable timeframe; we have clarified this limitation in the manuscript.

      We have incorporated relevant prior work (e.g., George and Ricci, 2026) demonstrating minimal Annexin V labeling prior to P6 and lack of PS externalization in TMC1/2 double knockout models.

      We have clarified that hearing thresholds for TMEM30B-HA and ATP8B1-HA lines were addressed in this study, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      We have softened statements regarding HA-tag insertion and clarified that, to our knowledge, localization and function are not disrupted, while acknowledging this as a potential limitation.

      We have revised the Methods section to clarify differences in fluorescence measurements across experiments.

      Public Reviews:

      Reviewer #1 (Public review):

      Figure1D.

      The authors should clarify how the qPCR data were normalized and specify the reference (housekeeping) genes used. This information is necessary to evaluate the robustness and comparability of the gene expression data.

      We thank the reviewer for this comment. qPCR data were normalized to GAPDH as the reference (housekeeping) gene. We have clarified this in the Methods section to ensure transparency and reproducibility.

      (2) Figure 1F.

      The lack of F-actin staining at the hair cell base raises the possibility that the permeabilization conditions may have limited antibody access to certain membrane regions. This is especially important given that the authors used a gentle permeabilization agent such as saponin to preserve membrane integrity. Because the authors conclude that ATP8B1 and TMEM30B are localized "almost exclusively to OHC bundles and the apical membrane, with minimal staining in the remaining plasma membrane," (line 128). Including co-labeling with a plasma membrane marker or more comprehensive F-actin visualization of lateral and basal regions would help ensure that the restricted localization is biological rather than technical. In the absence of such controls, the localization claim may be somewhat overstated and should be tempered accordingly.

      We thank the reviewer for this important point. The image shown represents a single z-slice from a larger stack, and the hair cell body lies outside the plane of this section. To clarify this, we revised the accompanying text.

      (3) Figure 7B.

      Although quantification of ATP8B1-HA intensity at the bundle appears similar between WT and Cib2 KO samples, the representative image suggests that some bundles lack detectable labeling. To better capture phenotype variability, it would be helpful to include an additional quantification showing the fraction or number of bundles with detectable ATP8B1-HA signal in Cib2 KO mice.

      We thank the reviewer for this suggestion. We have clarified the quantification of the fraction of hair cell bundles with detectable ATP8B1-HA and TMEM30B-HA signal per field of view. Although the representative images may give the impression that some hair bundles lack staining, this is due to changes in ATP8B1-HA and TMEM30B-HA distribution within the cell body. In all cases, detectable ATP8B1-HA and TMEM30B-HA signal remained present in the hair bundles.

      (4) Lines 346-349

      The manuscript suggests that IHCs lack stereocilia-enriched P4-ATPases. However, this conclusion is not directly supported by the presented data. The authors should either provide supporting localization or expression data for other P4-ATPases or soften the statement to indicate that no stereocilia-enriched P4-ATPases were detected under the conditions examined.

      We agree with the reviewer and have revised this statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      Recommendations:

      (5) The authors convincingly demonstrate that TMEM30B loss results in ATP8B1 mislocalization. While not essential to the central conclusions, examining TMEM30B localization in ATP8B1 KO hair cells would clarify whether this interdependence is reciprocal, as described for other P4-ATPase-CDC50 complexes.

      While we agree that this experiment would provide valuable information, performing it would require generation of a compound mouse line carrying both the TMEM30B-HA allele and the ATP8B1 knockout allele. This work is beyond the scope of the current revision and cannot be completed within a reasonable timeframe.

      (6) Lines 359-374. The discussion of Annexin V labeling is careful and balanced. This paragraph would benefit from referencing other studies that showed minimal Annexin V labeling in healthy P6 organ of Corti, reinforcing that robust PS externalization in the present study is pathological rather than developmental.

      We thank the reviewer for this suggestion and have incorporated relevant prior work, including George and Ricci (2026), which demonstrates minimal Annexin V labeling prior to P6 and further supports our interpretation.

      (7) Lines 392-399.

      The proposed feedback model linking MET activity and ATP8B1-TMEM30B localization is compelling. The discussion could be strengthened by noting that in TMC1/2 double knockout hair cells, PS externalization is not observed, consistent with the idea that flippase activity becomes critical specifically when scrambling occurs. The mislocalization observed in Cib2 KO hair cells further supports the coupling between TMC-mediated scrambling and flippase-mediated membrane restoration.

      We agree and have revised the text to include that TMC1/2 double knockout hair cells do not exhibit phosphatidylserine externalization, supporting the idea that flippase activity becomes critical in the context of scrambling.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. It would be good to know, for each knock-in model (TMEM30B, ATP8B1, ATP8A1, ATP8A2, and ATP11A), whether the HA-tagged protein is causing any issues with the mice and particularly with hearing (ABRs). Are these mice normal? Can they hear? These data are missing.

      We thank the reviewer for raising this important point. In this study, we focus on TMEM30B-HA and ATP8B1-HA mouse lines, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      Both TMEM30B-HA and ATP8B1-HA mice are viable and exhibit normal breeding and ageing. We have included adequately powered ABR measurements of both TMEM30B-HA and ATP8B1-HA which indicate wild-type–like hearing thresholds.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential mislocalization causing any functional phenotypes? (ABRs of point 1). I find it surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out). Given that this manuscript will likely become foundational, and that there is evidence that at least two of the other flippases are involved in hearing loss, it would be good to provide more information about the mice and HA-tagged proteins in the other knock-ins (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA). Depending on the data available for the knock-ins, the authors may want to discuss these scenarios and soften the statement indicating that inner-hair cells may lack flippase activity altogether.

      We appreciate this concern. To our knowledge, the HA tag does not appear to disrupt localization or function of the tagged proteins. However, we agree that this cannot be fully excluded. We have therefore softened our conclusions about IHC flippases and clarified that additional flippases (ATP8A1, ATP8A2, ATP11A) are under investigation and will be described in a separate study.

      (3) Expression of ATP8B1 at P0 (Figure 1D), when there should not be protein in outer hair cells yet seems high. Does this mean that other cells in the cochlea also express ATP8B1? Is this a concern?

      We thank the reviewer for this observation. We interpret the elevated ATP8B1 transcript levels at P0 as reflecting transcription that precedes detectable protein accumulation in OHC stereocilia. While expression in other cochlear cell types cannot be excluded, we did not detect ATP8B1-HA immunolabeling outside hair cells in the knock-in model.

      (4) Fluorescence scales in Figure 6 B and D and Figure 7 B and D are very different. So are the values for WT. One would expect that the WT would be similar in all cases (at least within the same compartments), given that the methods section indicates that "All images were collected using identical acquisition parameters, including zoom and laser power, across genotypes". If WT shows such variability, how can we compare?

      We appreciate the need for clarification. Identical acquisition parameters were maintained within each experiment used for direct comparison (e.g., within a given panel). However, different panels (e.g., Figures 6B vs. 6D) were acquired on different days using different imaging settings.

      We have revised the Methods section to explicitly state this and clarify that comparisons are intended only within panels, not across experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 42: When discussing TMC similarity to TMEM16 scramblases, it may be helpful to mention that some TMEM16 family members (TMEM16A and B) function as ion channels, highlighting the dual ion/lipid functionality within the superfamily. The similarity to TMEM63/OSCA ion channels and lipid scramblases could also be noted. The fact that TMC, TMEM16, and TMEM63/OSCA belong to the same superfamily would provide a broader context. I also suggest referencing the work that initially suggested this relationship: (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0192851)

      We have included this citation and expanded on this discussion in the revised version.

      (2) Line 45: Consider including the recent Cryo-EM structure of CeTMC2 (PNAS, 2023), which provides structural insight into TMC-lipid interactions.

      We have included this citation in the revised version.

      (3) Line 62: For precision, consider removing "calcium-activated," as caspase-activated scramblases also disrupt membrane asymmetry.

      For precision, we have removed “calcium-activated” in the revised version.

      (4) Line 71: Clarify whether this refers to "fusion of membranes" or "cell fusion."

      We have clarified this statement to mean cell-cell fusion.

      (5) Line 85: Consider citing studies showing constitutive PS externalization in TMC1 mutant mouse models linked to deafness.

      We have added a citation to show that constitutive PS externalization is linked to deafness (Ballesteros and Swartz, 2022, and Beurg et al. 2025).

      (6) Line 93: TMEM30C is not discussed. A brief comment on its expression or relevance in hair cells would provide completeness.

      We have added a brief statement regarding TMEM30C and cited prior work describing its expression pattern (Osada et al. 2007).

      (7) Figure 1A: Use distinct colors for the P4-ATPase and CDC50 subunit rather than a rainbow scheme to improve clarity.

      We have retained the original color scheme in this panel.

      (8) Figures 3C-D and 5C-D: Increase legend symbol size for clarity. Update Y-axis labels to "Number of OHCs/100 μm" and "Number of IHCs/100 μm." Correct "um" to "μm."

      We changed the legend to improve the presentation of these panels to be more legible and changed the measurement to μm.

      (9) Figures 3F, 5F, 5H: Add scale bars.

      We have added scale bars to these figures.

      (10) Figure 7: The confocal images (A, C) show the bundle on top and cell body below, but the quantification (B, D) is in the opposite order. Reorganizing the panels for consistent orientation would improve clarity.

      We have reorganized the panels to improve clarity.

      (11) ABR measurements: Please specify the sex of the mice tested or clarify whether both sexes were included.

      We have included both male and female mice in this study as there were no differences in hearing function. We have added this clarification to the methods section under hearing tests.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1A, the panels show CDC50. I would either change to TMEM30B or mention in the caption that TMEM30B is also known as CDC50 as labeled in the figure.

      We have changed CDC50 to TMEM30B.

      (2) Figures 1F and 1G are missing scale bars.

      We have added scale bars to these figures.

      (3) Figures 2 C, D, and 5 C, D - difficult to tell what's what in the legend. Perhaps make symbols larger in front of WT P17, KO P17, etc.?

      We changed the legend to improve the presentation of these panels.

      (4) Text under "TMEM30B is required for hearing and OHC maintenance". There is a difference in phenotype between the TMEM30B (Figure 5C) and ATP8B1 (Figure 3C) knockouts that is not discussed, as apical and middle cells seem to be okay. Should this be discussed?

      We appreciate this observation. We have elected not to expand the discussion of these regional differences because apical and middle hair cells also undergo degeneration at later ages (after P30), suggesting that the observed differences primarily reflect the timing of degeneration rather than distinct underlying mechanisms.

      (5) In the discussion text, under "Why do ATP8B1/TMEM30B-deficient OHCs die?", "Tmc1/2 or Cib2" should probably be "TMC1/2 or CIB2" or "Tmc1/2 or Cib2"

      We have changed this to read TMC1/2 or CIB2.

      (6) The methods section states "..., whereas non-significant comparisons are not shown." However, non-significant p values are shown in Figures 7B and D (bottom panels).

      We have removed the nonsignificant comparisons from Fig 7B and D to be consistent with the rest of the paper.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved.

      The integration across biological scales and models is a major strength. The state dependency analysis provides additional, strong support. The methodology and statistical approaches used are thoughtfully presented and rigorously applied.

      In particular, this paper provides a strong response to the findings from Gliske et al, Nat Commun 2018. This study utilized long-term data analysis to uncover low rates of FRs detected from most recording sites, suggesting spurious detections, although FRs were concentrated within seizure onset areas.

      We fully agree with this comparison. Although we had already cited this paper, we now further emphasize this observation in the Discussion:

      “Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence.”

      Weaknesses:

      The authors clearly aimed to use a statistical rather than a mechanism-based approach in this work. However, the paper's framing of true fast ripples as oscillatory events with stochastic fast ripples considered as confounders does not take prior investigations into biological mechanisms, particularly prior studies that point to an important role for stochastic fast ripples in some contexts. Incorporating recognition of these mechanisms would strengthen the manuscript and provide a more complete and nuanced characterization.

      Some examples from the literature:

      Eissa et al, eNeuro 2016, a paper that closely parallels this manuscript but took a mechanistic rather than statistical approach, showed that fast ripples can arise from population paroxysmal depolarizations - a key feature of epileptiform discharges - as temporally clustered, jittered population firing, with FRs appearing in LFP or EEG due to summated postsynaptic potentials (which are slower than action potentials and can generate signals in the high gamma range).

      Foffani et al., 2007, Neuron, and Ibarz et al., 2010, J Neurosci, argue that FRs are pseudo-oscillations created by jittered neuronal populations in the setting of altered spike timing.

      Smith et al., 2020, Sci Rep, contrasts FR characteristics in different regimes, i.e., intact inhibition early in a seizure vs. implied collapse of inhibition after recruitment. Schlingloff et al., 2025, J Neurosci, reported analogous findings in an animal model.

      We agree with the reviewer that even stochastic events may be of biological importance and an increase in stochastic events will occur when there is an increase in synchronisation and excitability, two properties of pathological cortex. We also don’t disagree that FRs can occur as distinct entities, although our work indicates that most are due to chance.

      To address this point, we have clarified our claims in the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      In addition, we expand on these points at various junctures in the Discussion. In particular, we reiterate our assertion that FRs may still be a useful biomarker, but that their interpretation should be moderated to reflect the fact that they often occur by chance:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      The computational model and subtraction approach provide a strong case for the random emergence of clustered activity in the high gamma band, given its assumptions. However, any such modeling effort needs to account for inhibitory activity, including impaired inhibitory function that is expected in epileptic brain regions, which has a strong modulating effect on excitatory firing and is thought to play a significant role in FR generation.

      We appreciate the reviewer’s concerns, but we believe that the impact of inhibitory interneuron activity on excitatory firing rates and synchronisation is incorporated indirectly into our simulations, while keeping our model as parsimonious as possible by not directly incorporating interneuron activity into our simulations. We have addressed this point in the Methods section:

      “Varying synchrony allowed us to test the impact, on the network, of inhibitory cells, which have been shown to favour synchrony (Bocchio et al., 2024; Cobb et al., 1995).”

      The shuffling procedure aims to preserve the power spectrum but randomizes high frequency phase (>200 Hz). However, this procedure removes biologically meaningful spike timing correlations, as well as structured cross-frequency coupling. The subtraction method thus likely underestimates the incidence of structured "distinct" FRs, while perhaps overestimating "chance" FRs due to biologically infeasible activity, making the statement that most FRs are due to chance correlation too strong.

      We appreciate this concern, which is especially important given that our results depend crucially on the validity of our shuffling procedure (as described in the Discussion). To address this issue, we have implemented an additional shuffling algorithm that preserves cross-frequency coupling (see last section of the Results, especially Supplementary Fig. 10f). This method showed no qualitative difference, compared with other alternative methods presented in Supplementary Fig. 10. These new results are described in the Methods section:

      “Last, we also implemented a method based on wavelet-IAAFT with preservation of cross-frequency coupling, since fast ripples are typically locked to low-frequency phase (Sheybani et al., 2019). The code detects the highest phase-amplitude coupling (PAC) in the original signal between [300-6000 Hz] for amplitude and several low-frequency bands ranging from 2-20 Hz, bandwidth of 3 Hz. PAC is computed using the modulation index (Tort et al., 2008). Then, in the shuffled signal under construction and during convergence testing of PSD (see above), the PAC between high-frequency part of the signal (300-6000 Hz) and the identified low frequency for phase is normalized to that of the highest PAC identified earlier.”

      The kainate findings underscore this point: the increase in the number of FR detections could be, as the authors state, an increase in chance clustering due to increased network excitability generally. However, the likelihood of a parallel increase in pathological FRs cannot be ruled out, given likely pro-epileptic alterations in spike timing and circuit function.

      We appreciate the reviewer’s point but wish to re-emphasise our interpretation of these findings – that the observed increase in the incidence of FRs occurs as a result of increased network excitability/synchrony, secondary to the pathological mechanisms of epilepsy. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-duration FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Weaknesses:

      I found only one serious weakness in this study, but it is one that is of importance. Although the original work on FRs was done in an experimental model of epilepsy, the field really became prominent when ripples and fast ripples were found first in microelectrode recordings of epileptic patients and then in the intracerebral EEG of such patients. Numerous studies have been performed since then, with a valuable meta-analysis including 700 patients (Wang Z, Guo J, van 't Klooster M, Hoogteijling S, Jacobs J, Zijlmans M. Prognostic Value of Complete Resection of the High-Frequency Oscillation Area in Intracranial EEG: A Systematic Review and Meta-Analysis. Neurology. 2024 May 14;102(9). Although the consensus at this point is that FRs are not the ideal and totally specific marker of epileptic tissue that many thought it could be, FRs are nevertheless much more frequent in epileptic tissue than in non-epileptic tissue and are a solid biomarker.

      We agree with the reviewer, and do not intend to challenge the role of FRs as a marker of the seizure-onset zone, and potentially the epileptogenic zone. Instead, the aim of this study was to address the question of whether FRs are generated by intrinsic pathological mechanisms, or whether they arise due to the chance co-occurrence of action potentials that follow different dynamics in epileptogenic parenchyma. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      It is also well established that they are much more frequent in NREM sleep than in wakefulness, as reported in the original paper of Staba et al (Staba RJ, Wilson CL, Bragin A, Jhung D, Fried I, Engel J Jr. High-frequency oscillations recorded in human medial temporal lobe during sleep. Ann Neurol. 2004 Jul;56(1):108-15., not mentioned in this paper) and in the study of Bagshaw et al (2009). In this last paper, using SEEG in various brain regions, the average rate of FRs in NREM sleep is about 6 times that in wakefulness. In the paper by Staba, with microelectrodes in mesial temporal structures, it is about twice. As a separate issue, the paper of Fraucher et al (Frauscher B, von Ellenrieder N, Zelmann R, Rogers C, Nguyen DK, Kahane P, Dubeau F, Gotman J. High-Frequency Oscillations in the Normal Human Brain. Ann Neurol. 2018 Sep;84(3):374-385), which is not quoted, found that, in an extensive sample, non-epileptic human tissue sampled with SEEG generated extremely rare FRs (an average rate of 0.04/min/channel, i.e. 1 every 25 min).

      The results above are mentioned because they do not fit with the data provided in the present study: FRs are much more frequent in NREM sleep than in wakefulness in human epileptic patients, and they are much more frequent (not 30% more, but many hundreds of percent more) in epileptic tissue than in non-epileptic human tissue. The fundamental phenomenon of interest is, I believe, the FRs in epileptic patients. The animal experiments, tissue studies, and simulations are models to study the human phenomenon. With respect to the modulation by sleep and the differentiation between epileptic and non-epileptic tissue, it seems that the systems studied in this paper are not good models of the human condition. The human results presented in the study only reflect wakefulness recordings, which is not the condition in which most HFO studies have been done and in which most HFOs occur. The authors refer to the study of long-term fluctuations in HFO rates by Gliske et al. (2018) to say that one has to be careful with the results regarding sleep, for example, Bagshaw et al (2009), but the clear predominance in of HFOs in NREM sleep has been observed by many studies. The cautions regarding fluctuations over extended periods also apply to the awake human data analyzed in this study. The study's conclusions regarding the generation of FRs are therefore questionably applicable to the human condition. I do not dispute their validity for the models and situations in which they were studied.

      We looked at this in more detail. Our simulations were intended to test how the incidence of FRs can vary with different parameters of network activity (neuronal count, firing rate, synchronization). Indeed, since their incidence is known to vary across regions and within regions and across states, we wanted to test how FRs are controlled by different factors. As such, we do not wish to draw firm conclusions about the observed sleep-wake changes in FR incidence in rodents, and how it relates to humans – evidence shows that pathological FRs in rodents do not display state-specific preferential occurrence (Ewell et al., 2019). We have added new text to the Abstract and Discussion to emphasize this.

      Abstract:

      “Our simulations showed that chance aggregation can generate fast-ripples and that their incidence changes depending on brain state, an observation that we confirmed in our rodent data.”

      We acknowledge that previous publications have reported higher rates during sleep, although with shorter recordings than in our rodent recordings (Staba, 2004: one night; Bagshaw, 2009: 10 min; Frauscher, 2018: 20 min – only sleep recordings). We have rewritten the part of the Discussion on the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009). Hence, FRs reflect and are highly susceptible to changes in network excitability.”

      Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high-frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies, i.e., 250Hz, well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out-of-phase fashion, or rather at random intervals, may contribute to a spectrum of HFOs ranging from 250-500Hz that are observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate, and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated beyond chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle, could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy, and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization, and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Weaknesses:

      (1) Some of the simulated FRs appear short in duration and may not meet standard detection and definition criteria, potentially influencing validity.

      We thank the reviewer for raising this important concern. To address this issue, we quantified and compared the duration of FRs in original and shuffled rodent data. Consistent with the reviewer’s suspicions, we found that FRs in shuffled signals are shorter than FRs in original signals. This is important because it shows that: (i) a longer duration should be considered a core feature of genuine FRs; and (ii) depending on the basal duration of FRs, the shuffling procedure will lead to different ratios of genuine to stochastic FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      And Discussion accordingly:

      “In addition, we show that long durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      (2) The neuronal culture approach does not directly test random insertion of action potentials, limiting interpretation.

      Neither the neuronal culture approach, the rat data or the human data directly test random insertion of action potentials. The insertion of random action potentials is only performed in the simulated data to test if FRs can arise from the chance insertion of action potentials. Once this was confirmed in the simulations, we then used the shuffling procedure in biological data to test if FRs are more frequent than expected by chance.

      (3) Sleep is treated as a homogeneous state in the rat dataset, without accounting for stage-specific differences in synchronization, which may affect the results and interpretation.

      We agree with the reviewer, but our primary aim was to answer the question of whether FRs can arise by chance. Although it was interesting to see that, in our longitudinal rodent data, the incidence of FRs varies across the sleep-wake cycle, any sleep-stage-specific changes are beyond the scope of this work.

      (4) The analyses conducted in human data lack direct comparison with sleep data.

      We agree that it would have been useful to investigate variations in the incidence of FRs across the sleep-wake cycle in human microelectrode recordings. Unfortunately, however, such sleep recordings were not available. Hence, while we cannot compare variations in FR incidence across brain states between humans and animal models, our conclusions that FRs arise mostly by the chance co-occurrence of action potentials still holds.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please indicate where corrections for multiple comparisons were used.

      P-values corrected for multiple comparisons are indicated by the accompanying phrase: “adjusted p-value”. We had previously omitted to mention this once in the Results, which we have now corrected.

      (2) Delta amplitude is likely sufficient for detecting sleep-wake transitions, but the beta/delta ratio is better supported in the literature. Do the results change if beta activity is incorporated?

      We have now computed the beta (15-40 Hz) to delta (0.5-4 Hz) ratio and find that this is closely correlated with delta across time. We have updated the Results accordingly:

      “Importantly, these findings were robust to the specific method used to detect FRs (Supplementary Fig. 4c) (Padmasola et al., 2024; Sheybani et al., 2019, 2018). Also our use of delta power to identify periods of presumed wakefulness and sleep was highly (negatively) correlated with an alternative method of using the beta-to-delta ratio across time (another marker of increased vigilance; (Fraigne et al., 2023), see Supplementary Fig. 4d).”

      Methods:

      “We further verified that delta power across time displayed similar fluctuations to beta (15-40 Hz) to delta power ratio, another marker of vigilance (Fraigne et al., 2023).”

      And we updated Supplementary Fig. 4d

      “(d) Beta to delta power ratio across time is superimposed over delta power across time. There is a strong (inverse) correlation between the two time-series (inset), which is confirmed by the correlation coefficient across animals (right).”

      (3) Figure 2's axis labeling with the 3D plots is hard to read.

      We have enlarged the font size.

      (4) The scaling of the histogram in Figure 3 is unclear.

      This was on omission. The scale has now been added to Figure 3.

      (5) There is a risk of overfitting in the regression model. Was cross-validation used?

      We have now repeated this analysis with cross-validation, without any qualitative impact on the results (e.g. the model still performs well above chance). We have updated the Methods:

      “To further confirm the performance of GBT, we used a cross-validation procedure where the GBT is trained on 80% of data and then tested on the 20% remaining. The procedure is repeated 1000 times and the r<sup>2</sup> is saved at each round. We repeated the analysis with randomization of the outputs across 1000 rounds and saved this null distribution r<sup>2</sup>. We then compared the performance against original data.”

      Legend of Fig. 3:

      “(e) Performance of the GBT classifier using cross-validation (training: 80% of data; test: 20% remaining) using original (orange) and shuffled (blue) data. The difference is significant (paired t-test, p<0.0001).”

      And Results:

      “Furthermore, using a cross-validation approach with 80% of the data as training set and the remaining 20% as the test set, we obtained a significantly higher explained variance than when outputs were shuffled across the 125,000 solution points (paired t-test, p<0.0001, Fig. 3e), […]”

      Reviewer #2 (Recommendations for the authors):

      Maybe I missed it, but I did not find the length of human data analyzed or how the sections were selected.

      Apologies for this omission. The methods have been updated accordingly:

      “Microwire signals were selected based on high signal-to-noise ratio, as reflected by the detection of ≥ 1 single unit. Duration of recordings was of (median, interquartile range) 10 min and 17 s [3-13 min] and number of electrodes per patient was 4.5 [2.75-8].”

      The authors use the term "virtual simulation", which I find odd. I think the simulation is very real in the sense that it simulates reality, and I do not understand how a simulation can be virtual.

      We have updated the manuscript accordingly.

      Reviewer #3 (Recommendations for the authors):

      Major Comments:

      (1) In Figure 1, the authors suggest that random insertion of action potentials in a signal is sufficient to yield FRs. However, the observed FRs shown in panel 1b (also in supplemental Figure 5) seem pretty short in duration and may not meet the mentioned criteria in methods that require at least 4 cycles and ".whose amplitude is 3 times that of the surrounding baseline..". Moreover, in panel 1b, it seems that the FR shows a candle-like appearance, which has often been associated with filtering of sharp transients. How did the authors validate that the detected FRs were "real" FRs?

      Given the very large amount of data, it was not possible to visually verify all FRs. However, FRs were detected with published methods (Roehri et al., 2016; Roehri et al., 2017; and Sheybani et al., 2018 for confirmation of 24-hour variability in rodents) that have subsequently been used in several publications.

      Regarding the candle-like appearance of the spectrogram, the Delphos algorithm precisely looks for isolated “islands” of increased power (see Roehri et al., 2018, Ann Neurol), thus excluding any candle-like appearance. Similarly, the detector in Sheybani et al. (2018) J Neurosci first detects candidate FRs but then excludes those that are associated with a peak in lower frequencies, thus also limiting the risk of detecting candle-like events.

      Regarding duration, we have compared the duration of FRs in original and shuffled rodent data and found that FRs in original signals are indeed longer. This makes duration a key feature to identify distinct FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      (2) In the context of neuronal cultures, it is unclear how it could be deducted that the result relates to chance incidence of action potentials considering that no random action potentials were inserted, but only random shuffling of the high frequency component of the signal was attempted "Hence, neural networks with limited complexity (Kim et al., 2020; Saglam-Metiner et al., 2024; Sanchez-Vives and McCormick, 2000; Timofeev and Chauvette) fail to generate FRs beyond that expected from the chance coincidence of APs, even after increasing network excitability."

      FRs arise from series of action potentials occurring at a delay corresponding to their oscillatory frequency (250-500 Hz). Simulations demonstrated that FRs can occur by chance. When the EEG is shuffled, the only FRs that remain are those occurring by chance, because those occurring as individual entities have been broken up. Hence, if the original EEG displays more FRs than the shuffled EEG, then it means that these additional FRs were generated as individual entities. We have improved the Results section to clarify this:

      “We hypothesized that if FRs arise purely from chance firing, then temporally shuffling these recordings while conserving their spectral properties (Supplementary Fig. 3) would disrupt any oscillatory structure, leaving only FRs that occur due to chance.] Any additional FRs in the original data, compared to the number of FRs in the shuffled EEG, should thus be assumed to be individual entities.”

      (3) In the rat dataset, sleep was treated rather homogenously, without accounting for the sleep stage that is characterized by different synchronization and firing. An analysis of different sleep stages would be valuable.

      Although we agree that it would be scientifically interesting, we believe that our claim – that the ratio of genuine to stochastic FRs changes across the sleep-wake cycle – would hold. Unfortunately, lack of EMG prevents us from performing reliable sleep scoring. However, we do now include an alternative method for differentiating sleep from wake using the beta-to-delta ratio, which was highly correlated with delta activity, supporting our previous approach. Please refer to Supplementary Fig. 4d for further information.

      (4) The authors found that chance aggregation was highest during periods of wakefulness. Analyses of human data also confirmed that FRs could occur by chance aggregation during wakefulness. However, a comparison with sleep data would further strengthen this finding.

      We fully agree, but unfortunately, we do not have sleep data using microwires. Although our central claim – that FRs can occur by chance clustering of action potentials – would hold, we agree that it would have been scientifically interesting to add sleep data.

      (5) The statistics section would benefit from addressing how normality was determined and power analysis, as well as the inclusion of the exact sample size for all experiments.

      With large sample sizes, ANOVA and linear mixed models are robust to non-normality. Given the large sample sizes of our data, we thus used ANOVA and linear mixed model. For tests with small sample sizes where normality was violated, we used non-parametric tests, indicated by their name, e.g., Wilcoxon test for Supplementary Fig. 3b.

      (6) Greater discussion on the implications of this study for proposed in-phase or out-of-phase FR generation mechanisms is suggested.

      We have added further discussion on this. In the aim to keep the Discussion short and impactful, we could not elaborate too much. We have synthetized other parts of the Discussion to keep it within the right length. Here is the additional part:

      “It has been argued that the very high frequency that can be obtained during FRs are due to out-of-phase firing of excitatory neurons (Foffani et al., 2007; Ibarz et al., 2010), which is also consistent with our concept of stochastic firing. The conceptual difference is the degree to which there is any underlying organization of this firing. We argue that in the majority of cases there is no organization, although a substantial minority cannot be explained on a stochastic basis.”

      (7) More explanation around why wakefulness may drive chance aggregation and the clinical relevance of it, as often presurgical epilepsy recordings are being evaluated during sleep.

      We have profoundly rewritten the Discussion regarding the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009).”

      Minor Comments:

      (1) Abstract, please include the frequency range of fast ripples explored in this study.

      The abstract has been updated accordingly.

      (2) Abstract, consider including the exact epilepsy model system in rats instead of "a rodent model of hippocampal epilepsy".

      The abstract has been updated accordingly.

      (3) Line 87, while Ylinen uses the term "high frequency oscillations" to refer to ripples up to 200Hz, which are different from the ones discussed here, better to rephrase or use another reference.

      The reference has been changed for Bragin et al. (1999), Epilepsia

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes is differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1, that are downregulated during differentiation, have predicted binding site enrichment at other switch genes, and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes the differentiation of aged OPCs. Viral expression of Bcl11a in Sox10-expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      We sincerely thank the reviewer for their positive assessment and for recognising the significance of our study, as well as the bioinformatics approach and tool developed as part of this work.

      Weaknesses:

      Although the PCA plots show distinct and reproducible global gene expression differences between the different isolated cell populations, the authors do not present a figure showing expression levels of typical stage-specific markers (e.g., Pdgfra, Pcdh15, C1ql1 for OPCs, Bcas1, Enpp6, Gpr17 for preOLs, Mobp, Mog, etc. for OLs) or confirm the absence of markers of other lineages (astrocytes, neurons, microglia, etc.). This makes it difficult to evaluate the success of their cell isolation strategy at different ages without reanalyzing the raw data.

      Thank you for this suggestion. We have presented markers expression in a new figure (Supplementary Figure 1) and included a description in the new Supplementary text.

      We observed elevated expression of Hes1 in OPCs as compared to both PreOL and OL, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (PMID: 19104146, PMID: 21167918).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs.

      One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra are not immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPC to OL, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparing PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared to be lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. However, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern OL > PreOL > OPC.

      We did not find any difference of astrocytes marker Gfap in those cell types comparison, suggesting similar level of unavoidable contamination which will not affect determination of differential gene expression. Regarding this please also see reviewer #2 major point 1.

      We now included this in the supplementary text:

      “Please see Supplementary Figure 1. We observed elevated expression of Hes1 in OPCs compared with both PreOLs and OLs, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (Brosnan et al, 2009; Ogata et al., 2011).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs. One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra may not be immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPCs to OLs, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparison of PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. In contrast, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern: OL > PreOL > OPC.

      We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      In the main text we have added the following text:

      “The expression patterns of cell-type-specific markers were consistent with their being distinct OPC, Pre-OL, and OL populations (Supplementary Figure 1, see Supplementary text for detailed description).”

      Please note that a detailed discussion of marker expression in the main text will disrupt the flow of the manuscript in manner we feel would detract from its clarity. We have therefore provided this discussion in the Supplementary Text.

      In addition, other publicly available datasets (e.g., the Barres lab bulk RNA-Seq datasets from PMID 25186741 or the Castelo-Branco lab single cell datasets from PMID 27284195) do not show downregulation of Bcl11a during OL differentiation as is described here - this apparent discrepancy is not discussed.

      Thank you for raising this point. We have now included new data as a Supplementary Figure 4. We performed RT-qPCR (reverse transcription followed by qPCR) to quantify Bcl11a expression and found that it was significantly lower in OLs than in OPCs, and significantly lower in aged OPCs than in young OPCs. These data were presented together with stage-specific markers.

      Regarding the comparison with PMID: 25186741: we extracted Bcl11a FPKM values from their dataset (GSE52564) and plotted, as shown in Author response image 1. We found that Bcl11a expression is downregulated during differentiation. However, the dataset contains only two replicates, and the SEM between the two OL replicates is very high, which may have contributed to the apparent lack of clarity. With such high SEM and only two replicates, the statistical power is poor, making robust statistical inference difficult.

      Author response image 1.

      Plotting of FPKM values of Bcl11a (obtained from GSE52564). mean+SEM shown along with individual data points. OPC: Oligodendrocytes progenitor cells, NFO: Newly formed oligodendrocytes, MO: myelinating oligodendrocytes. Dotted red line: linear regression line.

      Regarding comparison with PMID 27284195: we contacted the Castelo-Branco laboratory, and they kindly provided us with the analysis shown below in Author response table 1. This analysis showed that Bcl11a expression is lower in myelinating oligodendrocytes (MOLs) compared with OPCs. The apparent discrepancy observed in the web interface is likely because MOLs are displayed separately by subtype in the online resource. In single-cell datasets, particularly earlier pre-10x datasets with relatively lower cell numbers and sparser transcript detection, visualisations such as violin plots or t-SNE plots can be difficult to interpret when expression is distributed across multiple subclusters. Therefore, directly examining the differential expression statistics, including fold-change and significance values, provides a clearer and more quantitative assessment of the expression change.

      Author response table 1.

      Bcl11a expression difference in MOLs vs OPCs (dataset: GSE75330)

      FC: fold change, p_val_adj: adjusted p-value.

      Therefore, our bulk RNA-seq and RT-qPCR analyses presented in this paper are consistent with the Barres laboratory bulk RNA-seq dataset (PMID: 25186741) and the Castelo-Branco laboratory scRNAseq dataset (PMID: 27284195).

      Reviewer #2 (Public review):

      Aging poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs) to differentiate and myelinate neuronal axons. Myelin abnormalities accumulate with age, and it is likely that the ability of OPCs to differentiate into myelinating oligodendrocytes becomes progressively impaired during aging, leading to inefficient turnover of damaged myelin and oligodendrocytes, as well as reduced adaptive myelination. Understanding the molecular mechanisms underlying the compromised capacity of aged OPCs is therefore critical for addressing age-related white matter decline.

      This study aims to decipher the intrinsic molecular changes that occur in aged OPCs. By profiling differentially expressed transcription factors (TFs) between young and aged OPCs, and by employing a novel bioinformatic tool to identify key TFs that undergo dynamic changes across distinct stages of OPC differentiation, the authors identify Bcl11a as a potential regulator. Bcl11a is highly expressed in young OPCs but markedly reduced in aged cells. Functional experiments further demonstrate that while Bcl11a does not affect OPC proliferation, it significantly promotes the differentiation of aged OPCs. Importantly, this effect is also observed in vivo following demyelinating injury in aged mice.

      While the study provides compelling evidence that BCL11A represents a limiting factor for OPC differentiation during ageing, the downstream targets and molecular mechanisms through which BCL11A exerts its effects are not directly addressed. As such, the work should be interpreted primarily as identifying a key regulatory node rather than a fully defined molecular pathway.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      We are grateful to the reviewer for their supportive comments and for highlighting the broader relevance of our computational framework beyond our specific subfield.

      Major Points:

      (1) MACS mouse anti-A2B5 microbeads are not OPC-specific and may also label astrocyte precursor cells or immature astrocytes. How do the authors justify this caveat? Could some of the claimed "OPCspecific" switch genes in fact be enriched in astrocyte lineage cells?

      We thank the reviewer for raising this important point. While anti-A2B5 is a well-established and widely used antibody for isolating OPCs, we nonetheless agree that no technique can isolate a specific cell type with 100% purity, and this also applies to OPC-specific isolation using a validated anti-A2B5 antibody.

      To check whether astrocyte contamination could be an issue in determining differential expression, and specifically whether the OPC population was affected by astrocyte contamination, we checked the relative expression and statistical significance of the astrocyte marker Gfap. We refer to our new Supplementary Figure 1 and Supplementary text. This suggests that no difference exists in Gfap levels when comparing OPC, PreOL and OL populations. Therefore, we contend that it is unlikely that the differential expression observed in any cell population is actually due to astrocyte contamination, or that the OPC population is selectively contaminated by astrocytes.

      We now included the following in the Supplementary text:

      “We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      (2) Overall, Figures 1 and 2 are not very informative in terms of biological insight. The authors should provide more detail in the main figures regarding the enriched gene sets associated with each of the Type 1-4 switch categories. For example, summarizing the top Gene Ontology terms for each switch type would greatly enhance interpretability.

      We agree that GO analysis can add further interpretability. We have now prepared a new Supplementary Figure 3A to summarise the significant top GO-term enrichment for switch Types 1– 4, for which gSWITCH-identified patterns are presented in Figure 1C. We also prepared a Supplementary Figure 3B to summarise the top significant GO-term enrichment for the 135 Type 3 switch genes affected in ageing, presented in Figure 2C. Please note that only 8 Type 4 genes overlapped with differentially expressed genes in ageing. Due to this small number, we could not identify any significant GO-term enrichment, and therefore this was not plotted.

      (3) A similar issue applies to Figure 3. The authors should explicitly specify the transcription factors in the main figure, particularly the 27 TFs identified through theENCODE/ReMap2 analysis.

      Thank you for raising this point. We have now prepared a new Supplementary Table 3, where we list 27 TFs and highlight, with light grey shading, the 5 TFs that overlapped with Type 3 switches.

      (4) Have the authors validated Bcl11a expression across different CNS cell types and between young and aged conditions using independent methods such as qPCR, immunofluorescence, or western blotting?

      Thank you for this suggestion. We performed qPCR and presented this data in a new Supplementary Figure 4. We found that Bcl11a expression is lower in OLs than in OPCs (Supplementary Figure 4A). We also observed reduced Bcl11a expression in aged OPCs compared with young OPCs (Supplementary Figure 4B). (see also response to Reviewer 1’s recommendations).

      (5) Regarding OPC aging, an open question is whether the reduced differentiation capacity of aged OPCs is an intrinsic property of the cells themselves or whether it results from prolonged exposure to an aging environment that induces non-cell-autonomous epigenetic or genetic changes, thereby rendering OPCs less efficient at differentiating. It would be helpful if the authors could expand on this point in the Discussion, with reference to relevant previous studies and experimental evidence.

      We thank the reviewer for suggesting this important aspect be discussed. We have now included the following paragraph in the discussion section:

      “The extent to which the reduced differentiation capacity of aged OPCs is intrinsically encoded within the cells themselves or induced by prolonged exposure to an aged tissue environment is an interesting question. Based on our previous work, we favour the view that loss of OPC function is primarily determined extrinsically since various manipulations of the aged environment such as heterochronic parabiosis (Ruckh et al., 2012), fasting and calorie restriction mimetics (Neumann et al., 2019), and niche biomechanics (Segel et al., 2019) can all alter the cell-intrinsic state, reverting aged cells to a ‘youthful state’. Significantly, when aged OPCs are transplanted into the neonatal CNS they proliferate and differentiate as if they were neonatal OPCs (Segel et al., 2019). The reversion of aged OPCs to a functional state by changes in their external environment necessarily operates through changes in cell intrinsic function, suggesting that the same intrinsic mechanisms could be targeted directly to restore declining OPC function—for example through epigenetic regulation of differentiation inhibitors (Shen et al., 2008) or overexpression of transcriptional regulators such as c-Myc (Neumann et al., 2021, Dimas et al., 2025).”

      (6) Do the authors observe a change in the number or density of OPCs between young and aged mice?

      Thank you for asking this important question. In 2002 we reported that there was no difference in the OPCS density between young adult and old adult rats, at least in the deep cerebellar white matter (Sim et al. 2002 - PMID: 11923409). We also refer the reviewer to Figure S1 of another previous study, published in Cell Stem Cell in 2019 (PMID: 31585093). We did not find any difference in OPC number between young and aged brains. Quantification was performed using FACS, where freshly isolated cells were stained with A2B5 (OPC marker), CD11b (microglia marker), and MOG (oligodendrocyte marker). Thus, we do not find any evidence for an age-related decline in OPC densities.

      (7) The in vivo characterization of Bcl11a overexpression using the AAV-based approach appears incomplete. Do aged mice overexpressing Bcl11a in Sox10⁺ cells exhibit reduced age-related myelin degeneration under baseline conditions? In the LPC model, do the authors observe differences in lesion size and/or remyelination efficiency?

      Again, we thank the reviewer for raising these interesting points. To assess whether Bcl11a overexpression in Sox10+ myelinating oligodendrocytes exhibit less age-related myelin degeneration would, we suspect, require long-term experiments. For this to be the case would require a role for Bcl11a in myelin maintenance – and interesting question but one we feel (and hope the reviewer agrees) is beyond the scope of the current study. We do not see any difference in lesion size (and would not expect the expression of elevated levels of Bcl11a to protect against the membrane-solubilising effects of LPC) but do see changes in remyelination efficiency as shown in Figure 6.

      (8) Are the authors presenting gSWITCH for the first time in this manuscript? Given that the gSWITCH framework is novel and central to the study, its conceptual contribution could be emphasized more strongly. A brief comparison with existing trajectory- or pattern-based methods-ideally in the main text around Figure 1-would help readers better appreciate its novelty.

      We thank the reviewer for this important suggestion. Yes, gSWITCH is presented for the first time in this manuscript as a new computational framework and web application. We agree that its conceptual contribution should be made clearer in the main text itself, although we explained its concept in detail in ‘Materials and Methods’ and in the supplementary Figure 2 (which was Supplementary Figure 1 in first version of this manuscript).

      We now included the following paragraph in the manuscript:

      “Existing computational tools such as Monocle (Trapnell et al., 2014), tradeSeq (Van den Berge et al., 2020) and maSigPro (Nueda et al., 2014) are highly valuable for identifying genes with dynamic expression changes across pseudotime or time-course data. gSWITCH addresses a different question. It does not aim to infer trajectories. It works with a user-defined ordered series of biological states or time points and asks a more specific question — does this gene show a statistically supported "switchlike" change in expression as cells move through these states, and if so, what shape does that change take? It combines GLM-based statistical testing with criteria that capture where a gene reaches its highest or lowest expression and whether its expression changes steadily in one direction across the ordered series. To our knowledge, no existing tool combines significance testing with this type of explicit, shape-based classification into discrete, interpretable switch categories. gSWITCH sorts genes into four biologically meaningful patterns, rather than producing only a ranked list of significant genes based on pairwise comparisons between multiple conditions or states. gSWITCH also flags which of these switch genes are transcription factors, making it easier to prioritise candidates for follow-up experiments.

      This biologist-friendly tool is freely available as a web application requiring no programming, works with experimental designs containing three or more stages or time points with at least two replicates per stage (no upper limit on either), and can be applied to bulk RNA-seq or to single-cell RNA-seq data aggregated as pseudobulk.”

      (9) The evolutionary analysis also appears somewhat disconnected from the rest of the study. Could the authors leverage available public datasets to test whether a similar Bcl11a expression trajectory is observed in human oligodendrocyte lineage cells?

      We thank reviewer for mentioning this. We would like to clarify that the evolutionary analysis was included to examine whether Bcl11a sequences across vertebrates, including humans, show evidence of selective constraint, meaning that the sequence has been preserved during evolution because changes in it are likely to be disadvantageous. This analysis was therefore intended to provide broader evolutionary support for the functional importance of Bcl11a, rather than to stand as a separate or disconnected component of the study.

      For this analysis, we included Bcl11a DNA and protein sequences from 23 vertebrate species, including humans. We refer the reviewer to the Methods section of this paper, under “dN/dS analysis”, for further details. To provide further clarity regarding the different species used in this study, we have now prepared a new Supplementary Table 4, listing the 23 species together with their DNA and protein sequence accession numbers for Bcl11a.

      We also added this sentence in the main text:

      “We included twenty-three vertebrate species, including humans (Supplementary Table 4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Given how central the isolated cells are to the subsequent analysis, the manuscript would be strengthened by a figure showing expression of stage and lineage-specific markers.

      Ideally, the authors would provide some sort of orthogonal experimental approach to confirm downregulation of Bcl11a during oligodendrocyte differentiation and loss with age (e.g., IF or RNAScope in conjunction with stage-specific markers in tissue, or western blot in culture).

      Thank you again. We have performed these. Please see the Reviewer #1 comment (above).

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1A: It should be 'anti-O4' instead of 'anti-04'.

      This is now corrected. Thank you.

      (2) Figure 1C: The authors should specify what the connecting lines indicate (e.g., gene sets or gene modules).

      Each coloured line represents one gene and connects its log2 fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualize gene-wise patterns of expression change across these cell states. For example, in Type 1, each line shows a pattern in which gene expression increases progressively from OPC to PreOL to OL, with the highest expression change observed in OLs: OL > PreOL > OPC.

      We now included the following line in the figure legend:

      “Each coloured line represents one gene and connects its log<sub>2</sub> fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualise gene-wise patterns of expression change across these cell states.”

      (3) Figure 2C: The authors should specify "DF genes" in the figure legend.

      Thank you for pointing this out. This was a typo: it was written as DF, but it should be DE (differentially expressed) genes. We have now corrected this in the figure and spelled out the abbreviation in the legend. Also, DE gene list is accessible through GEO accession: GSE303317. This also mentioned in the figure legend as:

      “DE: Differentially expressed. DE gene list is accessible through GEO accession: GSE303317.”

      (4) Figure 4C & Figure 5B: the title for the y-axis of the bar graph is confusing. The authors should specify what "#" indicates. Does it represent the counts? What are the thresholding criteria to judge whether an Olig2 cell is MBP-positive or not? It's unclear what the unit is here for the 0-100 scale.

      We apologise for the confusion. We used ‘#’, which is a common notation in mathematical and quantitative contexts, to denote counts, so you are correct. We now mentioned in the legend: “The symbol “#” indicates cell count.”

      We counted the number of MBP+OLIG2+ cells, divided this by the total number of OLIG2+ cells, and expressed the value as a percentage. For greater clarity, instead of writing #MBP+/#OLIG2+, we have now written #MBP+OLIG2+/#OLIG2+.

      Regarding the 0–100 scale, the unit of the Y-axis is percentage, as stated in both figure legends.

      The criterion for classifying an OLIG2+ cell as MBP+ was morphological: an OLIG2+ nucleus, shown in white, had to be surrounded by MBP+ staining, shown in red. Cells meeting this criterion were counted as MBP+OLIG2+ cells. Manual counting was performed blinded to sample identity.

      (5) Figure 6B: To discriminate from IF staining, the authors should use italic'Plp1' to indicate the RNA in situ results.

      Thank you for pointing this out; we have now corrected it.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewer’s for their thoughtful comments that have significantly strengthened the paper. Below, we have outlined our responses to both the public reviews and recommendations.

      In addition to the alterations to the manuscript based on the reviews, during our review of the data analysis we uncovered some small errors that we have now corrected. In looking back over the image registration, we identified three animals whose olfactory bulbs did not register properly and one with poor cell counting in the telencephalon. To account for these issues, we imputed the missing data using an iterative soft-threshold singular value decomposition (described on lines 779-783 of the updated manuscript). This update had little impact on the results. We also identified a small error in how we determined ‘unique’ and ‘overlapping’ edges in the network analysis (Figure 8). In the previous analysis we had incorrectly noted that all ‘unique’ edges did not have an overlapping confidence interval with the two other networks (i.e., the networks for evading freezers, freezers, and non-reactive). Instead, the ‘unique’ edges in the prior version of the manuscript did not have an overlap with at least one other network. We have now updated the analysis so the reader can distinguish between edges that are truly ‘unique’ versus those with ‘1 overlapping confidence interval’ or ‘2 overlapping confidence intervals’ with other networks. As before, this update and change to the analysis does not materially affect the results or conclusions.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The neural analysis part is very comprehensive. Figure 5 and Figure 6 are independent but complement each other very well. They together support that the cerebellar system is the key brain component for a freezing response. Their extreme focus on high-level analyses, however, came at the expense of biological intuitions. I suggest adding some figure panels and result/discussion paragraphs to help with that aspect.

      Thank you for the suggestion. We have made extensive edits to the manuscript to include additional discussion and biological intuition. Specifically:

      We added a supplemental figure (Figure S6-2) that has scatterplots showing how cfos levels vary with the different behavioral contrasts. Although the PLS analysis is multivariate, this univariate analysis should help give readers a better intuition of how the behavior relates to brain function.

      We have also rewritten the results sections for both the PLS analysis (lines 303-361) and network analysis (lines 396-437) to incorporate more of a discussion about the biological context of different regions identified. Thank you for this suggestion, we feel that this significantly strengthens the biological interpretation of the data for the reader.

      Reviewer #2 (Public review):

      (1) My first concern relates to the claim in the abstract that "We found that fear memory behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers".

      In my opinion, the "freezing" aspect is well supported as being both triggered by the CAS and for memory effect upon re-exposure to the tank, but I am less convinced about the "evasive" behaviour. In Figure 2, it appears that "evasiveness" is generally not increased in both the Exposure or Memory phases for many groups, and in Figure 5, it appears that "evasiveness" is expressed by nearly 50% of the fish in the pre-exposure condition before CAS addition and in all phases in the vehicle condition. Therefore, it appears that most of the expression of this behaviour is independent of any memorybased effect.

      We thank the reviewer for this suggestion and we agree that this line in the abstract was unintentionally misleading. We have now altered this line in the abstract (lines 34-36) to read:

      “We also found that that behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers with the evading freezer and freezer groups most clearly associated with memory formation.”

      On the larger point of the inclusion of evasion as part of the fear response, we believe this is warranted for the following reasons: (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      (2) My second concern relates to the claim in the abstract that "background strain and sex influenced how fish respond to CAS, with males more likely to increase evasive behaviors than females and the TU strain more likely to be non-reactive."

      My understanding, based on the introduction and on the methods, is that it is likely important that the CAS be prepared from conspecifics of the same strain and sex, and for this reason, they prepared different CAS specific for each strain and each sex. Therefore, the "CAS" that is applied is necessarily different for each condition, and I am concerned about if the differences observed could relate more to variation in the quality, purity, concentration, etc. of the specific CAS samples for different groups, rather than their reactivity to the substance or their ability to form memories based on such experiences.

      The CAS was prepared by mixing extracts from all four strains and both sexes (so 8 fish per batch). Thus, all the fish were exposed to the same CAS mix derived from the same donors. This is described in the methods (lines 626-629). However, to ensure that this is clear to readers, we’ve now included a line indicating this in the results section (lines 123-124).

      (3) My third concern relates to the interpretation of the cFos data.

      As I mentioned above, I feel as though the behavioural analysis is perhaps more complex than is warranted via the inclusion of evasiveness, and I wonder if the conclusions from the experiments would be simpler if analyzed only from the perspective of freezing.

      We agree that the freezing response is driving the majority of the neural cfos response that we are seeing (e.g., Figure 6A-C). However, we feel that the network analysis (Figure 8) justifies the distinction between freezers and evading freezers. This is because the brain networks for these two groups (freezers and evading freezers) are quite distinct, even though these groups both have the same levels of freezing behavior (Figure 4). This stark difference in patterns of neural activity suggests the brain of a freezer and an evading freezer are engaging with the world in two distinct ways that is worth noting. We’ve updated the abstract to make this point clearer (abstract: lines 39-48) and discuss the biological interpretations of patterns of brain activity unique to evasion or evading freezers in more depth (lines 303-361; lines 396-437).

      Reviewer #3 (Public review):

      (1) The three-day contextual fear paradigm, as implemented - one CAS pairing on day 2 followed by a single recall test on day 3 - inevitably conflates acquisition and long-term memory, making it impossible to know whether strains like TU truly recall the association poorly or simply learn it more slowly. For example, given that TU fish extinguish fear faster than AB or TL strains in extended protocols, they may simply require additional or repeated CAS pairings to achieve the same asymptotic performance. To disentangle learning kinetics from recall strength, the assay could be revised to include multiple acquisition trials (e.g., conditioning on two or more consecutive days) with an immediate post-conditioning probe to assess acquisition independent of consolidation, and continuous measurement of freezing and evasive behaviors across each trial to fit learning curves for each strain. Such refinements - even if on a subset of the strains - would reveal whether "non-reactive" phenotypes reflect genuine recall deficits or merely delayed acquisition.

      We thank the reviewer for this thoughtful comment. We agree that it is difficult to disentangle acquisition from consolidation. Indeed, the TU fish do appear to have lower levels of freezing in response to the CAS (Figure 2A), supporting the idea that reduced performance at memory day could be due to some sort of deficit at acquisition. However, pursuing a detailed examination of strain dependent differences in fear memory acquisition versus consolidation is beyond the scope of the current paper where we primarily focus on individual differences in behavior. Nonetheless, we have included this important point in the discussion (lines 470-471).

      (2) My second major question is with respect to Figure 3 panel B. This is a complex figure, and I can understand the gist of what the authors are attempting to show, but it is difficult to understand as it is. Can this be represented in a way that is clearer and explained a bit more easily?

      We agree that this figure is one of the more complex in the paper. However, we’ve struggled to come up with a better way to present it. We have improved the presentation based on other reviewer comments by making the vehicle and CAS groups more easily distinguishable by using open versus closed circles. We’ve also included additional interpretations of the data in the results, which we hope will help guide readers through this figure better (lines 208-223).

      (3) The brain mapping is by far one of the most interesting aspects of this study, and the methods that the group used are interesting. The brain mapping, however, relies on generating "contrasting" groups (Figure 6A), and I was not clear as to how these two groups were formed. Could the authors elaborate a bit?

      These contrasting groups (contrast 1, contrast 2) arise analytically from the partial least squares (PLS) analysis; they are not defined by the experimenter. In brief, PLS is a multivariate technique that identifies latent variables that capture axes of maximal covariation between two datasets: behavior and brain activity. As an analogy to a more widely known technique, principal components analysis (PCA) uncovers axes of maximal variance within a single dataset. PLS, in contrast, simultaneously analyzes the covariation in two datasets. The contrast groups in Figure 6A represent the behavioral weights of the latent variables that capture the most covariance, which illustrates how the four behaviors load onto these top two contrasts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The c-Fos analysis in Figure 5 is very comprehensive and convincing, but lacks intuitive presentations. In my understanding, the increase in c-Fos expression in red areas means increased freezing behavior for Contrast 1 for the PLS analysis? Do you have representative c-Fos expression images between different groups of fish?

      We decided not to include a representative cfos image because the data is derived from a large number of fish (N=87) and thus it can easily be cherry-picked to choose images that match the narrative. Instead, to more accurately capture the breadth of the data while providing a more intuitive presentation, we have included an additional supplemental figure that includes scatterplots of scaled cfos data against behavioral scores for each of the two contrasts (S6-2). We believe this more fully and accurately captures the relationship between behavior and brain activity. We included six different example brain regions and scatterplots for cfos activity against behavioral scores for contrasts 1 and 2, demonstrating a range of relationships. However, we should note that PLS is a multivariate technique, and so this univariate analysis does not fully capture the subtleties of the PLS analysis. Nonetheless, we think this will help give a more intuitive interpretation of the data to readers. We have also referenced this additional data in the manuscript (lines 307-309). We thank the reviewer for this excellent suggestion that improves the ability of readers to understand the paper.

      (2) Also related to Figure 5, the result section only describes the PLS statistics and does not try to describe the biological interpretation. Do the authors think the c-Fos expression directly represents lowlevel behavior, such as swimming, or a high-level behavioral state or learning? Maybe different areas mediate different aspects?

      For example, the medullary locomotor areas, which are usually highly correlated with swimming in terms of neural activity, seem to have higher c-Fos expression in freezing fish. I'm not saying this shouldn't be the case. c-Fos expression in this area was not elevated in larval fish during OMR in Shainer et al., 2023, indicating that it doesn't linearly reflect neural activity. But discussing a bit of intuition on the connection between c-Fos expression and biological process, rather than just saying "the cerebellum could regulate emotional states", would help us guide through this highly complex analysis.

      We have now added more interpretation of the data in both the PLS and network analysis sections (lines 303-361 and lines 396-437). Again, thank you for this excellent suggestion. This helps make the biological interpretation of the data clearer.

      Minor points:

      (1) Figure 2B titles: please write "memory" on the right side.

      We considered writing ‘memory’ on the right-hand side, but we thought this may add confusion because it would not apply to both graphs in the row. The left-hand graphs are the responses during ‘exposure’ and the right-hand graphs are the responses during the ‘memory’ phase. This is indicated by the titles above the left and right-hand sets of graphs.

      (2) Figure 2C: needs legend lines.

      We have now moved the legend lines from the top of the graphs to below the graph to make them more visible to readers.

      (3) Line 187: "aggregated" data.

      This has now been changed to ‘aggregated’ (now line 194).

      (4) Line 371: I'm not sure what "Beyond" means.

      We have now significantly changed this part of the paper and we no longer use the word ‘beyond’ here.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding point (1) in the Public Review:

      I would encourage the authors to consider whether this study might be better focused exclusively on the freezing behaviour, which does appear to be reliably expressed during CAS exposure and in the memory phases, and would significantly simplify the subsequent analyses of neural activity, and perhaps may lead to a more coherent conclusion.

      As noted in our response to the public review, we appreciate this suggestion, but we have decided to keep the inclusion of the evasive behavior. This is because (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are very distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      A more minor concern related to the analyses in Figure 2: in the figure legend, it is stated that "*-P < 0.05 compared to vehicle treated fish via t-tests". How are the authors dealing with the multiple comparisons problem? Would something like an ANOVA not be more appropriate?

      Thank you for bringing this point up. We did not initially correct for multiple comparisons because we considered each of these experiments across sex and strain separate since we did not compare across strains. However, the way we’ve grouped the data together in figure 2 makes it appear as if they are one large experiment. To alleviate any concern about multiple testing, we have now corrected for multiple comparisons using the false discover rate (FDR) correction. The statistics in the figure and captions have now been updated.

      (2) Regarding point (2) in the Public Review:

      If the authors agree with my concern regarding potential variability in the CAS samples, I would suggest either testing for differences among strains using the same batch of CAS, or including and explaining this caveat in the text.

      As noted in our response to the public review, the CAS was the same for all the fish. Each batch was derived from 8 donor fish, one fish from each strain and sex (described in lines 123-124 of the results and lines 626-629 of the methods).

      (3) Regarding point (3) in the Public Review:

      I feel like the standard in the field for such conclusions would be after

      (a) Direct analyses of the activity states in these areas. I was surprised not to see a direct analysis of the cFos stainings in the cerebellum relative to freezing behaviour, for example, ideally in a different animal cohort.

      The PLS analysis does relate activity in the cerebellum (and other brain regions) to specific behaviors via the the behavioral contrasts (Figure 6A). We believe this approach (instead of dividing fish into ‘high and low freezers’) is a more powerful way to leverage the data from all the animals tested (87 fish). However, we appreciate that the interpretation of the PLS analysis is not as intuitive as seeing scatterplots or bar charts comparing neural activity. For this reason (and in response to a comment from reviewer 1), we have included as a supplementary figure (Figure S6-2) scatterplots showing how standardized c-fos activity varies with the behavioral scores from the contrasts identified from the PLS analysis. Given that contrast 1 weights heavily in the positive direction on freezing, these figures can essentially be read as looking at cfos activity as a function of freezing levels. What can clearly be seen is that for regions of the cerebelleum (E.g., the LCa and CC) there is a clear positive relationship between cfos activity and the behavior scores for contrast 1.

      (b) Some kind of manipulation of the brain area resulting in the relevant behavioural modification.

      We completely agree with the reviewer. However, at the moment, we do not have the tools to do this in adult zebrafish. It is something we’re actively working on.

      Of course, I appreciate that such experiments might not be possible or feasible, and in which case I would suggest adjusting the claims accordingly and highlighting the caveats to their interpretations.

      We have incorporated the caveat that we have not directly altered neural activity into the discussion (lines 542-543) and adjusted how we discuss our findings in the abstract (lines 39-41) to more accurately represent the type of evidence we provide. Hopefully we’ll be able to do so in the near future!

      MINOR CONCERNS:

      (1) In Figure 3, how is the end of a behavioural epoch defined? I am surprised to see that you consider transitions between the same behavioural state. How does erratic swimming -> erratic swimming differ from a longer single epoch of erratic swimming? In general, I find this analysis confusing, and I am not sure if it adds significantly to the message of the paper.

      Thank you for this question as it prompted us to realize we were missing this in our methods section. We have now updated the methods to include how we calculated the behavioral transitions (lines 644-650). In short, we used a 750 ms behavioral epoch time that corresponds to the size of the sliding window we used for the random forest model.

      We have also updated the description of this analysis in the results to indicate the main finding from it (lines 207-223). In brief, the main finding is that exposure to CAS results in longer bouts of evasive behavior without increasing its frequency. Whereas CAS induced freezing arises from both longer bouts and likelihood of occuring. While we agree that this is a relatively minor finding in the paper, one of our goals is to provide as comprehensive analysis of fear behavior as possible to help guide future researchers interested in using fish for understanding different aspects of fear-related behaviors.

      (2) In the PLS analyses, two measures of evasion are used: evasion time, and evasion as a percent of active behavior. I don't understand the justification for both of these being used rather than one. Again, my overall recommendation is to reduce the focus on the analysis of evasion behaviour, but if you do not choose to do this, I think the rationale of how both measures are used and why needs explanation.

      We chose to incorporate two different measures of evasion throughout the study because the high levels of freezing in some animals results in little opportunity to express other behaviors (like evasion). Thus, to better capture what fish may be doing in the absence of freezing (i.e., when they are active) we also calculate the amount of active time spent performing evasive behaviors (instead of normal swimming). We have now included an explanation for this earlier in the results section when we first use this metric (lines 149-152).

      (3) In the methods, I don't understand this: "Animals that were assigned the wrong sex were removed from data analysis, as well as its paired fish (< 2%)".

      We determine the sex of fish when we set them up for dual housing. However, we occasionally make errors in sex determination. To ensure we properly sexed the fish, at the end of experiments, we euthanize the fish and check for the presence of eggs. If we incorrectly assigned the sex to a fish, they are removed from the experiment alongside the other fish they were dual housed with. This is because we want to ensure all fish are housed in the same way (i.e., a male fish with a female fish).

      (4) How was this determined differently from the first time, resulting in exclusion?

      After experiments, fish were euthanized and we checked for the presence of eggs (line 599-601). We’ve now added a line in this other part of the methods referring back to where we describe this (lines 676678).

      Reviewer #3 (Recommendations for the authors):

      Here are some minor concerns and errors found in the manuscript:

      (1) For Figure 2B and Figure 3B, can the group make the lines solid and dotted? The circle or triangle designation is difficult to see, and since the crux of the figure depends on comparing Veh and CAS, it would be easier to see if the lines were altered.

      Thank you for this suggestion. Instead of making the lines solid and dotted, we decided to make both the CAS and vehicle group circles and then have open and closed circles. We believe this solves the issue of being able to distinguish these groups and makes the data more readable.

      (2) Figure 2C: It appears that the line colors in the legend are missing.

      We have moved the line colors below the graphs to make them more obvious.

      (3) Figure 8A: Same thing here - could the text be enlarged? It's really difficult to make out each node, and when I zoom the text becomes pixelated. This is an important figure and one that will likely be referenced, and making it clear would be helpful.

      This one is difficult. We have made the network images as large as would fit on a page. We have now uploaded vectorized versions of the images so that they do not become pixelated when zooming in. As part of our supplemental materials we also include a cystoscope file that can be explored in greater depth as well.

      (4) The paper is really well written: I found a few typos, though:

      (a) Line 529: "Institutional Cara and Use Committee" should be "Institutional Animal Care and Use Committee" (Change cara to care and add animal).

      (b) Line 274: "hybdridization" should read hybridization.

      Thank you for catching these typos. They have now been fixed.

    1. Author response:

      We thank the editorial team and all three reviewers for their time and attention to detail in reviewing our manuscript. We are particularly grateful for comments recognising the “direct practical value” and “comprehensive analysis” in our work (Reviewer #2) as well as its “thorough and methodical” approaches (Reviewer #1).

      We also appreciate the reviewers’ constructive comments to improve our manuscript, which we plan to address in a revised version. Specifically, we plan to more explicitly acknowledge some of the limitations of our study, including clearer highlighting of the experiments performed only in females (Reviewer #1); the inability of our experimental design to control for chloride concentration (Reviewers #1 and #2); the limitations of transgene induction using the AGES system in adult flies (Reviewer #2); and the rationale for using different concentrations of auxin across different experiments (Reviewer #3). We also note that Reviewers #1 and #2 have provided additional recommendations beyond the public reviews (largely relating to helpful ways to clarify our text and more explicitly acknowledge limitations), which we also plan to address.

      In addition, we plan to perform additional experiments to address specific points raised by reviewers in both their public reviews and additional recommendations. In the first instance, we plan to follow Reviewer #3’s suggestion to characterise induction in the brain using 10mM auxin, as well as Reviewer #2 and #3’s suggestions to explore feeding behaviour of flies fed auxin within our experimental setup.

      We look forward to submitting an improved manuscript guided by the reviews, with the aim of strengthening this “well-needed validation” study that also “highlights critical caveats that will support future studies” (Reviewer #2).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1C1) The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions, then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      This is an interesting point raised by the reviewer.

      First, we need to clarify an important point: we do not make a well-stirred assumption. Samples can be drawn and plated from any region of space however small and that region’s population can be quantified using our method. The stirring only occurs after we collect a sample in order to dilute the contents and pour the solution homogeneously over the plate.

      As such, learning multiple independent species is possible and not impacted by the dilution (“well-stirred” assumption). In the new first paragraph of the methods section, we made it clear that this assumption concerns the dilution process. REPOP is designed to recover the true underlying heterogeneity in species abundance (even from limited data) by leveraging a Bayesian framework that remains valid regardless of whether species are independent or correlated.

      If the method is applied to multiple species as currently implemented, REPOP can recover the marginal distribution of each species, provided that the species are either selectively cultured or produce sufficiently distinguishable colonies on the same plate. To demonstrate this, we have added a new Results subsection with a synthetic two-species example in which the species abundances are correlated across samples.

      However, in order to learn the joint distribution and capture correlations between species within samples, the method would need to be extended. At present, in Eq. 5 we sum the likelihood over all values of n, using a data-driven cutoff (twice the largest naïvely estimated count times the dilution factor). Extending this to multiple species adding up to (n<sub>1</sub>,n<sub>2</sub>), while retain the generality of the method, would require quadratically scaling memory with this cutoff in the population number. For this reason while we comment on this in the new paragraph in the conclusion, it is not implemented as part of REPOP.

      Reviewer #2 (Public review):

      (R2C1) A more thorough discussion of when and by how much estimated microbial population abundance distributions differ from the ground truth would be helpful in determining the best practices for applying this method. Not only would this allow researchers to understand the sampling effort necessary to achieve the results presented here, but it would also contextualize the experimental results presented in the paper. Particularly, there is a disconnect between the discussion of the large sample sizes necessary to achieve accurate multimodal distribution estimates and the small sample sizes used in both experiments.

      That is a great suggestion from the reviewer. To address it, we expanded Appendix B. We know report (1) the relative error in the estimated means (as already done for Fig. 4 formally 3), and (2) the Kullback-Leibler (KL) divergence between the reconstructed and ground-truth distributions. These metrics will are show as a function of the size of the dataset, for the examples in Fig 3. enabling a direct assessment of how the sampling effort affects the precision of the inference.

      That said, we now highlight in the Conclusion that, by explicitly modeling the dilution process within a Bayesian framework, REPOP extracts the maximum information available from each individual sample at a given sample size. This strategy therefore enables more accurate inference with fewer measurements, which is particularly important in applications such as plate counting, where data acquisition is labour-intensive.

      Reviewer #3 (Public review):

      (R3C1) While the study is promising, there are a few areas where the paper could be strengthened to increase its impact and usability. First, the extent to which dilution and plating introduce noise is not fully explored. Could this noise significantly affect experimental conclusions? And under what conditions does it matter most? Does it depend on experimental design or specific parameter values? Clarifying this would help readers appreciate when and why REPOP should be used.

      We agree with the reviewer that this is an important point, and we expanded Appendix B to include a quantitative analysis using simulated data (Fig. 3, formely 2), reporting both relative error and KL divergence as a function of dataset size. This complements our response to R2C1 clarifying when REPOP offers the greatest benefit.

      In addition, we will expand the discussion on how modeling dilution noise becomes essential when learning population dynamics. In particular, we emphasize? the role of Model 3, especially relevant when working with multiple plates and approaching the asymptotic regime; an aspect that was alluded to in Fig. 3 but not fully explored.

      (R3C2) Second, more practical details about the tool itself would be very helpful. Simply stating that it is available on GitHub may not be enough. Readers will want to know what programming language it uses, what the input data should look like, and ideally, see a step-by-step diagram of the workflow. Packaging the tool as an easy-to-use resource, perhaps even submitting it to CRAN or including example scripts, would go a long way, especially since microbiologists tend to favor user-friendly, recipe-like solutions.

      In the new paragraphs of the introduction, we made clear that REPOP is written in Python (PyTorch), installable via pip, and designed for ease of use. We are also expanding the tutorials to include clearer guidance on data formatting and common workflows. The new workflow figure (Fig 2) better illustrates the full process.

      (R3C3) Third, it would be great to see the method tested on existing datasets, such as those from Nic Vega and Jeff Gore (2017), which explore how colonization frequency impacts abundance fluctuation distributions. Even if the general conclusions remain unchanged, showing that REPOP can better match observed patterns would strengthen the paper’s real-world relevance.

      We thank the reviewer for this interesting suggestion. We agree that applying REPOP to additional existing datasets would make REPOP’s relevance clearer. However, the Vega and Gore datasets lack the information required. REPOP requires the plate count measurement process to be specified, including the dilution factors used for each measurement. Furthermore, we can leverage on additional information about the experimental procedure when the colony cutoffs and dilution schedules used are reported. Without the dilution factors, the likelihood connecting the observed colony counts to the underlying population size is not possible. We hope this clarification will help make future datasets made available publicly more useful for purposes of uncertainty propagation.

      (R3C4) Lastly, it would be helpful for the authors to briefly discuss the limitations of their method, as no approach is without its constraints. Acknowledging these would provide a more balanced and transparent perspective.

      We agree with the reviewer. We have added two new paragraphs to the conclusion highlighting important current constraints and future development directions of the framework. In particular, we now discuss that, in its present implementation, REPOP focuses on the population distribution that maximizes the posterior, rather than returning posterior uncertainty over the reconstructed distributions themselves. We also note the computational demands of the method, making GPU acceleration highly beneficial and more complex multi-population inference computationally challenging. This discussion synthesizes points raised throughout our response to R1C1 and the reviewers and provides a more balanced perspective on the current scope of the method.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to some neural changes across sessions. However, several aspects of our behavioral design and results argue against extinction or repeated nonreinforced retrieval as the primary drivers of the observed effects. Importantly, discrimination ratios remained stable or increased across time rather than progressively diminishing as would be expected under extinction (this new analysis will be added to the resubmission). Nevertheless, we will address this important point in the Discussion and explicitly acknowledge that repeated retrieval may contribute to some component of the observed representational changes.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      We thank the reviewer for raising this important point, which was also noted by the other reviewers. To address this issue, we will implement Reviewer 3’s suggested Generalized Linear Model (GLM) analysis using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. Because freezing behavior varies across trials whereas stimulus identity is fixed, this approach will allow us to dissociate their respective contributions to neuronal activity. If, after accounting for freezing behavior, responsive neurons continue to exhibit graded coding consistent with inferred threat value, this would strengthen the interpretation that the identified ensembles reflect generalization gradients related to aversive value rather than freezing behavior alone. Otherwise, we will adjust the conclusions according to the interpretation that freezing itself drives the generalization gradients.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We will include measures of registration quality in the resubmission.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We will correct correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      Thank you for this suggestion. We will clarify the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      We will redo the Statistics of Figure 3 to take into account the following variables: group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30).

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term “inference” may overstate the cognitive processes engaged by the current task. Accordingly, we will revise the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We will correct the language and use “threat value”.

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for these important points, which we address in further detail below.

      Ongoing learning during test sessions: The reviewer correctly notes that unreinforced test presentations may constitute extinction-learning trials and that some neural changes across days could therefore reflect ongoing learning rather than spontaneous ensemble reorganization. However, new analyses indicate that extinction is unlikely to be the primary driver of our findings. Discrimination ratios do not decay over time; instead, they either sharpen or remain stable across sessions (new analyses to be included in the resubmission). These results argue against robust extinction as the primary source of the neural changes observed across sessions. This interpretation is also consistent with the strength of our conditioning protocol, which used 10 CS+ shock pairings and 10 CS− no-shock pairings specifically to minimize extinction across repeated testing sessions. Nevertheless, we acknowledge that the current design cannot fully dissociate time-dependent consolidation from retrieval-induced plasticity, and we will explicitly discuss this limitation in the revised Discussion.

      Stability reflecting behavioral consistency: We agree this alternative cannot be fully excluded. However, the cluster stability analyses assess identity at the level of response profile across all four frequencies, not response magnitude alone. Tone-selective clusters, which also show consistent behavioral correlates (firing rate correlates with threat-value, Fig. S8), do not show equivalent profile stability, suggesting that the stability of graded clusters is not simply a consequence of behavioral consistency. This point will be added to the Discussion in the resubmission.

      Language of "inference" and "integration": The reviewer is correct that responses to novel tones are consistent with graded stimulus generalization. We will substantially revise the manuscript to replace "inference" and "integration" with more precise language describing graded frequency generalization gradients.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific and constructive criticisms. We will revise the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      The hypothesis will be rewritten as: "We hypothesized that generalization to tones acoustically similar to the CS+ and CS− depends on the emergence of stable ensembles encoding threat value, and that population-level response similarity across stimuli would correlate with the degree of behavioral fear generalization, consistent with prior work in auditory cortex [1]."

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      The summary statement will be rewritten: "Our results show that stable cortical sub-ensembles preserve the emotional content of the fear memory over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity in response to tones associated with threat correlates with the degree of behavioral threat generalization."

      (c) 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We will adopt the reviewer's suggested rewording: "In CS+15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization."

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we will follow Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis will allow us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we will revise and appropriately adjust our conclusions.

      In addition, we will revise the section heading and surrounding text to remove the terminology of “learned and inferred valence.” Instead, the findings will be described more conservatively as: “PL population responses reflect behavioral generalization to auditory stimuli following discriminative fear conditioning.”

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      The reviewer is correct that controls do not form threat associations; however, these animals still could respond differentially to distinct frequencies, something that is not reflected in the data. We will correct the section indicating that distinct neutral frequencies do not produce graded responses: "graded responses require associative learning" will be retained but reframed simply as: "graded frequency-dependent population responses were absent in animals that did not receive fear conditioning." The concluding statement of the paragraph will be rewritten as: "PL population responses reflect behavioral generalization to acoustically similar stimuli following discriminative conditioning," in line with the reviewer's suggestion.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We respectfully disagree with the reviewer’s interpretation. While repeated testing cannot be entirely excluded as a contributing factor, several lines of evidence suggest that it cannot fully account for our observations.

      Regarding extinction: discrimination ratios between CS+ and all other frequencies either remained stable or increased over time (new analysis included in resubmission), indicating that animals continued to discriminate threat value across the testing period rather than showing the progressive suppression expected under extinction — the opposite of what we observe.

      Regarding the recruitment of new neurons: repeated non-reinforced tone exposure would be expected to produce stimulus-specific adaptation — characterized by reduced, less discriminative neural responsiveness and flatter tuning profiles [2]— not the progressive sharpening we observe. The same would be expected if these neurons represent or are associated with new extinction learning.

      Finally, sharpening of generalization gradients during repeated within-subjects testing has been reported previously [3], suggesting that successive exposures may promote more precise discrimination in some cases. Consistent with this, discrimination learning has also been shown to narrow or sharpen fear generalization gradients rather than broaden them [4], supporting the idea that discriminative conditioning enhances stimulus specificity during testing. Although we cannot exclude the possibility that more extended training could eventually broaden the generalization gradient, under the training parameters and temporal window used in our study, the data support a progressive sharpening of the gradient over time. In the revised Discussion, we will present systems consolidation as the primary interpretive framework and further elaborate on why repeated testing is unlikely to account for the full pattern of behavioral and neural findings reported here.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      The GLM analysis described in our response to reviewers 1 and 3 will directly address the contribution of freezing. We will report these results in the resubmission and revise the interpretive language in the manuscript accordingly.

      However, regarding the analysis of population vector similarity, we need to clarify a point of confusion. The reviewer states “Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained”. The similarity vectors were calculated by correlating activity across all tone presentations within each testing day, not only the first two presentations. In Fig. 4, “Early” and “Late” refer to the order of a tone within a trial, which we will clarify more explicitly in the resubmission. Notably, repeated-measures analyses did not reveal any effect of the time variable (Fig. 4e,f), indicating that similarity across tone presentations remained high for tones associated with high threat value. Importantly, our data showed no evidence that responses to 11 kHz or 15 kHz in the CS15 group, or to 3 kHz in the CS3 group, exhibited extinction-like patterns at either the behavioral or neural level. Therefore, the persistence of high population similarity across time provides additional evidence against extinction as the primary explanation for our findings.

      We will remove the word "determines" from the manuscript, as our data cannot conclusively establish a causal relationship.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We will make this correction.

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We thank the reviewer for raising this point. The phrase "integrated learned valence across tones" refers specifically to a subpopulation of neurons that respond to all four frequencies in a graded manner, with response magnitude scaling according to threat value. This is distinct from tone-selective neurons, which respond preferentially to a single frequency. The neurons responding to all tones in a graded manner are present only in conditioned animals and not in no-shock controls, demonstrating that their graded response profile is shaped by associative learning.

      We agree, however, that the phrase "integrated learned valence" is unnecessarily opaque and we will replace it with more precise language: these neurons will be described as showing graded frequency-dependent responses whose magnitude scales with threat value. We believe this subpopulation represents a genuinely novel finding that complements the behavioral generalization literature by identifying a specific neural substrate for the generalization gradient within PL.

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      As stated in our previous responses, in the resubmission we will determine the contribution of freezing. If we find that freezing predicts graded neural responses, we will adjust the language of the manuscript.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting the lack of clarity in this passage and agree that the original phrasing was insufficiently precise. What we intended to convey is that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. Therefore, we suggest that neurons recruited into the active population after conditioning — likely frequency-selective neurons — contribute to the graded population responses through changes in their firing-rate activity, which is modulated by threat value (Fig. S8). We will rewrite this passage in the resubmission to make this interpretation explicit rather than leaving it to the reader to infer.

      Regarding the reviewer's suggestion that the characteristics of newly recruited neurons more likely reflect learning from non-reinforced exposures during repeated test sessions, we respectfully maintain that this interpretation is difficult to reconcile with two aspects of our data. First, graded-response neurons are absent in no-shock controls that are exposed to nonreinforced repeated testing. Second, as detailed in our responses to previous points, the progressive sharpening of population responses over time is inconsistent with what would be expected from repeated non-reinforced exposure, which would more plausibly produce broader or flatter tuning profiles.

      We agree that the phrase "shaped by associative processes" was ambiguous and will replace it with explicit language clarifying that we refer to fear conditioning as the associative process driving the emergence of graded responses, rather than any learning occurring during the test sessions themselves.

      (10) The following points all relate to the Discussion and reiterate many of the points above. 

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We will modify the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We will incorporate new analysis (GLM) to better address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We thank the reviewer for this comment and agree that both phrases require clarification or revision.

      By "graded representational axis" we intended to convey that PL population activity varies systematically as a function of stimulus similarity to the conditioned tone — that is, population responses are not categorical but scale continuously with spectral proximity to the CS+. We agree this was not clearly stated and will revise the manuscript accordingly.

      Regarding recurrent connectivity, we agree with the reviewer that nothing in our data directly measures or manipulates connectivity between neurons. This statement was intended as a speculative interpretive hypothesis in the Discussion, motivated by the established literature linking strong recurrent connectivity in prefrontal circuits to stable population-level representations [5]. However, we acknowledge that invoking it in this context, without direct evidence, risks overstating our conclusions. We will revise this sentence to make its speculative nature explicit and ground it more carefully in the cited literature rather than presenting it as an inference from our own data.

      In summary, we will ensure our conclusions will be restricted to population-level coding of learned threat value and its generalization across auditory frequencies. We will revise the relevant passages in the Discussion to ensure that speculative interpretations regarding emotional memory content are either removed or clearly flagged as speculative hypotheses.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this point. Our interpretation that these neurons reflect pre-existing PL sensory-driven properties was based on the observation that tone-selective responses were present in control animals that never received conditioning, consistent with prior reports of sensory responsiveness in PL cortex ([6, 7]. Because these responses emerge from the first time we expose mice to the intermediate frequencies, they cannot be explained by repeated exposure. Moreover, we did not observe progressive refinement, emergence of discrimination-like changes, or suppression of responding to non-reinforced tones in control mice. This difference between conditioned and control animals indicates that repeated tone exposure alone is not sufficient to produce the observed dynamics — associative learning is necessary. We therefore maintain that the tone-selective responses of these neurons reflect pre-existing sensory-driven properties of PL cortex that are present independently of conditioning history.

      In summary, we thank the reviewer for suggesting clarifications to our interpretation, for raising the possibility that freezing behavior may contribute to graded neural responses, and for raising the question of whether repeated tone exposure may contribute to the properties of neurons recruited after conditioning. In the revised manuscript, we will include additional analyses to better dissociate the contributions of freezing behavior and tone identity, clarify passages that were insufficiently precise, and include a paragraph in the Discussion addressing potential alternative explanations alongside our own interpretation of the data.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded-response ensembles may partially reflect freezing-related activity rather than mnemonic or salience-related representations of the conditioned stimuli themselves. In the revision, we will acknowledge that prior work has identified PL neurons that encode freezing independently of stimulus identity or associative content. Furthermore, we will implement the reviewer’s suggested generalized linear model (GLM) approach using inferred spiking activity derived from the Ca2+ signals. Specifically, we will include both stimulus identity and freezing behavior as predictors. Because freezing varies across trials whereas stimulus presentation is fixed, this analysis will allow us to dissociate the relative contributions of stimulus-related versus freezing-related activity to the graded neuronal responses. We thank the reviewer for this excellent suggestion.

      If graded stimulus coding remains significant after accounting for freezing behavior, this would strengthen the interpretation that these ensembles encode learned salience or associative properties of the stimuli rather than behavioral output alone. Conversely, if freezing explains a substantial proportion of the variance, we will revise our interpretation accordingly.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such holographic two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      To provide an indirect link between ensemble organization and behavior within the constraints of the current dataset, we will examine inter-individual variability in the revised manuscript. Specifically, we will test whether the proportion of neurons participating in stable graded-response ensembles versus dynamic stimulus-specific ensembles predicts individual differences in freezing behavior and fear generalization across retrieval sessions. If animals with a higher proportion of stable graded-response neurons show stronger discrimination and less generalization to non-conditioned tones, this would strengthen the association between ensemble organization and behavioral outcome, while remaining correlational in interpretation.

      We will modify the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the associative nature of our conclusions.

      References

      (1) Aschauer, D.F., et al., Learning-induced biases in the ongoing dynamics of sensory representations predict stimulus generalization. Cell Rep, 2022. 38(6): p. 110340.

      (2) Kato, H.K., S.N. Gillet, and J.S. Isaacson, Flexible Sensory Representations in Auditory Cortex Driven by Behavioral Relevance. Neuron, 2015. 88(5): p. 1027–1039.

      (3) Vervliet, B., et al., Generalization gradients in human predictive learning: Effects of discrimination training and within-subjects testing. Learning and Motivation, 2011. 42(3): p. 210–220.

      (4) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (5) Mante, V., et al., Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature, 2013. 503(7474): p. 78–84.

      (6) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (7) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive and mechanistic analysis of how tsetse flies feed on blood across a wide range of host skin types. The authors combine detailed anatomical characterization of the feeding apparatus with quantitative measurements of mechanical properties, probing forces, and blood uptake, complemented by experiments using artificial skin. They show that tsetse flies do not rely on extreme forces or uniquely specialized structures, but instead on subtle and highly efficient structural and mechanical adaptations (such as the toothed labellum and coordinated proboscis movements) to achieve effective blood pool feeding. The study successfully moves beyond descriptive anatomy to a quantitative, functional analysis that explains how feeding is accomplished across diverse substrates.

      Strengths:

      A major strength of the work is the impressive integration of multiple complementary approaches. Advanced imaging tools provide a convincing three-dimensional view of the proboscis, labellum, and associated structures, while direct force measurements and blood intake quantification place these observations on a solid quantitative footing. The use of artificial skin with different mechanical properties is particularly powerful, as it allows structure-function relationships to be tested under controlled and reproducible conditions. Together, these datasets provide strong and coherent support for the authors' central conclusions. The quantitative treatment of feeding mechanics represents a significant advance over largely descriptive prior work by others (e.g., Gibson W et al 2017) and establishes a valuable mechanistic insight for studying blood feeding in insect vectors more broadly.

      Weaknesses:

      The study focuses almost entirely on uninfected flies and does not address how infection might alter feeding mechanics or performance. Previous work has shown that trypanosome infection can affect salivary gland function and feeding time (Van Den Abbeele et al 2010), and even cause damage to mouthparts, all of which can influence feeding behavior and efficiency. While this does not detract from the technical quality or the core findings of the study, a more explicit discussion of these biological variables would help place the results in a broader transmissionrelevant context and clarify how generalizable the conclusions are to natural infection settings.

      We thank the reviewer for this important comment. While our study focused on uninfected flies, we agree that parasite infection may influence feeding performance and should therefore be considered when assessing the broader relevance of our findings. Previous studies have shown that trypanosome infections can alter salivary gland physiology and saliva composition (Van Den Abbeele et al., 2010; Matetovici et al., 2016). In addition, transcriptomic analyses suggest that infection with T. congolense may affect the molecular and physiological state of the proboscis (Awuoche et al., 2017). However, there is currently no direct evidence that these changes translate into fundamental alterations of the mechanical properties or function of the mouthparts themselves, which were the primary focus of our study. We have now expanded our manuscript to discuss this (lines 508-520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, leading to increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Overall, this is an outstanding and carefully executed study that will have a significant impact on the fields of vector biology and parasite transmission.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an impressively detailed, multidisciplinary analysis of the mechanics of blood feeding in Glossina spp. Combining SEM, CLSM, µCT, FIB-SEM, macro-videography, and quantitative force measurements, the authors characterize the structures and biomechanics of attachment, proboscis deployment, tissue penetration, and blood uptake. They also examine interactions with diverse host-type substrates, from human skin equivalents to cow, deer, and lizard skin, and integrate these with force measurements to quantify penetration and retraction dynamics.

      The work's key conclusion is that the tsetse fly does not rely on any single exceptional morphological innovation, but rather uses a suite of subtle structural features and retractive forces to feed efficiently across diverse hosts. This result is novel, insightful, and evolutionarily compelling. Overall, this is a strong manuscript that combines methodological sophistication with biological relevance. It should be of high interest to researchers studying vector biology, biomechanics, parasite transmission, and vector-host interactions.

      Strengths:

      (1) The combination of SEM, CLSM, µCT, and FIB-SEM provides an unusually comprehensive anatomical characterization of the tsetse feeding apparatus.

      (2) The direct measurement of proboscis penetration and retraction forces across diverse substrates is highly original and fills a major knowledge gap in vector-host interaction mechanics.

      (3) The study bridges morphology, mechanics, behavior, and host tissue properties, which strengthens the overall conclusions.

      (4) Imaging of trypanosomes within the hypopharynx and surrounding tissue during feeding provides new information about parasite delivery mechanisms.

      Main Comments:

      (1) The authors conclude that feeding versatility arises from the sum of subtle adaptations. This interpretation is reasonable, but it would help to sharpen which findings most robustly support this statement. For example, the relative similarity of proboscis forces across skin types is compelling evidence that the proboscis is broadly tuned rather than specialized. The observation that tsetse targets softer interscale regions on lizard skin suggests behavioural selectivity, not morphological specialisation. It would strengthen the discussion to highlight which data most directly refute the hypothesis of a unique specialization.

      We thank the reviewer for this comment. To address this point more explicitly and to sharpen the interpretation of our findings, we have expanded the final conclusion in the Discussion (lines 528544):

      "Ultimately, the objective of this study was to investigate how tsetse flies can feed on a seemingly random selection of animals with highly diverse skin structures. In our detailed anatomical studies and force measurements, we did not identify a single dominant trait that explains the fly's feeding versatility.

      Instead, our results indicate that this capability emerges from the combined effect of multiple, more subtle traits. In particular, the proboscis generates broadly similar penetration forces across a wide range of skin types, suggesting a generalised mechanical mechanism rather than hostspecific optimisation. The intricate architecture of the labellum and the strong retractile forces during probing likely contribute to efficient penetration and blood pool formation across heterogeneous substrates. Behaviourally, tsetse flies further increase feeding success by flexibly targeting mechanically favourable sites, such as the softer interscale regions on lizard skin, rather than relying on specialised morphological adaptations.

      This composite strategy likely reflects evolutionary fine-tuning that enables the broad host range of tsetse flies. By allowing efficient blood feeding across diverse vertebrate hosts, this versatility may also have facilitated the ecological success and transmission opportunities of African trypanosomes.”

      (2) A central finding is that retraction forces exceed penetration forces across substrates, implying that backward pulling is a key component of wound creation. However, the biological interpretation could be deepened. Specifically, do the authors believe retraction serves primarily to enlarge the pool-feeding site? How does this compare mechanically to mosquito fascicle oscillation or other blood-feeding arthropods (especially other flies such as those in the tabanidae family)? Could retraction forces contribute to anchoring or resisting host grooming behaviors?

      The stronger retraction forces observed during probing indeed suggest that backward pulling is not a passive withdrawal, but likely an active component of tissue disruption. As discussed in the manuscript (lines 475–482), we interpret these repeated pullback movements, together with the outward-facing prestomal teeth of the everted labellum, primarily as a mechanism to enlarge the feeding lesion and improve access to blood, consistent with the blood pool feeding strategy of tsetse flies. To make this more clear, we have added a half sentence to line 482 "..., thereby creating a larger blood pool for feeding."

      We also already compare this mechanism to mosquito feeding mechanics in the discussion (starting from line 487). In mosquitoes, high-frequency fascicle oscillations are thought to reduce insertion resistance and facilitate minimally invasive capillary feeding. Although we also observed oscillatory movements during tsetse feeding (Video 4), the underlying mechanical strategy appears fundamentally different. In contrast to the mosquito’s system optimized for delicate penetration, the tsetse proboscis appears adapted for forceful tissue disruption during pool feeding. Notably, the oscillations observed in tsetse flies seem to occur during active blood uptake rather than initial tissue penetration. Consequently, the functional role of these oscillations in tsetse flies remains unclear. We have now addressed this more specifically in the discussion (lines 490-495):

      "Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown."

      When looking at other species, stable flies (Stomoxys) may represent a particularly relevant comparison because they employ a similar penetration mechanism and are pool feeders with prominent prestomal teeth (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). In contrast, tabanids employ a different mouthpart architecture with rasping/cutting structures but without comparable prestomal teeth. Whereas mosquito mouthparts have been described as functioning like a syringe, and we compare the tsetse proboscis to a saw, tabanid mouthparts have been likened to scissors (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). Although tabanids are also known to inflict substantial tissue damage, it remains unclear whether their feeding movements produce retraction-dominated force patterns comparable to those we observed in tsetse flies.

      Lastly, we agree that the elevated resistance generated during retraction may contribute to withstanding host defensive behaviour such as shake-off responses. Structurally, the orientation of the prestomal teeth and the architecture of the everted labellum could provide temporary anchoring during feeding, as we have already briefly discussed in the manuscript (lines 458– 461). However, while stronger anchoring may increase feeding stability, it could also increase the risk of injury to the fly if detected by the host. Compared to other pool-feeding flies such as stable flies, tsetse flies have been reported to respond more readily to host defensive behaviour (Schofield and Torr. A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002). We therefore currently consider anchoring to be a possible secondary function but lack direct experimental evidence to assess its practical importance.

      (3) The study analyzes a diverse set of substrates, which is a strength. However, some caveats deserve explicit discussion. Human skin equivalents and dermal equivalents lack the full mechanical complexity of real skin (e.g., innervation, perfusion, tension). Frozen or ethanol-stored samples, particularly reptile skin, may also exhibit altered mechanical properties compared to live tissues. These limitations do not undermine the findings but should be explicitly acknowledged as they influence the interpretation of absolute force magnitudes.

      The reviewer raises a valid point regarding the interpretation of absolute force magnitudes across the measured substrates. We have therefore added a clarifying statement to the discussion (lines 469-476):

      "When interpreting absolute force magnitudes, it is important to bear in mind that our samples do not fully recapitulate physiological conditions. Skin explants and skin equivalents may behave differently to skin under active perfusion and native tissue tension, as may our fixed and frozen animal skin samples. Nevertheless, comparative force measurements revealed consistent biomechanical signatures across substrates, suggesting that the observed force patterns reflect fundamental aspects of the feeding mechanism that are likely relevant in vivo.”

      (4) The SEM and FIB-SEM images showing trypanosomes in the hypopharynx and surrounding tissue during penetration are visually striking and suggest rapid dispersal. It would be helpful to connect these observations more clearly to the kinetics of parasite deposition and whether mechanical tissue laceration is likely to increase inoculation efficiency. Without conducting additional experiments, the authors could discuss whether these findings support or modify existing models of salivary-gland-derived parasite release.

      We have now expanded the Discussion to clarify that our observations of trypanosomes in the hypopharynx are consistent with the established model of salivary-gland-derived parasite release during probing and feeding, in which infective metacyclic trypanosomes are delivered with saliva into the host tissue. Furthermore, the presence of trypanosomes beyond the immediate feeding canal supports rapid parasite dispersal following inoculation, as described in previous work (Reuter et al., 2023). In this context, the tissue laceration generated by the tsetse proboscis may facilitate local parasite distribution by creating a larger, mechanically disrupted feeding lesion. However, our data provide high-resolution structural snapshots and were not designed to quantify deposition kinetics or inoculation efficiency. We therefore refrain from concluding that mechanical laceration increases transmission efficiency and instead view this as a plausible consequence that should be tested directly in future work. Specifically, we have added this paragraph to the discussion (521-527):

      "Overall, our observations of trypanosomes within the fly's hypopharynx, labial gutter, and host tissue are consistent with the established model of salivary-gland-derived parasite release during probing and feeding. Their presence beyond the immediate feeding canal is consistent with rapid local dispersal following inoculation, as described previously (Reuter et al., 2023). This process may be facilitated by the extensive tissue disruption caused by the tsetse mouthparts, although this hypothesis will require direct experimental testing."

      (5) The authors demonstrate that tsetse attachment abilities fall within the range of generalist insects and are far lower than those of obligate ectoparasites. However, the manuscript could discuss how attachment forces relate to the tsetse's ecological context, e.g., whether their attachment is generally brief, whether host shaking strongly selects for grip strength, etc. Is there evidence that other Glossina species or tabanids with different host preferences show variation in attachment performance? This would broaden the relevance of the findings.

      Tsetse flies are obligate blood feeders, but host contact is typically brief and frequently interrupted by host defensive behaviour. As a result, selection may favour rapid and efficient feeding rather than exceptionally strong attachment. This interpretation is supported by Schofield and Torr (A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002), showing that tsetse flies experience more feeding interruptions than the stable fly Stomoxys calcitrans, despite completing successful blood meals in less time. These differences are consistent with life-history theory (Anderson and Roitberg. Modelling trade-offs between mortality and fitness associated with persistent blood feeding by mosquitoes. Ecology Letters, 1999), which predicts that long-lived species with low reproductive rates, such as tsetse flies, should be less willing to risk injury by persisting on a host than shorter-lived, more fecund species. Against this background, our finding that tsetse attachment forces fall within the range reported for generalist insects, appears biologically plausible. Their attachment performance needs to be functionally sufficient for brief feeding events rather than maximized for prolonged host retention.

      We are not aware of comparative biomechanical data on attachment performance across different Glossina species or tabanids. We agree that such comparative studies would be valuable to test whether differences in host preference and feeding ecology correlate with variation in attachment capacity.

      (6) In video 4, could the authors clarify whether the observed maxillary vibrations are hypothesized to reduce penetration resistance or serve another function?

      The vibrations of the maxilla specifically appear during active blood uptake rather than during initial tissue penetration, suggesting they are linked to the ingestion phase. Whether they serve a mechanical function, such as facilitating blood flow or preventing canal occlusion, or represent a passive consequence of pump activity, remains unclear. We consider this an open and interesting question that warrants dedicated investigation.

      We have therefore clarified that the functional significance of these oscillations remains unresolved to date (lines 490-495). This reads: “Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown.”

      Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The authors aim to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely, the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!), so researchers interested in trypanosomes, tsetse flies, or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figures 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips, explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and the probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      We agree that “jigsaw” may be too specific and not fully appropriate in this context. We have therefore replaced it in the manuscript with the more general term “saw.”

      The impressive aspect of this paper is the range of imaging techniques (CLSM, SEM, uCT, FIB SEM), the quality of the images, which attests to the obvious care taken with sample preparation. The biomechanical analysis, especially the penetration analysis, is impressive. Finally, the paper is clearly written and presented; it was a very easy read and, overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis, but that is not a prerequisite for sharing it. Perhaps the least convincing parts are the imaging of the flexible versus rigid parts of the structures, which is based on the amount of resilin (flexible) and chitin-protein (stiff), based on their autofluorescence. It seems odd that the joints would be less blue (stiffer) in Figure 1i, or what the blue structures correspond to in Figure 6B-D.

      Our analysis is based on established CLSM approaches that use exoskeleton autofluorescence as a proxy for relative differences in cuticular composition and material properties (Michels & Gorb, 2012; Michels et al., 2016). In the tarsus, the observed differences in inferred stiffness are relatively subtle, with most regions exhibiting broadly comparable material properties. This becomes particularly evident when compared with the proboscis, where the contrasts in cuticular composition are much more pronounced (Figure 6). We also note that locally stiffer regions at joints are not unexpected, as stiffness gradients in arthropod joints can provide mechanical support and help constrain the direction of movement. Importantly, our images show a flexible, ring-like blue region directly at the articulation, surrounded by slightly stiffer material. We therefore interpret this pattern as a combination of a flexible hinge region and adjacent supporting structures that together enable controlled joint motion.

      The blue structures in Figure 6B–D correspond to flexible regions of the furca (f). Because this spring-like cuticular element undergoes substantial configuration changes during labellar eversion, the presence of highly flexible regions is consistent with its proposed mechanical function.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further experiments or analyses are suggested. However, the Discussion would benefit from briefly acknowledging how trypanosome infection can alter feeding behavior and mouthpart function, based on prior work, to place the mechanical findings in a more biologically relevant transmission context.

      We thank the reviewer for this suggestion. Previous studies have indeed shown that trypanosome infection can alter tsetse feeding behavior, primarily through changes in saliva composition. Van den Abbeele et al. (2010) demonstrated that infection with T. brucei significantly impairs the anti-haemostatic activity of tsetse saliva, resulting in prolonged prefeeding probing and therefore extended feeding times. These findings are consistent with earlier observations by Jenni (1980), who reported increased probing frequency in infected flies.

      Jenni (1980) also proposed that these behavioral changes might be linked to altered mechanoreceptor function. However, Van den Abbeele et al. (2010) argued against this interpretation for T. brucei, noting that this parasite does not colonize the mouthparts where these mechanoreceptors are located. Taken together, the available evidence suggests that the observed changes in feeding behavior are mediated primarily through altered interactions with host blood rather than through direct effects on the mouthparts themselves.

      It should be noted that this conclusion is specific to T. brucei. Other tsetse-transmitted trypanosome species, such as T. congolense, do colonize the proboscis. However, a comparative study examining flies infected with T. brucei, T. congolense, or T. vivax found no significant effects of infection on feeding behaviour relative to uninfected controls (Moloo, 1983). To our knowledge, there is currently also no direct evidence that any trypanosome species alters the physical properties or mechanical function of the mouthparts, or causes damage that would directly affect feeding performance. We have added this paragraph to the Discussion (lines 508520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, including increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Reviewer #2 (Recommendations for the authors):

      Several figures (particularly SEM-based ones) contain very dense labeling. Consider providing simplified overviews or annotated "orientation guides" in figure supplements to improve navigability for readers unfamiliar with proboscis anatomy.

      We thank the reviewer for this helpful suggestion. While we agree that orientation aids can be valuable, we have decided not to include additional simplified overview figures, as we consider that introducing separate schematic summaries could potentially complicate rather than improve navigation of the structural detail. We therefore rely on consistent labelling within the existing figures and detailed captions to guide interpretation.

      The manuscript uses appropriate non-parametric tests, but could benefit from reporting effect sizes and indicating sample sizes on all plots.

      Sample sizes are reported in the figure legends, Methods section, and Supplementary material for all experiments. We agree that reporting effect sizes can be informative and will consider this in future studies. However, because the primary objective of the statistical analyses in the present work was to support comparisons between experimental conditions rather than to estimate effect magnitudes, and because the figures are already information-dense, we therefore decided not to further modify the graphical presentation in this revision.

      Reviewer #3 (Recommendations for the authors):

      (1) P5 L111. Perhaps indicate these knobs on the image Figure 1S). I assume these are the structures visible under the pointer labelled spa? Maybe highlight some of the worn areas in Figure 1G.

      The knob-like structures in Figure S1 are highlighted in green and we have now revised the figure description from:

      “…showing fine crests on the underside and surface modifications (green) on the upper side.”

      to:

      “…showing fine crests on the underside and knob-like surface modifications (green) on the upper side.”

      Regarding Figure 1G, the purpose of the panel is to illustrate the contrast between deformed spatulae (Figure 1G) and intact spatulae (Figure 1H). We therefore chose to retain the original presentation, as we feel that additional markings would not substantially improve interpretation and could obscure structural details. We hope that the direct comparison between the two panels provides sufficient visual guidance.

      (2) P9. The frictional force (and P38/39) has the units of N (kg.m/Sexp2). The safety factor is this force divided by the weight of the fly? So are there units (Kg/sexp2) or are these not shown? Perhaps this is a convention.

      The safety factor is defined as the ratio of the total frictional force to the fly’s weight force (m·g), where m is body mass and g is gravitational acceleration. Since both quantities are express in Newtons (kg·m·s<sup>-2</sup>), the safety factor is dimensionless.

      We agree that the terminology in the original manuscript may have been ambiguous, as “body weight” is sometimes used colloquially to refer to body mass. To avoid confusion, we have revised the text to explicitly refer to weight force and now define the safety factor as the total friction force divided by weight force (mg, where m is body mass and g is gravitational acceleration). We have clarified this in the main text, the Figure 2 legend, and the description of Supplementary Material 1.

      (3) P10 Figure 2G & H. It is not very clear...are these the data, the average of all readings across all surfaces in E and F? If so, why is this value useful...how does it add to what is already shown?

      The figures 2G and 2H summarize the friction forces (G) and safety factors (H) across all tested substrates, based on the values from the male (B, E) and female (C, F) datasets. The purpose of these panels is to provide an overall comparison between sexes independent of substrate type. While this information can also be inferred from the substrate-specific plots, the sex-separated presentation does not make the absence of an overall sex difference immediately obvious. Figures 2G and 2H therefore serve as concise summary plots highlighting this result.

      (4) P12. For the nonspecialist, it might be useful to draw a cartoon showing the organisation of the labium, labrum and the hypopharynx...this is visible in Figure 4i but not in the dissected proboscis and labellum ....only the labium as the labrum doesn't extend this far?

      To clarify the anatomical arrangement in the dissected specimen, we have added the following statement to the Figure 4 legend (lines 224–226):

      “In an intact fly, the labrum would be positioned within the empty groove of the labium visible in J; however, it is absent in this dissected preparation.”

      (5) P17 legend to Figure 5. Include the Lm abbreviation in the legend, and maybe a close-up of the rsp teeth?

      We have added “lm, labellum” to the Figure 5 legend (line 250), as this abbreviation was previously missing. Panel J is a close-up of the rasping teeth.

      (6) F3S and Video 3. Are the images in B and C taken from the FIB SEM video images? It is not clear. A small legend descriptor for video 3 would be helpful.

      The images in Supplementary Figure 3B and C are reconstructed from the same FIB-SEM dataset shown in Video 3, but they are displayed in a different orientation. This is indicated schematically in Supplementary Figure 3A, which illustrates the viewing plane used for the reconstruction.

      We already included the following legend for Video 3 (lines 1146–1149):

      "Video 3: FIB-SEM of the tsetse labellum. Sequential cross sections reveal internal ultrastructure progressing from near the tip of the labellum downward. Data were acquired on a Crossbeam 540 (Zeiss) with the EsB detector in continuous milling mode."

      To improve clarity, we have now added a sentence to the video legend linking the figures to the video: (lines 1149-1150)

      “Reconstructed images from this dataset are shown in Figure 5A and Supplementary Figure 3B and C.”

      In addition, we have now explicitly cross-referenced Video 3 in the legends of Figures 5 and S3 to make the connection clearer for the reader.

      (7) Figure 7. These are amazing images, especially G-I.

      Thank you for this positive feedback, we appreciate it.

      (8) P24. It is really good to see that there is a difference in force penetration for full skin v dermal...this deserves a comment.

      We agree and have revised the text accordingly. We replaced:

      "Human skin substrates required the lowest penetration forces, with 0.97 mN for full-thickness skin equivalents, 0.67 mN for dermal equivalents, and 0.85 mN for native skin explants (Figure 8C, D)."

      With this (lines 363-367):

      "Human skin substrates showed the lowest penetration forces, with dermal equivalents requiring less force (0.67 mN) than full-thickness skin equivalents (0.97 mN), reflecting the additional mechanical resistance of the epidermal layer absent in dermal-only constructs. Native skin explants fell intermediate at 0.85 mN (Figure 8C, D)."

      (9) P26 Figure S5. Panel c, there seems to be a big scatter in the drinking time. Was there an outlier?

      Indeed, the observed scatter is due to a single fly with an unusually long drinking time of 184.44 seconds, which is approximately six times the median duration. We have verified the underlying data and found no indication of a measurement error; the value therefore remains included in the analysis. The data for the plots in Supplementary Figure 5 are also available in Supplementary Material 3.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

      Comments on revised version.

      The manuscript overall is significantly improved and authors addressed majority of my concerns. The addition of the computational vertex model (Figure 7) as well as Atg1 RNAi (Figure 4) to inhibit cell fusion provide more mechanistic insight to their study. However, the analysis of Atg1 RNAi wound assay falls short as it does directly measure changes in syncytium frequency nor size to confirm that cell fusion is reduced. The authors should quantify the number of nuclei per syncytium over the 2hr wound healing period as performed for WT in Figure 1C. It would have been ideal if they could have also performed the Act-GFP spreading assay in WT and Atg1 RNAi strains to determine if Act-GFP movement is dependent on cell fusion as purposed. At the least, further quantification of Atg1 RNAi phenotype is warranted to support their conclusions.

      In response to the reviewer's comment, we have repeated the analysis of Fig 1C and generated a new panel, Fig. 4C, which is directly comparable to the control and shows that syncytial size is dramatically reduced in the Atg1 knockdown area. Unfortunately, we cannot perform the second analysis of actin-GFP spreading in the Atg1 knockdown cells because we need Gal4 for labeling individual cells and for knocking down Atg1, and we can't do both at the same time.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Comments on revised version:

      The authors have extended their original manuscript by adding two key parts. First, they show a role of Atg1 in mediating cell fusion (Figure 4). Second, they provide additional evidence for a contribution of radial border fusions to wound closure through its effect on tissue fluidity and through computational modelling (Figure 7).

      This new version of the manuscript is greatly improved and provides significant new insights into the role of syncytia in aiding wound repair. There are just a few minor, yet important, additions needed to back up Figure 4 which should not require new experiments.

      Minor but important points:

      The authors show a role of Atg1 in mediating syncytia formation in Figure 4. However, since the Pnr>+ side of the wound closes slower than the non-Pnr side (control side), a few additions to this figure would be important and should not require additional experiments.

      (1) The authors should show, similar to the data shown in Figure 4D of the wound radius over time for control versus Pnr>Atg1RNAi, also the same type of data for control versus Pnr>+.

      The data the reviewer requests is available in our bioRxiv manuscript, in Fig. 6B (Hua, Krystofiak, Pumford, Page-McCaw, and Hutson, https://doi.org/10.64898/2026.05.31.728998). These experiments were all done at the same time. As you can see, the difference in closure rate is quite subtle in control wounds.

      (2) Since Pnr>+ also slows down wound healing, albeit to a lesser extent than Pnr>Atg1, the authors should also show an extra graph that provides evidence that Pnr>Atg1RNAi reduces syncytia formation more than Pnr>+ does. E.g. Two graphs could be added that show individual cell size at 4 or 5h post wounding for control versus Pnr>Atg1RNAi as well as for control versus Pnr>+ and also another graph with the same data but comparing cell size between Pnr>+ and Pnr>Atg1RNAi. Otherwise, if the expected minimum cell size for a syncytium is easy to estimate, a graph could be added that shows the percentage of cells that are above this threshold (e.g. above 100 square micron) for control versus Pnr>Atg1RNAi and control versus Pnr>+ and Pnr>+ versus Pnr>Atg1RNAi.

      In response to this comment and the comment from reviewer 1, we have now added new Fig. 4C, which addresses the reviewer's question about the comparative frequency of fusion in pnr>Atg1RNAi and pnr>+. These graphs show that Atg1 knockdown significantly reduces the size of syncytia.

      Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches. Some clarifications are needed as described below.

      Major suggestions:

      (1) Introduction, page 4. The examples of developmental syncytia formation of invertebrates and vertebrates are confusing. The authors may want to make the examples clear and add additional examples. Currently, readers may assume that C. elegans cell fusions occur only in the hypodermis - other structures can be mentioned like the vulva, pharyngeal muscles, glia, tail. In addition, the authors may want to add injury-induced fusions like the C. elegans' PLM and PVD neurons (Ghosh-Roy et al., 2010; Newman et al., 2015; Oren-Suissa et al., 2017).

      We appreciate the suggestions and have included the additional examples of C. elegans vulva and PLM and PVD neurons. We are limiting ourselves to those because we don't want to focus too heavily on C. elegans examples, as that's not the direction this paper is heading.

      (2) In cases where it is not clear whether fusion has occurred or whether mononucleated cells were ejected from the leading edge, membrane markers can be used. Page 6. Lines 96-99. The authors may want to use a membrane marker like RFP-PH driven by the epithelial cell promoter.

      At this point in the manuscript, we are introducing syncytia and are not concerned yet with their origin. 

      (3) Pages 8-10. The authors may want to clearly explain that apical junctions shrinking is a post fusion event. That the apical shrinking is caused by the expansion of fusion pores and the migration of apical junctions towards the basolateral domain. This is something that was clearly shown during physiological epidermal cell-cell fusion in C. elegans by Mohler et al., 1998 and 2002. A cartoon showing the process of cell-cell fusion, pore expansion and apical junction dynamics would make the manuscript much clearer.

      Apical shrinking cannot be caused by the "migration of apical junctions towards the basolateral domain" because that is not what we observed -- rather, we observed labeled adherens junctions remaining at the apical surface while the area they enclose becomes smaller (shrinks). Further, despite close reading of the Mohler papers, it is not clear how similar the apical shrinking events of this manuscript are to the fusion events described there. Finally, we do not want to include a schematic describing this process because that would suggest certainty that we do not have. Unlike in C. elegans development, wound-induced cell fusion is stochastic, not stereotyped; with cells that display apical shrinking, the fusion partner of a labeled cell is difficult to identify because it is often not a neighboring cell. These factors make it difficult to describe this process in detail, but we have sufficient data to conclude that these are indeed cell fusion events.

      (4) Page 9. Line 170. "...as these cells represent fusion initiation events (fusion pore) but were unable to productively stabilize and expand the site of fusion and so returned to the diploid state." The authors may want to make clear that this is an assumption that needs to be tested. Live imaging using a membrane marker may resolve whether a reversible fusion pore was generated.

      Thank you for the suggestion; we have updated this text to make it clear that this is an interpretation.

      (5) Page 11. It is not clear whether Atg1 is directly required for cell fusion, or that autophagy is required for efficient cell fusion or both Atg1 and autophagy participate in the fusion process.

      Our data show that Atg1 is required for cell fusion. The work that inspired this experiment, Kakanj et al 2022, concluded from their more comprehensive studies that the process of autophagy was required. We have clarified the text.

      (6) Page 12. Line 235. "Indeed, we observed that several hours after wounding, the entire leading edge was occupied by syncytia." This observation is based only on the adherens junction marker. Can they test basal cell membrane marker? Is it possible that the mononucleate cell in the leading edge is under the two syncytia?

      Unfortunately, there are not good basal markers -- the recently reported basal spot markers also label adherens junctions. Nonetheless, we are confident that the mononuclear cell is not under the syncytia because we image Z-stacks and thus can detect cell overlap.

      Recommendations for the authors:

      Reviewer #3 (Recommendations For The Authors):

      Minor suggestions:

      (1) Figure 1. The authors may want to add an image immediately after laser ablation of the actual wound and the area around the wound. Add an arrow to mark the wound.

      With this wounding modality, the extent of the wound is unclear for ~30 min. As we reported in O'Connor et al, PLoS One, 2021, there is a gradient of damage emanating out from the center of the wound, and cells with greater amounts of damage die while those with less damage repair and survive. Immediately after laser ablation, very little visible damage is evident by 120ctn-RFP and Histone-GFP (the markers in Fig. 1) until the cells die and the surrounding cells respond.

      (2) Page 6. Line 86. "A mitotic tissue utilizes cell-cell fusions during wound repair." replace "during wound repair" with "after wound induction" since in this section the authors do not show that this process is part of wound repair.

      Thank you for the suggestion - we reworded this heading to remove "wound repair".

      (3) Page 6. Line 92. The authors may want to be consistent with the terms used in the text and in the figure - His2GFP in the text versus Histone GFP in the figures.

      Thank you for the suggestion, we have revised for consistency.

      (4) Figure 1 - supplement figure 1D. The "v" of Div panel moved below D.

      Thank you, we have corrected it.

      (5) Figure 1 - supplement figure 1G. add "i" to second Gii to make it Giii.

      Thank you, we have corrected it.

      (6) Page 24. Figure 1H legend. 3 or 4 wounds?

      Thank you for catching this error - 4 wounds.

      (7) Page 7. Line 124. "GFP mixing always preceded border breakdowns (n=11)" instead of "always" use "in all observed cases".

      We have made this change.

      (8) Figure 2. Switch the writing "Apical Shrinking: Nuclear Transfer" since apical shrinking represented in panel 2A and Nuclear Transfer in panel 2B. If this description applies only to panel 2B, make it clear.

      We consider this heading to apply to panels A and B together (as they show the same sample, just different channels).

      (9) Figure 2C. Is ActinGFP a cytoplasmic GFP driven by actin promoter or Actin-bound GFP? Cytoplasmic GFP versus membrane-cortex GFP?

      It is a transgene expressing an actin-GFP fusion protein, as noted in the key reagents table and discussed in Fig. 8. We corrected the manuscript to ensure it is always referred to now as Actin-GFP in the text, figures, and legends.

      (10) Video 3 - Impressive movie!

      Thank you!

      (11) Page 9. Line 155. "In both these cells, as the cell lost its basal volume, cytoplasm moved laterally to join the neighboring syncytia." It seems that the apical shrinking cells' cytoplasm joined the neighboring syncytia even before.

      Because both indicated cells (yellow and white arrows) and the neighboring syncytium are all labeled with GFP, it is not possible to determine precisely when the cells' cytoplasm joined the syncytium.

      (12) Page 9. Line 158. "...but fusions associated with apical shrinking occurred later and were more numerous." Did the fusion occur later or the apical shrinking itself as was mentioned before and shown in Figure 2F?

      We have changed the wording, as for many apical shrinking events we cannot tell exactly when the fusions were initiated.

      (13) Page 25. Figure 3A legend. What is the meaning of morphological fusion? Border breakdown and apical shrinking? The authors may want to define it.

      We have defined it now in the legend.

      (14) Page 26. Figure 3B-C legend. "Panel C shows that apical shrinking fusion and border-breakdown fusion occur at similar distances from the wound." It seems that fusion by apical shrinking mostly occurs within 60-70 um from wound center and fusion by breakdown occurs equally at all distances up to 80 um.

      We don't disagree with your comment, but we feel the dataset is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (15) Page 9. Line 165. "...but infrequently (n=3) with GFP mixing and no subsequent cell fusion..." Does this mean that there were GFP mixing without border breakdown or apical shrinking?

      Yes, that is correct. We assume that in this case a fusion pore opened and then closed again. We have added a phrase to clarify.

      (16) Page 9. Line 175. "...the spatial distribution of fusing cells that shrank vs. lost borders was similar (compare Figures 1G and 2E)." Even though the visual comparison suggests similar spatial distribution, the carefully quantified distribution in figure 3C suggests more fusion by shrinkage at 60-70 um from wound center of the 5 tested wounds.

      As we noted to comment 14, we feel the data set is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (17) Figure 3. The shown pies sum the results from 5 wounds. It would be interesting to add a graph comparing the percentage of fused and persisted cells per wound to see the variability, if exists.

      Unfortunately, the number of fused/persisting cells in each wound is greatly affected by the heat-shock conditions that generate the labeled clones; even the ratio of these fates would be heavily influenced by noise because the numbers are small in each animal. Further, the frequency of fusion is determined by the wound size as shown in Fig. 1. Because of these variables, such data could be easily misinterpreted.

      (18) Figure 3 - figure supplement 1D. Even though it was mentioned that the duration of some border breakdown is finished within minutes it is worth comparing it with shrinking duration on one graph.

      Unlike apical shrinking, it is difficult to identify exactly when border breakdown concludes, so this data is difficult to compare. We have provided several examples of border breakdown in the manuscript that give an overview of the process.

      (19) Video 1 is not mentioned in the main text.

      Thank you for catching that omission. We now refer to it in the first paragraph of the results.

      (20) Figure 4B. The difference between the treated group and the control group is unclear. Add arrows.

      We have added some arrows to Fig. 4B.

      (21) Figure 4C. For consistency use percentage for both border breakdown and shrinking cells.

      In response to the reviewer's comment, we now provide the consistent metric of number of lost borders and number of shrinking cells.

      (22) Page 11. Did the authors try other wound types (e.g. mechanical/chemical wounds)? May other wound causes besides laser ablation result in different response? This may help to answer whether there is a causation between syncytia formation and speed wound closure.

      There are reports of puncture and pinch wounds inducing cell fusion. Perhaps the reviewer is suggesting that we might be able to identify a wounding method that does not induce cell fusion and then compare the rate of wound closure. However, another type of wound would probably inflict different amounts of cell damage and so would be hard to compare. Overall, we think the half-and-half system of comparing responses on the two sides of the wound is the best, most controlled comparison.

      (23) Figure 4F-G. It was mentioned that there is less syncytia formation in Atg KD cells, however the difference in cell area between control, WT and Atg KD is not obvious. The authors may want to mark the dots that represent syncytia to distinguish them from mononucleated cells.

      The point we are trying to make (now Fig. 4G-H) is that cell area is related to distance moved, regardless of how cell area is determined. We do not have the ability to count nuclei in the control sides (nuclei are labeled only on the pnr side), and further, we have reported separately (White et al, 2024) that there is a limited amount of endocycling in these cells, which should also increase area.

      (24) Figure 5G. y axis. The authors may want to change "small cells" to "mononucleate cells".

      We changed it to "unfused cells" which is the term we used in the legend. In these wounds we were unable to visualize nuclei.

      (25) Page 12. Line 243. (Figure 5D,G) instead (Figure 5D).

      We changed it to read (Figure 5D,G).

      (26) Page 12-13. Lines 241-246. The description of "mononuclear cells removed" and "syncytia outcompete unfused cells" may be clearer if explained here as mononuclear cells joining the syncytium by cell-cell fusion.

      Here we are describing a different phenomenon - not that fusion is removing all the smaller cells but rather that the syncytia are faster/better/more effective at wound closure than the smaller cells. This is illustrated in Fig. 5Cii-Ciii.

      (27) Page 13. Line 261. "Thus, there were about five-fold more tangential borders lost to fusion than radial" Is this conclusion also true when analyzing each wound individually?

      This is a reproducible finding, that there is more fusion across tangential borders than across radial borders. The ratio of tangential-border loss: radial-border loss for each wound is as follows:

      wound 1, 63:12

      wound 2, 39: 11

      wound 3, 44:8

      wound 4, 50:8

      (28) Page 31. Figure 6 - figure supplement 1 legend, Line 652. Make "B)" bold.

      Done.

      (29) Page 14. Lines 270-275. If there is an advantage to radial fusion versus cell intercalation for wound closure speed, how do the authors explain that the percentage of radial fusion is lower than the percentage of intercalation? (Figure 6D) How does the wound affect the molecular level (fusogen expression?) of the surrounding cells? 

      We expect that radial fusion specifically reduces the need for intercalation at the leading edge, as shown in Fig. 6C. Both would speed closure, however, as any increase in cell area will allow more efficient redistribution of resources such as actin and will also reduce the total number of junctions needing to be remodeled as the wound closes. Since we don't know the fusogen, we can't say how the wound affects its distribution.

      (30) Page 14. It is not clear where the experimental data ends and the model starts. For example, in line 276, it would be clearer to describe the "tissue fluidity as measured" or is it more precise to write instead "as estimated/calculated". The fusion between observations and model is confusing and maybe this should be unfused.

      This text, referring to the analysis in Fig. 7A, B, is not a computational model but rather a quantitative analysis of tissue fluidity as measured by a pre-existing metric, the shape index. This is experimental data. The computational model begins in the next paragraph, accompanying Fig. 7C, D. We edited the language slightly in this paragraph to clarify.

      (31) Figure 8 versus Figure 2C. Actin-bound GFP versus cytoplasmic GFP? Both mentioned as Actin GFP. Make it clear.

      They are indeed the same thing, actin protein fused to GFP, as described in the text and legend, and they are labeled identically.

      (32) Figure 8Biii, Div. Nice presentation of signal distribution between the cells.

      Thank you!

      (33) Page 30. Figure 8G legend. Lines 623-628. Is the shown mean profile plot based on specific images shown in Fi and Fv or just the cells represented there? Since Fi is a single z slice and Fv is maximum intensity projection which are not comparable.

      In response to the reviewer's question, we reanalyzed the image. Fig. 8G compares Z-projections.

      (34) Page 15. Line 303-304. "Tangential border fusions allow resources from distant cells to be mobilized to the wound edge." Does not this leading-edge actin localization happen in radial fusions close to the region of the wound?

      Fusions along radial borders, as shown in the top panel of Fig. 6A, would not offer the opportunity to move actin from distant cells to cells nearer to the wound.

      (35) Page 15. Did the authors test any predictions from the simulations of the model experimentally?

      This isn’t so much a predictive model as an exploratory model that addresses one question: is it plausible that the presence of syncytia can speed closure by reducing the need for intercalations, even if the syncytia have no other special properties. The only prediction would be that inhibiting fusion would slow wound closure.

      (36) Page 17. It would be interesting to discuss the following questions: (A) Is autophagy required for fusion. (B) Is Atg1 required for epithelial cell fusion? (C) Is autophagy required for wound repair? Are any of the combinations correct (A&B, A&C, B&C, A&B&C)

      The role of autophagy in wound-induced cell fusion was thoroughly explored in the 2022 EMBO J paper from Maria Leptin's lab, "Autophagy-mediated plasma membrane removal promotes the formation of epithelial syncytia" by Kakanj et al. We merely knockdown a gene they discovered to be important for wound-induced epithelial fusion, Atg1, as one means of investigating how syncytia contribute to wound closure. Our results don't add to their findings, and the role of autophagy is not what we want to focus on in our Discussion.

      (37) Page 19. Line 381-384. "If N represents the number of cells that fused, our results suggests that syncytia can apply up to N times more actin to the leading edge; considering that we observed syncytia with dozens of nuclei, this could represent a significant enhancement of actin at the leading edge. Increased actin might explain the ability of syncytia to outcompete diploid cells at the leading edge." To enhance this suggestion, the authors may want to compare actin signal in the leading edge of different size syncytia.

      We thought a lot about this experiment because reviewer 2 asked for it in the previous round of review, but as we said then, we can imagine too many caveats to the interpretation to make it worthwhile.

      (38) Page 22. Line 447-448. "...fusion would act the fastest after wounding because there is no need for DNA replication." There may be a potential need for protein (fusogen) synthesis.

      The timing of fusion, which we report here begins within 10 minutes after wounding, suggests that if there is a fusogen, it is already present in the cells before wounding.

      (39) Page 34. Line 716. Add "C" to "29{degree sign}".

      Done.

      (40) Page 38. "Wound closure analysis" part. Can the wound closure be visualized using brightfield?

      The scar also impedes imaging through bright-field microscopy.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This valuable study addresses the effects of selection on aggression on fitness and life-history trade-offs in Drosophila melanogaster. However, the evidence presented is incomplete and does not support the claims proposed in the study of increased survival of highly aggressive males at the expense of reproductive success and shorter mating duration. The main limitation of the study is the choice to use males from only one aggressive Drosophila line in combination with Canton-S females, that do not allow disambiguation between nonaggression-related factors, such as hybrid vigor and aggression-related factors influencing mating and lifespan.

      We would like to clarify the points raised in the eLife assessment.

      The report states that we relied on a single line of hyper-aggressive males tested with Canton-S females, and implies that Bully and Cs have not co-evolved. This is a misunderstanding: Bully flies were derived from Cs population. Thus, Bully and Cs have co-evolved. In addition to the Bully A line presented in the main figures of the manuscript, we replicated several of our findings with a second independent selected line, Bully B. Results from courtship assays involving both Bully A and Bully B couples males and females were presented in Figure Supp1. We apologies for not having made this more explicit in the original manuscript, which we will correct. These experiments should alleviate the concerns from the reviewers; they demonstrate that our conclusions are supported by two independent hyper-aggressive lines, and these include assays with selected male and female flies.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      We thank the reviewer for recognizing these strengths.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated to non-coevolved Cs females with Cs males mated to coevolved Cs females.

      We thank the reviewer for this comment, which made us realize that we had not sufficiently highlighted some of our experiments. The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs. As originally described by Penn et al. (2010), highly aggressive “Bully” lines were generated through selective breeding from Canton-S males that consistently won aggressive encounters. After 34–37 generations, stable Bully lines were established. Thus, 1) Bully and Cs flies have co-evolved and 2) the selection applied was male-specific. Independent selection replicates produced distinct lines, including Bully A and Bully B. Previous studies only characterized Bully A (Penn et al., 2010; Chowdhury et al., 2017), but our work includes both Bully A and Bully B (Fig. S1).

      The rationale for pairing Bully or Cs males with Cs females (with which both male types co-evolved) follows the approach used by Dierick et al. (2006), who investigated how the male-specific selection for aggression affected courtship and mating behaviors by testing them with standard Canton-S females. This design allows to isolate the effects of male genotype and behavior on courtship and mating outcomes, avoiding confounding effects from female behavioral changes.

      We initially compared selected Bully pairs (Bully males × Bully females) (Fig. S1) with Cs pairs and observed similarly shortened mating durations in both Bully × Bully and Bully × Cs matings (Fig. S1, Fig. 1F and G). Thus, the reduction in mating duration arises specifically from Bully males. We therefore chose to use Cs females as a standard background to assess the consequences of male-specific selection for aggression on reproductive behaviors.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males; the longer lifespan and performance on the former is interpreted as a trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      We appreciate this comment, which again stems from a poor explanation from our part about the origin of the Bully line in the original manuscript. The Bully flies were derived from the same original population as the Cs line. Hybrid vigor typically arises when crossing individuals from distinct populations, which is not the case here as both Bully and CS come from the same population.

      To further support our conclusions, we conducted additional experiments using progeny from within-line crosses (Bully males × Bully females) and results revealed the same phenotype: the progeny of these flies also exhibited significantly longer lifespans than Cs males x Cs females progeny. This finding argues against hybrid vigor as the main explanation for the observed phenotype, since both the Bully and Cs crosses result in inbreeding, yet give longer lifespan in Bully. We will include these additional longevity data (currently not included in the manuscript) to strengthen our results and reinforce our interpretation.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause differences in measured behavioral outcomes.

      We thank the reviewer for raising this important point regarding causality. One way to establish a causal link between differences in CHCs observed in Bully and Cs flies and the corresponding behavioral outcomes would be to experimentally manipulate CHC profiles. For instance, one could perfume oenocyte-less males with the compounds found in higher abundance in Bully flies, then perform behavioral assays to assess causality. We agree that such experiments would be highly informative in determining the functional roles of specific CHCs elevated in Bully males. However, this approach is technically challenging, as the perfuming technique must be optimized to transfer precise amounts of each compound. For example, this method can be used to gradually perfume flies to assess dose–response behavioral effects, whereas matching exactly the natural concentrations found in individuals, especially given inter-individual variability, remains difficult.

      We considered conducting such experiments during our study but did not pursue them for these technical reasons. Nevertheless, we can include a statement in the Discussion acknowledging this as an important future direction to test the causal relationship between CHC variation and behavior.

      Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbons (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding.

      Paradoxically, Bully males live longer and remain fertile at older ages when CS males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations.

      Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits, thus offering new entry points for understanding life-history evolution.

      We thank the reviewer for this positive summary of our work.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in female reproductive tracts after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits. The study will be of interest to Drosophila biologists working on sexual selection, behavioral evolution, and aging.

      We thank the reviewer for recognizing the integrative design and mechanistic contributions of our study.

      Weaknesses:

      The weaknesses are primarily conceptual rather than procedural. The generality of the findings is uncertain, as selection appears to be represented by only one (and a second closely related) Bully line, limiting conclusions about selection responses versus line-specific drift or founder effects. The causal link between aggression selection and increased longevity is not established: the data show a correlated shift but do not identify mechanisms underlying lifespan extension. In several places, the manuscript uses causal language (e.g., that selection 'influences' longevity or mating strategy) where association would be more accurate; this should be toned down to avoid overstatement. Ecological relevance is also not addressed, since laboratory conditions may bias the balance between costs and benefits of aggression compared with variable natural environments. Addressing these points would strengthen both the impact and clarity of the study.

      (1) Generality of findings and potential line effects

      We agree that our results presented in the main figures of the manuscript relied mainly on one Bully line (Bully A). To address potential line-specific effects, we replicated key courtship experiments with another independent line, Bully B, selected in parallel from the same Canton-S stock but through distinct selection replicates. The results obtained from Bully B closely matched those from Bully A, suggesting that the observed phenotypes are consistent consequences of aggression selection rather than random drift or founder effects.

      (2) Causality versus correlation

      We concur that some sentences in the manuscript could overstate causal interpretations. We will revise the text to clearly distinguish correlation from causation and to avoid implying direct causal relationships where data only support association.

      (3) Ecological relevance

      We appreciate this point. Our experiments were performed under controlled laboratory conditions, which may not fully capture the ecological contexts shaping the costs and benefits of aggression. We will acknowledge this limitation and expand the Discussion to consider how environmental variability could modulate the fitness trade-offs associated with aggression in natural populations.

      We thank both reviewers for their constructive feedback, which will help us strengthen the rigor and clarity of the manuscript. We believe that the additional results and revisions will satisfactorily address their concerns.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The major weaknesses raised by the reviewers, namely the flies used in the study (CsxBully compared to CsxCs) where the effect on lifespan could be explained by hybrid vigor and the use of only one Bully and one Cs line that does not allow to link unambiguously the observed effect to the selection for aggression, should be addressed by a different experimental design and additional lines to exclude the effect of non-aggression related factors.

      We thank the Reviewing Editor for these comments.

      (i) Experimental design and hybrid vigor:

      Hybrid vigor typically arises from crosses between genetically divergent populations. In our study, Bully lines were derived from Canton-S background and are not therefore not genetically distant from controls. To directly address this concern, we included new data from Bully × Bully pairs (Figure 1), using independently selected Bully lines. These experiments reproduce the key aggression and courtship phenotypes observed in Cs × Bully assays, indicating that the effects are not attributable to hybrid vigor.

      (ii) Use of additional selected lines:

      We now include data from two independently selected lines (Bully A and Bully B), both derived from Cs, which show consistent behavioral phenotypes. This supports the conclusion that the observed effects are associated with selection for aggression rather than line-specific artifacts. We note that generating such lines is time- and labor-intensive, and only a few laboratories have established aggression-selected lines in Drosophila melanogaster (e.g., Penn et al., 2010; Dierick et al., 2006; Edwards et al., 2006). Accordingly, we have revised the manuscript to explicitly acknowledge this limitation and to frame our conclusions in terms of association rather than causation.

      Reviewer #1 (Recommendations for the authors):

      I can't see any way to interpret the data using CsxCs vs CsxBully comparisons.

      We thank the reviewer for this important point. This concern appears to arise from the assumption that Bully and Cs represent genetically distinct or non-coevolved populations. However, Bully lines were directly derived from Cs and therefore share a common genetic background. We have clarified this point in the Introduction (lines 104-106) and Results (lines 132-137).

      Importantly, we now include additional data showing that key phenotypes, including reduced mating duration, are also observed in Bully × Bully pairings and across independently selected Bully lines (new Figure 1). These results demonstrate that the observed effects are driven by the male genotype and do not depend on the female background.

      Because selection for aggression was applied specifically to males, we used Cs females as a standardized background to isolate male-specific effects while minimizing variability arising from female genotype or behavior. This rationale is now explicitly stated in the Results (lines 163-166) and at the beginning of the Discussion (lines 326-330). This experimental design allows interpretation of male-specific effects, and the observed differences cannot be attributed to cross design artifacts or hybrid vigor.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) Several passages currently imply causality, whereas the data support correlations between selection and trait differences rather than direct causation. This overstatement also appears in section subheadings within the Results, such as "Hyper-aggressive males display reduced mate-guarding efficiency, without compromising female fertility." Please consider toning down the wording by replacing active causal verbs with more neutral phrasing. Additionally, it would be important to include a clear, explicit sentence in the Discussion acknowledging this caveat, as the existing phrase "is associated with changes in reproductive traits" does not fully convey this nuance.

      We thank the reviewer for this important comment. We have revised the manuscript throughout, including Results subheadings, to replace causal language with association-based phrasing. We also rephrase the first sentence of the Discussion to clarify that our conclusions are correlational (see line 320).

      (2) Figure 1 would benefit from a simple schematic of the behavioral paradigm and the arena, since the authors' arena design minimizes manual handling; a cartoon would help readers quickly grasp the assay flow and the conditions under which interactions occur.

      We thank the reviewer for this helpful suggestion. We have added a schematic to Figure 1 illustrating the behavioral paradigm and arena design. Additional details are provided in the Materials and Methods (Trannoy et al., 2015). This improves clarity and accessibility of the experimental design.

      (3) In Figures 2A-B and A'-B', the higher post-mating UWE in Bully males is intriguing, but these panels do not actually measure the refractory period. It would be helpful to include 'latency' in the first UWE after mating in these swapped-female conditions. This could also be repeated with pheromone-standardized (cVA/CHC-equalized) decapitated females to disentangle effects of female pheromone load from male sensory perception. In addition, a baseline courtship control (naive males with decapitated virgins) is necessary to test whether Bully males simply have a lower threshold for initiating courtship.

      We thank the reviewer for this suggestion. The referenced panels are now shown in Figure 3. We quantified post-mating courtship latency; however, latencies were very short across conditions, and no differences were observed between genotypes. We therefore revised the text to interpret these results in terms of post-mating courtship motivation rather than refractory period. The baseline courtship control with decapitated virgins is provided in Fig 2G. These changes clarify the interpretation of post-mating behavior and address the reviewer’s concerns.

      (4) Related to my above point, the results in Figure 2B-B′ raise the possibility that Bully males have reduced perception or neural sensitivity to anti-aphrodisiac pheromones deposited by CS males, which could account for their elevated post-mating courtship; the authors might consider experiments that directly test male sensory responsiveness to these cues or mention this possibility in the Discussion.

      We thank the reviewer for this point. We performed additional assays to test males’ sensory responsiveness using binary choice assays and measured the time spent performing UWE towards decapitated females versus males. These results were added in Figure 3-Figure Supp 1, and indicate that both Cs and Bully males displayed courtship preferentially towards females, providing a control for sensory perception.

      (5) In multiple figure panels, virgin and mated females are depicted with the same symbols, which makes interpretation confusing.

      Thank you for pointing this out. We have updated the figure panels to use distinct symbols for virgin and mated females to improve clarity.

      (6) For Figure 4, it would be helpful to provide standalone KM curves for Bully versus CS males, including a separate panel for isolated (never-mated) males, and present mating counts in a separate panel while reporting survival models that incorporate mating frequency (or use it as a time-dependent covariate). Although Figures 4C-D report median survivals, full KM plots and an isolated-male curve are important since mating itself elevates mortality and can otherwise confound intrinsic lifespan differences.

      Thank you for this important point to improve clarity of this figure. We have reorganized this figure (now Figure 5) to now, present survival curves first (isolated and group-housed males), followed by lifetime mating counts. This reorganization separates survival from mating activity and addresses the potential confounding effect of mating on lifespan. For the lifetime mating data, we used bar plots rather than curves to better visualize individual mating events.

      (7) The following sentences overstate the results and imply causality; consider toning down: "These findings suggest that 5-P and 5-T might contribute to promoting remating in females that have previously mated with Bully males (Figure 2F' and G'). Given that Bully males also showed higher levels of both 5-P and 5-T compared to naïve Cs males (Figure 3B), it is likely that the elevated levels of these compounds observed in females result from their transfer during mating."

      We thank the reviewer for this important point. We have rephrased these sentences to remove causal language and instead describe associations between CHC profiles and behavioral outcomes. In particular, statements implying that 5-P and 5-T promote remating or are directly transferred during mating have been revised to reflect correlational evidence only (see lines 253-256).

      (8) It is not entirely clear how aggression was quantified in each generation, what proportion of males were selected to breed, and whether the findings generalize beyond a single Bully line (Figure Supplement 1 shows data from a closely related Bully line). Without independent replicate lines or sham-selected controls, it remains difficult to rule out drift or line-specific artifacts, and this limitation should be explicitly acknowledged.

      We thank the reviewer for this important point. We have clarified the aggression selection procedure by adding methodological details from Penn et al., including how aggression was quantified and how breeders were selected (lines 104-106 and 131-137). Briefly, independent selection replicates were initiated from the same Canton-S population, generating three lines (Bully A, B, and C), which were maintained separately.

      To address generality, we now include data from multiple lines. In particular, a new Figure 1 presents aggression and courtship phenotypes across Bully A, B, and C, and key behavioral results are consistent across independent lines.

      We acknowledge that additional independent lines would further strengthen generality; this limitation is now explicitly stated in the Discussion (lines 325-326).

      These additions clarify the selection procedure and support that the observed phenotypes are associated with aggression selection rather than line-specific artifacts.

      Minor Comments:

      (1) Exact sample sizes for every experiment should be included in the main figure legends.

      Thank you. We have added the number of replicates in each figure legends.

      (2) In Figure 1-Supplement 1, the orientation for depicting mating success is reversed compared to Figure 1, which is a bit jarring; it would be clearer to keep the orientation consistent with the main figure.

      Thank you. We have incorporated the results initially presented in Figure 1-Sup 1 into a new Figure 1 with additional results, and have taken into account reviewers’ comment.

      (3) For multivariate analyses, I suggest including important details such as group sample sizes, p-value, and the percent variance, etc., in the figure legend rather than keeping this only in Supplementary Table S1.

      We have inserted these details directly into the figure legends for clarity.

      (4) Why was the food cup used for arenas where decapitated virgins were used in mating assays?

      Thank you for pointing this. We now have clarified the experimental procedure in the M&M of the revised manuscript (lines 465-469).

      (5) For cartoons in Figure 2, the current yellow background makes it very difficult to distinguish flies drawn in yellow or green. Please adjust to a higher-contrast background or add darker outlines so that the cartoons are clearly legible.

      Thank you. We have increased the contrast of the female bodies to ensure the cartoons are clearly distinguishable (now figure 3).

      (6) Addition of line numbers in the manuscript would be helpful during the review process.

      Line numbers have been added throughout the manuscript.

      (7) I noticed a few typos in the manuscript. For example, in the Introduction, "seminal fuids" should be corrected to "seminal fluid." In the Discussion, the phrase "CHCs profiles compared those" requires a "to" before "those." Please carefully review the manuscript for similar errors.

      Thank you for pointing this out. We carefully reviewed the manuscript for typos and corrected all identified errors.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the editors and reviewers for your careful evaluation of our manuscript and for the constructive recommendations that have helped us improve the rigor, clarity, and balance of the study. We are pleased that the reviewers recognized the potential value of linking red light exposure to SIRT4 downregulation, fatty acid metabolism, H3K9 acetylation, and attenuation of ageing-related phenotypes. We have revised the manuscript extensively in response to the reviewers’ comments.

      In particular, we have clarified the wavelength specificity of the red-light response, reanalyzed and more cautiously interpreted the omics data, improved the presentation and quantification of semi-quantitative experiments, revised statistical reporting, corrected gene/pathway annotations, toned down mechanistic claims where direct evidence was insufficient, and expanded the Discussion to integrate recent evidence on red-light-induced fatty acid oxidation and AMPK/ACC signaling. We also added a dedicated limitations paragraph addressing the use of female mice, the absence of a complete in vivo wavelength-control and source-blocked sham cohort, and the need for future direct metabolic flux and isolated mitochondria studies.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, is sometimes approached in unconvincing ways, and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      We would like to thank the reviewer for the careful and patient examination of the issues identified in our manuscript. The poor quality of some of the Western blot bands in Figure 4 may have been caused by inappropriate electrophoresis conditions during the Western blot experiments. In the revised manuscript, we will optimize the electrophoresis conditions to obtain higher-quality protein bands and update the quantitative data. Regarding the quantification format, we believe that heatmaps provide a more intuitive representation of trends in protein expression across different treatment groups. This approach more accurately reflects the results of our biological replicates than simply analyzing the significance of differences in the grayscale values of protein bands. For the analysis of transcriptomic data, we will conduct a more detailed analysis of signal pathway enrichment and the identified differentially expressed genes to ensure that predicted genes are excluded from our current results and redundant data presentation is removed.

      Regarding additional experimental controls, such as incorporating experimental data under blue light treatment conditions as a control for red light. While exploring the optimal red light irradiation dose at the cellular level, we simultaneously conducted experiments on the effects of blue light irradiation at the same dose on keratinocyte activity. The results indicated that as the blue light irradiation dose increased (0–160 J/cm<sup>2</sup>), the keratinocyte activity exhibited a dose-dependent decline. This indicates that blue light is phototoxic to keratinocytes. The relevant experimental results have already been published in our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). Taken together with the data from our study, this demonstrates that the anti-ageing effects of red light reported in the current manuscript are indeed driven by red light.

      Reviewer #2 (Public review):

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to the current study.

      Herrera et al. demonstrate that red light increases basal, ATP-linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript, i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focuses on the SIRT4-MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. The authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long-lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilised cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation) aligns perfectly with Herrera´s critique.

      Long-term metabolic effects-Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12-24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al.'s results would not only acknowledge independent, corroborating evidence but would also allow the authors to position their SIRT4-centric mechanism within a broader, emerging understanding of red-light photobiomodulation.

      We would like to thank the reviewer for providing us with constructive suggestions for discussion. Our results showed that under red light conditions, both glycolipid and lipid metabolism were activated in keratinocytes, and cellular metabolic flux increased. The activation of lipid metabolism directly led to an increase in metabolism-associated H3K9ac and drove the upregulation of anti-ageing-related genes; we believe this is key to the anti-ageing effects of red light. Mechanistic analysis combining proteomics and acetylation proteomics revealed that red light significantly downregulated SIRT4 expression and increased the acetylation of MCD, a protein regulated by SIRT4 that governs cellular fatty acid oxidation rates. Through validation using cell-level knockdown and inhibitors, we confirmed that SIRT4 inhibition exerts anti-ageing effects in vitro and that inhibiting MCD function under red light conditions suppresses H3K9ac. These results establish the role of the SIRT4-MCD signalling axis in mediating the anti-ageing effects of red light.

      The study by Herrera et al. included a substantial body of validation data confirming the role of red light in promoting fatty acid oxidation, providing robust empirical support for our research. Furthermore, Herrera et al. revealed that red light-induced fatty acid oxidation depends on AMPK and ACC phosphorylation. This mechanism of red-light photobiomodulation may refute the notion that its bio-regulatory effects rely solely on the action of mitochondrial cytochrome c oxidase. Furthermore, together with our study revealing that red light exerts anti-ageing photobiomodulatory effects via the SIRT4-MCD signalling axis, these findings independently confirm that red light regulates cellular fatty acid oxidation, thereby demonstrating the pivotal role of activated fatty acid oxidation in the bio-regulatory effects of red light. In the revised manuscript, we will include a discussion on the potential link between the red light-driven downregulation of SIRT4 and the phosphorylation of AMPK/ACC. This will be of positive value in elucidating how SIRT4 exerts its anti-ageing effects by regulating lipid metabolism, as well as in explaining the possible mechanisms by which red light downregulates SIRT4.

      Recommendations for the authors:

      Summary of Major Revisions

      Changes made in the revised manuscript:

      (1) Added a clearer explanation of why the 625-635 nm red-light regimen was considered the active intervention and how the available blue-light data from our previous work support wavelength-dependent effects on keratinocytes.

      (2) Revised the language describing inflammatory regulation. We now avoid presenting red light as producing a uniform anti-inflammatory effect and instead describe selective remodeling of ageing-associated inflammatory and SASP signatures.

      (3) Improved figure presentation and quantification for immunofluorescence, metabolite, and western blot assays; clarified image-analysis regions, replicate numbers, and normalization procedures.

      (4) Reanalyzed transcriptomic, proteomic, and acetyl-proteomic datasets with appropriate multiple-testing correction and corrected erroneous pathway/gene annotations in metabolic gene panels.

      (5) Replaced overly strong causal wording with more conservative language, especially regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4 localization, PPARα immunofluorescence, and direct fatty acid oxidation flux.

      (6) Expanded the Discussion to incorporate Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), highlighting convergence between the SIRT4-MCD model and AMPK/ACC-dependent fatty acid oxidation.

      (7) Corrected typographical, nomenclature, and figure-legend inconsistencies throughout the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) Wavelength specificity and need for a non-red-light control

      As a reader, one is left wondering whether the effects are due to red light specifically. An important control would have been to irradiate mice and cells with another light wavelength, such as blue light.

      We agree that wavelength specificity is a critical issue for interpreting photobiomodulation studies. In the revised manuscript, we have clarified that the anti-ageing and metabolic effects described here apply specifically to our 625-635 nm red-light regimen, rather than to visible light in general. We have also added a discussion of our previously published blue-light experiments, in which keratinocyte viability decreased in a dose-dependent manner across the same 0-160 J/cm<sup>2</sup> dose range (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). These data indicate that blue light and red light produce distinct biological outcomes in keratinocytes. Because high-dose blue light was cytotoxic under comparable cellular conditions and because the present study was designed to investigate the long-term effects of red light in aged mice, we did not perform prolonged in vivo blue-light irradiation as an ageing intervention.

      Changes made in the revised manuscript:

      Clarified in the revised Introduction and Discussion that the conclusions are specific to 625-635 nm red light under the irradiation parameters used in this study. (Lines 92 to 94, Lines 1146-1149)

      Added text summarizing the published blue-light comparison data from our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1), including the dose-dependent decline in keratinocyte activity after blue-light irradiation. (Lines 90 to 92, Lines 1146-1149)

      Added data on the wavelength range of the red light used in this study. (Lines 529 to 531, Fig S1a)

      (2) Complexity of inflammatory effects

      The manuscript repeatedly emphasizes anti-inflammatory effects, yet some cytokines such as IL-18, Ccl2, TNF-α, Ccl2, and IL-8 appear increased. This suggests that the effects may be more complex than presented and may require additional readouts or stronger statistical power.

      We unanimously agree that the inflammatory response to red light should not be described as a simple, uniform suppression of all cytokines. We have demonstrated that changes in the levels of the senescence-associated secretory phenotype (SASP) at the cellular level and in skin tissue following red light treatment not only indicate that red light-induced metabolic activation can reduce the age-related inflammatory baseline, but also reveal red light-driven short-term reparative effects or stress-related cytokine responses. We consider this to be consistent with the findings, and the downregulation of NF-κB-related signalling observed in skin tissue following periodic red light irradiation of aged mice further supports the conclusion that red light alleviates the age-related inflammatory baseline. In fact, in our previous study, we did observe that red light treatment promoted increased levels of the cytokine Ccl2, which plays an important positive role in rapid wound healing (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). In the revised manuscript, we have reworded the relevant Results and Discussion sections to indicate that red light remodels ageing-associated inflammatory signalling rather than globally reducing every inflammatory mediator. We have also toned down statements suggesting that red light ‘reverses’ or ‘suppresses’ inflammation where the underlying data support a more selective effect.

      Changes made in the revised manuscript:

      Replaced broad terms such as “anti-inflammatory effects” with more precise wording such as “remodeling of ageing-associated inflammatory signaling” where appropriate. (Lines 541 to 543, Lines 573 to 574, Lines 893 to 895)

      Expanded the Discussion to explain that red-light-induced metabolic activation may simultaneously reduce senescence-associated inflammatory tone while allowing transient reparative or stress-related cytokine responses. (Lines 1135 to 1142)

      (3) Figure clarity, semi-quantitative methods, western blot quality, and inconsistent band patterns

      Many differences are difficult to see or require orthogonal validation. Some tissue-specific signals and western blots are difficult to judge. Several western blots are of poor quality, and multiple markers show inconsistent band profiles across experiments, including SIRT4 in Figure 5.

      We thank the reviewer for highlighting these technical and presentation issues. We have reviewed the semi-quantitative data and revised the presentation of the figures to improve their interpretability. For Western blot experiments, we optimised the electrophoresis and transfer conditions, replaced low-quality representative images where possible, and updated the semi-quantitative results. Furthermore, regarding the lack of clarity in the SIRT4 protein band, we have conducted repeat experiments and updated the main text to include a clearer image of the band. The issue with the annotation of the protein location was in fact due to an oversight during the data analysis process; we have carried out a detailed review and provided the uncropped full-length Western blot images for all experiments in the Supplementary Materials for the reviewers’ scrutiny. Finally, we would also like to point out that factors such as sample origin, protein extraction, electrophoresis conditions, antibody exposure time, and potential non-specific detection may all contribute to differences in band patterns. At present, the core conclusions regarding SIRT4 are supported by multiple lines of evidence, including mRNA analysis, immunofluorescence, Western blotting of bands at the expected sizes, and SIRT4 knockdown experiments, rather than being based solely on any single semi-quantitative Western blot result. We therefore believe that the conclusions drawn from the data presented in the revised manuscript are equally convincing.

      Changes made in the revised manuscript:

      Replaced or improved low-quality Western blot panels and updated quantitative analyses in revised Figures 3-5 and associated supplementary material. (Fig 3m, Fig 4, Fig 5f)

      Clarified the normalization approach for H3K9ac/H3 and target/loading-control comparisons, and the use of Actin or H3 as appropriate loading controls. (Lines 283 to 293)

      Bands with nonspecific profiles were excluded from quantitative conclusions and the manuscript conclusions no longer depend on those ambiguous signals.

      Revised the Results (Repeat the experiment to update the low-quality Bands) to avoid overstating changes that are not clearly visible or not supported by statistical analysis. (Fig 4n and r)

      (4) Choice of pharmacological agents and need for genetic strategies

      The choice of drugs in Figure 4 is puzzling. More specific and widely used inhibitors could be used to block PI3K/Akt or mTOR, and natural agonists such as insulin or EGF could be used. Genetic strategies should complement these observations.

      We agree that pharmacological perturbation experiments should be interpreted with caution. In this study, our criteria for selecting inhibitors were based on transcriptomic and proteomic analyses; we sought to determine how the most direct inhibition of red light-activated signalling pathways would affect H3K9ac levels. In the revised manuscript, we have clarified the rationale for the compounds used and have reduced the causal weight assigned to these inhibitor/agonist experiments. These data are now presented as supportive evidence that red light is associated with metabolism-related signalling changes, rather than as definitive proof that PI3K/Akt/mTOR is the primary upstream mechanism. We have also emphasised the genetic SIRT4 knockdown experiments as a more direct mechanistic test for the SIRT4-centred part of the model. We acknowledge that additional experiments using more selective inhibitors, physiological agonists such as insulin or EGF, and genetic perturbation of PI3K/Akt/mTOR components would be valuable for future studies.

      Changes made in the revised manuscript:

      Revised the text describing pharmacological experiments to distinguish supportive pathway modulation from direct causal evidence. (Lines 787 to 789)

      Added a limitation and future direction noting that genetic perturbation of PI3K/Akt/mTOR and physiological pathway activation with insulin or EGF would strengthen the model. (Lines 1190 to 1194)

      (5) Serum NADH measurement

      In Figure 1t, the authors measure serum NADH. NADH is poorly detectable in serum or plasma, and changes may reflect blood-cell lysis during collection rather than circulating NADH.

      We appreciate this technical concern. We have revised the manuscript so that serum NADH is no longer used as a central mechanistic readout. We now treat this measurement only as an exploratory indicator of systemic redox-related changes and explicitly acknowledge that serum or plasma NADH is vulnerable to artifacts from blood-cell disruption during sampling. The mechanistic interpretation has been shifted toward cellular and tissue measurements, including intracellular NADH/NADPH/GSH, ATP, acetyl-CoA, fatty acid uptake, and H3K9ac, which are more directly relevant to keratinocyte metabolic remodeling.

      Changes made in the revised manuscript:

      Removed serum NADH from the main causal argument linking red light to metabolic flux and H3K9ac.

      Placed greater emphasis on cell-based metabolite assays, tissue acetyl-CoA, and H3K9ac measurements as the main metabolic-epigenetic evidence. (Lines 582 to 587, Fig 1s)

      (6) Direct assessment of glycolysis and fatty acid oxidation

      The authors propose that red light increases glycolysis and fatty acid oxidation, but this could be assessed directly rather than through surrogate measures.

      We agree. Our current data include multiple metabolic readouts, including glucose and fatty acid uptake, ATP, NADH/NADPH/GSH, triglycerides, fatty acids, pyruvate, lactate, acetyl-CoA, and MCD-dependent changes; however, these assays are not equivalent to direct flux measurements such as Seahorse extracellular flux analysis, isotope tracing, or etomoxir-sensitive respiration. We have therefore revised the wording throughout the manuscript to distinguish metabolic remodeling and fatty-acid-oxidation-related signatures from direct measurements of fatty acid oxidation flux. We also incorporated the recent independent work by Herrera et al., which directly measured oxygen consumption and demonstrated red-light-induced fatty acid oxidation in keratinocytes. This external evidence supports the biological plausibility of our SIRT4-MCD model while making clear which aspects are directly measured in our study and which are inferred.

      Changes made in the revised manuscript:

      Replaced overstrong language such as “red light increases fatty acid oxidation” with “red light promotes PPAR-α-related fatty acid metabolism pathway” where direct flux data were not measured in our experiments. (Lines 882 to 883)

      Expanded the Discussion to integrate direct FAO evidence from Herrera et al. and to place the SIRT4-MCD axis within a broader red-light metabolic framework. (Lines 1169 to 1189)

      (7) Incorrect annotation of metabolic genes in Figure 3e

      Acss2, Aldh3b1 and Aldh3a1 are not glycolytic enzymes, Aldh3a3 does not appear to exist, and several enzymes classified as FAO are fatty acid synthesis enzymes. This questions the interpretation of the data.

      We thank the reviewer for identifying these annotation errors. We have rechecked the gene names and pathway assignments in the transcriptomic analysis and corrected the metabolic gene panels. We have confirmed that Acss2 is an acetyl-CoA synthase involved in the metabolism of acetate to acetyl-CoA. The Aldh family genes, meanwhile, are associated with aldehyde metabolism and detoxification. We have made the corresponding adjustments in the manuscript. The incorrectly listed Aldh3a3 entry has been removed. Furthermore, we have categorised genes involved in fatty acid metabolism as ‘fatty acid metabolism-related genes’, rather than grouping them all under the FAO category. These revisions have significantly improved the accuracy of the metabolic interpretation.

      Changes made in the revised manuscript:

      Reannotated Figure 3e and the corresponding Results text to correct glycolysis, TCA cycle, pentose phosphate pathway, and fatty acid metabolism categories. (Fig 3d)

      Removed the erroneous Acss2 and Aldh family genes. (Fig 3d)

      Revised the metabolic model to avoid using incorrectly grouped genes as evidence for direct fatty acid oxidation. (Lines 707 to 709)

      (8) Transcriptomic analysis and implausible volcano-plot p-values

      The transcriptomic analysis raises concerns. For example, the volcano plot in Figure 3d appears incorrect, with -log<sub>10</sub>(P-value) around 300 despite n=3 biological replicates.

      We thank the reviewer for pointing out this important issue. We have reopened the transcriptomic data and found that the extremely high -log<sub>10</sub>(P-value) in the original volcano plot were caused by the automatic replacement of very small P-values—generated during the differential expression analysis—with zero in the tabular data. To avoid misleading visualisations, we have regenerated the volcano plot using Q-values in place of the original P-values. Differentially expressed genes were defined as those with a Q-value < 0.05 and |log<sub>2</sub> fold change| > 1. Furthermore, for visualisation purposes only, the upper limit for q-values was set to 1 × 10<sup>-50</sup> for values below 1 × 10<sup>-50</sup>. This adjustment does not affect the statistical classification of differentially expressed genes but prevents over-interpretation of extremely small values. The revised volcano plots and legends have been updated accordingly. To avoid any potential misinterpretation arising from these updates, the updated volcano plots are presented in the supplementary materials.

      Changes made in the revised manuscript:

      Reanalyzed transcriptomic data using appropriate multiple-testing correction and revised the volcano plot. (Fig S3a)

      Corrected the y-axis transformation and removed implausible -log<sub>10</sub>(P-value) presentation. (Fig S3a)

      Updated Methods to specify the statistical workflow for transcriptomic differential expression and pathway enrichment. (Supplementary materials Lines 46 to 52)

      Moved the analysis of metabolic pathways based on transcriptomic data to the supplementary material, thereby reducing the reliance of the conclusions on transcriptomic data (Fig S3b).

      (9) Need to tone down mechanistic claims regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4, and PPARα

      The mechanisms proposed must be toned down. PI3K/Akt/mTOR should not be called glycolytic pathways, the link to red light or cytochrome c oxidase is vague, SIRT4 reduction requires mitochondrial counterstaining, and PPARα appears cytosolic after SIRT4 knockdown.

      We agree and have substantially revised the mechanistic language. PI3K/Akt/mTOR is no longer referred to as a ‘glycolytic pathway’; instead, it is described as a metabolism-related signalling axis that may influence glucose uptake, growth and nutrient-responsive metabolism. We have also toned down statements attributing red-light effects directly to cytochrome c oxidase, as our study primarily examines downstream metabolic and epigenetic remodelling rather than direct photoreceptor activation. With regard to SIRT4, we have revised the text to avoid interpreting changes in SIRT4 immunofluorescence alone as evidence of altered mitochondrial abundance or mitochondrial localisation. The conclusion is now based on a combination of SIRT4 mRNA levels, western blot bands of the expected size, immunofluorescence trends, and SIRT4 knockdown phenotypes. With regard to PPARα, we have re-examined the PPARα antibody used for the cellular immunofluorescence experiments. In the original Figure 5p, we mistakenly used a PPARα antibody (PPARα, Abclonal, A25296) that is only suitable for Western blot (WB) experiments; we believe this was the cause of the mislocalisation of the fluorescent signal; Consequently, we conducted new experiments using a PPARα antibody (PPARα, Abclonal, A22887) specifically designed for cellular immunofluorescence. The relevant experimental data have been corrected in the manuscript.

      Changes made in the revised manuscript:

      Replaced “PI3K/Akt/mTOR glycolytic pathway” with “The PI3K-AKT signalling pathway is involved in the regulation of glucose metabolism” throughout the revised manuscript. (Lines 701 to 705, Lines 809 to 810, Lines 813, Lines 1193)

      Reduced mechanistic certainty around cytochrome c oxidase and framed it as a possible upstream photoreceptor rather than an experimentally proven mechanism in this study. (Lines 827 to 830, Lines 842 to 848)

      Repeat the PPARα immunofluorescence staining experiment. (Fig 5p)

      (10) Need for isolated mitochondria experiments and red/blue light comparison of mitochondrial respiration

      If the effect of red light relies on mitochondrial cytochromes, additional proof would be needed, potentially using isolated mitochondria and comparing how red and blue light influence respiration capacity.

      We agree that isolated mitochondria experiments would be an important way to test direct mitochondrial photoreception. Because the present study was designed around cellular and in vivo metabolic-epigenetic remodeling, we did not perform isolated mitochondria irradiation experiments. To address this concern, we have toned down statements implying direct cytochrome activation and revised the Discussion to distinguish between direct mitochondrial photoreceptor models and downstream metabolic reprogramming. We also added a future direction proposing isolated mitochondria or permeabilized-cell experiments comparing red and blue light effects on respiration, ATP-linked OCR, maximal respiration, and FAO-dependent respiration. The revised manuscript now emphasizes that our data support a downstream SIRT4-MCD-H3K9ac mechanism after red-light exposure, while the proximal photophysical event remains to be fully defined.

      Changes made in the revised manuscript:

      Added discussion of the need for isolated mitochondria, permeabilized-cell, and wavelength-comparison respiration experiments. (Lines 1194 to 1197)

      Reviewer #2 (Recommendations for the authors):

      (1) Statistical reporting, post-hoc tests, normality/equal-variance tests, exact p-values, and FDR control

      The manuscript states that one-way ANOVA followed by Tukey or Dunnett tests was used, but it does not consistently specify the post-hoc correction for each figure. Normality and equal-variance tests are not reported, p-values are shown only as asterisks, and FDR control is not mentioned for transcriptomics and proteomics.

      We agree that the statistical reporting needed to be more complete. We have revised the Statistics and reproducibility section and the figure legends to specify the statistical test used for each experiment, the post-hoc correction applied after ANOVA, the number of independent biological replicates, and the definition of error bars. Where multiple comparisons were performed, we now state whether Tukey’s or Dunnett’s correction was used. Regarding P-value presentation, we have retained the use of asterisks in the figures as visual indicators of statistical significance to maintain figure readability. For transcriptomic, proteomic, and acetyl-proteomic analyses, we have revised the Methods section to state that multiple-testing correction was performed using the Benjamini–Hochberg false-discovery-rate procedure. Adjusted P values or Q values were used for differential-expression and pathway-enrichment analyses. These revisions clarify the statistical workflow and strengthen the reproducibility of the study.

      Changes made in the revised manuscript:

      Revised the Statistics and reproducibility section to define statistical tests, post-hoc corrections, assumption checks, and multiple-testing correction. (Lines 509 to 522)

      Updated relevant figure legends to include n values, statistical tests, post-hoc corrections, and definitions of significance symbols. (Lines 515 to 516)

      Added FDR control details for RNA-seq, proteomics, acetyl-proteomics, and pathway-enrichment analyses. (Lines 450 to 455)

      (2) Figure clarity and quantitative analysis of fluorescence, JC-1, metabolite, and western blot data

      Several figures lack clarity or appropriate quantification. Figure 1i-j H3K9ac quantification should be based on whole-image or multiple fields; Figure 2e JC-1 should include red/green ratio quantification; Figure 2k-p metabolite data should include absolute concentrations; Figure 3j needs appropriate loading controls.

      We appreciate these specific suggestions and have revised the figure presentation accordingly. For H3K9ac immunofluorescence in skin sections, we have clarified the anatomical region quantified and performed a more objective quantification using multiple fields/regions per section rather than relying on a visually selected dashed area. The dashed regions in the representative images were made clearer and the quantification criteria were added to the Methods and legend. For JC-1 staining, the bar chart on the right-hand side of the mitochondrial membrane potential fluorescence image in Figure 2e shows the quantitative data for the red/green fluorescence ratio obtained from independent experiments; compared with providing only a representative image, these data offer a more easily interpretable quantitative measure of mitochondrial membrane potential. To avoid any potential misunderstanding, we have corrected the vertical axis. For metabolite assays, we clarified normalization to cell number or protein content and revised the data presentation to include absolute or normalized concentrations where available, rather than relying solely on fold changes with variable y-axis scaling.

      Changes made in the revised manuscript:

      Revised Figure 1i-j quantification using multiple fields/regions per mouse section and improved dashed-region visibility. (Lines 431 to 440, Fig 1i and j)

      Corrected the vertical axis of the quantitative data for the JC-1 red/green fluorescence ratio. (Fig 2e and Fig S2a)

      Updated metabolite panels and/or source data to include absolute or protein-normalized values where available, and standardized y-axis interpretation. (Fig 2k-p and Fig 4s and v, Given the diversity of intracellular fatty acid and triglyceride species, absolute quantification based solely on absorbance measurements would be technically challenging and may not accurately reflect the content of each molecular component. Therefore, we presented the changes in fatty acid and glycerol levels as percentage-normalized relative absorbance values, which allowed consistent comparison among the experimental groups.)

      (3) Experimental design limitations: sex of mice and sham control

      Only female C57BL/6 mice were used, although aging and metabolic responses can be sex-dependent. The thermal-control argument lacks a true sham control in which mice are placed in the same apparatus with the light blocked at the source.

      We agree with these points. We have added a section on limitations stating that all aged mice used in this study were female, and that sex-dependent responses to red light, SIRT4 regulation, metabolism and skin ageing should be investigated in future studies using both male and female cohorts. Furthermore, regarding the design of the non-irradiated control group: although the control mice underwent the same depilation and routine procedures, they did not receive red light irradiation. However, we also acknowledge that establishing a sham-irradiated control group with light shielding would allow for stricter control of factors such as restraint, contact with equipment and procedural stress. However, given that this experiment involved a continuous cyclic photoperiodic treatment lasting two years, we were unable to supplement the study with a control experiment involving only red light shielding. Nevertheless, based on the fact that we observed only minimal changes in the mice’s skin temperature following red light irradiation, we believe that the primary factor driving the alleviation of the skin ageing phenotype in the mice remains red light-induced.

      Changes made in the revised manuscript:

      Added a limitation noting that the study used female C57BL/6 mice only and that sex as a biological variable should be addressed in future studies. (Lines 1198 to 1201)

      (4) Textual errors, nomenclature inconsistencies, and ChIP-qPCR normalization

      Several textual errors and inconsistencies should be corrected, including Pparg1a/Ppargc1a, Ricotr/Rictor, Pi3k/PI3K, Sirt4/SIRT4 protein nomenclature, and the use of RPL30 normalization in ChIP-qPCR without showing that RPL30 is unchanged.

      We thank the reviewer for their careful reading. We have corrected the typographical errors and standardised gene and protein nomenclature throughout the manuscript and figure legends. Specifically, Ppargc1α has been corrected to Ppargc1a, Ricotr to Rictor, and the capitalisation of PI3K has been standardised. We now use Sirt4 for the mouse gene and SIRT4 for the protein, applying the same convention to other genes and proteins. For ChIP-qPCR, we have revised the Methods and Results sections to describe normalisation against input and IgG controls more clearly, and to specify the role of the RPL30 locus as an internal control. We have also included data in the Supplementary Materials showing relative enrichment of H3K9ac in the RPL30 promoter region in PAM212 cells before and after red light irradiation; the results indicate that H3K9ac enrichment at the RPL30 locus remained stable across treatment groups after normalization to input DNA and correction against IgG background. This result indicates that the use of RPL30 as an internal control in ChIP-qPCR experiments is feasible.

      Changes made in the revised manuscript:

      Corrected Ppargc1a, Rictor, PI3K, Sirt4/SIRT4, and related nomenclature throughout the manuscript.

      Supplement the experimental results on the effect of red-light irradiation on the level of H3K9ac enrichment at the RPL30 locus in keratinocytes. (Lines 339-352, Lines 539 to 541, Fig S1d)

      (5) Additional Revision Addressing the Public Review and Herrera et al.

      The reviewer suggested integrating the recent study by Herrera et al. showing that 660 nm red light stimulates mitochondrial fatty acid oxidation in keratinocytes through AMPK-dependent phosphorylation of ACC, without changing electron transport chain complex expression. The reviewer also noted that these findings may complement the SIRT4-MCD axis and challenge a cytochrome-c-oxidase-only model of photobiomodulation.

      We are grateful for this constructive suggestion. We have expanded the Discussion to incorporate Herrera et al. and to place our SIRT4-MCD-centered mechanism within the broader emerging model of red-light-driven metabolic remodeling. Herrera et al. provide direct oxygen-consumption evidence that red light enhances fatty acid oxidation in keratinocytes and that this effect involves AMPK/ACC signaling. This is highly complementary to our data, in which red light decreases SIRT4, increases acetylation of MCD, promotes fatty-acid-metabolism-related signatures, elevates acetyl-CoA, and increases H3K9ac. In the revised Discussion, we propose two nonexclusive models: red light-induced SIRT4 downregulation may converge with AMPK/ACC-dependent relief of fatty acid oxidation, or the two pathways may represent parallel reinforcing mechanisms that together enhance lipid metabolic flux.

      Changes made in the revised manuscript:

      Added a paragraph discussing Herrera et al. in the revised Discussion. (Lines A1169 to 1189, Lines 1206 and 1208)

      Revised the conceptual model of red-light photobiomodulation to emphasize downstream metabolic reprogramming rather than direct cytochrome c oxidase activation alone. (Lines 827 to 830, Lines 843 to 844)

      Added future directions to test whether red-light-induced SIRT4 downregulation causally affects AMPK/ACC phosphorylation and FAO-dependent respiration. (Lines 1169 to 1189)

      We again thank the editors and reviewers for their thoughtful and constructive comments. The revised manuscript now provides a more rigorous and balanced presentation of the evidence, distinguishes direct measurements from inferred metabolic flux, corrects pathway annotations, improves figure quantification and statistical transparency, and places the SIRT4-MCD-H3K9ac mechanism within a broader framework of red-light-induced fatty acid metabolic remodeling. We believe these revisions substantially strengthen the manuscript and clarify both the significance and the limitations of our findings.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to.

      This is indeed a caveat of our study and we discuss this issue now in lines 557-566. We refrained from testing necessity to specific US on purpose as we knew that another research group was focussing on this question in parallel (Weber et al., 2023, also published in eLife) and therefore rather focussed on complementary experiments. That study also includes a rather deep discussion about the inputs to individual DANs. We briefly refer to this discussion (lines 442-444) but decided to not go into detail to avoid too much overlap.

      Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      We thank the reviewer for raising these points and included a brief discussion of both scenarios in lines 469-474.

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning?

      We do not know of such a stimulus. Based on our data, a US that only activates a single DAN should only make safety memory – however, the available data of which US activates which DAN is very limited in larvae and currently no such US is known. We discuss this point briefly in lines 557-563 and 572-574.

      The authors provide a detailed behavioral analysis of locomotion behavior; however, the detailed analysis seems unnecessary for that dataset. Modulation of speed and bending rate has been described before with simpler methods (specifically for MBONs). The revealed locomotion phenotypes probably affect larval locomotion during memory recall with light activation, thus the authors should show that larvae are potentially able to move during light-on memory tests.

      We expanded our locomotion analysis of the innate and learned odor preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal and the existing locomotion phenotypes are not correlated to their olfactory choices.

      We do not agree that the locomotion analysis is unnecessary. Modulations of speed and bending have been described for MBONs but to our knowledge not for DANs. It is not trivial at all that DANs and MBONs cause the same behavioral modulations (see, for example, this adult study: Mohammad et al., 2024 Plos Biol). There is extremely limited knowledge about the motoric effects of dopaminergic neurons in larvae - we therefore find it important to describe our results in detail. We added some further rationale of why we think it is crucial to explore the functions of DANs for learning and movements together (lines 81-86).

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      We had decided to put the controls into the supplement in some cases to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (previously Fig. 7) but decided to keep other figures unchanged as we feel that the current design best fits the purpose of each figure. We provide a figure-for-figure rationale in our response to the recommendations for the authors.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      We thank the reviewer for the suggestion. Although we agree that such a discussion would be interesting, we hesitate to expand on this topic, as the discussion is already quite long and our study does not contribute any new data to clarify the molecular dopaminergic mechanism.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      We understand that this choice was not clearly explained and expanded on our rationale in lines 244-251.

      Reviewer #3 (Public review):

      Weaknesses:

      The larvae exhibit directed locomotory action to express punishment or safety memory. If the larvae did not move, we would not be able to assess memory function. Hence, functional activation of DANs could result in one action, which seems like two different functions of memory expression and locomotion. It can also be argued that activation of DANs represents a teaching signal to the KCs, and then eventually, downstream of the MBONs, it results in locomotion modulation. Hence, the seeming functional diversity could be a function of different downstream neuronal pathways and not molecular context-dependent diversity inside dopaminergic neurons. The authors should address this possibility or point out the fallacy in the above argument.

      We thank the reviewer for raising this issue. To the first point, we expanded our locomotion analysis of the innate and learned odour preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal. In addition, the existing locomotion phenotypes in these experiments were not correlated with the animals’ olfactory choice. This makes it unlikely that the changed locomotion directly determines our observation during the olfactory experiments.

      We do agree that it is possible that both the preference after learning and the changed locomotion could come through the same dopaminergic mechanism via diverse downstream pathways. We cover this hypothesis in Fig. 10G and address this question briefly in lines 580-582.

      The finding that activation of TH-GAL4 conveys aversive valence and R58E02-GAL4 conveys appetitive valence seems redundant (Figure 6). I understand they say this in the context of locomotion. However, they may not have mentioned similar findings in adults. In adults, artificial activation of DANs covered by the same GAL4 lines acts as aversive and appetitive teaching signals for memory formation. These references should be cited appropriately in the results and discussion if not currently included.

      We thank the reviewer for this comment and tried to include the relevant adult literature (see e.g. lines 351-356 and 583-604). In particular, we added a quite detailed discussion about a paper published after our initial submission that performed similar experiments for the adult PAM-DANs Lozada-Perdomo et al., 2025, iScience).

      We do not agree, however, that the experiments in Fig. 7 (previously Fig. 6) are redundant. Recent studies in adults found no correlation between the rewarding/punishing effects and the innate valence a given dopaminergic neuron induces (Rohrsen et al., 2021, bioRxiv; Mohammad et al., 2024, PLOS Biol; Lozada-Perdomo et al., 2025, iScience). To our knowledge, no such studies have been carried out in larvae so far. Therefore, we think that it is not only important to test it but that the respective results compared to the results in adults are of relevance for the readership.

      The evidence for the role of dopamine (Figure 7) can be bolstered by using other available RNAi lines against TH. A valium20 vector-based shRNA line is recommended. The current evidence is based mainly on non-specific pharmacological intervention with 3IY.

      We agree to this caveat and made it transparent now in lines 387-389 and 401-403. We nevertheless chose, for the time being, to not include further experiments to address this point in the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Activation of specific or multiple DANs seems to increase naïve odor preference (f1 or TH). Is this due to locomotion defects in TH - how does the odor preference develop over time? Can this increased odor preference explain the safety learning - where they also approach the odor stimulus?

      The reviewer is right that in presence of light, we see increased odor preference both innate and after unpaired training – theoretically, that could be the same effect. However, this would not explain why we see the same increased preferences after unpaired learning with all driver strains but increased innate preferences only for some of them. Moreover, when activating TH, we also see increased odor preference in absence of light after unpaired training but not innately. We therefore think that these are independent effects.

      We also include a new Fig. 6 providing additional information, including the development of preference over time, and addressing the question whether the modulations in locomotion can explain differences in odor preference.

      Locomotion behavior was assessed in 30s light-on periods - was the behavior different from the memory test or naïve preference test which had light on for 3 (or 2.5) minutes - which light intensity was used for the data shown in Figure 5.2? How did the larvae move in the high light concentration/ or with ATR - did they not show memory due to impaired locomotion?

      Light intensity for all odour preference and learning experiments with ChR2-XXL was 100 µW/cm<sup>2</sup> (except Fig. 2 – S2C), i.e. equivalent to the experiments in Fig. 5 – S1 and Fig. 9 – S2 (weak light). We replaced Fig. 5 – S2 with a new expanded Fig. 6 analysing the locomotion of our experiments shown in Fig. 2F, 3F and Fig. 2 – S1F. We show in this figure that the larvae can move relatively normal and that the locomotion effects are weaker than in our experiments with 30s light periods. Unfortunately, we do not have videos available for all experiments and therefore cannot make a similar analysis for Fig. 2 – S2C and D when we used strong light or ATR feeding. The experimenters did not notice impaired locomotion during the experiment and the animals did show normal odour preferences similar to those shown in Fig. 3 – but the preferences were the same after paired and unpaired training, resulting in zero Memory Scores. Therefore, we do not think that the locomotion prevented the memory expression.

      In several experiments, even genetic controls seem to show learning with blue light activation. Thus, the light itself seems to activate DANs. The authors should discuss these effects and explain what this could mean for the findings. The light stimulus might not just activate the specific DAN that expresses the optogenetics, but also additionally other DANs which respond to light.

      We thank the reviewer to point this out and point out this caveat in lines 159-165.

      The authors speculate about the function of potential MB circuits - the DAN-MBON circuit is not well described so far and might be required for the US in test memory recall. A straightforward experiment to investigate the involvement of this circuit in punishment or safety memory recall would be to block dopamine receptors in the MBON.

      We very much agree to this suggestion, but believe these experiments are beyond the scope of the current study. We therefore decided to not perform these experiments for the current paper.

      Reviewer #2 (Recommendations for the authors):

      (1) As self-explanatory as the figures are, it would be interesting to also see controls in some of them. For example, in Figure 5, the effect size graph (Figure 5C) clarifies to an extent the difference between control genotypes and the experimental genotypes. It would be nice to see the results of genetic controls in Figures 5A and 5B instead of in Figure 4 - supplement 2.

      We originally decided to put the controls into the supplement to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (originally 7). For Fig. 4 and 5 specifically, we decided to keep the current layout because each serves a different purpose: Fig. 4D-L, Fig. 5 – S1 and S2 present the actual data with all genotypes that were made in parallel and therefore can be compared directly. Fig. 4 – S2 aims to visualize the effect of the light by comparing all controls across all experiments, normalized to the same starting value. Fig. 5A and B aim to compare the shape and effect size of activating DANs on top of the effect of the light - therefore, we subtracted the controls in each experiment from the experimental group. We think that adding the controls’ behaviour to Fig. 5A and B would undermine the aim of this figure.

      (2) It is a bit unclear why bending and tail velocities were the parameters chosen to compare between groups while in most cases they were of lower relative importance according to Figure 4 - supplement 1. Elaborating on this would strengthen the differences in behavior and also the claims of this study.

      We thank the reviewer for the suggestion and tried to make our choice clearer. Please see our answer to the respective part of the public review.

      (3) In adults, it has been shown that the same DAN can encode opposing valence depending on whether it was activated before or after odor presentation. Discussing the importance of temporal order of stimulus processing would bolster the results regarding paired and unpaired training in Figure 3.

      We thank the reviewer for this very good suggestion – also in larvae, this temporal function has been described. We discuss these observations in relation to our results in lines in 481-494.

      Reviewer #3 (Recommendations for the authors):

      Toshima et al., as stated in the public reviews, have done an admirable job with this manuscript. Below are specific suggestions that could improve the manuscript. It is, of course, up to the authors to decide which ones to attend to.

      Treat controls consistently. In Figure 2 and others, parental controls are not pooled, but in Figure 3, for odor preference, controls are pooled.

      We agree that the same things should be treated in the same way throughout a study and normally adhere to this principle. We nevertheless made an exception for Fig. 3 only because its goal is to provide a post-hoc analysis across several replications of experiments, some of which included genetic controls, others not (from Fig. 2, Fig. 2-S2 and S3). Due to relatively small effect sizes and high variability in odor preferences, to answer the question of paired and unpaired learning, we need higher sample sizes than each individual experiment provided. We therefore decided to pool all “equivalent” data across all these experiments. We do agree that this is a suboptimal approach but hope the reviewer can understand the rationale behind it. We explained our rationale clearer now (lines 180-183).

      I prefer to see all data points in a graph. It is more transparent than the box plots. Also, could you note why the data median is preferable to show over the mean?

      Although we in principle agree to the notion that presenting all data points is more transparent, we opted against it as it makes some graphs harder to read in particular with high sample sizes – in some of our figures, we have hundreds of data points per group. We explain our choice, including for using the median, in the method section (lines 835-839).

      Please undertake another round of language editing to handle spelling errors, etc. Use consistent British/American English.

      We thank the reviewer for their suggestion and tried our best to fix any spelling and grammar errors.

      I urge the authors to move beyond the false dichotomy of 'p' value statistics to using the statistical framework of estimation statistics for data analysis. I understand switching from familiar statistical analysis in such a late manuscript stage is very difficult. However, the authors can consider the estimation statistics framework in subsequent studies. https://www.estimationstats.com is a good starting point for biologists to get to know a framework that has been extensively worked on and is arguably a more 'honest' way of analyzing data. Disclaimer: I am not associated with the above website.

      We agree that the p-value has problems and are aware of the estimation statistics framework. We had considered applying it here, but we decided against switching to a completely different statistical framework for a research project that was ongoing since several years. However, we are sincerely considering it for our current research projects.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) A major question that remains is why the mutations have so much more detrimental effect in MutL (100-fold lower k<sub>cat</sub>/K<sub>M</sub>) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      We agree that the quantitative effects of the mutations differ between MutL and GyrB. However, we do not think that this difference argues against conservation of the catalytic mechanism. The trends of the mutational effects are highly consistent between the two enzymes. In both proteins, replacement of the conserved catalytic glutamate with Ala (E29A in aqMutL and E48A in aqGyrB) abolished ATPase activity and ATP binding, whereas replacement with the isosteric amide residue (E29Q and E48Q), which preserves hydrogen-bonding capability but lacks proton-accepting capacity, retained ATP binding and measurable ATPase activity. Likewise, substitutions of the second acidic residue (E32Q in aqMutL and D51N in aqGyrB) also retained substantial ATPase activity despite the loss of proton-accepting capability. Most importantly, simultaneous substitution of both acidic residues (E29Q/E32Q in aqMutL and E48Q/D51N in aqGyrB) completely abolished ATPase activity in both enzymes. Therefore, although the magnitude of the activity reduction caused by the individual mutations differs between MutL and GyrB, the qualitative pattern is essentially identical. We therefore think that the proposed catalytic mechanism is conserved, while the quantitative differences likely reflect differences in the local catalytic environment rather than differences in the underlying mechanism.

      (2) The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the acidic residues and their mutants is not shown.

      We have added Supplementary Fig. S2, which shows the 2F<sub>o</sub>–F<sub>c</sub> electron density maps around residues Glu48/Asp51 (or their substituted residues) in the wildtype and all mutant structures. The following sentences have been added in the revised manuscript:

      “To show that the introduced substitutions were unambiguously supported by the crystallographic data, electron density maps around residues 48 and 51 are shown in Supplementary Fig. S2.” (p. 4 line 198-200 in the revised manuscript)

      Reviewer #2 (Public review):

      (1) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of the effect (or likely effect) of disease-linked mutations in GHKL ATPases would have strengthened this study.

      We agree that extending the analysis to additional disease-associated variants in other human GHKL ATPases would further strengthen our understanding of the conserved catalytic mechanism and its clinical relevance. However, we believe that such a comprehensive analysis is beyond the scope of the present study, which focuses on establishing the fundamental catalytic mechanism shared between MutL and GyrB. We consider systematic functional and structural analyses of disease-associated variants across the GHKL ATPase family to be an important direction for future research. We have now added a statement to the Results and Discussion section to acknowledge this limitation and highlight this future perspective:

      “Although we focused here on pathogenic variants in the MutL homologs MLH1 and PMS2, extending similar structural and biochemical analyses to disease-associated variants in other human GHKL ATPases will be important for evaluating the generality and clinical relevance of the conserved catalytic mechanism proposed in this study.” (p. 6 line 303-306 in the revised manuscript)

      (2) In MLH1, the E37K mutation completely abolishes ATPase activity, but the corresponding mutations in aqMutL, aqGyrB, and PMS2 do not. It remains unclear why E37K in MLH1 leads to complete loss of activity, as the authors propose that water molecule positioning via the first acidic residue, as well as ATP lid stabilisation and associated conformational changes, should still be possible.

      We agree that the complete loss of ATPase activity caused by the MLH1 E37K variant cannot be explained solely by loss of the catalytic carboxylate. However, we note that the corresponding aqMutL E32K variant analyzed in this study also exhibited essentially no detectable ATPase activity, indicating that this phenotype is not unique to MLH1. It can be thought that the severe defect of the lysine variants arises not merely from loss of the acidic side chain but from charge reversal. We have clarified this point in the Results and Discussion sections:

      “In contrast, the E37K mutation in the MLH1 NTD completely abolished the ATPase activity under our assay conditions (Fig. 5B and Table 1) unlike the corresponding glutamine substitutions, which retained substantial residual ATPase activity in aqMutL, aqGyrB, and PMS2 NTDs. A similar complete loss of ATPase activity was also observed for the E32K mutant form of the aqMutL NTD. These observations suggest that the severe defect caused by the lysine substitution cannot be attributed simply to loss of the catalytic carboxylate. Instead, introduction of a positively charged side chain (charge reversal) is likely to perturb the local electrostatic environment. Structural characterization of the MLH1 E37K and aqMutL E32K mutant forms will be required to clarify the molecular basis of this severe functional defect.” (p. 6 line 281-289 in the revised manuscript)

      (3) The authors do not examine ATP binding in the E32 mutants of aqMutL NTD and the D51 mutants of aqGyrB, or AMPPNP binding of the NLH1 and PMS2 mutants. Hence, the relative contributions of the acidic residues to ATP binding and hydrolysis remain partially unclear.

      We performed additional ATP-binding experiments using the aqMutL NTD E32A and aqGyrB NTD D51A mutant forms. Both mutant forms exhibited ATP-binding activities comparable to those of the corresponding wildtype forms. These results support our conclusion that the second acidic residue primarily contributes to ATP hydrolysis rather than ATP binding, whereas the first acidic residue plays dual roles in ATP binding and catalysis.

      Although we agree that nucleotide-binding analyses of the MLH1 and PMS2 variants would be informative, these experiments were not feasible because the equilibrium dialysis assay requires high protein concentrations, which we were unable to obtain for the recombinant human MLH1 and PMS2 N-terminal domains.

      We have incorporated these new data into the Results and Discussion section:

      “In contrast to the E29A mutant form of the aqMutL NTD, the E32A mutant form exhibited ATP binding ability comparable to that of the wildtype form (Supplementary Fig. S1A), indicating that Glu32 does not contribute to ATP binding.” (p. 3 line 143-145 in the revised manuscript)

      “The D51A mutant form of the aqGyrB NTD retained ATP binding ability comparable to that of the wildtype form, indicating that Asp51 is not required for nucleotide binding (Supplementary Fig. S1B).” (p. 4 line 183-185 in the revised manuscript)

      (4) The ATPase assays for PMS2 and MLH1 (Figure 7 and Table 1) were performed with purification/solubility tags still present. Hence, it cannot be ruled out that these tags influence the measured activities.

      We thank the reviewer for raising this important point. We agree that the possible influence of the purification/solubility tags on the absolute ATPase activities of the PMS2 and MLH1 NTDs cannot be completely excluded. However, the wild-type and mutant forms for each homolog were analyzed using identical constructs under the same experimental conditions. Therefore, the affinity/solubility tags are unlikely to affect the relative comparisons of the mutational effects. Furthermore, because the affinity tags are located at the N terminus and are distant from the ATPase active site, they are unlikely to directly perturb the catalytic center.

      (5) The authors suggest that the two-acidic-residue mechanism proposed in this study could be shared among several GHKL ATPase families, yet they also state that the hydrogen-bonding network was not observed in MutL and MORC family proteins. This raises doubt about how conserved the mechanism is, e.g., in MutL and MORC proteins.

      We thank the reviewer for this insightful comment. Our proposed mechanism is based on the cooperative catalytic roles of the two conserved acidic residues, namely the involvement of the first acidic residue in ATP binding and nucleophilic water positioning and the role of the second acidic residue in proton abstraction. In contrast, the Glu48–Gln340 hydrogen-bonding interaction described in aqGyrB was proposed only as a structural feature that may modulate the contribution of the first acidic residue to ATP binding. It is not an essential component of the catalytic mechanism proposed in this study. Therefore, the absence of this particular hydrogen-bonding network in the currently available structures of MutL and MORC proteins does not argue against conservation of the catalytic mechanism itself.

      Recommendations for the authors:

      Reviewing Editor Comments:

      One of the structures (Crystal Structure of the E48A variant) has relatively poor statistics in the PDB validation report. Please improve this structure.

      We performed additional refinement of the E48A crystal structure. This resulted in a clear improvement in the overall model quality, with the Ramachandran favored residues increasing from 93.4% to 95.4%, the percentage of side-chain outliers decreasing from 6.1% to 1.4%. The refined structural model has been used throughout the revised manuscript, and the updated refinement statistics are provided in Table 2.

      Reviewer #1 (Recommendations for the authors):

      Please show conventional density maps (e.g., sigmaA weighted 2fo-fc maps).

      This comment is closely related to Comment (2) in the Public Review by the Reviewer #1. In response, we have added Supplementary Fig. S2, which presents conventional σA-weighted 2F<sub>o</sub>–F<sub>c</sub> electron density maps around the catalytic acidic residues in the wild-type and mutant aqGyrB structures.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding the analysis of clinical variants, it would be informative to note that the second allele is lost before tumor growth in Lynch syndrome.

      “Therefore, these variants might contribute to the development of Lynch syndrome by weakening the ATPase-driven regulatory functions of MutL.” (p. 6 line 280-281 in the original manuscript) has been changed to:

      “In individuals carrying these germline variants, subsequent loss or inactivation of the remaining wildtype allele would leave only the ATPase-defective MutL protein, thereby compromising mismatch repair and promoting tumorigenesis.” (p. 6 line 296-298 in the revised manuscript)

      (2) P. 4, in the paragraph "Conserved roles of two acidic residues of aqGyrB in ATP hydrolysis", the E48Q mutant retains approximately one third of the WT activity, not one quarter as stated in the text (Table 1). Additionally, later in the article, the D51 mutant is reported to retain approximately one-sixth (~17%) of the WT activity, rather than ~25% as written.

      We thank the reviewer for carefully identifying these inconsistencies. The text has been corrected to accurately reflect the data presented in Table 1: “…one third of the wildtype activity” (p. 4 line 180) and “…retaining ~16%...” (p. 4 line 186 in the revised manuscript)

      (3) P. 6, lines 275-276, this sentence should be rephrased for clarity, as the authors note at the end of page 5 that not all members of the GHKL ATPase family possess this second acidic residue.

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis across the GHKL ATPase family.” in the original manuscript has been changed to:

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis among some members of the GHKL ATPase family.” (p. 6 line 292 in the revised manuscript)

      (4) P. 9, in the "Data Accessibility Statement", the PDB code 23UY is missing. This entry corresponds to the crystal structure of the D51A mutant of aqGyrB NTD and should be included.

      The Data Accessibility Statement has been revised to include the code 23UY. (p. 9 line 454 in the revised manuscript)

      (5) P. 14, the table should be labelled "Table 2. Data collection and refinement statistics for the aqGyrB NTDs", rather than "Supplementary Table 2", to ensure consistency with how it is cited in the main text.

      The table title has been corrected from "Supplementary Table 2" to "Table 2”. (p. 4 line 198 in the revised manuscript)

      (6) It is difficult to determine from the figures whether the magnesium ion is positioned equivalently in aqMutL and aqGyrB. Did the authors observe any differences in ion positioning?

      To facilitate direct comparison of the catalytic Mg<sup>2+</sup> ion between the aqMutL and aqGyrB NTDs, we have added Supplementary Fig. S3, which shows a structural superimposition of the ATPase active sites of the two proteins:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of the conserved Asn, AMPPNP, and surrounding water molecules. Neither Glu48 of aqMutL nor Asp51 of aqGyrB directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (7) The authors should discuss the interaction between the aqGyrB NTD, Mg<sup>2+</sup>, and ATP during the binding step. In the case of the E48A mutant, where ATP binding is lost, does E48 directly establish contacts with Mg<sup>2+</sup>, or is another residue involved (with conformational changes preventing this interaction)?

      Our structural analyses indicate that Glu48 does not directly coordinate the catalytic Mg<sup>2+</sup> ion. Instead, as shown in Supplementary Fig. S3, the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. We have clarified this point in the Results and Discussion sections:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. Neither Glu48 nor Asp51 directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (8) Figures 1 and 3: Use ribbon representation and no shadows, at least for the inset panels, to enhance clarity and interpretability.

      We have revised Figures 1 and 3 by displaying the protein structures in ribbon representation and removing shadows from the inset panels.

      (9) Combine Figures 1 and 2, and combine Figures 3 and 4.

      Following the reviewer's recommendation, we have combined the original Figures 1 and 2 into a single figure and the original Figures 3 and 4 into another single figure.

      (10) Figure 5: Zoom in further and remove shadows. The current panels are not very effective in highlighting how ATP is bound by the different protein variants.

      Figure 5 has been revised by increasing the magnification of the ATP-binding sites and removing shadows from the structural renderings.

      (11) Figure 8. Add a scale bar to show evolutionary distance.

      We thank the reviewer for this helpful suggestion. To provide information on evolutionary distances while preserving the clarity of the main figure, we have added a new Supplementary Fig. S5 showing the same phylogenetic tree with branch lengths proportional to the inferred evolutionary distances and an evolutionary distance scale bar. Figure 6 has been retained in its simplified form with equal branch lengths to facilitate visualization of the ancestral-state reconstruction, and we have clarified this distinction in the Materials and Methods section:

      “For visualization purposes, branch lengths were not scaled and were displayed with equal lengths in Fig. 6. The corresponding phylogeny with branch lengths proportional to the inferred evolutionary distances is provided in Supplementary Fig. S5.” (p. 9 line 429-432 in the revised manuscript)

    1. Author response:

      We would like to thank the editor and reviewers for their thoughtful and constructive feedback. We appreciate the time and care devoted to reviewing our manuscript, as well as the recognition of rigorous experimental design, the technical challenges involved in conducting an fMRI study with two sensory modalities and two tasks in both deaf and hearing participants, and the value of the findings for understanding crossmodal plasticity and cortical organisation in deafness. We are encouraged by the overall assessment of the study, and appreciate the suggestions for strengthening the manuscript. Below, we provide a summary of how we plan to address the reviewers’ comments in our formal revision of the manuscript:

      (1) Additional analyses

      (a) We will incorporate behavioural performance measures into the relevant analyses to disentangle potential behavioural contributions to the observed effects.

      (b) We will calculate the noise ceiling value for each of the RSA analyses. 

      (c) We will conduct a correlation analysis between RDMs of auditory and control regions, to investigate the similarity between these computations and whether this is influenced by sensory experience.

      (d) Regarding the suggestion to conduct whole-brain searchlight analyses, we respectfully do not believe that this approach would address the primary research question of the study, namely whether and how representations within auditory cortex differ between deaf and hearing individuals. Our central hypotheses specifically concern representational content within predefined auditory cortical regions, making the ROI-based approach the most appropriate and sensitive method for testing these questions.

      Furthermore, the searchlight approach would require adequately powered group comparisons at the whole-brain level. Given the challenges associated with recruiting deaf native signers participants and the resulting sample size, we do not believe the study is sufficiently powered to draw reliable conclusions from this analysis. We will further clarify this rationale in the revised manuscript.

      (2) Presentation of the results

      We will revise the presentation of the findings to better guide the reader through the analyses and facilitate interpretation of the figures. In particular, we will ensure that significant effects, interactions, and their relationship to the corresponding figures are described more explicitly throughout the manuscript.

      (3) Revision of the discussion

      Following the reviewers’ feedback, we will revise the Discussion to more clearly distinguish between results that directly support a conclusion and hypotheses that remain speculative.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding, the position-selective organization of visual response structures, spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An eventtriggered analysis suggests that retinotopic coding shapes both "top-down" (DNinitiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxellevel visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly-matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients, not only dATN-initiated ones, is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely-sampled designs.

      Weaknesses

      (1) The nature of negative pRFs requires more scrutiny

      The entire interpretive framework depends on treating negative pRFs in the DN as genuine position-selective neural responses (suppression). However, negative BOLD signals are well known to arise from non-neural sources, specifically, vascular stealing (where activation in nearby tissue diverts blood from adjacent voxels) and macrovascular draining vein effects that produce spatially displaced signal inversions. These concerns are amplified at 7T, where T2*-weighted GEEPI carries substantial macrovascular weighting. The DN and dATN interdigitate extensively in the posterior cortex, often within millimeters. A negative pRF in a DN voxel adjacent to a positive dATN voxel could, in principle, reflect the hemodynamic shadow of its neighbor rather than an independent neural response.

      The spatial dispersion control (matched vs. random pRFs have similar cortical distribution) is valuable but addresses long-range confounds, not local hemodynamic crosstalk. The reliability of sign and center position across runs is reassuring but does not exclude a vascular origin, as vascular architecture is itself stable across sessions. I would encourage the authors to test whether the matched-vs-random effect survives exclusion of voxels near large pial vessels (identifiable from T2* contrast or the venograms available in the NSD). These analyses would not be dispositive, but they would meaningfully strengthen the neural interpretation.

      The reviewer raises an important concern about the interpretation of negative pRFs in the DN, namely that spatially specific negative BOLD responses could, in principle, reflect local vascular effects rather than genuine position-selective suppression. The reviewer suggests excluding voxels near large vessels to address this issue.

      Based on the reviewer’s suggestion, we repeated the pRF matching analysis excluding any voxels within 3mm of a major vein, as identified using the time-of-flight (TOF) MR venography included in the NSD. This analysis therefore tests whether the retinotopically specific DN–dATN coupling persists after removing voxels most likely to be affected by vascular signal.

      Excluding these voxels did not impact our results: we found preferential coupling according to response valence and center position, with stronger correlation between matched +DN and +dATN voxels (t(6) = 6.054, p < 0.001), and a more pronounced negative correlation between matched -DN and -dATN voxels (t(6) = -5.0448, p < 0.01). We have added these results to the supplemental figures (Fig. S7), and also added to the text (Pg. 8). Together with the run-wise reliability of pRF sign and position, and the persistence of the matched-versus-random effect after vessel exclusion, this analysis supports the interpretation that negative DN pRFs reflect structured, spatially specific responses rather than a vascular artifact.

      “Finally, to rule out any possible influences from vascular stealing (i.e. the shunting of blood into active tissue from nearby regions), we repeated the matching analysis after excluding any voxels within a 3mm radius of a major vessel (Fig. S7; see Methods). Both matching effects remained after excluding vascularly susceptible voxels (+DN x dATN: t(6) = 6.054; p < 0.001; -DN x dATN: t(6) = -5.0448; p < 0.01).”

      (2) Amount of retinotopic mapping data and choice of pRF pipeline

      The NSD includes 6 runs of retinotopic mapping (~5 minutes each; 3 baraperture, 3 wedge/ring). The authors use only the 3 bar-aperture runs (~15 minutes total per subject) and fit their own pRFs using AFNI's 3dNLfim procedure, rather than using the pRF estimates provided as part of the NSD release (which were fitted using the analyzePRF toolbox with all 6 runs).

      Fifteen minutes of bar data is quite limited for reliable voxel-wise pRF estimation, especially in regions far from the early visual cortex, where signal-to-noise is inherently lower. Standard recommendations for robust pRF mapping in higherorder regions generally suggest substantially more data. The variance-explained threshold is close to the noise floor by design, meaning that a non-trivial number of the "retinotopic" DN voxels may be poorly estimated. Given that the core analyses depend on both the sign and the center position of these pRFs, the limited data is a significant concern.

      The authors do not explain why they chose to re-fit pRFs rather than use the NSD-provided estimates. If the motivation was methodological (e.g., the NSD pRF pipeline does not readily yield signed amplitude, or the bar-only fits were judged more appropriate for detecting negative responses), this should be made explicit. If the NSD-provided pRFs can reproduce the key findings, this would substantially increase confidence in the results. If they cannot, that divergence itself would be important to understand. I would ask the authors to address this choice and, if feasible, to report whether the core results replicate using the NSDprovided pRF estimates and/or whether using all 6 runs of retinotopy data changes the findings.

      The reviewer raises two related concerns: first, that the amount of retinotopic mapping data available in the NSD may be limited for estimating voxel-wise pRFs in higher-order cortical regions; and second, that we re-fit the pRF model using AFNI rather than relying on the pRF estimates provided with the NSD release. We appreciate the opportunity to clarify both points. We agree with the reviewer that more travelling bar data would be preferable and would likely yield more robust model fits, particularly in higher-order regions with lower SNR. This is a limitation of our paper that we now acknowledge in the discussion section. However, we do not think that more data would fundamentally change the pattern of our results for the following reasons.

      First, we implemented a novel data-driven approach to derive a threshold for thresholding significant pRF fits (a noise floor). Importantly, our noise floor estimation yields a conservative threshold (R<sup>2</sup> > 0.14), which is greater than both our previous work characterizing cortical pRFs (Steel et al. 2024: R<sup>2</sup> > 0.08) and other work exploring visual responses in the default network (Klink et al. 2021: R<sup>2</sup> > 0.05; no threshold: Szinte and Knapen 2020; Knapen 2021).

      Second, the key pRF features used in our analyses – response sign and centre position – were reliable across retinotopic mapping runs. This reliability is important because our central matching analysis depends on voxel-wise estimates of both response valence and visual-field position.

      Third, the matching analysis asks whether pRF parameters estimated from the retinotopic mapping task predict functional coupling measured during independent resting-state scans. Noisy or unstable pRF estimates should weaken this relationship, because they would degrade the accuracy of voxel-wise matching. Thus, parameter instability would be expected to obscure retinotopically specific coupling rather than systematically produce the observed matched-versus-random effects.

      To the reviewer’s question about our decision to re-fit the pRF model using AFNI, the reviewer is correct that this was motivated by the requirements of our analysis: we re-fit the pRF estimates using AFNI because it allows for both positive and negative signed amplitudes. The pRF model fits provided with the NSD do not allow bivalent amplitude estimates. We have made this decision clearer in the text, reproduced below (Pg. 4-5; Pg. 14-15).

      “We chose to re-fit the data using a simple Gaussian approach as implemented in AFNI to allow for both positive and negative signed amplitudes.”

      “The limited amount of pRF mapping task data included in the NSD posed a challenge for establishing reliable visual response estimates. Here, we addressed this issue by developing a novel thresholding method to establish robust voxel-wise model fits. Among voxels that passed this empirical threshold, we observed a significant correlation in voxel-wise estimates of centre position and visual response amplitude. In addition, our pRF matching results were based on the relationship between the voxel-wise estimates of centre position and response amplitude with resting-state fMRI – a completely independent measure. Crucially, noisy estimates of pRF parameters would obscure this relationship and make our results less likely. Therefore, despite the relatively limited pRF mapping data available, unstable pRF estimates are unlikely to drive our results.”

      (3) pRF model adequacy for the Default Network

      The isotropic Gaussian pRF model was developed for and validated in early and mid-level visual cortex, where it captures the dominant spatial selectivity of neuronal populations. In DN voxels where the model explains comparatively little variance, it is less clear that the model is capturing the right quantity.

      Specifically, the negative pRFs could conceivably be described by a model with a dominant suppressive surround (e.g., a difference-of-Gaussians model), in which what appears as a "negative pRF" in the standard model is actually the surround component of a center-surround mechanism whose center is poorly resolved. This distinction matters: a genuine inverted code (negative center response) implies a qualitatively different computation than inherited surround suppression from nearby visual cortex.

      The authors should consider discussing why the standard model is sufficient for the questions asked, or ideally, testing whether the sign distinction survives under alternative pRF model specifications.

      We appreciate the reviewer’s comment about the limitations of a single gaussian pRF model. We chose the single gaussian model as a direct extension of prior work from our lab and others (Steel et al., 2024, Klink et al., 2022, Szinte and Knapen, 2021). We agree that a negative response in this model could, in principle, reflect a more complex spatial profile, such as a dominant suppressive surround. However, adjudicating among alternative pRF models would require more retinotopic mapping data than are available in the NSD, particularly for higher-order cortex. Thus, we feel that it is outside the scope of the current work. We now address this limitation in our discussion (Pg. 15).

      “Relatedly, here we used a single gaussian model, consistent with prior work on negative visual responses in memory systems (31, 33, 34). However, other models of visual receptive fields might offer further insight into the DN’s visual responsiveness, such as double gaussian models of surround suppression (65) or compressive summation (66). Future studies might directly compare different visual models to further refine the computations underpinning visual responses in the DN.”

      (4) Interpreting resting-state transients as top-down vs. bottom-up The event-triggered analysis labels high-amplitude DN pRF activations as "topdown events" and dATN activations as "bottom-up events." This is a reasonable inference given experience-sampling studies showing that rest involves alternation between internal and external attention, but it remains an inference. Without concurrent experience sampling, eye-tracking, or physiological monitoring, we cannot establish that a spontaneous DN transient reflects memory retrieval or internally-directed thought rather than a global arousal fluctuation. Similarly, dATN transients during rest could reflect covert shifts of spatial attention to remembered or imagined locations rather than bottom-up processing per se. I would ask the authors to soften this framing or to discuss what additional data would be needed to validate the top-down/bottom-up attribution.

      The reviewer raises an important concern about the strong interpretation of elevated BOLD activity detected in the DN and dATN as top-down and bottom-up events. We agree that the limitations of fMRI in our current data prevent these strong claims about the origin of these signals. We have therefore softened this framing throughout the manuscript, and we now refer to these events as DN-driven and dATN-driven. We think that this more directly describes the analysis: events were defined by transient high-amplitude activity in DN or dATN pRFs, respectively.  

      (5) The "retinotopic code" vs. "visual field bias" distinction The paper uses the language of a "retinotopic code" throughout and correctly distinguishes this from a "retinotopic map," noting that DN voxels do not form a continuous topographic representation on the cortical surface. This distinction deserves greater emphasis. In vision science, retinotopic maps carry computational significance through their topographic continuity and relationship to cortical wiring. A distributed collection of voxels with coarse visual field preferences but no cortical topography is a fundamentally different organizational feature. Recent reviews have drawn an explicit distinction between retinotopic maps and visual field biases (Groen, Dekker, Knapen & Silson, TiCS 2022), and the present findings may be more accurately characterized as the latter. Perhaps the authors think that the distinction is merely a signal-to-noise distinction, in which case I would invite them to clearly speak to this interpretation. In any case, this is not a criticism of the findings themselves, but clarity on this point would prevent conflation of two different organizational principles and would help position the work for both the vision and network neuroscience communities.

      The reviewer raises a valuable point about the distinction between a retinotopic code, a retinotopic map, and a visual field bias, and we are happy to add discussion of this topic to our manuscript.

      Our results show that the DN does not exhibit a continuous retinotopic map in the sense used in early visual cortex. Rather, our results suggest a distributed voxel-level code for visual-field position: individual DN voxels show reliable spatial preferences, and these preferences predict retinotopically specific functional coupling with dATN voxels. This voxel-level organization is analogous to other distributed spatial codes, such as head-direction coding in retrosplenial cortex, where spatial variables are represented by population activity without requiring a topographic map on the cortical surface. This differs from a coarse visual-field bias, including preferential responses to the contralateral visual field, although we do also observe such biases. We have added text unpacking this important distinction to the Discussion (Pg. 15-16):

      “Prior work has emphasized the visual response bias in regions where voxel-wise retinotopic responses lack a map-like organization(35); overall, the DN does exhibit this kind of bias. However, our results show that the voxel-scale activity underpinning this bias reflects the latent connectivity of those voxels. Thus, we adopt the term “retinotopic coding”, because this voxel-scale coding scheme exists without a map-like organization on the cortical surface. For example, rodent and bat head direction cells are not laid out in a literal ring, but the population code of these neurons forms a ring manifold(68, 69).”

      Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that restingstate activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with the sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks, and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g., see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis, and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN-positive pRFs. The authors do show that the positive pRF voxels have abovechance consistency across runs, again providing evidence that there are true positive voxels in this set, but perhaps a stricter criterion (such as having consistent negative fits across runs) would provide more targeted identification of the DN voxels with true retinotopic sensitivity.

      We thank the reviewer for giving us the opportunity to discuss this important decision. We agree with the reviewer that a stricter R<sup>2</sup> criterion could result in more targeted pRF identification. Motivated by the reviewer’s suggestion, we repeated the cross-region pRF matching analysis across multiple R2 thresholds.

      The retinotopic matching effects were not dependent on the original threshold. In fact, we found that the pRF matching effects are enhanced as the R<sup>2</sup> value increases (Fig. S5). This pattern suggests that any false-positive voxels admitted near the original threshold would dilute, rather than drive, the observed matched-versus-random effects. We have added text to the results highlighting this finding (Pg. 7):

      “In contrast, DN voxels that responded positively to visual stimulation (DN positive pRFs, +pRFs) had a positive correlation with the dATN (mean correlation = 0.284±0.152, t(6) = 4.96, p = 0.0025), while DN voxels with systematic negative responses to visual stimulation (DN negative pRFs, -pRFs) were anti-correlated with the dATN (mean correlation = -0.21±0.149, t(6) = -3.75, p = 0.0094). This relationship was further strengthened by adopting more conservative R<sup>2</sup> thresholds up to 0.30 despite the overall number of included voxels decreasing, suggesting that this effect is not driven by false-positive voxels at the edge of our threshold criteria (Fig. S5).”

      (2) The claim that "opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning" is not well supported. The fraction of DN voxels with negative pRFs is small: 9.42% of DN voxels have pRFs, and 58.77% are negative, so about 6% of DN voxels have negative pRFs. The fact that any DN voxels have negative pRFs is notable, but the authors do not provide evidence that these 6% are driving the overall behavior of the DN. They do show (e.g., in Figure 2B) that negative and positive pRFs have opposing influences, but the overall correlation with dATN does not look similar to the negative pRF connectivity. I'm also unsure whether "opponency" is a reasonable description for two networks that are "independent (i.e., not correlated)" in this analysis.

      The reviewer raises an important point about whether negative DN pRFs should be described as driving the overall DN–dATN relationship. We agree that this language was too strong. Negative pRFs constitute a small subset of DN voxels, and our analyses show that this subset has a distinct pattern of functional coupling with the dATN, not that it explains the global relationship between the DN and dATN as a whole.

      We have therefore revised the manuscript to avoid implying that negative DN pRFs drive overall DN–dATN opponency. Instead, we now frame these voxels as an important retinotopically tuned subpopulation nested within broader network dynamics. Specifically, our results show that visually responsive DN voxels are not homogeneous: positive and negative DN pRFs show opposing patterns of coupling with dATN pRFs, and these interactions are strengthened when voxels share visual-field preferences. This suggests that a small but structured subset of DN voxels may provide a route for retinotopically specific communication between internally and externally oriented networks, without implying that this subset determines the mean activity pattern of the entire DN:

      “Spontaneous DN and dATN activity during rest is uncorrelated at the network level. However, voxel-scale functional coupling across networks is shaped by the latent visual field preferences of individual voxels in each network, as measured during independent retinotopic mapping.” Abstract (Pg. 2)

      “This result shows that voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity. Specifically, the DN and dATN activation is independent during rest. However, at the voxel-level, specific sub-groups of DN voxels have distinct coupling patterns with the dATN that depends on the valence of voxels’ visual responses. DN and dATN voxels with positive visual responses show a positive relationship during rest, and a notable subset of DN voxels with negative visual responses display the canonical opponency with dATN voxels. This suggests that retinotopic coding may be a mechanism that enables visual information to be exchanged between these large-scale brain systems. Specifically, opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning.” Results (Pg. 7)

      These findings offer a multi-scale account of neural communication, in which interactions among sub-populations of voxels with shared tuning preferences are nested within macro-scale network dynamics. Nesting multiple neural codes might enable ongoing computations within a larger brain system (e.g., attending to internal mental states within the DN during memory recall), while simultaneously allowing for the sharing of fine-grained representations across brain systems (34). Discussion (Pg. 14)

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study. First, is 13 TRs a typical length of time that activity is elevated during one of these events? Second, the top-down and bottom-up terminology is perhaps too loaded and not well-justified; if the negative pRFs in the DN reflect a meaningful coding system, then couldn't low (rather than high) activity indicate a top-down event?

      We thank the reviewer for these helpful suggestions. To the best of our knowledge, there is not currently a widely agreed-upon time window for performing event-based fMRI analyses. We chose a 13 TR time window to balance between sufficiently capturing BOLD signal related to the chosen event while also minimizing influence from other signal fluctuations, based on the procedure adopted in Gordon et al. (Nature, 2023) and Mitra et al. (J. Neuro Phys, 2014), which considered temporal relationships among brain regions over comparable timescales. In our analysis, this window considered 6 TRs (9.6s) on either side of the detected event, which we felt comfortably captures the peak BOLD signal that would result from an impulse at the event time, and responses that may reflect upstream activity leading into it.

      The reviewer has raised an additional comment about the terms “top-down” and “bottom-up.” These concerns were shared by Reviewer 1. Based on these comments, we have adopted the terms “DN-driven” and “dATN-driven”, which we think aligns more closely with our analysis approach.

      (4) The framing of this paper relative to the authors' past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems"), could be improved. The existence of negative pRFs in the DN and a functional relationship between these pRFs and the sensory pRFs have already been described in prior work. My understanding of the primary novelty here is that this paper examines resting-state data, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

      We appreciate the opportunity to clarify the novel aspects of our paper. The reviewer correctly identifies the extent of prior work, which identified -pRFs in regions of the canonical default network (Szinte and Knapen, 2021; Klink et al., 2022) and characterized the local interactions between adjacent perceptual and mnemonic regions (Steel et al., 2024). Our current work builds upon these findings in two key ways.

      First, we explore the effect across individually-defined whole brain networks. While the DN and dATN are often adjacent, these networks are spatially discontinuous and are comprised of distinct sub-regions (e.g. in prefrontal cortex). Whether retinotopic patterning of activity would persist in distributed networks could not have been extrapolated from our prior work. We think that finding will be of broad interest to the community studying perception and memory systems, because it offers a mechanistic account of how information is read in/out of memory.

      Second, here we considered whether spontaneous activity across networks would be structured by a retinotopic code. Our previous work characterized activity during tasks that depended on visual information: either scene perception or mental imagery. While the prior work was an important first step, it left open the possibility that retinotopic coding may only be relevant in visual tasks. By demonstrating that the retinotopic coding structures voxel-specific coactivation during rest, which entails no overt visual demands, we provide evidence that retinotopic features are a general, mode-agnostic code between regions.

      (5) The definition of the default mode (DN) in this study aligns with past research, but the definition of the dorsal attention network (dATN) seems at odds with standard terminology. For example, the authors cite Fox et al. 2006, which depicts the dATN as including regions such as IPS, FEF, SMA, and MT+. Here, however, the "dATN" seems to be primarily lateral and ventral visual cortex (e.g., Figure S5). The exact location of these sensory pRFs is not critical to the authors' claims, but this labeling seems incorrect, and the motivation for defining/selecting the sensory network in this way is not described.

      We thank the reviewer for this insightful comment and their careful consideration of our network definition.

      Our method for network identification, and the topography of the resulting networks, are broadly consistent with more recent conceptualizations of the DN and dATN (e.g. Du et al. 2024, Gordon et al. 2017, Braga and Buckner 2017). Relatedly, because we defined brain networks based on the unique connectivity patterns of each individual participant, we expect them to differ from previous group-level network descriptions. The increased resolution of the 7T data in the NSD may also result in greater departure from prior definitions compared to previous work done at 3T.

      Further study into dATN differences between group-level 3T, individualized 3T, and individualized 7T networks could be a valuable future direction, but this is outside the scope of this work.

      Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (the relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the key claims are not fully supported by the data presented. There is also a concern about over-interpretation of the results. Key issues:

      (1)  The authors claim that retinotopic coding scaffolds the interaction between DMN and dATN. However, retinotopically tuned voxels account for a mere 9% of DMN voxels. So this appears to be a major overstatement. For instance, the statement that "these findings would position retinotopy as a unifying framework for brain-wide information processing" is not justified given the presented data.

      We appreciate the reviewer’s concern about the framing of our conclusion, which was shared by reviewers 1 and 2. In response to these comments, we have revised our paper to more accurately reflect the observed data. Specifically, we focus on the specific sub-populations of voxels within the DN and dATN that show retinotopic responses, and we have removed references to explaining the overall pattern of activity across networks.

      (2) Given that positive pRF voxels in DMN positively correlate with dATN voxels and negative pRF voxels in DMN negatively correlate with dATN voxels, there is a concern that these results could be contributed to by imprecise brain network parcellations. E.g., could some of the positive pRF voxels in DMN be erroneously assigned to DMN and actually belong to one of the other task-positive networks? There is insufficient validation of network parcellation to put this worry to rest, especially since it depends on ICA, which has a degree of arbitrariness built in.

      We thank the reviewer for the opportunity to clarify our method for network definition.

      Precision functional mapping is a growing field with many methods for defining personalized functional networks for each individual. Because the NSD resting-state data is relatively high resolution, we chose an approach designed to improve the stability of voxel-wise network assignment: Multi-Session Hierarchical Bayesian Modeling approach (Kong et al. 2019; Du et al. 2024). This approach enhances stability of network assignment by including a group-based prior and accounting for both within- and across-subject variability. This approach is more stable than ICA, and, because this approach leverages a prior, there is less concern about arbitrary or idiosyncratic network definitions.

      However, it is still common for network assignments to have lower confidence around the borders between networks. Yet, we also do not think border misassignment is likely to explain the present results for two reasons: first, while DN and dATN nodes are sometimes adjacent, there are many regions where they are spatially distant, such as the IPS for dATN and the lateral temporal lobe for DN. Second, the DN pRFs do not appear to cluster selectively along DN–dATN borders, suggesting that they are not simply misassigned dATN voxels.(Fig. S3) Therefore, we think voxels on the edge of these networks are unlikely to drive the effects observed here (see Fig. S3).

      (3) The claim that retinotopic coding is intrinsic to the DN network is not supported by rigorous analysis and results. The analysis here has many arbitrary factors, including: the threshold of the 99th percentile of resting-state distribution; the designation of DN as "top-down" and dATN as "bottom-up"; the definition of "anti-matched" voxels instead of using randomly selected voxels; and the statistics being paired between matched and anti-matched voxels instead of using comparisons to baseline. Overall, I do not think that the result supports the conclusion that retinotopic coding in DN is intrinsic instead of being bottomup-driven, given the very high threshold (99%) used and the fact that many other networks could also send bottom-up input to DN. Furthermore, the idea that bottom-up inputs only occur when the dATN (or any other RSN)'s spontaneous BOLD activity is above a certain threshold is a huge and unvalidated assumption.

      The reviewer raises several interesting concerns about decisions in our event-detection analysis. Here, we clarify the rationale for several analytic choices:

      (1) The 99th-percentile threshold was chosen to identify sparse, high-amplitude events while minimizing contamination from smaller ongoing fluctuations.

      (2) The other reviewers also noted a concern with the top-down/bottom-up terminology. We have revised these terms to DN-driven and dATN-driven, which we think reflect our approach more accurately.

      (3) We used anti-matched rather than randomly selected voxels because the full event-by-voxel randomization procedure was computationally intractable at the network level.

      (4) We did not understand the reviewer’s contention about activation baseline, but we would welcome clarification.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) The reliability analysis (Figure 1D) notes that dATN negative pRF amplitude was not reliable above chance in 2 of 7 participants. This could be discussed more prominently as it suggests that negative pRFs may not be stable features in all networks, which tempers the generality of the sign distinction as a fundamental organizational property.

      We thank the reviewer for raising this point of clarification. It is true that 2/7 participants did not show reliable negative pRFs in the dATN. However, the majority of participants show stable negative pRFs, and even in these 2 participants, the negative result does not indicate that negative pRFs would not be stable in those individuals with additional data. 

      Based on the reviewer’s comment, we have added emphasis to this point, but we do not feel that this warrants greater discussion in the paper. 

      Both positive and negative pRF amplitude was reliable in the DN for all subjects. In the dATN, positive amplitude pRFs were reliable in all participants, and negative amplitude pRFs (which constituted a small proportion of the overall pRFs in this network) were reliable in 5/7 participants. For the remainder of the paper, we only consider positive pRFs in the dATN. Importantly, pRF center position was highly reproducible across runs of pRF data in the dATN and DN in all subjects (Fig. 1D).  Pg. 5

      (2) The paper would benefit from situating the findings more explicitly within the cortical gradient framework (Margulies et al., 2016), which predicts that DN regions have maximally abstract, transmodal codes. The present findings complicate this view productively and deserve to be "situated" within that ongoing debate.

      We agree that the gradient framework is interesting, and we have added discussion of Margulies to our paper. (Pg. 16)

      Relatedly, the DN is considered a transmodal hub for cortical processing, where disparate sensory and motor processes converge (59, 75) The DN’s position at the cortical apex implies connections with and influence over unimodal cortical areas. However, the mechanism for liking unimodal and transmodal networks had been unknown. Prior work posited that sensory coding in transmodal areas might serve this function (31, 35), and our data provide direct empirical support for this account: specific visually-responsive voxels provide an input/output interface linking perceptual and memory systems. This complements work delineating specific affective and effective subregions within the DN that link the DN to other brain areas (76). Thus, while the DN may be “distant from input” (28), these data suggest that it is not disengaged from sensory processing.

      (3) It would be informative to know whether the *proportion* of negative vs. positive pRFs differs between DN-A and DN-B, given their distinct functional roles.

      Despite the functional specialization of DN-A and DN-B, and the slightly higher mean proportion of negative pRFs in DN-A (61% vs 56%), we found no statistically significant difference in the proportion of negative pRFs across the two networks (t(6) = 0.888, p = 0.409).

      (4) Low N is inherent to the densely-sampled NSD design, and the within-subject consistency is a strength. Nevertheless, with 6 degrees of freedom, the precision of specific quantitative estimates (e.g., that 58.77% of DN pRFs are negative) is uncertain, and the authors should be cautious about the generalizability of these point estimates.

      The reviewer raises a concern about the inclusion of specific levels of decimal place in our statistical reporting. We do not think that this is a major issue with the paper, but we are willing to change if the reviewer feels strongly.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1C could use an explicit legend (I believe it is following the color convention from the bar plots in 1F?). Also, for consistency, it would be helpful to make all the colormaps in 1F correspond to the bars (i.e., change the dATN colormap to go white->green).

      We thank the reviewer for this suggestion, and we have added explicit labels to Fig. 1C

      (2) Providing a scatter plot, in which each dot is a voxel and the x and y axes are the pRF amplitude estimates in different runs, could help provide evidence that there are voxels with pRFs that have consistently negative amplitudes across runs. This would also go beyond the binary consistency analysis in Figure 1D to show that the magnitudes of the amplitude estimates are also consistent.

      We thank the reviewer for this suggestion. We feel that the binary consistency conveys sufficient information. Because the analysis is done using pairwise correlation, how the scatter plot would reflect the three-way consistency is not clear. 

      (3) For understanding how the overall correlation between DN and dATN could be driven by voxel populations with opposing effects (e.g., Figure 2B), it would be useful to show how the +pRF and -pRF voxels compare to other voxels within the DN. For example, are these the voxels with the strongest negative and positive correlations with dATN, or are there many other DN voxels (among the 90% that do not have pRFs) that also have similarly-strong dATN correlations?

      The reviewer offers a very interesting suggestion. Based on the reviewer’s suggestion, we have refocused our paper on the particular subpopulations of +/- pRFs in the DN, rather than on an explanation for the overall pattern of correlation between the DN and dATN. Because our revised framing focuses on the properties of these retinotopically defined voxel populations, rather than on explaining whole-network DN–dATN coupling, we have not added this additional analysis. We have revised the relevant text to avoid implying that these pRF subpopulations drive the overall network-level relationship.

      (4) Initially, the baseline comparison pRFs for the matched pRFs are labeled "random" pRFs, which seems misleading; these are closer to "mismatched"/"anti-matched" pRFs since they are selected from the 1/3 that are farthest away. Then the comparison switched to using the anti-matched pRFs that are the 10 very farthest away, though I didn't understand the rationale that "the large number of pRFs made the random matching procedure impractical" - in what way is the number of pRFs larger in this analysis? Having a more consistent baseline (e.g., just using the 10 anti-matched pRFs the whole time) would be easier to interpret.

      We thank the reviewer for this suggestion. We have compared the results between the randomly-sampled bottom ⅓ matched versus the 10 worst matched, and the pattern of results is identical (the effect is strongest in the 10 worst matched). Therefore, we include the bottom ⅓ matched in the main text as a more conservative test of this effect. We are happy to include this as a supplemental figure if the reviewer feels it is essential. 

      (5) In the past, I have only seen the terminology "bootstrapped" to refer to sampling with replacement from the data sample, producing samples/statistics that are centered on the observed data. Here (lines 704-708), the sampling is coming from the null distribution of randomly-chosen voxels, and therefore the term "bootstrapped" would not apply (and could just be replaced with "null").

      We have revised this terminology in the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract and Discussion should be significantly toned down. E.g., the claim that "These findings challenge the prevailing view of global DN-dATN antagonism" is not really supported by the data provided. The claim that "retinotopic coding underpins the dynamic coordination of perception and thought" is also unsupported by the presented data.

      We have revised the manuscript in light of this comment.

      (2) Line 233-235: The null statistical result cannot support the claim reached here. Correlation analysis or Bayesian statistics should be used.

      We have revised the manuscript in light of this comment.

      (3) Line 250-254: Comparison to baseline should be used, in addition to comparing matched and random voxels.

      We agree that baseline comparisons can be useful in event-triggered analyses. However, for the pRF-matching analysis discussed here, the critical question is whether shared visual-field preference influences resting-state functional coupling between DN and dATN voxels. For this question, we believe that the appropriate baseline is the coupling observed for pRFs that do not share visual-field preferences. We therefore compare retinotopically matched pRFs to randomly matched pRFs drawn from the same networks. 

      (4) Line 271: "not" is missing.

      We have revised the manuscript in light of this comment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs), they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across over 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While conservation maps are valuable resources, the manuscript lacks functional validation in congener species, limiting claims about broad applicability across related genomes/species.

      The approach also failed to validate developmental CREs. None of the candidates from combined ATAC and conservation filtering drove reporter expression matching endogenous patterns. The authors appropriately hypothesize technical limits (low expression) or biological factors (long-range enhancers, shadow enhancers).

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific), where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      We thank the reviewer for their comments and valuable feedback.

      Reviewer #1 (Recommendations for the authors):

      (1) Standardize terminology and introduce acronyms properly:

      I suggest standardizing technique names throughout the manuscript (e.g., ATAC-seq rather than ATACseq, RNA-seq rather than RNAseq, ChIP-seq rather than ChipSeq). Please also introduce technical terms with their full name at first use, such as 'Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq)' rather than just the acronym.

      Corrected.

      (2) Correct author name:

      On page 15, "Lo brutto" appears to be misspelt. Please review and fix throughout the text and the corresponding reference in the bibliography if this is the case.

      Corrected.

      (3) Revise the title to reflect dual contributions:

      The paper delivers two key advances: (a) ATAC-seq atlas plus functional CRE discovery in P. hawaiensis, and (b) innovative low-coverage sequencing for conservation mapping across congenerics. Current title highlights only the first. Consider modifying the title to better reflect both.

      The title reflects our overall objective, without highlighting one of the two approaches (advances) in particular. We would like to keep this concise title and invite readers to read about the two approaches in the abstract.

      (4) Emphasize the combined power of the approach in the Discussion and Conclusions:

      A significant innovation of this manuscript is integrating low-coverage comparative genomics with ATAC-seq to prioritize functional CREs in a large non-model genome. The abstract highlights this well, but the Discussion and Conclusions could better emphasize the power of this pipeline over ATAC-seq alone. The authors could also add 2-3 sentences quantifying cost savings versus traditional assemblies and reiterating how this complements ATAC-seq for efficient CRE prioritization in non-model species.

      We have added the following text in the Discussion to describe complementary contributions of ATAC-seq and sequence conservation to CRE discovery: "Previous efforts to identify cis-regulatory elements in Parhyale relied on reporter constructs carrying a few kb of sequences upstream of selected target genes, an approach that has worked well in animals and plants with relatively small genomes. As presented earlier, however, this approach was often unsuccessful in Parhyale: from tens of reporters tested, only four robust native drivers had so far been identified (refs). The present work adds 5 new drivers to that collection, including ones with ubiquitous, neuron- and muscle-specific activities. This result comes from combining information on genome-wide chromatin accessibility and evolutionary conservation profiles.

      We cannot at this point distinguish the relative contributions of chromatin profiling and sequence conservation to CRE discovery, because these sources of information were not tested separately. At minimum, we can state that (1) ATAC-seq profiles serve to identify robustly the promoters of candidate genes that are active in a particular cellular context (cell type and stage) and (2) coupling this information with sequence conservation narrows down candidate promoters and distant CREs by a factor of 4 to 10, since only a fraction of ATAC-seq peaks show sequence conservation (Figure 3D). This represents a great improvement in our ability to select candidates to test by transgenesis, the most labour-intensive step in the process."

      Further, we have added this text to explain the advantages and cost savings of our low coverage sequencing strategy: "Our strategy of mapping regions of sequence conservation by direct mapping of short sequence reads across species is much more accessible than conventional strategies that rely on genome assembly. The latter require much higher sequence coverage (> 50x) from multiple libraries, long-read sequencing or other scaffolding methods, and complex bioinformatic pipelines to assemble large genomes. Moreover, these approaches are often compromised by high levels of polymorphism found in natural populations. We estimate that our approach is 5- to 10-fold cheaper than assembly-based methods, even excluding labor costs."

      (5) Improve figure readability:

      The authors could improve figure readability by introducing a schematic representation of the specimens and more references in Figures 4, 5 and Supplementary Figure 5. A schematic representation would help non-experts in Parhyale understand what they are looking at. Some figures might also benefit from improved color contrast (e.g., Figure 3 has very similar orange/red colors; black dots on a dark grey background are hard to distinguish).

      We have added additional labels and explanations in the legends of Figures 4, 5 and Supplementary Figure 5, which we think will make the images more intelligible to the readers. In Figure 3 we modified the colouring in panels B and C to improve contrast.

      (6) Quantify reporter validation efficiencies:

      The authors should add a summary table/plot (e.g. n surviving, n fluorescent) or label the figures (n of specimens showing that pattern/n of specimens that do not show the pattern) with exact numbers to explicitly illustrate the observation. For example, Supplementary Table 3 contains excellent data for the putative developmental CREs tested, but the main text lacks equivalent quantification for successful reporters.

      This information is already provided in Table 3.

      (7) Discuss the rapid evolution of developmental CREs:

      The failure to validate developmental CREs using conserved candidates may also reflect the rapid evolution and turnover of developmental enhancers, which can erode detectable sequence conservation over these phylogenetic distances. As a result, functionally relevant elements may have been excluded during candidate selection. It may be worth discussing this possibility alongside the proposed long-range and shadow enhancers hypothesis.

      We added the phrase “or the rapid evolution of these enhancers leading to low sequence conservation" in the relevant part of the Discussion.

      Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next-generation sequencing technology and the decrease in price of short-read sequencing.

      We thank the reviewer for their comments and valuable feedback.

      Two major weaknesses are:

      (1) The novelty of the approach and its advantages should be more explicitly stated.

      (2) The authors do not discuss in depth the strength of using a combination of two methods rather than either of the two, especially considering that previously known CREs do not overlap with conserved sequences.

      We have added two paragraphs at the start of the Discussion to address the reviewer's comments 1 and 2 more explicitly (see response to reviewer 1, comment 4).

      Previously known CREs do include some conserved sequences, see Suppl. Figure 7.

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the two above-mentioned weaknesses, the authors should address the following minor issues:

      (1) It is difficult to refer to particular regions of text without line numbers.

      Spelling of ChIPseq is inconsistent in the Introduction.

      Spelling corrected. (Sorry for not including line numbering, we'll try to remember next time.)

      (2) "6.4% of reads from P. aquilina, 4.1% of reads from P. darvishi, and 0.46 % of reads from P. plumicornis could be mapped unambiguously to the P. hawaiensis genome" seems quite low for closely related species. Do the authors expect such low rates?

      Neutral nucleotide substitution rates in multicellular animals are in the order of 1 per site per 100 million years (e.g. https://pubmed.ncbi.nlm.nih.gov/12949132/) or a little lower (e.g. https://pubmed.ncbi.nlm.nih.gov/11792858/, https://pubmed.ncbi.nlm.nih.gov/34049492/). With the evolutionary times separating P. hawaiensis from P. aquilina/darvishi and P. plumicornis estimated at roughly 50 and 180 million years (2x25 and 2x90 million years, respectively), we expect a large fraction of neutrally evolving nucleotides in these genomes to have changed. We performed the read mapping using bowtie2, which requires a ~20 nt long perfect match with the reference sequence. We were therefore not surprised to obtain such low rates of read mapping. In fact, these low mapping rates (long divergence times) are important for islands of sequence conservation to stand out.

      (3) "Of these, 37% are found in introns, 54% in intergenic regions, and 1% overlap with promoters (TSS), marking regions that evolve at a lower rate than surrounding non-coding sequences". The authors explain in the Methods why they omit exons, but in the Results and Discussion, it is not stated. In addition, discussing the conservation with exons would be helpful, and the % in exons should be compared to non-coding regions.

      We have added "Of these, 8% are found in exons, likely reflecting conservation in protein-coding sequences".

      (4) "Two of the 7 reporters we tested, named neuro5 and neuro6," if I understood correctly, neuro5 and neuro6 are CREs, however, they are named quite ambiguously, and their names can be mistaken for gene names.

      Indeed, neuro5 and neur6 are the names of the CRE reporters. We have now added the names of the corresponding genes ("carrying CREs associated with the genes αTub and Cdk5α, respectively"). The gene names are also given in Table 3.

      (5) Why was single-end sequencing done for E24?

      We now explain this in the Methods: "Sequencing was carried out on an Illumina NextSeq 500 sequencer; we carried out single-end 76 bp sequencing for the first sample we generated (E24), and then switched to paired-end 76 bp sequencing for the other samples, because this leads to more specific read mapping."

      (6) Syntax related to in-line references should be double checked as the following sentences are broken by parentheses, e.g., "updated in (Almazán et al. 2022))".

      Corrected.

      (7) Could the authors discuss the P. aquilina genome size, which was estimated to be 3-times less than P. hawaiensis? Considering that in their phylogeny these two species are closest, it is quite surprising that they have such differing genome sizes. Do you expect it to be true? If yes, what could be the reason?

      As we explain in the manuscript, our estimates of genome size were obtained by dividing the total number of nucleotides sequenced by the estimated genome coverage, for each species. This method could overestimate genome sizes if there was a significant fraction of contaminating DNA in the preps, or a high degree of sequence variation that would prevent efficient mapping to BUSCO genes (both would underestimate genome coverage), but we find no evidence of this when we estimate the genome size of P. hawaiensis (see manuscript). We used the same method to estimate genome size in all four Parhyale species and have no reason to think that the method would be biased in one species and not in others. We therefore think that we have comparable estimates of genome size for the 4 species and the size difference is real.

      Variations in genome size can be driven by changes in the fraction of repetitive sequences found in a genome. We therefore checked the proportion of repetitive elements in each Parhyale genome using dnaPipeTE (https://github.com/clemgoub/dnaPipeTE). Based on this method (which likely underestimates the repetitive genome content) we find that the genomes of P. hawaiensis, P. aquiline, P. darvishi and P. plumicornis contain 31%, 22%, 18% and 39% of repetitive sequences, respectively. These figures do not fully account for the differences in genome size (particularly since P. darvishi appears to have even fewer repetitive sequences than P. aquiline). We therefore hesitate to add this very preliminary analysis to the manuscript.

      Of note, such rapid change in genome size is not unprecedented: in fruit flies genome size can vary more than 3-fold in species that have diverged over about 30 million years (https://elifesciences.org/articles/66405).

      (8) Wording "and found a genome coverage of 5.8x, corresponding to a genome size of 3.0 Gbp instead of 3.6 Gbp" is confusing and unclear as to what the authors exactly did here.

      We modified the sentence: "As a control, we followed the same procedure for P. hawaiensis, for which genome size is known (ref), and found a genome size of 3.0 Gbp instead of 3.6 Gbp (with a genome coverage of 5.8x)."

      (9) In the figures and supplementary figures, the genome browser screenshots should also include tracks of macs2 called peaks (those in narrowPeak format).

      Each ATACseq and sequence conservation track has its own set of peaks; we think that adding more tracks would overcrowd the figures. All the tracks (including called peaks) are provided as genome-browser-readable files in Supplementary Data files 1-3, so readers should be able to explore the data and reconstruct the figure panels without much effort.

      Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as a very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Reviewer #3 (Recommendations for the authors):

      I have no specific comment for the authors. While in my opinion the study has two limitations (as described in the public review), these are clearly acknowledged and properly discussed in the manuscript.

      The manuscript is extremely well written. It has been a great pleasure to read it. Excellent job!

      Thank you!

    1. Author response:

      The following is the authors’ response to the original reviews.

      Overview of Revisions

      We thank the editors and reviewers for their constructive and insightful comments, which have substantially improved the manuscript. We have carefully addressed every point raised. The major revisions include:

      (1) Methods 2.1: completely reorganized to clarify the allotetraploid genome structure of Phragmites australis, the rationale for single-chromosome-anchored microsatellite markers, the maximum distinguishable alleles per ploidy level, and the conservative Ploidies(mydata) <- 4 setting in polysat.

      (2) Methods 2.3 / Results 3.3 / Discussion 4.4: clarified common garden sample sizes, added Cohen's d effect sizes, and acknowledged the correlational nature of the lineage-level comparisons.

      (3) Introduction: added a new paragraph on the eco-evolutionary significance of gene flow in mixed-ploidy systems, and another paragraph emphasizing the novelty of integrating SDMs with physiological and common garden experiments.

      (4) Discussion 4.1 and 4.4: reframed all "polyploidy-driven" language to "polyploidy-associated", explicitly acknowledging that ploidy is confounded with genetic background and was not experimentally manipulated.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1-P1) Inadequate explanation of allele dosage for ploidy levels

      Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      We appreciate this comment and have substantially revised Methods 2.1. The key clarifications are:

      (A) Marker specificity: All 42 microsatellite markers were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2). This confirms that each marker amplifies a locus specific to one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from one subgenome), while in octoploids (autopolyploid derivatives with four copies of the same chromosome) each marker detects up to four alleles.

      (B) Conservative ploidy setting: We set Ploidies(mydata) <- 4 for all samples in polysat because the exact ploidy of many samples could not be confidently assigned a priori. This uniform treatment is conservative: for actual tetraploids, the two unobserved "copies" are scored as null; for actual octoploids, all four detected alleles are accommodated.

      (C) Dosage estimation: Allele dosage was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), which applies stutter correction, amplification bias correction, and ploidy-optimized dosage calling, not inferred from allele presence/absence alone.

      Methods 2.1, second paragraph onward

      “Phragmites australis has a base allotetraploid genome. All 42 microsatellite markers used in this study were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2), confirming that each marker amplifies from one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from the target subgenome) while the homologous region from the other subgenome is not amplified (Saltonstall, 2003). In Asia, the prevalent octoploids are most likely autopolyploid derivatives of tetraploids, carrying four homologous copies of the same chromosome and thus capable of up to four distinguishable alleles per locus (Liu et al., 2022; Wang et al., 2024). Hexaploid individuals are rare and occur primarily in contact zones, likely originating from inter‑lineage hybridization (Wang et al., 2024).

      In practice, the 42 selected markers very rarely produced more than four alleles in any single individual (Table S2), consistent with a ploidy ceiling of octoploid. Because the exact ploidy of many samples could not be confidently assigned a priori (ploidy was inferred from a combination of chloroplast haplotype, geographic origin, and flow cytometry from prior studies; Lambertini et al., 2020; Liu et al., 2022), we consistently set Ploidies(mydata) <- 4 in the polysat R package (Clark & Jasieniuk, 2011), treating every individual as having four homologous copies. This uniform treatment is conservative: for an actual tetraploid (two copies per locus), the two unobserved "copies" are simply scored as null (missing data) in the dosage matrix; for an actual octoploid, all four detected alleles are accommodated. The allele dosage itself was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), not inferred from allele counts alone.”

      (R1-P2) Common garden setup and sample sizes unclear

      The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      We agree that the original description was insufficiently detailed. We have made three modifications:

      (A) Methods 2.3: clarified that each population contributed one rhizome segment planted in one pot (one biological replicate per population per site), and emphasized that the key inference rests on the direction and consistency of differences across four climatically distinct sites rather than on significance at any single site.

      (B) Results 3.3: added Cohen's d effect sizes (verified from the raw data), showing that where lineage differences are present, they are biologically substantial (d = 1.10–1.43 at three of four sites).

      (C) Discussion 4.4: added a paragraph acknowledging the limited number of populations per lineage and the need for future confirmation with a larger panel.

      Methods 2.3

      “(1) A previously published common garden experiment (Song et al., 2021) conducted in 2017 across Jinan (36.43°N, 117.45°E) and Panjin (41.20°N, 122.02°E), using CN (n = 11) and FEAU (n = 9) lineages, for which we determined the haplotype information of all samples; (2) A new common garden experiment established in 2021 across Qingdao (36.36°N, 120.69°E) and Shanghai (30.20°N, 121.29°E), with CN (n = 9) and FEAU (n = 8) lineages (Table S3). Each rhizome segment (2–3 buds per segment, one segment per population) was transplanted into an individual 20 L pot, yielding one biological replicate per population per site. Although the number of populations per lineage is modest, the key inference rests on the direction and consistency of lineage differences across four climatically distinct sites rather than on the statistical significance at any single site.”

      Results 3.3

      “Similarly, plant height was significantly greater in the FEAU lineage than in the CN lineage in Jinan, but not in the other common gardens (Figure 3C). Effect sizes for total biomass were large in Jinan (Cohen's d = 1.10), Panjin (d = 1.12), and Qingdao (d = 1.43), but negligible in Shanghai (d = 0.37), confirming that the lineage differences, where present, are biologically substantial.”

      Discussion 4.4

      “We also acknowledge that the common garden experiments, while replicated across four climatically distinct sites, involved a limited number of populations per lineage (9–11 CN and 8–9 FEAU), which constrains our ability to fully separate lineage-level effects from population-level variation. The consistent direction of biomass differences across three of four sites, supported by large effect sizes, nonetheless provides robust evidence for a lineage-level performance advantage that merits further confirmation with a larger, more geographically representative panel of populations.”

      (R1-P3) How allele dosage is determined

      Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

      We agree that a self-contained description is warranted. We have rewritten the relevant paragraph in Methods 2.1 to describe the three core steps of the SSRSeq V1.1 pipeline (Cui et al., 2022): (i) stutter correction based on empirically estimated slip ratios; (ii) amplification bias correction across alleles of different repeat lengths; and (iii) ploidy-adjusted allele dosage calling. The pipeline's source code and full documentation are available at https://github.com/ccoo22/SSRseq_count.

      Methods 2.1

      “Microsatellite genotyping was performed using the SSRSeq V1.1 pipeline (Cui et al., 2022; https://github.com/ccoo22/SSRseq_count). Briefly, the pipeline takes the per-locus per-sample read count table generated from high-throughput sequencing and processes it through three core steps fully described in Cui et al. (2022): (i) stutter correction, which reallocates a fraction of reads from each allele to its adjacent repeat class based on empirically estimated slip ratios; (ii) amplification bias correction, which normalizes read counts across alleles of different repeat lengths using locus-specific bias coefficients; and (iii) allele dosage calling, which selects the maximum number of alleles consistent with the specified ploidy (four in this study) and assigns integer dosages (0–4) by comparing corrected read ratios to a ploidy-adjusted threshold optimized to minimize both allelic dropout and false positives. The final output is a genotype matrix with integer allele dosages for all samples and loci, which was used directly as input to the polysat R package for subsequent population genetic analyses.”

      Reviewer #2 (Public review):

      (R2-P1) Polyploidy has no causal evidence; confounded with genetic background

      First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted.

      We fully acknowledge this critical limitation and thank the reviewer for this important critique. We have revised the manuscript at four locations to reframe all claims from "polyploidy-driven" to "polyploidy-associated" and to explicitly state that ploidy is confounded with lineage identity and was not experimentally manipulated.

      Abstract (last sentence)

      “These results demonstrate that climate change interacts with intraspecific variation among polyploidy-associated lineages, manifested through differences in thermal tolerance, biomass production, and asymmetric gene flow, to drive potential lineage replacement within a native range”

      Introduction (polyploidy paragraph)

      “The octoploid FEAU lineage is distinguished from its tetraploid relatives not only by ploidy level but also by its distinct evolutionary history, genomic background, and geographic origin. Polyploidy has been shown in other systems to generate genetic novelty, alter gene expression, and enhance physiological stress tolerance (Bureš et al., 2024; Cheng et al., 2021; Kolář et al., 2017; Van de Peer et al., 2017), potentially pre-equipping polyploid lineages to occupy new geographical ranges and endure environmental shifts (Cheng et al., 2021; López-Jurado et al., 2019). The FEAU lineage's superior thermal tolerance and biomass are consistent with such polyploidy-associated effects, although ploidy is correlated with, rather than experimentally separable from, the broader genetic identity of each lineage.”

      Discussion 4.1 (title and opening paragraph)

      “Our findings demonstrate that the octoploid FEAU lineage of P. australis possesses greater heat tolerance and biomass production than the tetraploid CN lineage. Under a high emission scenario (SSP5-8.5), the projected suitable habitat for the FEAU lineage expands by 18.6%, while the CN lineage exhibits a much smaller relative increase. Several non-mutually-exclusive mechanisms could explain these lineage-level differences, including increased gene dosage from whole-genome duplication, divergent selection histories, and/or standing genetic variation in thermal tolerance loci unlinked to ploidy (Bures et al., 2024; Cheng et al., 2021; Van de Peer et al., 2017). Our data cannot fully partition these factors, but the strong association between lineage identity and both physiological performance and projected range dynamics highlights the importance of incorporating intraspecific lineage information into ecological forecasts, regardless of the ultimate causal mechanism.”

      Discussion 4.4 (limitations paragraph)

      “Crucially, ploidy was not experimentally manipulated in this study; it is inherently confounded with the distinct evolutionary history and genomic background of each lineage. While the observed thermal tolerance and biomass differences are consistently associated with the octoploid FEAU lineage, we cannot formally exclude the possibility that these traits are driven by genetic factors independent of ploidy per se. Future studies using experimental approaches that can partition ploidy effects from lineage-specific genetic effects, such as common gardens with synthetic polyploids or transcriptomic analyses comparing gene expression dosage responses, are needed to strengthen causal inference (Wei et al., 2020). Similarly, the common garden results should be interpreted as lineage-associated rather than ploidy-causal performance differences. The potential role of admixture in facilitating the adaptive introgression of heat tolerance alleles also warrants deeper investigation (Suarez-Gonzalez et al., 2018).”

      (R2-P2) SDMs treat lineages as homogeneous entities

      Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models... the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species"—thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species.

      We acknowledge this important limitation and agree that it deserves explicit discussion. While disaggregating the species into three major genetic lineages is a step forward from species-as-monolith approaches, within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible given the broad geographic ranges of the CN and FEAU lineages. We have added a new paragraph in Discussion 4.4 to address this point.

      Discussion 4.4 (new paragraph)

      “We also recognize that our SDM approach, while disaggregating the species into three major genetic lineages, still treats each lineage as a homogeneous entity. This simplification parallels—albeit at a finer scale—the species-as-monolith assumption that we critique in the Introduction. Within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible, particularly given the broad geographic ranges of the CN and FEAU lineages. By modelling each lineage as a uniform group, our projections may overestimate the precision of range forecasts and underestimate the evolutionary potential of standing variation within lineages (Chardon et al., 2020). Future frameworks that incorporate trait variation at multiple hierarchical levels (population, lineage, ploidy) will be necessary to capture both the adaptive potential and the ecological constraints that shape species' responses to climate change.”

      (R2-P3) Asymmetric introgression and thermal tolerance lack context in Introduction and Discussion

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species.

      We agree that the evolutionary significance of asymmetric gene flow was underdeveloped. We have added two substantial new passages:

      (A) Introduction: a new paragraph explaining the dual role of gene flow in climate adaptation: introgression of adaptive alleles vs. asymmetric introgression as a mechanism of gradual lineage replacement. This paragraph explicitly connects genome dosage differences (octoploid vs. tetraploid) to the natural directionality of backcrossing.

      (B) Discussion 4.2: a new paragraph extending the discussion of asymmetric introgression into its long-term evolutionary consequences, including the potential erosion of the CN lineage's genetic distinctiveness and the risk of losing cold-adapted alleles under future climate volatility.

      Introduction (new paragraph)

      “Gene flow between lineages of differing ploidy can play a dual role in climate adaptation. Introgression may introduce adaptive alleles (e.g., heat tolerance loci) into a recipient lineage, facilitating its persistence under warming (Suarez-Gonzalez et al., 2018). Conversely, if introgression is asymmetric, such that one lineage's genome is disproportionately represented in admixed populations, it can drive a gradual but systematic shift in genetic composition within the contact zone—effectively functioning as a mechanism of lineage replacement without requiring complete competitive exclusion (Bartolić et al., 2024; Zohren et al., 2016). In mixed-ploidy systems, genome dosage differences create a natural directionality in backcrossing: hybrids tend to backcross more frequently with the high-ploidy parent (Bartolić et al., 2024). In the present study, we test whether such a bias exists between the octoploid FEAU and tetraploid CN lineages and examine its consequences for future distribution under climate warming.”

      Discussion 4.2 (new paragraph)

      “From an evolutionary standpoint, asymmetric introgression can erode the genetic distinctiveness of the minority lineage (CN) while enriching the majority lineage (FEAU) with alleles that may have been locally adapted in the CN genomic background. This could reduce the species' overall evolutionary potential, even if the FEAU lineage itself thrives—because cold-adapted alleles from the CN lineage, which may be valuable under future climate volatility (including extreme cold events), risk being diluted or lost (Exposito-Alonso et al., 2022). The directionality of introgression is also not fixed; it could shift if environmental conditions alter hybrid fitness or if the demographic balance between lineages changes. Long-term genomic monitoring of the CN–FEAU contact zone will be essential to determine whether the asymmetric gene flow documented here represents a transient phase or a persistent trajectory toward genomic homogenization.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      (RE-1) Explain how ploidy level is inferred

      We invite the authors to clearly explain how the level of ploidy is being inferred (Reviewer #1).

      We agree that the rationale for ploidy assignment and its relationship to allele counts needed greater clarity. This has been addressed by the comprehensive revision of Methods 2.1 described in response to R1-P1 (Part 1, Reviewer #1 Public Reviews). The revised text now presents a complete step-by-step logical chain: (i) P. australis has an allotetraploid base genome; (ii) all 42 markers map to a single unique chromosome in the reference genome, confirming that each marker amplifies from only one subgenome; (iii) tetraploids therefore show at most two distinguishable alleles per locus, while octoploids (autopolyploid derivatives of tetraploids) show up to four; (iv) because many samples lacked independent ploidy confirmation, we uniformly set Ploidies(mydata) <- 4 in polysat as a conservative treatment that accommodates both tetraploids (two observed copies + two null) and octoploids (four observed copies).

      See the full revised text under R1-P1 (Part 1) above.

      (RE-2) Common garden results are correlational

      We note that the findings of the common garden experiment, although interesting, are mostly correlational (not causative) and rely on a relatively small sample size and confound lineage isolation and adaptive differentiation (both Reviewers).

      We fully acknowledge this limitation. Because all octoploids belong to the FEAU lineage and all tetraploids to CN, ploidy and lineage identity are inherently confounded. This has been addressed by the four-part revision described in response to R2-P1 (Part 1, Reviewer #2 Public Reviews). Specifically:

      The Abstract now frames the findings as "polyploidy-associated" rather than "rooted in polyploidy."

      The Introduction now explicitly states that ploidy is correlated with—but not experimentally separable from—the broader genetic identity of each lineage.

      The Discussion 4.1 title was changed to "Polyploidy-associated thermal tolerance" and the opening paragraph now presents multiple non-mutually-exclusive mechanisms rather than asserting a causal role for polyploidy.

      The Discussion 4.4 now includes an expanded limitations paragraph acknowledging that ploidy was not experimentally manipulated and that common garden results should be interpreted as lineage-associated rather than ploidy-causal.

      In addition, the Methods 2.3 and Discussion 4.4 revisions described in response to R1-P2 (Part 1) address the sample size concern by clarifying the experimental design and adding a dedicated acknowledgement of the limited population replication.

      See the full revised text under R2-P1 and R1-P2 (Part 1) above.

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) "Large morphological traits"

      Line 85. Large morphological traits. Does this mean physically large? Or higher values of some trait.

      We agree the original phrasing was ambiguous. We have replaced "large morphological traits" with explicit trait descriptions.

      Introduction

      “The octoploid FEAU lineage exhibits greater shoot height, larger leaf size, and thicker stems (K. Chen et al., 1993; Guo et al., 2025; Liu et al., 2021b, 2026; Yin et al., 2024), along with stronger salt tolerance and higher thermal tolerance”

      (R1-R2) Why 2 allele copies expected for a tetraploid

      Line 119. Not clear why this allele copy number is expected. A tetraploid can have up to 4 unique alleles (e.g., ABCD), not two.

      This comment arises from the same conceptual gap addressed in R1-P1. The key point is that P. australis is an allotetraploid with two subgenomes, and our 42 markers each map to a single unique chromosome (one subgenome). Therefore, the marker only amplifies the two homologous copies from that subgenome, giving at most two distinguishable alleles. A true autotetraploid would indeed show up to four alleles—but that is not the genomic architecture of P. australis. The revised Methods 2.1 (see R1-P1 in Part 1) now explicitly explains this logic.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R3) Theoretical expectation of allele number vs. ploidy

      Line 127-129. As above, this is unclear and not what we expect theoretically. If there is a reasonable number of alleles, there should be a maximum of 4 for tets, 6 for hex, and 8 for octs.

      Same point as R1-R2. The reviewer's expectation (4 for tetraploids, 6 for hexaploids, 8 for octoploids) is correct for autopolyploids with markers that amplify all homologous copies. The discrepancy arises because P. australis is an allotetraploid and our markers are single-chromosome-anchored (each amplifying from only one subgenome). The revised Methods 2.1 now clarifies this distinction explicitly.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R4) Typo "makers"

      Line 136. Should be 'markers' not 'makers'

      We have performed a full-text search and corrected all instances of "makers" to "markers" in the manuscript.

      Full-text search and replace.

      (R1-R5) Reason for removing markers with >4 alleles

      Line 142. The reason for the removal of more than 4 alleles is not clear. What about hexaploids and octoploids? They can carry 6 or 8 alleles, respectively.

      We agree the original text did not adequately justify this quality-control step. In our study, the maximum expected distinguishable alleles (given the allotetraploid genome and single-chromosome-anchored markers) is two for tetraploids and four for octoploids. The observation of five or more alleles in multiple individuals is therefore diagnostic of multi-locus amplification (the marker amplifying more than one genomic locus), not of high ploidy. This is a quality-control filter, not a ploidy assignment criterion.

      Methods 2.1

      “During genotyping, eleven markers (including four of the five multi-mapping markers) were removed because more than ten samples exhibited more than four alleles per sample at these loci. Because the maximum number of distinguishable alleles expected under our ploidy model is two (tetraploid) to four (octoploid), the observation of five or more alleles in multiple individuals indicates that these markers amplify more than one genomic locus, rendering them unsuitable for dosage-based genotyping. This filtration is a quality-control step, not a ploidy assignment criterion.”

      (R1-R6) Unclear sample sizes in common garden

      Line 240. Unclear sample sizes. If these are the numbers, they are a very low level of replication expected for a common garden experiment.

      Same point as R1-P2 (Public Review). Please see the full response under R1-P2 in Part 1, where we have (A) clarified the experimental design in Methods 2.3, (B) added Cohen's d effect sizes in Results 3.3, and (C) acknowledged the sample size limitation in Discussion 4.4.

      Fully addressed by the three-part revision in R1-P2.

      Reviewer #2 (Recommendations for the authors):

      (R2-A1) De-emphasize polyploidy

      De-emphasize polyploidy, as it's not manipulated in the study and is entirely confounded with the genotypes of the distinct lineages, and the putative links between polyploidy and heat tolerance are circumstantial and lacking in a clear mechanism.

      We agree fully. This has been addressed comprehensively across four locations in the manuscript (Abstract, Introduction, Discussion 4.1, Discussion 4.4). See the full response under R2-P1 (Part 1) for the revised text at each location.

      Fully addressed by the four-part revision in R2-P1.

      (R2-A2) Emphasize the novelty of combining SDMs with experiments

      Emphasize the novelty of combining SDMs with experiments (or, if I'm not up on the literature and they are more common, explain how they have been used to make new insights).

      We appreciate this suggestion and agree that explicitly stating the novelty of our integrative approach strengthens the manuscript. We have added a new paragraph at the end of the Introduction.

      Introduction (end, before "Here, we integrate population genomics…")

      “Studies that combine species distribution models with physiological or common garden experiments remain surprisingly uncommon (but see López-Jurado et al., 2019). Such integration is essential for transforming correlative SDM projections into mechanistically grounded predictions. In the present study, we adopt this integrative approach: common garden experiments directly test growth performance under controlled conditions, heat-tolerance measurements identify the specific physiological thresholds (T<sub>crit</sub>, T<sub>50</sub>) underlying lineage-specific climate responses, and SDMs project how these experimentally documented differences translate into spatial dynamics under future warming. By linking experimental data with spatial forecasting, we move beyond correlative climate matching toward a trait-based understanding of how intraspecific variation shapes species' future distributions.”

      (R2-A3) Elaborate on the importance of gene flow

      Elaborate on the importance of gene flow and the potential connections between gene flow and evolving species (or lineage) geographical limits.

      This has been addressed by the two new paragraphs described under R2-P3 (Part 1)—one in the Introduction on the dual role of gene flow in climate adaptation (introgression of adaptive alleles vs. asymmetric introgression as a mechanism of lineage replacement), and one in Discussion 4.2 on the long-term evolutionary consequences of asymmetric introgression.

      Fully addressed by the two-part revision in R2-P3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Many thanks to the three reviewers and the editors for their thoughtful comments and careful evaluation of our manuscript. These are fair, consistent and largely expected comments. On behalf of my co-authors, we provide this response to the public reviews to summarize the main issues raised and the corresponding revisions we have made in the revised manuscript.

      (1) The main consistent comment from all three referees was that our single-nucleus RNA-seq data should be further validated. The reviewers differ in the detail of exactly what they think should be validated, but collectively these comments referred to validation of: (1) the identified cell types, (2) pathways inferred from trajectory analysis, (3) differentially expressed genes between plucked and control conditions across the four sampled time points, and/or (4) inferred ligand–receptor pairs from the cell–cell communication analysis.

      We believe that we are on strong footing for some of these points because of extensive work we’ve done in the past in the cichlid fish model.

      In the references cited in the manuscript and highlighted below (References 1, 10, 11, 29, 30, 31), we tally 29 figures with 273 individual figure panels presenting histology, in situ hybridization, and immunohistochemistry featuring genes expressed in cichlid (replacement) teeth. Most of these genes are markers of dental competency and/or indicative of regenerative potential.

      In addition, in multiple of these papers, we use pharmacology to manipulate the role of key pathways (Hh, BMP, Wnt, Notch) in cichlid tooth development and replacement. Validation of cell types in the present study therefore draws on these published data in cichlids (and other vertebrates), as well as on an unbiased comparative approach, SAMap, which identifies homology between cichlid and mouse dental cell types based on shared gene expression.

      In short, experiments to validate cell types and pathways active in cichlid teeth have been published and are referenced herein. We recognized, however, that these references (some of which include Gareth Fraser as an author, when he was a postdoc in my group; for Reviewer 2) were cited primarily in the Introduction, rather than in the Rationale/Methods or Results sections. We have therefore clarified these connections in the revised manuscript (line 173-74).

      We have not validated nor analyzed functionally the ligand-receptor pairs we inferred from cell-cell communication analysis. This work is beyond the scope of the current paper, and we now state more clearly that these inferences represent hypotheses to be tested in future studies, although many of these ligand–receptor pairs have been noted in other tooth-related publications cited in the manuscript.

      (2) The biggest weakness of our manuscript, noted by referees, is that we do not provide serial histology to accompany our snRNA-seq time course after plucking. We previously described this as a limitation in the “Study limitations and future direction” section of the Discussion, but we have now strengthened this discussion. In particular, we more explicitly acknowledge that we do not directly document the histological progression of tissue responses across the plucking time course or the degree of tissue damage caused by the plucking paradigm at each sampled time point.

      In the “study limitations” section, we note both issues 1 and 2 and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      (3) Reviewers also asked about the presence and interpretation of stromal cells in our snRNA-seq data. In response, we re-examined the mesenchymal compartment and added additional analyses to better characterize stromal/mesenchymal populations and their inferred trajectories in the revised manuscript. This includes a revised Figure 4, revised text around Figure 4 and revised Supplementary Figures.

      (4) Multiple (minor) suggestions for clarification in text and figures have been adopted throughout the revised manuscript, figure legends, and supplemental materials.

      Overall, we do not anticipate that further reviewer engagement will be necessary, and we believe that editorial review of the revised manuscript should be sufficient.

      References cited in the manuscript, highlighted here:

      (1) Fraser, G. J. et al. An Ancient Gene Network Is Co-opted for Teeth on Old and New Jaws. PLoS Biol. 7, e1000031 (2009).

      (10) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. Common developmental pathways link tooth shape to regeneration. Dev. Biol. 377, 399–414 (2013).

      (11) Bloomquist, R. F. et al. Developmental plasticity of epithelial stem cells in tooth and taste bud renewal. Proc. Natl. Acad. Sci. 116, 17858–17866 (2019).

      (29) Streelman, J. T., Webb, J. F., Albertson, R. C. & Kocher, T. D. The cusp of evolution and development: a model of cichlid tooth shape diversity. Evol. Dev. 5, 600–608 (2003).

      (30) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. A periodic pattern generator for dental diversity. BMC Biol. 6, 32 (2008).

      (31) Bloomquist, R. F. et al. Coevolutionary patterning of teeth and taste buds. Proc. Natl. Acad. Sci. 112, (2015).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell-type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Weaknesses:

      The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data.

      We thank Reviewer #1 for the thoughtful and positive assessment of our study, including the recognition that our snRNA-seq time course provides insight into cellular trajectories, cell-type-specific gene expression, and inferred intercellular signaling during accelerated tooth replacement in cichlid fish. We also appreciate the reviewer’s central concern that the manuscript would be strengthened by additional experimental validation of the single-nucleus RNA-seq data.

      As summarized above and discussed in more detail in our point-by-point responses below, we have clarified how the present cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. We have also added analyses demonstrating reproducibility across biological test subjects and consistency of representative differentially expressed genes between paired plucked and control samples. Finally, we have strengthened the Study Limitations section to more clearly state that future spatial transcriptomic, histological, and functional validation experiments will be important next steps.

      Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues studied the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single-nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data, including the inference of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear takeaway message for biologists not necessarily interested in tooth replacement.

      Conclusion:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g., referring to "cell types" as "putative cell types").

      The work should have a high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow-up studies to further study this process. The work could also have a strong impact through the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish, such as zebrafish, to test function during tooth replacement.

      The work is unique and interdisciplinary, and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic and environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

      We thank Reviewer #2 for the careful and constructive evaluation of our manuscript and for highlighting the novelty of the cichlid tooth-plucking paradigm, the temporal design of the snRNA-seq experiment, and the computational analyses used to infer cellular trajectories and signaling interactions during accelerated tooth replacement. We also appreciate the reviewer’s comments regarding validation of inferred cell types and interpretation of signaling pathways.

      In response, we have revised the manuscript to clarify that our cell-type annotations are supported by marker-gene expression, previously published cichlid tooth studies, and an unbiased comparative approach, SAMap, which relates cichlid and mouse dental cell types based on shared gene-expression structure. We have also clarified that inferred ligand-receptor interactions represent computational hypotheses rather than functionally validated mechanisms. In addition, we revised Figure 6 and Figure 7A to improve the readability and interpretation of inferred signaling results, and we edited the relevant text and figure legends to make these results easier to follow. These points are addressed in greater detail in the point-by-point responses below.

      Reviewer #3 (Public review):

      Summary:

      This is an interesting paper. The process of tooth exfoliation and replacement in vertebrates remains an intriguing and fascinating subject of inquiry. As the scientists noted, there are no mammalian models that can be used to examine signaling pathways in real time.

      Strengths:

      This work integrates in vivo and high-resolution transcriptomics. The study confirms previous findings and emphasizes the need for additional research into the processes that drive the restoration of missing teeth for future therapeutic uses.

      Weaknesses:

      I disagree with the use of the phrase "plucking". Instead, the authors use tooth extraction or tooth removal, which is clinically more correct for the procedure they are doing.

      The inspiration for our ‘plucking’ experiment is work done in the hair follicle model (lines 73-74). Because cichlid teeth are so numerous, are very small, and lack dental roots, this is an accurate description of the procedure. We opt to retain the phrasing.

      The title is rather broad and appears to be more appropriate for a review than an original research work. I would advise specifying the species under research and/or the sort of damage model used in the transcriptome analysis.

      We opt to retain the title.

      It's uncertain whether the findings are exclusively based on regeneration. The presence of tooth remnants, as well as unintended harm to surrounding tissues, may have triggered repair mechanisms, thereby biasing the current data. How did the authors handle this issue? The oral cavity was under severe manipulation, increasing the inflammatory stimuli, a situation that does not take place in physiological exfoliation.

      In the revised manuscript, we have more clearly acknowledged that our plucking paradigm may induce tissue damage and repair-associated responses in addition to accelerated tooth replacement. We have strengthened the Study Limitations section to state that we do not directly document the histological progression of tissue responses across the time course or the degree of damage caused by plucking at each sampled time point. One caveat, however, is that bone remodeling and immune response is likely triggered on the ‘control’ side of the jaw also, just not to the same degree as after plucking.

      The authors indicated the use of microCT analysis; however, no such information appears in the main text. In fact, this manuscript lacks anatomical information. It is required to conduct histological examinations of the regenerated teeth at various time points.

      microCT data were included as a Supplemental Figure to demonstrate the dental formulae of our chosen species; but we did not characterize post-plucking recovery using this technique (see above summary and below point-by-point comments).

      Although the current findings confirm previously found and verified signaling pathways, the absence of functional data lends uniqueness to this work.

      In the revised manuscript, we also clarify that, while our transcriptomic analyses identify candidate cell states, pathways, and signaling interactions associated with accelerated replacement, the functional roles of these inferred pathways remain to be verified in future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Figure 1 should include representative H&E staining images comparing the left control side and the regenerated side at 7 days post-plucking. This would provide important histological context for the regeneration process and help readers better interpret the molecular findings.

      This would indeed be valuable information, but we did not carry out histology of paired control vs plucked jaws to accompany our pulse-chase and dissections for single-nucleus isolation. This comment is similar to that below about validation of what is happening on plucked vs. control jaw halves and is the first limitation we discuss in the “study limitations and future directions” section (from line 538).

      (2) Each tooth position consists of a functional tooth, a replacement tooth, and the dental (successional) lamina. On the control side, the successional lamina contains teeth at different developmental stages, analogous to the mammalian bud, cap, and bell stages. Can the snRNA-seq analysis distinguish among tooth families at different developmental stages, as well as the individual components within a single tooth unit? Clarification of this point would enhance the developmental interpretation of the dataset.

      No, our approach does not distinguish among teeth at different stages, nor among teeth in even vs odd positions that tend to be synchronized in replacement cycles. Theoretically, one could do this by dissecting individual teeth and pooling by tooth stage, but we did not.

      (3) Identifying successional lamina cells is critical, and the authors report putative SL cells within the VEE cluster. However, the stromal cells surrounding the successional lamina are also known to play important roles in tooth regeneration. Can the authors further annotate and characterize stromal cell populations in the snRNA-seq dataset? Additional analysis of these supporting cells would strengthen the conclusions regarding epithelial-mesenchymal interactions.

      We thank the reviewer for this insightful suggestion. In response, we further characterized mesenchymal subpopulations and included these new analyses in the revised manuscript (updated Figure 4 and Supplemental Figure 8). Specifically, pseudotime and CellRank analyses identified a mesenchymal subpopulation enriched for Twist1, Dnmt1, and Runx2, which we interpret as putative dental ectomesenchyme (DEM) based on the established roles of these genes in odontogenic mesenchymal development and differentiation, as well as their reported expression in mouse and human tooth single-cell transcriptomic studies. Notably, this putative dental ectomesenchymal population resides within the broader dental follicle compartment identified in our dataset (Figure 4A-C).

      To further assess supporting stromal populations, we examined the expression of established stromal marker genes, including Lum, Col6a3, Aspn, and Vegfc. These markers were broadly restricted to mesenchymal populations and showed strong enrichment overlapping the newly identified putative dental ectomesenchymal region (Supplemental Figure 8B). Consistent with these observations, differential expression analysis identified additional DEM-enriched genes that substantially overlap canonical stromal markers, including extracellular matrix-associated genes, supporting a close transcriptional relationship between the putative dental ectomesenchyme and the surrounding stromal mesenchymal compartment (Supplemental Figure 8C). Together, these findings refine the mesenchymal landscape surrounding the putative successional lamina and support the presence of a specialized stromal microenvironment associated with tooth regeneration.

      Consistent with this interpretation, our CellChat analysis identified significantly increased interactions between the dental ectomesenchyme and cycling ameloblast populations on the plucked side at Day 0 (Supplemental Figure 8D). These interactions were enriched for signaling pathways including SEMA4, EPHB, SLIT and SPP1, all of which have established roles in tissue remodeling, extracellular matrix organization, and regenerative processes. (Supplemental Figure 8E). Because Day 0 contained sufficient biological replicates and cell numbers for robust statistical comparison, we focused our interaction analyses on this time point. Collectively, these additional analyses provide a more comprehensive characterization of the stromal compartment and further support the conclusion that a specialized dental ectomesenchymal population actively participates in epithelial-mesenchymal communication during the earliest stages of tooth regeneration. So, in total, Figure 4 was revised, the text on lines 281-309 was revised, and Supplemental Figure 8 was added.

      (4) The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data. The authors should validate the expression of major signature genes using RNAscope or immunostaining, ideally comparing regenerated samples with the left-side control. Such validation would significantly enhance the robustness of the conclusions.

      We did not validate up- or down-regulation of differentially expressed genes in intact tissue, owing in part to (1) the complexity of this experiment, (2) the fact that the majority of DEGs, or ‘major signature genes’ have been observed to be expressed in dentitions generally, and often by us in previous work on cichlid teeth, and the fact that (3) independent biological replicates were strongly consistent in the direction of effects (see below). In the “study limitations” section, we note this issue and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      Minor Points

      (1) In Figure 1, the color scheme used in the schematic drawing (Figure 1A) should match the corresponding structures shown in Figure 1B to improve clarity and consistency.

      We appreciate the reviewer’s thoughtful suggestion regarding the color consistency between the schematic (Figure 1A) and the fluorescence images (Figure 1B). However, the color scheme in the schematic (Figure 1A) was intentionally selected to maximize accessibility, particularly for readers with color vision deficiencies, and therefore differs from the magenta and green fluorescence channels used in Figure 1B. In the fluorescence images, the magenta and green colors reflect the native display colors used for the Alizarin Red and Calcein labeling channels in the pulse-chase experiment. Directly matching the schematic colors to the fluorescence images could reduce the visual contrast between key anatomical structures and compromise accessibility for some readers. We have therefore retained the current color scheme in Figure 1A while ensuring that the corresponding structures are clearly identified through consistent labels and annotations across both panels.

      (2) The abbreviation for successional lamina (SL) should be defined upon first use in the Introduction.

      We thank the reviewer for catching this omission. We have now defined the abbreviation “successional lamina (SL)” upon its first appearance in the Introduction.

      (3) Regarding biological replicates, the authors should provide data demonstrating the consistency and reproducibility across replicated samples.

      We thank the reviewer for this suggestion. To demonstrate the consistency and reproducibility across biological test subjects, we have added analyses summarizing sequencing quality metrics, test subject contributions, integrated clustering, and representative differential gene expression across individual samples (see Figure S4, panels C, D & E and Author response image 1). Panel A shows that nuclei from different biological test subjects are well integrated across clusters rather than segregating by sample origin. Finally, Panel B presents representative differentially expressed genes from multiple cell populations, demonstrating consistent expression differences between paired plucked and control samples across biological test subjects.

      Author response image 1.

      (A) UMAP embedding of dental nuclei. Each point represents a single nucleus, colored by test subject. (B) Representative differentially expressed genes show consistent expression differences between plucked and control samples across biological replicates. Paired boxplots of average gene expression for representative differentially expressed genes from multiple cell populations at Days 0, 1, 3, and 7. Each point represents one biological replicate (test subject), with paired plucked and control samples connected by dashed lines. The y-axis shows average gene expression, and the x-axis indicates the experimental condition. These representative examples illustrate the consistent direction of differential expression across biological replicates, supporting the reproducibility of the single-nucleus RNA-seq dataset.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1: Can the panels to the right of panel B be labeled? It's not clear what these six images are showing, so giving them letters and explaining briefly in the legend what the point of each panel is would clarify. "Right, example of individually classified teeth" - can the authors elaborate on what each tooth is an example of (i.e., how each tooth shown was classified"?) For clarity, the graphs in panels C and D should have the y-axes labeled

      We thank the reviewer for this helpful suggestion. In response, we revised the Figure 1B legend to clarify the classification criteria used for dye incorporation analyses and to better describe the representative fluorescence images. Specifically, teeth positive for both Alizarin and Calcein were classified as pre-existing old teeth, whereas teeth positive only for Calcein were classified as newly formed teeth. We additionally clarified that the images to the right of panel B show representative individually classified teeth, with the top row representing pre-existing old teeth and the bottom row representing newly formed teeth. We also added y-axis labels to panels C and D to improve figure clarity and readability.

      (2) Figure 2 legend: should "the cell type" instead be "the putative cell type"? Without validation for all cell types, it seems adding some sort of qualifier is in order here. Can the authors comment further on examples of validation from other studies? For example, Gareth Fraser has published numerous studies that show Pitx2 expression marking dental epithelium in different fish, yet none of these older papers are cited.

      Identification and validation of cell types make use of multiple published datasets in cichlids (for markers matched to mouse), as well as an unbiased computational approach (SAMap) that draws homology between cichlid and mouse dental cell types, based on shared global patterns of gene expression. There is perhaps a philosophical debate to be had about the validity of ‘cell types,’ generally, but our data are validated using two methods. We edited the text in lines 167-177 to clarify, including citing references to our own work (these studies include Gareth Fraser as an author, when he was a postdoc with Streelman).

      (3) Figure 6 is extremely complicated. Can any portions of rows or columns in these tables be highlighted in the figure to help the reader follow the proposed signaling interactions highlighted in the text?

      We thank the reviewer for this helpful suggestion. To improve the readability of Figure 6 and better guide readers through the dynamic signaling patterns described in the text, we revised the figure by visually highlighting the key sender-receiver interaction regions discussed in the Results. Specifically, we annotated the interactions involving mesenchymal subpopulations and alveolar bone (OST) signaling toward CYC-AMB at Days 0 and 7, mesenchymal signaling toward NK/T cells at Day 1, and epithelial cross-talk centred around ES-2 at Day 3. These visual annotations allow readers to more readily identify the signaling interactions highlighted in the text and relate them to the corresponding regions of the interaction heatmaps.

      (4) In Figure 7A, what does the black font indicate (if grey is up in control and red is up in plucked)? I'd guess not up in either, which then makes it unclear whether the sets in black are different or why they are being presented.

      We thank the reviewer for pointing out this ambiguity. In Figure 7A, blue and red labels indicate signaling pathways identified by CellChat as condition-specific, with blue representing pathways detected only in the control condition and red representing pathways detected only in the plucked condition. In contrast, pathways shown in black represent signaling pathways detected in both conditions but exhibiting significant differences in inferred communication probability between conditions. Thus, the black labels denote shared signaling pathways whose activity differs significantly between control and plucked samples, rather than pathways unique to either condition. We have revised the figure legend to clarify this distinction and improve interpretability.

      Reviewer #3 (Recommendations for the authors):

      (1) I encourage the authors to offer information on the histological differences between teeth during physiological and accelerated replacement. I'm curious if the eruption's accelerated rate has any effect on the mineralization of those teeth.

      We did not examine the histology of individual teeth, and so can’t comment on differences in mineralization.

      (2) The findings section contains multiple sentences that should be moved under material and techniques.

      We expect the reviewer is referring to paragraph lines 104-114, which was a tricky paragraph to place in the manuscript. In the end, we believe it represents important context necessary to interpret findings (which could be missed if moved to ‘methods’) and so we’ve chosen to keep this paragraph in its place.

      (3) It would be useful to include a table showing sample distribution by experimental design.

      We thank the reviewer for this suggestion. Sample distributions across experimental conditions, time points, biological test subjects, and identified cell populations are already provided in Supplementary Table 1. To improve clarity and accessibility, we have revised the table legend to more explicitly describe the experimental design and sample annotations represented in the table.

      (4) The writers did a nice job with the graphics in Figure 8; however, the schematics in Figure C are difficult to follow and are not adequately discussed anywhere. Please note that this text may be of great interest to the dentistry community, including clinicians, and that a clear and succinct explanation of the schemes at the end would be quite beneficial.

      We thank the reviewer for this helpful suggestion. We have revised the Figure 8 legend to more clearly explain Panel C as a summary schematic of inferred cell–cell communication events associated with accelerated tooth replacement after plucking. The updated legend clarifies that the pathway labels in Figure 8C summarize results directly from Figure 7A: red pathway labels indicate plucked-only signaling events, corresponding to pathways shown as full red bars in Figure 7A, while black pathway labels indicate signaling interactions detected in both plucked and control conditions but showing significant differences in interaction probability between conditions. Panel C also includes a cell-type legend at the bottom to identify the relevant cell populations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related, either singly or in combination, tyrosine phosphatases in this conditional, stage development gene knockout. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Public Reviews:

      Reviewer #1 (Public review):

      Barré et al. investigated the role of Shp1 and Shp2 in megakaryocytes (MKs) and platelets by conditional knock-out of Shp1, Shp2, or both under the control of the Gp1ba promoter. Deletion of Shp1 and Shp2 in MKs and platelets was almost complete. The Shp1/Shp2 double knock-out mice displayed macrothrombocytopenia and increased bleeding, whereas the single knock-outs did not show significant defects. Platelet function was aberrant in DKOs, but not in single knock-outs, and so was ligand-induced signaling, particularly Syk phosphorylation.

      Megakaryocyte maturation was impaired in Shp1/Shp2 DKO mice. Ligand-induced signaling was impaired in Shp2 knock-out and DKO. Ex vivo formation of platelets and in vivo maturation of MKs were impaired in DKO mice. Pharmacological inhibitors of Shp1 and Shp2 had largely similar effects as observed in the single knock-outs. The authors conclude that Shp1 and Shp2 have synergistic functions in the MK/platelet lineage, and that Shp2 may be a potential therapeutic target in myeloproliferative neoplasms.

      Strengths:

      The data clearly show effects of the Shp1/Shp2 double knock-out on MKs and platelets.

      Weaknesses:

      There appears to be a discrepancy between the results with the Shp2 single knock-out and the Shp2 inhibitor: the Shp2 knock-out does not affect MKs and platelets, except Erk1/2 signaling, whereas the Shp2 inhibitors appear to affect MK function.

      This work is interesting and may have potential from a therapeutic point of view.

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

      We thank the reviewer for recognizing the therapeutic potential of our findings.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Barré et al. investigate the roles of the phosphatases Shp1 and Shp2 in the megakaryocyte and platelet lineage using genetic depletion in mice. By employing Gp1ba-Cre-based models, the study builds on the authors' previous work and addresses some limitations associated with earlier Pf4-Cre approaches. The authors report relatively mild alterations in megakaryocyte and platelet parameters in mice lacking either Shp1 or Shp2 alone, whereas combined deletion of both phosphatases results in macrothrombocytopenia, mild bleeding, and impaired GPVI-dependent platelet aggregation accompanied by reduced Syk phosphorylation. The functional platelet defects are linked to reduced expression of GPVI and integrin α2, while thrombocytopenia is associated with impaired megakaryocyte maturation, reduced ploidy, defective proplatelet formation, and altered TPO-dependent Ras/MAPK signaling. Similar effects on megakaryopoiesis are also observed in vitro following treatment with newly developed Shp2 inhibitors.

      Strengths and Weaknesses:

      The study addresses an important biological question and presents a substantial dataset that could contribute to a better understanding of Shp1 and Shp2 function in platelet biology. However, several aspects of data presentation and interpretation would benefit from additional clarification. In particular, while the authors conclude that single genetic deletion or pharmacological inhibition of Shp1 has a limited impact and that the major phenotypes are specific to combined Shp1/2 deletion or Shp2 inhibition, some of the data suggest more nuanced effects that may warrant further discussion.

      We thank the reviewer for raising this point. The manuscript is being revised accordingly, including highlighting the potential role of Shp1 in megakaryopoiesis and thrombopoiesis under steady-state and stressed conditions, requiring more detailed investigation.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased bleeding primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Mpl signaling and downstream Ras/MAPK pathways, including ERK1/2. Given the key role of these pathways in human diseases such as myeloproliferative neoplasms and the challenges associated with modulating such a central pathway, identification of a specific regulator of Mpl signaling poses intriguing questions for future studies on clinical applicability.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1baCre as opposed to the Pf4-Cre system and demonstrate both the dose- and stagedependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from TPO-mediated signaling.

      Weaknesses:

      While the experiments have been thoughtfully designed and carried out, there is limited background explanation on relatively complex or niche pathways/mechanisms, such as the relationship between P-selectin, CRP, and PAR4p; the interactions between SFK, Syk, GPVI, and CLEC-2; and TPO, MPL, ERK1/2, AKT, and STAT3, which, while likely intuitive to experts in their respective fields, may be less obvious to a reader approaching this manuscript with a global interest in megakaryopoiesis/thrombopoiesis and thus detract from the impact of the findings.

      We thank the reviewer for raising this point. The manuscript is being revised to better explain the rationale and molecular mechanisms linking these pathways and functions.

      With regard to the science itself, some of the conclusions feel premature based on the available data.

      (1) The section "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets" is challenging to follow for those not well-versed in ITAM signaling and associated pathways, and may take additional outside reading to follow the conclusion that Syk-dependent signaling is modulated downstream of GPVI and CLEC-2 based on lack of change in Src p-Tyr418, especially considering that Src p-Tyr418 was previously introduced as a measure of SFK rather than Syk. In the introduction, Shp1 is specifically mentioned as a negative regulator of the ITAM/Syk/phospholipase pathway. However, in Figure 4Ai and Bi, Syk phosphorylation/activation in Shp1 knockout cells did not appear to be different from Shp2 knockout cells, and is lower than the control, which is surprising for a negative regulator. It is also not clear why, in the section (Figure 4A-B), there is reduced Syk activation in Shp1 and Shp2 single knockout cells upon CLEC2 stimulation (but apparently not with CRP) when there was no difference in response to CLEC2 (but a difference in response to CRP) in the previous section (Figure 3A, C).

      We thank the reviewer for raising these important points. The manuscript is being revised accordingly, including clarifying the roles of SFKs, Shp1 and Shp2 in the ITAM-Syk-PLCγ2 signaling pathway.

      Briefly, SFKs are essential for phosphorylating ITAMs, allowing SH2-dependent docking of Syk. Reduced reactivity of Shp1/2 DKO platelets to CRP and collagen is likely due to downregulation of the ITAM-containing GPVI-FcR γ-chain complex and integrin α2 subunit, and concomitant reduction in Syk phosphorylation.

      However, the marginal albeit significant reduction in Syk phosphorylation downstream of CLEC-2 in Shp1 and Shp2 KO platelets was not determined and was insufficient to impact CLEC-2-mediated platelet aggregation under the conditions tested.

      Differences in the stoichiometry and docking of Syk to phosphorylated GPVI-FcR γ-chain and CLEC-2 likely contribute to the differences in platelet reactivity and Syk phosphorylation downstream of the two receptors in the absence of Shp1 and Shp2.

      (2) In the section "Reduced Tpo signaling in Shp1/2-deficient MKs," only Western blot data for (p)ERK1/2, AKT, and STAT3 are presented before concluding that decreased ERK1/2 activity is a mechanistic explanation for thrombocytopenia seen in the Shp1/2 doubleknockout condition. Such a statement would benefit from additional experiments, such as protein or transcriptional levels of ERK1/2 targets specifically relevant to megakaryopoiesis, such as ETS, FOS, and JUN, to assess the consequences of decreased phosphorylated ERK1/2.

      We thank the reviewers for these constructive comments. Further experiments are being planned to determine the biological and transcriptional consequences of reduced ERK1/2 phosphorylation during megakaryopoiesis and thrombopoiesis.

      (3) Suggesting that "inhibiting Shp2 will not have any bleeding consequence in patients" and that Shp2 may be a therapeutic target in myeloproliferative neoplasms when none of these studies have been carried out in a human model is a bold conclusion. There are no data presented on, for example, whether Shp2 inhibition can help reverse the MPL/JAK/STAT pathway in the setting of gain-of-function mutations specifically associated with myeloproliferative neoplasms.

      This conclusion is being tempered in the revised manuscript. Genetic- and pharmacological-based approaches will be used to establish the therapeutic potential of inhibiting Shp1 and Shp2 in mouse models of MPN, including Jak2 gain-of-function mice. Bleeding and thrombotic complications of inhibiting Shp1 and Shp2 will be explored as part of these studies.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Altogether, we feel that this is an important study for those in the fields of hematology or signal transduction. Your important study characterizes the roles in late megakaryopoiesis and platelet biogenesis of single or combined conditional deletion of two tyrosine phosphatases, Shp1 and Shp2. Strengths include technical advances in single and combined deletions, the somewhat surprising results of synergy between the two phosphatases, focusing on the critical stage of late megakaryopoiesis, and clinical implications in bleeding diathesis.

      Weaknesses are mostly minor, but the numerous points raised by reviewer 3 need to be addressed and typographical errors corrected. Further discussion should include the relevance or dissimilarity in megakaryopoiesis and platelet biogenesis between murine and human blood health and disease. Since SHP1 is a negative regulator and SHP2 is a positive activator, additional discussion about how they coordinate and fine-tune ("nuanced") signal transduction in TPO- or GPVI-induced signaling in an explicitly stated pathway.

      We invite you to respond to the critiques and submit a revised manuscript.

      Sincerely,

      Seth Corey, MD MPH

      We thank the editor for the positive evaluation of our study and for highlighting its relevance to the fields of haematology and signal transduction. We have carefully addressed all comments raised by Reviewer 3 and corrected typographical errors throughout the manuscript.

      As suggested, we expanded the Discussion to better address the relevance of our murine findings to human megakaryopoiesis and platelet biogenesis. While our study relies on mouse models, key components of TPO/MPL signaling and platelet production are conserved between mice and humans, although differences in megakaryocyte maturation dynamics and platelet biology are acknowledged and now discussed.

      We also clarified the coordinated roles of Shp1 and Shp2 in signaling. Although Shp1 generally acts as a negative regulator and Shp2 as a positive mediator of signal transduction, our results suggest that they function in a complementary manner to optimize signaling downstream of TPO/MPL and GPVI pathways, thereby ensuring appropriate regulation of late megakaryopoiesis, platelet production and activation.

      These additional considerations have been incorporated into the revised manuscript to provide a clearer conceptual framework for how Shp1 and Shp2 cooperate to regulate platelet biogenesis.

      Reviewer #1 (Recommendations for the authors):

      (1) The effects of the Shp1/Shp2 DKO are clear, but the effect of the Shp2 single knock-out is less clear on all parameters that were tested. The exception is ERK1/2 phosphorylation, which was reduced in the Shp2 knock-out as well as the Shp1/Shp2 DKO. Why do the authors conclude that Shp2 may be a potential therapeutic target, while the data show that knock-out of Shp1 and Shp2 is required for the observed effects?

      We agree that the most pronounced phenotypes were observed in the Shp1/Shp2 DKO. However, Shp2 single knock-out consistently reduced ERK1/2 phosphorylation, indicating that Shp2 contributes to MPL downstream signaling in megakaryocytes. The absence of a strong phenotype in Shp2 single knock-out may be due to residual Shp2 protein. However, given the established role of the Shp2–ERK pathway in megakaryopoiesis and the observation that pharmacological Shp2 inhibition significantly affected MK ploidy, proplatelet formation, and ERK1/2 phosphorylation, our data support a contribution of Shp2 to these processes and suggest it as a potential therapeutic target.

      (2) Inhibitors of Shp1 and Shp2 had largely similar effects as Shp1 and Shp2 single knock-outs, respectively. The effect of Shp2 knock-out on MK ploidy is not clear, cf. Figure 5Ai (no effect) and Figure 5Aii (reduction, which is not significant), whereas a clear and significant effect was reported for the Shp1/Shp2 DKO. In contrast, in Figure 7Ciii, the Shp2 inhibitors SHP099 and RMC-4550 clearly affect MK ploidy and the percentage of MKs forming proplatelets. The discrepancy between the effect of Shp2 knock-out and Shp2 inhibitors suggests that the inhibitors may affect other targets. The authors should consider using the Shp2 inhibitors on the Shp2 knock-out to prove or disprove that the effects of the Shp2 inhibitors are mediated exclusively by Shp2.

      Pharmacological inhibition does not necessarily phenocopy genetic deletion. The allosteric Shp2 inhibitors used in our study (SHP099 and RMC-4550) stabilize Shp2 in an inactive conformation and inhibit its catalytic activity, whereas Ptpn11 deletion results in complete loss of the Shp2 protein, including both catalytic and scaffolding functions. These mechanistic differences may lead to distinct biological outcomes and could explain the discrepancy observed between Shp2 knockout and inhibitor treatments.

      (3) Since the most profound effects were found in the Shp1/Shp2 DKO, it would be interesting to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/Shp2 DKO.

      We thank the reviewers for these constructive comments. Further experiments are indeed being planned to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/2 DKO.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Additional details on the strategy used to isolate megakaryocyte progenitors from mouse bone marrow would improve clarity, including sorting approach, gating strategy, and assessment of population purity.

      We thank the reviewer for this suggestion. We have now expanded the Methods section to provide a more detailed description of the strategy used to isolate megakaryocyte progenitors from mouse bone marrow.

      Briefly, bone marrow cells were first enriched for hematopoietic progenitors and stained with antibodies against lineage markers and megakaryocyte-associated markers. Megakaryocyte progenitors were then isolated by flow cytometric sorting based on established surface marker combinations, including c-Kit and CD41 expression. The gating strategy excluded lineage-positive cells and debris before selecting the progenitor population of interest.

      (2) Platelet GPVI expression appears reduced not only in Shp1/2 double-knockout mice but also, to some extent, in single Shp1- or Shp2-deficient models. A more detailed quantitative comparison and discussion would be helpful.

      We thank the reviewer for this observation. Although the most pronounced reduction in GPVI surface expression was observed in Shp1/Shp2 double knock-out platelets, minor variations may appear in the single knock-out models. To address this, we performed additional statistical analyses comparing WT platelets with each single knock-out genotype. These analyses did not reveal any significant statistical differences in GPVI expression between WT and either Shp1- or Shp2-deficient platelets, indicating that the apparent variations fall within the range of biological variability.

      (3) The aggregation traces shown in Figures 3A and 3B would benefit from clarification regarding their representativeness relative to the corresponding quantitative analyses.

      We thank the reviewer for this comment. The aggregation traces in Figures 3A and 3B represent experiments selected from independent replicates included in the quantitative analysis. The figure legends have been revised to clarify that these traces are representative of the experiments summarized in the quantification panels, which include data from multiple independent mice.

      (4) In several experiments, statistical significance may be influenced by differences in sample size across genotypes (e.g., Figures 2Ci, 3Ai, 3Di, and 6Ai). Using comparable numbers of replicates would strengthen the interpretation.

      We appreciate the reviewer’s attention to statistical rigour. The differences in sample size between genotypes reflect the availability of animals from the different breeding cohorts. Importantly, all statistical analyses were performed using appropriate tests that account for unequal sample sizes. The observed differences remain consistent across independent experiments.

      (5) The rationale for assessing only P-selectin exposure following CRP and PAR4p stimulation is not fully explained. Including integrin αIIbβ3 activation, or clarifying its exclusion, would provide a more complete assessment of platelet activation.

      We thank the reviewer for this suggestion. P-selectin exposure was used as a primary readout because it provides a robust measure of α-granule secretion downstream of GPVI and PAR signaling. Integrin αIIbβ3 activation was not assessed in these experiments because platelet aggregation assays were performed in parallel, which already provide a functional readout of integrin activation, as aggregation requires αIIbβ3 engagement. Nonetheless, we agree with the reviewer that direct measurement of integrin activation (e.g., fibrinogen binding) would provide complementary information and will be considered in future studies.

      (6) Figure 3Dii is described as an aggregation assay, although it appears to report P-selectin exposure; this distinction should be clarified.

      We thank the reviewer for identifying this inconsistency. Figure 3Dii reports indeed P-selectin exposure measured by flow cytometry, rather than platelet aggregation. We have corrected the description in the Results section.

      (7) The suggestion of compensatory extramedullary hematopoiesis based on splenomegaly would be strengthened by immunophenotypic analysis of splenic hematopoietic progenitor populations.

      We appreciate this important suggestion. In the current study, the evidence for possible compensatory extramedullary hematopoiesis is mainly based on the splenomegaly observed in Shp1/2 DKO mice. We agree that detailed immunophenotypic analysis of splenic hematopoietic progenitors would provide additional mechanistic insight; however, this was beyond the scope of the present study, which focuses on the intrinsic role of Shp1 and Shp2 in the megakaryocyte and platelet lineage. We have therefore revised the Discussion to present this interpretation more cautiously and to indicate that further studies will be required to determine whether splenic hematopoiesis contributes to compensatory platelet production in this model.

      (8) In Figure S3, differences in platelet recovery kinetics among genotypes appear evident. Clarification of the statistical tests used to assess these differences would be useful.

      We thank the reviewer for this comment. Platelet recovery kinetics were analyzed using two-way ANOVA with appropriate post hoc tests. No statistically significant differences between genotypes were observed. These details have been added to the Methods and figure legend for clarity.

      Reviewer #3 (Recommendations for the authors):

      Overall, the manuscript suffers from multiple typographical and grammatical errors that distract from the data being presented.

      We have carefully revised the manuscript to correct typographical and grammatical errors throughout, improving clarity and readability.

      (1) Figure S1: I believe this should be referenced in the first paragraph of the results section.

      We have now referenced the Supplemental Figure S1 in the first paragraph of the results section as suggested.

      (2) Figure 2A: Although the individual points for the replicates are informative, they do make it difficult to appreciate the SEM, and to my eye it appears that, for example, there may not be a difference between Shp2 and Shp1/2 or that there may be a difference between Shp1 and Shp1/2 in (ii), as Table S2 suggests. In other words, it seems that the increased MPV (as well as the leukocyte phenotype) may be driven by the knockout of Shp2; are there statistical analyses that could be performed to show that the increased MPV is specific to the double knockout?

      We thank the reviewer for this comment. Despite the slightly higher MPV observed in Shp2 single knockouts, statistical analysis using one-way ANOVA, which is appropriate for comparing means across multiple independent groups, and taking all individual data points into account, revealed no significant differences between Shp2 or Shp1 single KO and the Shp1/2 DKO.

      (3) Figure 2Bi: Is this missing a statistical significance bar, or was there no significant difference in cumulative bleeding time between the conditions? If the latter, this should be clarified in the main text (although the specific sentence regarding bleeding time only claims "mildly prolonged," the preceding sentence indicates "significant increase in bleeding").

      Thank you for this comment. There was no statistically significant difference in cumulative bleeding time between the groups. We have now modified the text accordingly to clarify this point and to indicate that, while bleeding time was not significantly different, blood loss was significantly increased in Shp1/2 DKO mice.

      (4) Figure 2Ci: What was the extent (statistically) of GPVI reduction in the Shp1 and Shp2 single knockout mice compared to the control? It seems that although there was no change in alpha2 expression in the single-knockout conditions, the contributions of Shp1 and Shp2 loss may be additive on GPVI (although I acknowledge that this is not necessarily borne out in Figure 3Ai).

      Thank you for this comment. After reanalyzing the data using an appropriate statistical test (one-way ANOVA followed by Tukey’s post hoc test), we found that GPVI expression is significantly reduced in both Shp1 and Shp2 single knockout platelets compared with controls. However, this reduction did not result in detectable functional consequences on platelet aggregation, as shown in Figure 3Ai.

      (5) Figure 3Ai: It seems that the individual replicates for the Shp1/2 double knockout cluster in two populations, extreme non-responders and arguably normal responders to CRP. Are there any biological or technical explanations for this?

      We thank the reviewer for this observation. We agree that the distribution of individual replicates in the Shp1/2 DKO group suggests the presence of two subpopulations, with some samples showing markedly impaired aggregation and others retaining near-normal responsiveness to CRP. While all experiments were performed under standardized conditions, subtle differences in platelet preparation, agonist sensitivity, or assay timing could also contribute to dispersion within this group. Importantly, despite this variability, the overall trend indicates a significant reduction in aggregation in the Shp1/2 DKO condition compared to controls, supporting a critical and partially redundant role for Shp1 and Shp2 in GPVI-mediated platelet activation.

      (6) "Aberrant functional responses of Shp1/2-deficient platelets": It may be helpful, in the last paragraph of this section, to briefly explain the relationship between P-selectin, CRP, and PAR4p. If short on space/words, the introduction likely does not need an explanation of platelet function and definitions of megakaryopoiesis and thrombopoiesis.

      We thank the reviewer for this suggestion. We have revised the last paragraph to clarify that P-selectin surface expression reflects α-granule secretion following platelet activation. We now specify that CRP activates platelets via GPVI signaling, whereas PAR-4 peptide signals through thrombin receptors, providing context for the differential responses observed in Shp1/2-deficient platelets.

      (7) "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets": Is there a cartoon figure panel that could be added to clarify how SFK (which, as an aside, is not defined as an acronym), Syk, GPVI, CLEC-2 receptor, Shp1, and Shp2 are interrelated? In addition to the comments left in the public review, I was perplexed by Figure 4Bi, as the band for the Shp1/2 double knockout condition appears to be stronger than the other 3 conditions, but this is not what is depicted in the bar graph on the right.

      We thank the reviewer for this helpful comment. We have now added a schematic cartoon (new Figure 8) to clarify the relationships between SFKs, Syk, GPVI, and the regulatory roles of Shp1 and Shp2. All acronyms, including SFK, are now defined at first mention to improve accessibility.

      Regarding Figure 4Bi, we appreciate this observation. The apparent discrepancy between the representative blot and the quantification reflects variability across experiments. The bar graph represents the average of independent replicates.

      (8) I would also recommend considering reshuffling the panels in Figure 4 so that the 2 assays measuring Syk phosphorylation and the 2 assays measuring Src phosphorylation are next to each other, as opposed to grouped by agonist. They should also be presented in the order of the text, which states that SFK activation was measured via Src before mentioning Syk (but the data are presented in reverse).

      We thank the reviewer for this suggestion. We have reorganized Figure 4 so that the panels measuring Src and Syk phosphorylation are presented together and, in the order, described in the text. The manuscript text has also been updated accordingly to match the revised figure layout.

      (9) GPVI overexpression experiments in these megakaryocytes or, conversely, Syk inhibition in control cells, to reverse or recapitulate the phenotype, respectively, may be additionally informative.

      We thank the reviewer for this suggestion. We agree that modulating GPVI or Syk activity could provide additional mechanistic insight. While these experiments were beyond the scope of the current study, we plan to explore GPVI overexpression and Syk inhibition in follow-up studies to further validate the pathway’s role in the observed phenotype.

      (10) "Reduced Tpo signaling in Shp1/2-deficient MKs": In addition to the comments left in the public review, I would suggest moving this section to after "Defective proplatelet formation and MK maturation in Shp1/2-deficient mice" so that the 2 sets of proplatelet and ploidy data are consecutively presented.

      We thank the reviewer for this helpful suggestion. We have now revised the manuscript accordingly by reorganizing both the text and figures. The ploidy and proplatelet formation data are now presented together in Figure 5, followed by the Tpo signaling data in Figure 6, improving the overall flow and clarity of the results section.

      (11) Figure 6Cii: Why does Shp1 add up to >100%?

      The reason the Shp1 bar exceeds 100% is due to how the data were quantified and normalized. Each segment represents the mean from separate experiments. Stacking these means can exceed 100% because the sum of averages is not equal to the average of the total.

      (12) Figure 7D: How do you reconcile these findings of impaired AKT phosphorylation with the addition of a Shp2 inhibitor but no change with Shp2 knockout (Figure 5C)? Would you attribute it to the residual Shp1 and Shp2 in the Cre-Lox MKs?

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study provides valuable evidence that hilar mossy cells play important roles in maintaining the structural organization of the dentate gyrus and regulating the maturation of adult-born granule cells. The evidence for the structural reorganization and for the accelerated dendritic maturation of adult-born granule cells is convincing: it rests on converging anatomical, viral tract-tracing, retroviral birth-dating, and electrophysiological measurements, with appropriate controls for viral spread, off-target CA3 expression, and axonal degeneration. Support for the study's broader interpretive claim - that the dentate circuit functionally compensates for mossy cell loss - is incomplete. That claim rests on two null results obtained under baseline conditions (home-cage cFos and PTZ seizure metrics) in small cohorts, without behavioral assessment and without a stimulus-driven activity readout, and the manuscript does not engage with published work showing that mossy cells regulate neural stem cell activation and are required for stimulus-evoked neurogenic and behavioral responses.

      Thank you and we agree with all points. We specifically focused on the structural aspects of dentate rearrangement, the impact of mossy cell loss/silencing on dentate neurogenesis, and dentate function at the circuit level. Although our home-cage cFos and PTZ susceptibility assays are limited in terms of their sensitivity, these assays were chosen to address the role of mossy cells in controlling overall dentate activity levels and seizure susceptibility, and our data demonstrate no dramatic changes in overall activity levels or increased/decreased seizure susceptibility.

      Strengths:

      (1) The study is technically rigorous and employs multiple complementary approaches, including selective genetic manipulations, viral tracing, immunohistochemistry, retroviral labeling of adult-born neurons, electrophysiology, and anatomical analyses. The comparison between complete mossy cell ablation and chronic synaptic silencing is particularly powerful, allowing the authors to examine the significant role of mossy cells in structural and functional organization in the dentate gyrus.

      (2) One of the most notable findings is the identification of a previously unrecognized collapse of the inner molecular layer following extensive mossy cell ablation. This observation substantially expands current understanding of dentate gyrus structural plasticity. The demonstration that adult-born granule cells undergo accelerated dendritic maturation after both mossy cell loss and silencing also provides important insight into how mossy cells regulate adult neurogenesis.

      We were also surprised by the inner molecular layer (IML) collapse, as disease models that produce mossy cell loss often involve granule cell axon (mossy fiber) sprouting (and maintained IML thickness) rather than IML collapse. It is unclear whether axon sprouting, reduced degrees of mossy cell loss, or other signaling pathways drive the differences between our selective ablation and translational disease models. We agree that the differential effects of mossy cell ablation and silencing on adult neurogenesis and proximal spine formation highlight the remarkable plasticity in this circuit and provide insights into both the functional and structural circuit roles of mossy cell inputs.

      Weaknesses:

      (1) The functional significance of the observed structural remodeling remains incompletely addressed. Mossy cells have been strongly implicated in pattern separation, spatial information, and emotional behavior, yet no behavioral analyses were conducted. Consequently, it remains unclear whether the dramatic anatomical changes observed following mossy cell ablation translate into meaningful behavioral alterations.

      Our functional assays were primarily focused on the circuit (synaptic) level, with additional assessment of how mossy cell manipulations affected overall dentate activity levels (as reflected by cFos expression). Our limited behavioral analysis focused on seizures, based on prior foundational work on the roles of mossy cells in seizures/epilepsy. We tested the hypothesis that seizure susceptibility might be markedly changed in the near absence of mossy cells, using a “threshold” dose of PTZ that is just above that required to produce seizures, and which can produce dramatically enhanced seizures in hyperexcitable mice. Alternative seizure assays (dose-response curves, continuous monitoring, different seizure-inducing protocols), measurements of granule cell activity in response to environmental contingencies, and the assessments of the response of the dentate stem cell pool to neurogenesis-enhancing stimuli might absolutely produce further insights into how functional mossy cell inputs control stimulus-related dentate activation and/or neurogenesis. Our resubmitted manuscript will clarify that the preserved basal level of dentate gyrus activity after mossy cell loss does not preclude altered activity-dependent activation in other settings or behavioral/learning changes. This could be uncovered with additional behavioral testing or seizure modeling, and is something that we expect to address in future studies.

      (2) The conclusion that the dentate gyrus exhibits remarkable homeostatic compensation is reasonable but remains indirect. Although cFos expression and PTZ-induced seizure susceptibility are unchanged despite altered E:I balance, the mechanisms responsible for maintaining network stability are not investigated. Additional analyses of inhibitory circuit remodeling or compensatory synaptic adaptations would strengthen this conclusion.

      We believe that there are many potential mechanisms that could explain how the nearly complete loss of a major population of dentate neurons is not accompanied by dramatic changes in overall activity levels. Although a fully comprehensive functional assessment of dentate circuit elements is prohibitive, we will undertake what we believe to be the highest-yield analyses in this regard. We propose to stain tissue for inhibitory circuit markers such as VGAT and PV, to determine whether mossy cell loss alters the density or localization of inhibitory synapses as well as circuit elements involved in feed-forward inhibition. We also plan to perform additional electrophysiological experiments to directly assay whether changes in feed-forward inhibition, overall synaptic inhibition (sIPSCs) and/or tonic inhibition might accompany functional mossy cell loss. We will incorporate the outcomes from these additional assays into a revised manuscript. This will shed light on whether inhibitory circuit remodeling also contributes to compensation after mossy cell loss, and hopefully provide additional insights relevant to translational disease models that involve mossy cell loss.

      Reviewer #2 (Public review):

      Summary:

      The authors examine how hilar mossy cells (MCs) influence adult-born dentate granule cell (abDGC) maturation and dentate gyrus (DG) structural integrity. Using both MC ablation and chronic functional silencing, they find that lacking MC inputs accelerates early abDGC maturation without altering mature cellular or intrinsic properties. MC silencing specifically decreased inner molecular layer (IML) spine density, whereas MC ablation led to IML collapse and an increased E/I ratio. However, neither intervention altered overall network excitability (measured via c-Fos and seizure induction) or seizure thresholds. These results advance our understanding of DG circuit plasticity during neurodegeneration.

      Strengths:

      (1) The side-by-side comparison of ablation vs. silencing provides a clear distinction between structural synapse loss and functional inactivation.

      (2) The multi-level analysis spanning structural anatomy, single-cell physiology, and network-level assays yields a rich, comprehensive dataset.

      Thank you for these positive assessments of our study.

      Weaknesses:

      (1) Measuring composite E/I ratios without parsing isolated EPSCs and IPSCs limits direct evaluation of MC-driven excitatory inputs. Furthermore, electrical stimulation in the IML likely recruits local interneuron axons directly alongside MC fibers, complicating the attribution of these responses solely to feed-forward MC circuits.

      We fully expect electrical stimulation of the proximal molecular layer to directly recruit local interneuron axons in addition to feed-forward inhibition. Thus, our experimental design did not distinguish between directly stimulated and feed-forward inhibitory circuits, and we were only able to conclude that mossy cell loss caused circuit rearrangement without clearly attributing the differences specifically to feed-forward mechanisms. To provide additional insights into the underlying changes, we plan to examine both inhibitory circuit structure (using immunohistochemistry) and function using assays designed to distinguish between directly stimulated vs. feed-forward inhibitory mechanisms, which will be incorporated into the revised manuscript.

      (2) The dramatic structural reorganization and IML collapse observed following MC ablation make it difficult to attribute changes in the E/I ratio purely to functional synaptic remodeling rather than physical circuit distortion.

      We actually consider physical circuit remodeling after MC ablation to be the primary explanation for the E/I ratio changes, in that the proximal translocation of MEC synapses following mossy cell ablation allows them to be electrically stimulated in the proximal molecular layer. Thus, the altered E/I ratio of proximal synapses after MC ablation largely represents the fact that we are stimulating proximal MEC inputs rather than mossy cell inputs (which are now absent). Our MML terminal stain (VGlut2) and MEC viral labeling support this interpretation, which we will clarify in the results and discussion of this data.

      (3) Layer boundary shifts following MC ablation complicate the interpretation of site-specific spine density (Figure 4); without accounting for IML collapse, classifying spine loss purely by traditional layer boundaries rather than proximal vs. distal dendrites may obscure local structural changes.

      We initially kept the classic nomenclature (IML vs OML) to avoid confusion for readers, and defined “IML” vs “OML” spines based on proximity to the inner and outer edges of the molecular layer (the innermost and outermost 40 µm; see Methods). Thus, in the setting of IML collapse after mossy cell ablation, these spines almost certainly occurred in regions innervated by the MEC (formerly “MML”). To avoid obscuring this aspect of the data, we will clarify this in the Results and Figure 4, making the proximal vs. distal designations clear.

      (4) The convulsive dosing protocol used for the seizure threshold test lacks the sensitivity required to reveal subtle changes in excitability.

      Our PTZ dose (40 mg/kg i.p.) is just above a dose (30 mg/kg i.p.) that almost never causes seizures in healthy mice in our hands, making it potentially able to detect seizure resistance. This 40 mg/kg dose causes short, limited seizures with a relatively consistent latency, and in other (unrelated) experiments, mice with genetic hyperexcitability have dramatically increased seizure duration and accelerated seizure onset (and sometimes mortality) at this dose, indicating that it is sensitive to at least some forms of increased seizure susceptibility. That stated, we agree that this single-dose PTZ protocol could miss subtle changes in dentate excitability or seizure susceptibility. These could be unmasked by a more detailed dose-response analysis or by other induced seizure assays; we will clarify this limitation in our manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses the analysis of connectomic and transcriptomic datasets to survey the anatomy and connectivity of neurosecretory cells in the Drosophila brain. While the connectivity analyses are convincing, the anatomical and functional data provided to verify cell type identity and paracrine signaling is incomplete. Once these aspects are improved, this study would be of interest to neuroscientists working on hormonal signaling in Drosophila and other animals.

      We thank the editor and reviewers for their assessment of our manuscript. We hope that the additional results in the revised manuscript addresses all of the concerns.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by McKim et al seeks to provide a comprehensive description of the connectivity of neurosecretory cells (NSCs) using a high-resolution electron microscopy dataset of the fly brain and several single-cell RNA seq transcriptomic datasets from the brain and peripheral tissues of the fly. They use connectomic analyses to identify discrete functional subgroups of NSCs and describe both the broad architecture of the synaptic inputs to these subgroups as well as some of the specific inputs including from chemosensory pathways. They then demonstrate that NSCs have very few traditional presynapses consistent with their known function as providing paracrine release of neuropeptides. Acknowledging that EM datasets can't account for paracrine release, the authors use several scRNAseq datasets to explore signaling between NSCs and characterize widespread patterns of neuropeptide receptor expression across the brain and several body tissues. The thoroughness of this study allows it to largely achieve it's goal and provides a useful resource for anyone studying neurohormonal signaling.

      Strengths:

      The strengths of this study are the thorough nature of the approach and the integration of several large-scale datasets to address short-comings of individual datasets. The study also acknowledges the limitations that are inherent to studying hormonal signaling and provides interpretations within the context of these limitations.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      Overall, the framing of this paper needs to be shifted from statements of what was done to what was found. Each subsection, and the narrative within each, is framed on topics such as "synaptic output pathways from NSC" when there are clear and impactful findings such as "NSCs have sparse synaptic output". Framing the manuscript in this way allows the reader to identify broad takeaways that are applicable to other model system. Otherwise, the manuscript risks being encyclopedic in nature. An overall synthesis of the results would help provide the larger context within which this study falls.

      We agree with the reviewer and have modified the subsection titles to highlight the main findings within those sections.

      We have also included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      The cartoon schematic in Figure 5A (which is adapted from a 2020 review) has an error. This schematic depicts uniglomerular projection neurons of the antennal lobe projecting directly to the lateral horn (without synapsing in the mushroom bodies) and multiglomerular projection neurons projecting to the mushroom bodies and then lateral horn. This should be reversed (uniglomerular PNs synapse in the calyx and then further project to the LH and multiglomerular PNs project along the mlACT directly to the LH) and is nicely depicted in a Strutz et al 2014 publication in eLife.

      We thank the reviewer for spotting this error. We have now modified the schematic as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to provide a comprehensive description of the neurosecretory network in the adult Drosophila brain. They sought to assign and verify the types of 80 neurosecretory cells (NSCs) found in the publicly available FlyWire female brain connectome. They then describe the organization of synaptic inputs and outputs across NSC types and outline circuits by which olfaction may regulate NSCs, and by which Corazon-producing NSCs may regulate flight behavior. Leveraging existing transcriptomic data, they also describe the hormone and receptor expressions in the NSCs and suggest putative paracrine signaling between NSCs. Taken together, these analyses provide a framework for future experiments, which may demonstrate whether and how NSCs, and the circuits to which they belong, may shape physiological function or animal behavior.

      Strengths:

      This study uses the FlyWire female brain connectome (Dorkenwald et al. 2023) to assign putative cell types to the 80 neurosecretory cells (NSCs) based on clustering of synaptic connectivity and morphological features. The authors then verify type assignments for selected populations by matching cluster sizes to anatomical localization and cell counts using immunohistochemistry of neuropeptide expression and markers with known co-expression.

      The authors compare their findings to previous work describing the synaptic connectivity of the neurosecretory network in larval Drosophila (Huckesfeld et al., 2021), finding that there are some differences between these developmental stages. Direct comparisons between adults and larvae are made possible through direct comparison in Table 1, as well as the authors' choice to adopt similar (or equivalent) analyses and data visualizations in the present paper's figures.

      The authors extract core themes in NSC synaptic connectivity that speak to their function: different NSC types are downstream of shared presynaptic outputs, suggesting the possibility of joint or coordinated activation, depending on upstream activity. NSCs receive some but not all modalities of sensory input. NSCs have more synaptic inputs than outputs, suggesting they predominantly influence neuronal and whole-body physiology through paracrine and endocrine signaling.

      The authors outline synaptic pathways by which olfactory inputs may influence NSC activity and by which Corazonin-releasing NSCs may regulate flight. These analyses provide a basis for future experiments, which may demonstrate whether and how such circuits shape physiological function or animal behavior.

      The authors extract expression patterns of neuropeptides and receptors across NSC cell types from existing transcriptomic data (Davie et al., 2018) and present the hypothesis that NSCs could be interconnected via paracrine signaling. The authors also catalog hormone receptor expression across tissues, drawing from the Fly Cell Atlas (Li et al., 2022).

      We thank this reviewer for the thorough assessment and for highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      The clustering of NSCs by their presynaptic inputs and morphological features, along with corroboration with their anatomical locations, distinguished some, but not all cell types. The authors attempt to distinguish cell types using additional methodologies: immunohistochemistry (Figure 2), retrograde trans-synaptic labeling, and characterization of dense core vesicle characteristics in the FlyWire dataset (Figure 1, Supplement 1). However, these corroborating experiments often lacked experimental replicates, were not rigorously quantified, and/or were presented as singular images from individual animals or even individual cells of interest. The assignments of DH44 and DMS types remain particularly unconvincing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to m-NSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retro-Tango samples, lending further support to our assignment of DH44 and DMS cell types.

      The electron micrographs showing dense core vesicle (DCV) characteristics (new Figure 2 Supplement 2E-G) are also representative images based on examination of multiple neurons. However, we agree with the reviewer that a rigorous quantification would be useful to showcase the differences between DCVs from NSC subtypes. Therefore, we have now performed a quantitative analysis of the DCVs in putative m-NSC<sup>DH44</sup> (n=6), putative m-NSC<sup>DMS</sup> (n=6) and descending neurons (n=2) known to express DMS across three datasets (FlyWire, BANC and maleCNS connectomes). For consistency, we examined the cross section of each cell where the diameter of nuclei was the largest. We quantified the mean gray value of at least 50 DCVs per cell. The individual who performed these analyses was blind to the neuron identity. Our analysis (new Figure 2 Supplement 2H-J) shows that mean gray values of putative m-NSC<sup>DMS</sup> and DMS descending neurons in FAFB and maleCNS are not significantly different, whereas the mean gray values of m-NSC<sup>DH44</sup> are significantly higher. This analysis agrees with our initial DH44 and DMS NSC subtype assignments. Nonetheless, given the similarity in morphology and synaptic connectivity of DH44 and DMS neurons, we have included the limitation on cell type assignment in the absence of molecular markers in the connectome datasets.

      The authors present connectivity diagrams for visualization of putative paracrine signaling between NSCs based on their peptide and receptor expression patterns. These transcriptomic data alone are inadequate for drawing these conclusions, and these connectivity diagrams are untested hypotheses rather than results. The authors do discuss this in the Discussion section.

      We agree with the reviewer that the novel paracrine pathways presented are untested hypotheses. However, there is a very high likelihood that a given NSC subtype can signal to another NSC subtype using a neuropeptide if its receptor is expressed in the target NSC. This is due to the fact that all NSC axons are part of the same nerve bundle (nervi corpora cardiaca) which exits the brain. The axons of different NSCs form release sites that are extremely close to each other. While the release sites in NSCs cannot be visualized in adult Drosophila connectomes (since these regions were not included in the sample prep), these have been mapped in the larvae and shown to be in close proximity (Hückesfeld et al., 2021: https://doi.org/10.7554/eLife.65745). Neuropeptides from these release sites can easily diffuse via the hemolymph to peripheral tissues (e.g. fat body and ovaries) that are much further away from the release sites on neighboring NSCs. We believe that neuropeptide receptors are expressed in NSCs near these release sites where they can receive inputs, not just from the adjacent NSCs, but also from other sources such as the gut enteroendocrine cells. Hence, neuropeptide diffusion is not a limiting factor preventing paracrine signaling between NSCs, and receptor expression is a good indicator for putative paracrine signaling. Consistent with this, several pathways highlighted in the plot (CRZ to CAPA, DH44 to Hugin and Hugin to DH44) have been anatomically and/or functionally validated previously (Zandawala et al., 2021: https://doi.org/10.1371/journal.pgen.1009425; King et al., 2017: https://doi.org/10.1016/j.cub.2017.05.089; Mizuno et al., 2021: https://doi.org/10.1111/dgd.12733). Additionally, a similar analysis was also employed to depict putative interactions between NSCs in larval Drosophila (Hückesfeld et al., 2021). We have now modified the caption for this figure to explicitly state these connections are putative. We hope that the putative pathways presented here will inspire future functional studies, and have also highlighted this outstanding question in the summary Figure 10.

      Reviewer #3 (Public review):

      Summary:

      The manuscript presents an ambitious and comprehensive synaptic connectome of neurosecretory cells (NSC) in the Drosophila brain, which highlights the neural circuits underlying hormonal regulation of physiology and behaviour. The authors use EM-based connectomics, retrograde tracing, and previously characterised single-cell transcriptomic data. The goal was to map the inputs to and outputs from NSCs, revealing novel interactions between sensory, motor, and neurosecretory systems. The results are of great value for the field of neuroendocrinology, with implications for understanding how hormonal signals integrate with brain function to coordinate physiology.

      The manuscript is well-written and provides novel insights into the neurosecretory connectome in the adult Drosophila brain. Some, additional behavioural experiments will significantly strengthen the conclusions.

      Strengths:

      (1) Rigorous anatomical analysis

      (2) Novel insights on the wiring logic of the neurosecretory cells.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript.

      Weaknesses:

      (1) Functional validation of findings would greatly improve the manuscript.

      We agree with this reviewer that assessing the functional output from NSCs would improve the manuscript. Given that we currently lack genetic tools to measure hormone levels and that behaviors and physiology are modulated by NSCs on slow timescales, it is difficult to assess the immediate functional impact of the sensory inputs to NSC using approaches such as optogenetics. However, since l-NSC<sup>CRZ</sup> are the only known cell type that provide output to descending neurons, we have functionally tested this output pathway using different behavioral assays (new Figure 8 and Supplements). Our analysis identifies a novel role for l-NSC<sup>CRZ</sup> and DNg27 neurons in female reproduction (based on the number of eggs laid).

      Recommendations for the authors:

      Reviewing Editor Comments:

      You will see that the reviewers found your work interesting and valuable, but had some suggestions for how revision could improve the manuscript. A common thread in the reviews is that functional speculations about the extracted circuits and paracrine signaling would benefit from revision, and would fit better in the Discussion, not Results section. Caveats could be more explicitly stated and language asserting functionality could be tempered. The reviewers were unanimous in their desire for a summary diagram or model.

      We thank the editor for these suggestions to improve the manuscript. We have now functionally validated some output pathways from l-NSC<sup>CRZ</sup>. We have also toned down the language regarding functionality where appropriate. Finally, we included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      Reviewer #2 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      The authors present connectomic analyses for NSCs identified in the FlyWire dataset. All of their connectomic findings would be strengthened by executing these same analyses in the freely available female hemibrain connectome (Scheffer et al. 2020; Plaza et al. 2022), thereby effectively increasing their sample size from one whole brain to three hemispheres. It is unclear why the authors chose only to focus on the FlyWire dataset.

      We thank the reviewer for this suggestion. We had performed a preliminary analysis using the hemibrain dataset. However, out of the 80 endocrine cells that we found in FlyWire, the hemibrain dataset lacks both the NSC subtypes in the SEZ (SEZ-NSC<sup>CAPA</sup> and SEZ-NSC<sup>Hugin</sup>) as well as l-NSC subtypes in the other hemisphere (l-NSC<sup>ITP</sup>, l-NSC<sup>DH31</sup>, l-NSC<sup>CRZ</sup>). In addition, a majority of the input synapses for all NSC are in the SEZ region which allowed us to classify the different NSC subtypes in FlyWire. Since this information is missing in the hemibrain dataset, we are unable to classify the m-NSC into the different subtypes (not shown). Therefore, we cannot perform a comprehensive analysis of input and output pathways of different NSC subtypes using the hemibrain dataset. To address this concern, we have repeated several analyses with two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome (Table 1, new Figure 1 Supplement 1, new Figure 2 Supplement 2, new Figure 3 Supplement 3, new Figure 6 Supplement 1, new Figure 7 Supplement 3). Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      The authors initially map assign NSC types based on anatomical locations and clustering of presynaptic connections and morphological features. Due to matching cell counts and similar soma locations, they find that DMS and DH44 types cannot be easily distinguished. The authors attempt to assign cell types to these two populations using two methods, neither of which are convincing as executed:

      (1) The authors attempt to distinguish the identities of the two populations by anatomically comparing presynaptic inputs in FlyWire to those observed with light microscopy using retrograde trans-synaptic labeling. Due to the lack of a genetic driver line for the DMS population, the authors could complete this only for the DH44 population. The authors present only one animal, at inadequate magnification to see the absence of distinguishing presynaptic neurons. The results would be strengthened by the presentation and quantification of multiple samples; without more than one sample, it is not possible to know how robust this finding is in this genetic driver line. The authors might also consider taking advantage of the widely-used template brain (Bogovic, 2020) to align their light micrographs of presynaptic inputs from the retrograde tracing, with the presynaptic skeletons from FlyWire and compare in a more quantitative and precise manner. The authors might also consider taking a similar approach using anterograde tracing (Talay et al. 2017) to label postsynaptic outputs. Given that postsynaptic outputs are fewer, so long as there are identifiable, distinct postsynaptic partners, it may be easier to distinguish the two populations with anterograde tracing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also provide a magnified image in this figure to highlight the absence of presynaptic neurons that distinguish m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>.

      We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to mNSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retroTango samples, lending further support to our assignment of DH44 and DMS cell types.

      As per this reviewer’s suggestion, we also aligned our retrograde tracing light micrographs to a template brain (Author response image 1). However, we were unable to quantitatively compare neurons in our light micrographs with neuronal skeletons from the connectome. This is because retroTango labels several neurons in the SEZ which obscures morphology of individual neurons needed for such comparisons. Additional experiments, where retro-Tango output is restricted to sparse populations of neurons using a Flp-out strategy, are needed to perform such quantitative analyses. These experiments are beyond the scope of this study since we now provide additional lines of evidence for cell assignments.

      Author response image 1.

      DH44 > retro-Tango presynaptic signal aligned to JRC2018 unisex template brain

      We appreciate the suggestion to use the anterograde tracing tool trans-Tango to distinguish mNSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>. There is very little synaptic output from m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> based on the FlyWire connectome. There is no synaptic output from both of these cell types if we use a threshold of 5 synapses for significant connections (new Figure 7). Using a threshold of 2 synapses for significant synaptic connections, 3 neurons are downstream of m-NSC<sup>DH441</sup> and 5 neurons are downstream of m-NSC<sup>DMS</sup> (not shown). Since these postsynaptic neurons are not bilaterally paired (we do not anticipate unilateral pathways), we don’t think that these connections are significant. Consistent with our analysis with the FlyWire connectome, we did not observe any significant post-synaptic signal with DH44 > trans-Tango (Author response image 2) even using flies raised at 21ºC which increases the synaptic strength during development. Since we do not have a GAL4 driver to specifically target m-NSC<sup>DMS</sup> , we could not perform similar trans-Tango analysis of m-NSC<sup>DMS</sup>.

      Author response image 2.

      DH44 > trans-Tango (left) and w<sup>1118</sup> > trans-Tango (right; control). Presynaptic neurons are labelled in green and post-synaptic neurons are in red. Representative images based on 5 samples.

      Our connectome analyses revealed that putative m-NSC<sup>DMS</sup> receive direct synaptic inputs from enteric neurons but m-NSC<sup>DH44</sup> do not. We used this information to perform another trans-Tango analysis using Gr43a-Gal4 which labels a subpopulation of enteric neurons (Miyamoto and Amrein, 2013: https://doi.org/10.4161/fly.27241) (Author response image 3).

      Author response image 3.

      Initiating trans-Tango from Gr43a neurons (green) does not label any postsynaptic neurons (magenta) in the pars intercerebralis (white arrow head), including those labelled by the DMS antibody (cyan).

      Unfortunately, initiating trans-Tango from Gr43a neurons did not label any post-synaptic neurons in the pars intercerebralis where m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup> soma are located. This could be due to a) low trans-Tango sensitivity or b) m-NSC<sup>DMS</sup> are downstream from other enteric neurons not captured by Gr43a-GAL4. In the absence of other broad enteric neuron drivers, we are unable to perform additional analyses.

      (2) The authors attempt to assign cell types by qualitatively assessing the darkness of dense core vesicles in these two populations. However, there is a presentation of only single planar images through three selected cells (a DMS-expressing descending neuron, DMS-expressing NSC, and DH44-expressing NSC) without any quantitative analyses of vesicle characteristics within or across NSC cell types. It is not possible for the reader to assess whether the darker vesicles constitute a real trend, or if these images are hand-selected to support their point. This piece of evidence would be more convincing if the authors demonstrate consistent vesicle characteristics within NSC type and differences across type. Moreover, such analysis of dense core vesicle features in cell types with distinct and known peptide expression would be broadly interesting.

      Given that NSC type assignment is a major contribution of the present paper, it is critical that the authors are clear about the remaining uncertainty in assigning cell types, so as not to propagate false certainty into future work.

      This comment has been addressed above, and we refer the reviewer to the new Figure 2 Supplement 2E-J.

      The authors suggest larger peptide release capacity from CAPA-producing NSCs based on their larger morphological features (Figure 1, Supplement 2), which is more speculative than certain. In Figure 1, Supplement 1 the authors demonstrate the capacity to visualize vesicles number and size in individual NSCs. Rather than speculate over larger peptide release capacity based on cell size, the authors could quantify these vesicle features, which are surely a better indication of peptide release capacities.

      We thank the reviewer for this comment. We agree that number of dense core vesicles within these and other neurons would be a better indicator of their peptide release capacity. We are performing these analyses on a brain-wide scale as part of another project. Therefore, we have removed the following speculative statement from the present manuscript:

      “But given their location, large size, and presumed large release capacity, we speculate that SEZ-NSC<sup>CAPA</sup> participate in global modulation of post-feeding physiology.”

      The authors provide an analysis of NSCs' synaptic inputs and outputs, but never mention whether NSCs are synaptically connected to each other. If connected, it would be very sensible to provide some analysis of synaptic connectivity between NSCs. If they are not connected, the authors should explicitly mention this in the main text, as it is relevant to the overall aim of this study.

      All NSCs are classified as endocrine cells in the FlyWire connectome. Hence, as shown in new Figures 3B and 7B, NSCs do not provide output to any endocrine cells (NSCs) using a threshold of 5 synapses for significant connections. Similarly, we do not see any synaptic connectivity between NSCs in the BANC dataset (new Figure 7 Supplement 3A-B). We do observe sparse connectivity between NSCs in the maleCNS dataset with a threshold of 5 synapses (new Figure 7 Supplement 3C-D), as well as in the Flywire connectome when the threshold is reduced to 2 synapses (new Figure 7 Supplement 2). However, we refrain from emphasizing on these connections because additional validation is required to rule out false positives in synapse predictions. Dense-core vesicles in NSCs can frequently be mistaken for synaptic T-bars during the prediction (unpublished observation).

      Although unlikely, NSCs could also form synapses with each other near their release sites and outside the brain volumes captured in all three datasets examined in this study. This limitation has been included in the discussion.

      There is no substitute for a good circuit wiring diagram; the motifs that are extracted in Figure 3H might be better appreciated if the reader was first presented with a well-formatted complete circuit diagram, which may then foreshadow the points made in the main text and in Figure 3H.

      We appreciate this suggestion. We now include a circuit diagram (new Figure 3G) to highlight the connectivity between NSCs and their presynaptic partners. The proportion plot (old Figure 3G) has now been moved to new Figure 3 Supplement 5.

      The authors provide extensive bar graphs showing synaptic input body IDs in Figure 3 Supplement 2, however they don't complete the same analysis for synaptic outputs (likely due to low numbers). Even so, it would be useful to the reader if the body IDs and cell types for both synaptic inputs and outputs were documented in a supplemental table. Providing such an inventory is aligned with the goals of this study.

      Only l-NSC<sup>unknown</sup> and l-NSC<sup>CRZ</sup> provide synaptic outputs in the FlyWire connectome (new Figure 7B-F). We have included bar graphs showing output from both these cell types at a single-cell level (new Figure 7G). Additionally, we have annotated all the NSC subtypes in FlyWire and BANC on Codex. Further exploring the inputs and outputs of NSC subtypes can be done interactively on Codex. For example, the search command “{upstream_union} cell_type == SEZ_NSC_CAPA” will retrieve all the neurons providing inputs to SEZ-NSC<sup>CAPA</sup>. As a quick search, this is more convenient than pasting individual body IDs from a supplementary table into Codex. All code outputs (csv files) containing this information are also available on Zenodo.

      In describing possible paracrine signaling, the authors write "Given the proximity of NSC axon terminations, it is extremely likely that a hormone released from a given NSC will influence the activity of other NSC types if its receptor is expressed in those cells." In the absence of functional experiments and/or information about spatial localization and/or peptide diffusion and the proximity of receptors to release sites, the expression patterns alone are insufficient to support this conclusion. Thus, the authors might consider removing the circular connectivity plots in Figure 7C, and Figure 7, Supplement 1A-H, and instead emphasize what can be concluded with certainty from the transcriptomic data (which are expression patterns of the hormones and receptors across NSCs and other tissue types). The authors might instead speculate over potential paracrine signaling between NSCs in the Discussion. Given that the authors describe paracrine signaling between NSCs as 'putative' in the abstract and main text, the authors will likely agree the legend for Figure 7 is misleading.

      This comment has been addressed above. We agree with the reviewer and have modified the figure legends (new Figure 9 and Figure 9 Supplement 1) to emphasize that the connections are putative.

      Should the authors keep these connectivity diagrams, it is important to reconsider their threshold wherein 50% of cells in a cluster must express a given hormone for it to be considered present in their analysis. It is entirely conceivable that there is real heterogeneity in hormone expression within the cluster, so it is surprising the authors have applied this artificial criterion.

      We thank the reviewer for presenting us with this option.

      We also apologize for the oversight in explaining our thresholding carefully. To minimize false positives, neuropeptides were subjected to a two-step filtering process. First, only those expressed in at least 50% of the cells within a given cluster were retained. Second, a composite expression score was calculated for each neuropeptide by multiplying its average expression by its percent detection. These values were normalized to the maximum observed signal across the dataset, and only hormone-cluster pairs maintaining a relative score of 0.25 or higher were included in the final analysis. This stringent filtering approach was implemented to focus the analysis on dominant neuropeptides and to exclude contamination from ambient RNA, which is common for neuropeptides (Allen et al., 2020: https://doi.org/10.7554/eLife.54074). To account for lower abundance of receptor transcripts, we used a more permissive threshold for receptors by retaining those expressed in at least 5% of the cells within a cluster. Unlike the neuropeptides, no secondary relative-score filtering was applied to the receptors to ensure that biologically relevant signaling targets were not prematurely excluded due to low transcript density. We have now revised our methods to explain these details.

      Importantly, we used this thresholding criteria to align previous anatomical studies with our single-cell expression analysis and filter out neuropeptides that likely represent contamination: 1) Transcript for leucokinin (Lk) neuropeptide is expressed in l-NSC<sup>ITP</sup> (but previous studies have not been able to detect this peptide in l-NSC<sup>ITP</sup> (Zandawala et al., 2018: https://doi.org/10.1371/journal.pgen.1007767). 2) Hugin is not expressed in m-NSC<sup>DMS</sup> (Oh et al., 2021: https://doi.org/10.1016/j.neuron.2021.04.028). 3) Ilp2 is not expressed in SEZNSC<sup>CAPA</sup> and m-NSC<sup>DMS</sup> (this study and various others). 4) ITP is not expressed in any m-NSC (Gera et al., 2025: https://doi.org/10.7554/eLife.97043.3). Based on these and other examples, we feel that our stringent criteria recover putative pathways that are likely functional while filtering obvious false positive. Nonetheless, we agree with the reviewer that some of these NSCs could represent heterogeneous clusters as has been shown recently for m-NSC<sup>DILP</sup> (Held, Bisen, Zandawala et al., 2025: https://doi.org/10.7554/eLife.99548.3). This heterogeneity could result in some authentic connections to drop out. However, our goal for this analysis was to not identify all the putative paracrine connections, but rather the strongest ones with the hope that it can inspire future functional studies.

      The data shown in Figure 2 would be easier to interpret and therefore more convincing with better use of insets, appropriate overlays of multiple markers, higher image magnifications, and quantification across samples. Specifically: In Figure 2C, authors show mCherry expression but it isn't clear where these cell bodies are located with respect to the image in Figure 2B. This is also true for Figure 2D. Insets in Figure 2B that correspond with regions shown in 2C and 2D would be helpful. In Figure 2, the authors do not provide cell counts across samples for all markers. Thus, it isn't clear how consistent these cell counts are across samples. In Figure 2E, authors claim that no Gr64f-positive cells innervate the NCC, yet there is clearly a GFP signal in the NCC region in the merged image. The authors should provide an additional marker or a higher magnification image to convince the reader that these projections are not in the NCC region.

      We thank the reviewer for these suggestions. To improve clarity, we have made the following changes:

      Figures 2C and 2D are based on different samples than the one shown in Figure 2B. But we have added dashed boxes in Figure 2B to indicate the regions shown in Figure 2C and 2D.

      Included sample sizes in Figure 2A and 2C. The rest of the panels are representative images based on at least 5 samples. This has been included in the methods.

      The cell count has been provided for m-NSC<sup>DILP</sup> for both markers in Figure 2A. The cell counts for m-NSC<sup>DMS</sup>, labelled using mCherry alone, has been provided in Figure 2C. We did not perform cell counting when using the membrane GFP reporter as it is difficult to accurately count overlapping cells (see dashed box labelled C in Figure 2B).

      We also provide a new supplementary file (new Figure 2 Supplement 1) showing cell counts for m-NSC<sup>DILP</sup> using different markers. Based on this, we can confidently conclude that adult Drosophila typically have more than 14 m-NSC<sup>DILP</sup>.

      We have corrected a typo in our label for Figure 2E: it should be Gr64a instead of Gr64f.

      We have modified the Figure 2E inset to show that the four pairs of Gr64a > myrGFP expressing corazonin cells do not project via the NCC (labelled with an arrow). We have also identified the four pairs of Gr64a neurons (Author response image 4 left panel) in the FlyWire connectome, which shows that these neurons do not exit the brain via the NCC.

      Author response image 4.

      Corazonin-expressing Gr64a neurons in the FlyWire connectome (left) and a light micrograph (right, same as in Figure 2E) showing Gr64a neurons (green) and corazonin neurons (magenta).

      Recommendations for improving the writing and presentation:

      Throughout the paper, the authors provide scant or, at times, no citations. Inadequate citation is as much an issue in the introduction as it is in the results and discussion sections. As such, the authors often do not provide a well-supported premise for the present work and/or do not place their findings and interpretations into the context of existing literature. Related, there is a predominance of references to the work of the authors themselves, often in place of citing earlier foundational work. Citations are nearly exclusive to the Drosophila literature, with the exception of the second paragraph of the introduction. This paper would be greatly improved with references to a broader literature.

      We have now added additional references to give credit to foundational work where appropriate. We have also included citations to non-Drosophila literature for more general statements in the introduction and discussion; however, we refrain from citing such studies in the results section to keep it focused.

      Figure 1 Supplement 1 is referenced after Figure 2 in the text. The authors might consider reassigning it as a supplement to Figure 2, which also uses imaging methodologies to distinguish NSC cell types.

      We agree with the reviewer and have reassigned the figures accordingly.

      Figure 3G is difficult to interpret, and its figure legend is brief and inadequate.

      As suggested by this reviewer, we have replaced this panel with a circuit diagram. The proportion plot (old Figure 3G) has now been moved to Figure 3 Supplement 5, and we have expanded the figure legend.

      The bar graphs in Figure 3 Supplements 2 and 3 would best benefit the reader if the x-axis labels are not simply body IDs, but also cell types or instances (if assigned in FlyWire).

      We appreciate this suggestion. Cell types are routinely updated on Codex while the root IDs remain static for v783. Therefore, we chose root IDs for these plots as they can be used to query Codex easily and reliably. We now provide all raw data as csv files on Zenodo used to make these plots. This includes cell types and other classifications.

      Minor corrections to the text and figures:

      Table 1 compares the observed numbers of NSC types in adult flies to those in larvae and those expected based on previous literature. The authors should cite the previous studies that support each of the expected or larval numbers, either within the table or in the table legend. It would also be appreciated if the expected numbers were cited in the main text.

      References for NSC numbers in larvae and expected numbers in adults are now included in Table 1.

      In describing the author's approach to analyzing synaptic connectivity by cosine similarity, authors cite their own previous work rather than the foundational study describing this approach or earlier studies that use it.

      We have now also cited Schlegel et al., 2021 (https://doi.org/10.7554/eLife.66018) who used a similar approach in the olfactory system.

      Reviewer #3 (Recommendations for the authors):

      (1) The observation that most gustatory inputs to NSCs are indirect (particularly for feedingrelated NSCs) is very interesting but lacks functional validation. I suggest that the authors conduct behavioural assays where specific sensory inputs are activated or silenced while monitoring outputs from NSCs. This could include optogenetics to stimulate or inhibit sensory neurons, or alternative feeding assays.

      We thank the reviewer for this insightful suggestion. We agree that the functional validation of gustatory-to-NSC pathways is a highly compelling direction for future research. However, we believe that behavioral assays, as suggested, pose significant interpretive challenges for the following reasons:

      NSCs primarily function by releasing hormones into the systemic circulation. Unlike classical neurotransmission, hormonal modulation typically operates on much slower timescales (minutes to hours). Consequently, acute activation of sensory inputs is unlikely to elicit immediate, quantifiable behavioral changes that can be specifically attributed to NSC activity.

      Most NSC classes are known to influence multiple physiological and behavioral processes simultaneously. Attributing a specific behavioral phenotype to a single NSC class following sensory stimulation would be confounded by these overlapping roles.

      Activating or silencing taste neurons will directly impact feeding behavior through canonical motor circuits, independent of the neuroendocrine system. In such a paradigm, it would be nearly impossible to isolate the specific "indirect" contribution of the NSCs to the observed behavior.

      While we agree that functional connectivity, such as optogenetic activation of taste neurons paired with calcium imaging (e.g., GCaMP) in NSCs, would be the ideal way to validate these inputs, we consider these extensive physiological experiments to be beyond the scope of this anatomical and connectomic study.

      Nonetheless, to address this important question, we now use a recently developed approach (Bates et al., 2026: https://doi.org/10.1101/2025.07.31.667571) based on linear dynamical modeling to estimate the influence of various sensory neurons (gustatory, olfactory, enteric, hygrosensory, etc.) on different NSC classes. Our analysis (new Figure 6 and Figure 6 Supplement 1) reveals that contents of consumed food (detected by enteric neurons) have a stronger influence on NSCs compared to inputs from external taste receptors.

      (2) Descending neurons appear to play a crucial role in regulating both motor and endocrine output. However, their functional contribution is only inferred from the connectomic data. The authors could perform functional activity manipulations (silencing or activating) of these descending neurons (for instance dMS descending neurons) to explore their role in behaviour. This could be tested with simple behavioural assays such as feeding or reproduction (i.e egg laying).

      We believe that there might be some confusion. DMS descending neurons (DNp32 cell type) used for dense-core vesicle comparisons with m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> (new Figure 2 Supplement 2) are different from the descending neurons (DNg27 cell type) that receive inputs from l-NSC<sup>CRZ</sup> (new Figure 7). We have indicated the cell type of DMS descending neurons in the text to clarify this. We have also functionally tested DNg27 (instead of DMS descending neurons suggested by the reviewer) using optogenetic and chemogenetic approaches for effects in feeding, food preference, starvation survival, egg laying and flight (new Figure 8 and Figure 8 Supplement 1). While we expected DNg27 to influence flight based on our connectome analyses, we do not see any phenotype in our free flight setup following DNg27 activation (new Figure 8 Supplement 1). However, this could be due to the split GAL4 driver used being very weak (new Figure 8 Supplement 2). This is also supported by the egg-laying assay where DNg27 inactivation only produces a phenotype after day 8 (Figure 8). Since we currently don’t have access to another driver to specifically target DNg27, we are unable to validate our results in the free flight setup using an independent driver.

      (3) The authors describe a sparse olfactory input pathway to NSCs, with emphasis on odours playing major roles. However, the physiological consequences of these connections are not explored in detail. Authors should use ORN/AL stimulation (e.g., using optogenetics) to explore how odour sensory pathways affect hormonal secretion in NSCs.

      We acknowledge the reviewer’s interest in the physiological consequences of the olfactory to NSC pathways identified in our study. While we agree that exploring how specific odors modulate neuroendocrine output is a logical next step, we believe that such experiments are currently unfeasible due to significant technical and biological constraints as highlighted above for taste neurons. Hence, we calculated the influence of olfactory receptor neuron activation on different NSC classes using an approach based on linear dynamical modeling (new Figure 6 and Figure 6 Supplement 1). Our analysis reveals that smell has a weaker influence on NSC compared to taste.

      (4) The authors present a large amount of nice yet complex data, which can be difficult to navigate through and is sometimes hard to follow. Consider adding more schematic diagrams to summarize the key pathways and interactions between NSC types and their inputs/outputs.

      We thank the reviewer for this suggestion. We have now included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Mast cells have previously been reported to play an important role in bacterial immune defense and act protectively in sepsis. However, many of these findings were based on studies using Kit mutant mice. In this study, the authors conducted a detailed investigation using mast cell-deficient Cpa3 Cre-Master mice. As a result, the authors found that the Cpa3 Cre-Master mice exhibited responses similar to wildtype mice in terms of bacterial immune defense. This suggests that the observed phenotype is not due to mast cell-dependent bacterial immune defense, but rather is associated with dysbiosis of the gut microbiota.

      Strengths:

      Mast cells have long been reported to play an important role in the protective response against sepsis, and their function in infection defense has been demonstrated. However, Kit mutant mice have been reported to exhibit impaired peristalsis, and several mast cell-specific genetically modified mouse lines have since been developed and examined in detail. This study presents an important finding by logically demonstrating that the exacerbation of sepsis in Kit mice is due to alterations in the gut microbiota, and that the phenotype previously thought to be mast cell-dependent was, in fact, not.

      In addition, the experiments were carefully designed using mice with matched genetic backgrounds. These findings underscore the importance of microbiota composition in interpreting immune phenotypes and highlight the need for cohousing controls in mutant mouse studies.

      A major strength of this work is the robustness of the CLP data, generated over eight years by three independent researchers across two institutions with large sample sizes, lending strong support to the conclusions.

      Weaknesses:

      The study assesses only a limited subset of gut bacterial species, leaving the extent to which E. coli expansion contributes to the observed phenotype unclear.

      We now performed 16S rRNA sequencing of cecal samples isolated from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Results are display in a new Figure 4. Our comparative analysis of the cecal microbial communities in Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice (Figure 4A+B) confirmed the expansion of E. coli (Enterobacteriaceae) that we had observed by CFU counts (Figure 3D). Furthermore, it revealed a dysbiotic shift marked by increased abundance of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, and Erysipelotrichaceae in KitW/Wv mice.

      None of these changes was observed when comparing the cecal microbiomes of Cpa3Cre/+ and Cpa3+/+ mice (Figure 4C+D), indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to the deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes that we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function. These new findings fully align with and further support our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 9-10 and discussed on pages 13-14.

      Moreover, in the cohousing experiments, there is no evidence provided to confirm successful microbiota normalization between groups.

      It is correct that we have no direct data to confirm microbiota normalization between groups after co-housing. We note, however, that co-housing is a generally accepted method for microbiota equalization or conversion (Caruso et al., Cell Rep. 2019, Ridaura et al., Science 2013, and reviewed in Moore et al., Clin. Transl. Immunol. 2016). In any case, Kit<sup>W/Wv</sup> mutants were made resistant to CLP by co-housing. Similar microbiota sequencing results between groups, while useful, would again only be correlative.

      A more detailed analysis of the microbial composition would be necessary to strengthen the reliability of the findings.

      See above the new data from 16S rRNA sequencing.

      It is also important to note that Cpa3-deficient mice exhibit not only mast cell depletion but also defects in basophils and T cells. These additional immunological alterations may counterbalance one another, potentially masking phenotypic changes and complicating interpretation.

      Regarding basophils in Cpa3<sup>Cre/+</sup> mice, compared to wild-type mice, basophils are reduced to about 40% of normal (Feyerabend et al., Immunity 2011). In Kit<sup>W/Wv</sup> mice, compared to wild-type mice, basophils are reduced to about 10% of normal. To our knowledge, there has been no phenotype reported in which a reduction in basophils compensates for the loss for mast cells. Given that Kit<sup>W/Wv</sup> mice have about threefold lower numbers of basophils and are highly susceptible to sepsis, there is no evidence that a reduction in basophils is protective in mast cell-deficient mice. On the contrary, mice that were normal for mast cells but had their basophils depleted were more susceptible to sepsis (Piliponsky et al., Nat. Immunol. 2019). Hence, basophils appear to be protective, and their reduction increases susceptibility. In light of these data and considerations, there is no evidence for a reduction in basophils to counterbalance the loss of mast cells in Cpa3<sup>Cre/+</sup> mice.

      Regarding T cells, there is no evidence, and there are no reports, that Cpa3<sup>Cre/+</sup> mice have defects in T cells (Feyerabend et al., Immunity 2011, Feyerabend et al., Cell Metabolism 2016). Cpa3 is weakly and transiently expressed early in the T cell lineage (Feyerabend et al., Immunity 2009; for expression levels in T cells versus mast cells, see Author response image 1). In summary, in contrast to the reviewer's claim, there are no known defects in T cell development or T cell functions in Cpa3<sup>Cre/+</sup> mice. We think the reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      Author response image 1.

      Generated from the Immgen database. Shown are RNAseq gene expression levels of diverse T-cell and mast cell populations.

      Furthermore, it remains to be determined whether the altered gut microbiota observed in KitW/Wv mice is a consequence of impaired intestinal motility, whether a similar phenotype is observed in KitW-sh/W-sh mice, and whether comparable results occur in SCF-deficient models. Addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.

      The purpose of our study was to verify or refute the key claim dating back to two 1996 Nature papers that mast cells play important roles against sepsis. We demonstrate here that this is not the case because mice without mast cells (Cpa3<sup>Cre/+</sup> mice) were as resistant to sepsis as wild-type mice. Hence, mast cells are not involved in the immunity against sepsis, and 'secondary factors' are not involved in this simple experiment (both groups of mice, wild-type and Cpa3<sup>Cre/+</sup> mice, were on the identical genetic background). Second, Kit<sup>W/Wv</sup> mice are also as resistant to sepsis as wild-type mice when confronted with the identical intestinal slurry. Therefore, Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. Hence, in our view, the underlying immunological question regarding the role of mast cells in sepsis has been conclusively addressed and answered by our data. We have changed the title to emphasize this central question.

      The reviewer now asks us to delve even deeper into Kit biology and in particular intestinal pathophysiology in this and other Kit or steel mutants. While we share his/her interest in such questions, we fully disagree with the statement that 'addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.' We do not intend to enter the field of gut physiology or its link to microbiota, all the more because any results would not affect the central conclusion of our manuscript.

      Given that KitW/Wv mice exhibit impaired peristalsis, is the observed increase in E. coli a consequence of this dysfunction?

      See above

      Previous studies with BMMC reconstitution experiments have indicated that mast cells are a source of TNF - how does this align with the current findings?

      It does not align well. It is possible that cultured and transplanted mast cells (BMMC) produce TNF. Given that we did not find a reduction in TNF levels in the peritoneal lavage or serum in mice without mast cells undergoing sepsis, under physiological conditions mast cell-derived TNF does not seem to have a measurable impact on total TNF levels.

      Reviewer #2 (Public review):

      Summary:

      This study presents a useful finding that the high susceptibility to CLP sepsis of Kitmutant mice is not due to mast cell deficiency, but to dysbiosis.

      However, the present data are insufficient and incomplete to support the conclusion, and would benefit from more rigorous approaches. With the mechanism part strengthened, this paper would be of interest to researchers on mast cell biology and mucosal immunology.

      We disagree with the view that our data are insufficient and incomplete. Our results demonstrate that mice lacking mast cells (Cpa3<sup>Cre/+</sup> mice) are as resistant to sepsis as wild-type mice, demonstrating that mast cells do not play a detectable role in immunity against sepsis. Additionally, we show that Kit<sup>W/Wv</sup> mice exhibit the same resistance to sepsis as wild-type mice when confronted with the identical intestinal slurry. This finding demonstrates that Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. These central data are both sufficient and complete, given that our data fully address the potential role of mast cells in sepsis. Our study aimed to investigate the role of mast cells in sepsis, not to examine the mechanisms of dysbiosis or associated pathological phenotypes in Kit-mutant controls. We have changed the title to make this point.

      Recommendations:

      (1) The authors showed that E. coli increases in the cecum of Kit-mutant mice, which causes high CLP susceptibility. However, they did not provide any evidence E. coli is responsible for the high susceptibility.

      We showed that E. coli CFUs were increased in the cecum of Kit-mutant mice, but we did not state that this causes CLP susceptibility. We wrote: 'Hence, Kit<sup>W/Wv</sup> microbiota contains high levels of E. coli, which may underlie the observed pathogenicity'. We demonstrated that intestinal slurry from Kit<sup>W/Wv</sup> mice is more pathogenic compared to intestinal slurry from wild-type mice. However, we did not search for or identify the bacterial species that causes this increased pathogenicity because we were addressing the role of mast cells in sepsis. We demonstrate an association of pathogenicity in sepsis experiments with cecal content of pathogenic bacteria (see also the new data on 16S rRNA sequencing). The same argument could be made for each bacterial species identified but this would be very complex experiments (both microbiologically and immunologically) given requirements for bacterial isolation, titration, and considerations of synergism. We therefore refrain from this undertaking.

      In the Figure 3 experiments, the authors administered the same number of cecal bacteria and did not show the number of E. coli after the administration.

      The samples were split and one aliquot was analysed by microbiology and the other aliquot was injected intraperitoneally. Fig. 3d shows the colony-forming units (for Lactobacilli and E coli) from aliquots of cecal slurry used in the intraperitoneal injection experiments shown in Fig. 3a-c. Hence, our data show the colony-forming units that were injected into the mice. It is unclear to us why this is not the key information rather than 'the number of E. coli after the administration'.

      The authors should provide evidence showing that depletion of E. coli decreases susceptibility.

      See response to point 1 above.

      (2) The author should provide direct evidence of dysbiosis by, for example, shotgun sequencing of cecal and fecal contents.

      We performed 16S rRNA sequencing of cecal contents and observed a dysbiotic shift towards an increase of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, Enterobacteriaceae, and Erysipelotrichaceae in Kit<sup>W/Wv</sup> mice compared to Kit<sup>+/+</sup> controls. None of these changes was observed when comparing the cecal microbiomes of Cpa3<sup>Cre/+</sup> and Cpa3<sup>+/+</sup> mice, indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function.

      These new findings fully align and further support with our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 10-11 and discussed on pages 13-14.

      (3) In case the authors find dysbiosis, they should analyze the mechanisms by which Kit mutation causes dysbiosis.

      We have no intention to further explore Kit biology and in particular the intestinal pathophysiology caused by the Kit mutation because any results would not affect the central conclusion of our manuscript (see title). The review process and the revision shall center on making the core of a paper as conclusive as possible, and not widen a paper by requests 'tangential to the main conclusion' (Kaelin Jr. Nature 2017).

      References:

      Caruso, R., Ono, M., Bunker, M. E., Núñez, G. & Inohara, N. Dynamic and Asymmetric Changes of the Microbial Communities after Cohousing in Laboratory Mice. Cell Rep. 27, 3401-3412.e3 (2019).

      Feyerabend, T. B. et al. Deletion of Notch1 Converts Pro-T Cells to Dendritic Cells and Promotes Thymic B Cells by Cell-Extrinsic and Cell-Intrinsic Mechanisms. Immunity 30, 67–79 (2009).

      Feyerabend, T. B. et al. Cre-Mediated Cell Ablation Contests Mast Cell Contribution in Models of Antibody- and T Cell-Mediated Autoimmunity. Immunity 35, 832–844 (2011).

      Feyerabend, T. B., Gutierrez, D. A. & Rodewald, H.-R. Of Mouse Models of Mast Cell Deficiency and Metabolic Syndrome. Cell Metab 24, 1–2 (2016).

      Kaelin Jr, W. G. Publish houses of brick, not mansions of straw. Nature 545, 387– 387 (2017).

      Moore, R. J. & Stanley, D. Experimental design considerations in microbiota/inflammation studies. Clin. Transl. Immunol. 5, e92 (2016).

      Piliponsky, A. M. et al. Basophil-derived tumor necrosis factor can enhance survival in a sepsis model in mice. Nat. Immunol. 20, 129–140 (2019).

      Ridaura, V. K. et al. Gut Microbiota from Twins Discordant for Obesity Modulate Metabolism in Mice. Science 341, 1241214 (2013).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) The study examines only a limited range of gut bacterial species, making it difficult to determine the specific contribution of E. coli expansion to the observed phenotype. A more comprehensive microbial profiling (e.g., 16S rRNA sequencing or metagenomics) would significantly strengthen the conclusions (e.g., Cpa3-mast cell deficient mice, Kit WWv, Kit W-sh/W-sh mice).

      As mentioned above, we now performed 16S rRNA sequencing of cecal contents from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Addressing microbiota in Kit<sup>W-sh/W-sh</sup> mice would not add information relevant for our paper.

      (2) The role of impaired peristalsis in KitW/Wv mice as a contributor to microbial dysbiosis and increased E. coli burden should be further explored. Complementary studies using KitW-sh/W-sh or SCF-deficient mice could clarify whether the observed microbiota changes are unique to the W/Wv model.

      We changed the title of the manuscript to emphasize our (unchanged) focus on the immunological role of mast cells in protecting against bacterial sepsis. The responses of Cpa3<sup>Cre</sup> mice clearly ruled out a role of mast cells to these infectious conditions. Our observation of altered microbiota in Kit<sup>W/Wv</sup> mice is consistent with their increased CLP susceptibility. We cite the known peristalsis deficit of Kit<sup>W/Wv</sup> mice as a possible explanation for the microbiota alterations. In the future, other investigators may find it interesting to elucidate the link between Kit mutations and dysbiosis. As stated further above, in our view, these additional questions and potential data have no bearings on the conclusions of our paper.

      (3) In the cohousing experiments, no data are provided to confirm whether microbiota normalization was achieved between groups. Including microbial composition data pre- and post-cohousing would improve the reliability of the interpretation.

      Cohousing made the susceptibility of Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice comparable. Detailed analysis of the extent of microbiota normalization would only make sense to ultimately determine specific taxa or combinations thereof that are responsible for the increased susceptibility of Kit<sup>W/Wv</sup> mice, a question that was never the goal of this study.

      (4) The use of Cpa3 Cre/+ mice introduces potential confounders, as these mice also have defects in basophils and T cells. Functional validation or additional models (e.g., Mas-TRECK or Mcpt5-Cre mice) could help isolate the mast cell-specific effects.

      See our detailed explanation above (Reviewer #1 Public review). It is incorrect to claim that Cpa3<sup>Cre/+</sup> mice have defects in T cells. The reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      (5) Clarification is needed regarding the role of mast cell-derived TNF. Given previous reports using BMMC reconstitution that implicate mast cells as a source of TNF, reconciling these findings with the current study's results would strengthen interpretation.

      We also disagree here. Experiments in normal unmanipulated mice are inevitably superior to Kit mutants after BMMC reconstitution which is an artificial system. The transplanted cells do not mature normally and don’t settle in their natural niches. We mentioned in the discussion that there is conflicting literature derived from different models and mice.

      But we do not share the expectation that results obtained with Kit-independent models, that differ from previous studies using Kit mutants, necessarily require reconciliation. The experiments are simply incomparable and the most physiologically relevant experiment will pave the way. Research on mast cell functions based on BMMC-reconstituted Kit mutants has meanwhile proven to be unreliable.

      (6) A clearer delineation between mast cell-dependent and microbiota-mediated mechanisms in the discussion sections would enhance readability and impact.

      We restructured the discussion and distinguished between Kit-dependent and mast cell-dependent phenotypes. We also discussed in detail the observed microbiota differences and their influences for the outcomes in the different sepsis experiments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Weaknesses:

      (1) The cutoffs the authors used to define "conditionally essential" mutants are not reported. The results also lack validation for lethality using a titratable system. It would be ideal to validate several genes in each dataset to determine cutoffs (i.e. 5-fold decrease in insertion mutants) for conditional lethality. It was not done (or described) here.

      We acknowledge that independent validation using targeted mutants would further strengthen the assignment of synthetic lethality. As the primary aim of this study was the genome-wide identification of genetic interactions associated with loss of fitness in several mutant backgrounds, such validation was beyond the scope of the current work. Our experiments identified hundreds of lethal combinations and we have six datasets; therefore, validation of these interactions is not feasible and is indeed not common for a publication using TraDIS to generate leads for the community to follow up on. However, we already validated some of the hits in our original submission and compared our TraDIS data to some known synthetic-lethal interactions. We have also revised the manuscript to describe all other loci as candidate synthetic-lethal interactions and have highlighted the need for future validation studies in the Discussion.

      Regarding the reviewer’s query on thresholds, candidate synthetic-lethal interactions were identified using the tradis_essentiality.R script within the BioTraDIS analytical framework independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call. We have clarified this point on line 923 in the Methods section as follows:

      “We ran the tradis_essentiality.R script within the BioTraDIS package independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study[95]). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices [30, 34]. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call.”

      (2) Also, two mutations that both make the cells sick could provide an additive effect (i.e. dapF and BamB), which doesn't necessarily mean the pathways are linked. The authors should revise their wording. They have not shown genetic linkage in some cases.

      We revised the text to address this on line 693. However, the bamC mutant demonstrates no significant fitness cost under any of the conditions tested in the manuscript. Therefore, if this is simply an additive effect then it is not clear how this occurs, especially in the case of the dapF, bamC double mutant, and we offer an alternative explanation in the Discussion based on interpretation of the literature.

      (3) Mutations throughout the manuscript are not complemented. It would be ideal to add complementation data to show the gene-phenotype relationship is specific.

      We thank the reviewers for highlighting this and have complemented the experiments for the bamB-DNA replication link observation as described in response to reviewer 3.

      (4) Also, I would argue the term "conditionally essential genes" should be replaced with "synthetically lethal". Strains were compared in the same conditions but with different genetic backgrounds.

      We take the reviewer’s point and revised the text throughout.

      Reviewer #2 (Public Review):

      Weaknesses:

      (1) An important control in any genetic interaction study is to do complementation tests to demonstrate that the phenotype observed is indeed due to the missing gene under analysis. Although the Keio library was designed to avoid polar effects, it is impossible to predict other undesirable effects of the deletions (hitting of a non-annotated sRNA or RNA stability effects, for example). Thus, before one can safely conclude that a proposed genetic interaction is real, complementation tests should be carried out. This seems particularly important in the case of a new and surprising interaction, such as that between bamB and DNA replication and repair genes.

      We thank the reviewers for highlighting this and have provided the complementation experiments for the bamB-DNA replication link observations.

      (2) Why not include the suppressor interactions in the work? There are probably plenty, and in principle, they should be as informative as the conditional essential (or synthetic lethal) ones. The only one highlighted in the paper is that between bamB and diaA, since it nicely fits with the synthetic lethal effects with initiation inhibitors seqA and hda. Even if the authors cannot make sense of the suppressor interactions, their inclusion in the paper should make the dataset richer and more valuable to the community.

      Due to the nature of the BioTraDIS pipeline, we focused on gene essentiality and so only picked up potential genes that are essential in the parent but become non-essential in the mutants. This misses observations such as that made for diaA, which we hypothesised and checked manually. The data are publicly available for readers to use for their own studies and we have included some notes in Supplementary Table 1 to explain the filtering process along with another tab including the filtered essential gene lists.

      (3) The enrichment analysis in Figure 2B deserves some clarification. What is the meaning of gene ratio? How can single genes of a pathway yield an enrichment signal? Why weren’t seqA and hda included in the DNA replication class in 2B?

      We thank the reviewer for highlighting this point and realise we did not include a section on this analysis in the Methods section. As such we have included a section on line 935. KEGG pathway enrichment analysis was performed on the conditionally essential gene sets for each mutant background using the enrichKEGG function from the clusterProfiler R package [PMCID: PMC3339379], with the whole E. coli K-12 BW25113 genome used as the background gene set. Gene ratio is defined as the proportion of genes within a given conditionally essential gene set that are annotated to a specific KEGG pathway. Enrichment significance was assessed using a hypergeometric test comparing pathway representation within each query gene set to the whole-genome background.

      SeqA and Hda were not included in the DNA replication enrichment category because the KEGG enrichment analysis was based on existing KEGG pathway annotations, in which these genes are not assigned to the DNA replication pathway despite their well-established roles in replication initiation control. Considering the revision of the results regarding DNA replication, we feel this does not warrant further changes.

      (4) The writing puts too much emphasis on demonstrating that bam lipoproteins and chaperones are specialized instead of fully redundant. However, I have the impression this is a long-settled conclusion in the field, as the manuscript itself describes at several points when reviewing the literature.

      We revised the manuscript throughout to reduce this emphasis.

      Reviewer #3 (Public Review):

      In this work, Bryant, et al. investigate genetic interactions between non-essential members of the outer membrane protein biogenesis pathway and other genes in the genome using a transposon-directed insertion sequencing (TraDIS) approach in E. coli K-12. The authors identify interactions with other components of the envelope including LPS, peptidoglycan, and enterobacterial common antigen biogenesis, and they tie these interactions to specific members of the outer membrane biogenesis pathway. Although many of these interactions are known and have been previously investigated in the field, the study provides several synthetic phenotypes that could be useful for further investigations.

      The strengths of the paper include their unbiased, TraDIS approach, and follow up on the interactions they observe. The interactions with genes of unknown function also are of interest as they may suggest experiments to find the functions of these genes. The largest weakness of this paper is the use of a gene deletion allele for bamB that is known to be polar leading to decreased expression of an essential gene. This largely invalidates all results related to DNA replication. In addition, it is a weakness that the paper does not adequately address its place in the field through discussion of existing results on the interactions they investigate.

      The bamB mutant used here has been widely used in several previous studies (Cox et al., 2017, Gunasinghe et al., 2018, Psonis et al., 2019, Storek et al., 2019, Ranava et al. 2021, Steenhuis et al., 2021, Thewasano et al., 2023) with no concern raised and so we appreciate the reviewer’s expertise here and that they highlighted this issue for us to address.

      We thank the reviewer for highlighting this issue, as we have now completed complementation experiments for the CRISPRi depletion experiments and found that expression of bamB from a pBAD plasmid does not complement the DbamB strain in which seqA or hda is depleted, but expression of der in this system does complement the phenotype. Therefore, we have revised the title and the text to remove discussion of this potential link to DNA replication. We have included the new results and revised the existing DNA replication related figures as new figures S6-S8 and included a brief discussion of this polar effect in lines 283-314. We are very grateful to the reviewer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This study aimed to determine whether bacterial translation inhibitors affect mitochondria through the same mechanisms. Using mitoribosome profiling, the authors found that most antibiotics, except telithromycin, act similarly in both systems. These insights could help in the development of antibiotics with reduced mitochondrial toxicity.

      They also identified potential novel mitochondrial translation events, proposing new initiation sites for MT-ND1 and MT-ND5. These insights not only challenge existing annotations but also open new avenues for research on mitochondrial function.

      Strengths:

      Ribosome profiling is a state-of-the-art method for monitoring the translatome at very high resolution. Using mitoribosome profiling, the authors convincingly demonstrate that most of the analyzed antibiotics act in the same way on both bacterial and mitochondrial ribosomes, except for telithromycin. Additionally, the authors report possible alternative translation events, raising new questions about the mechanisms behind mitochondrial initiation and start codon recognition in mammals.

      Weaknesses:

      All the weaknesses I previously highlighted were adequately addressed.

      We thank the reviewer for carefully considering our revision.

      Reviewer #3 (Public review):

      Summary:

      Recently, the off-target activity of antibiotics on human mitoribosome has been paid more attention in the mitochondrial field. Hafner et al applied mitoribosome profilling to study the effect of antibiotics on protein translation in mitochondria as there are similarities between bacterial ribosome and mitoribosome. The authors conclude that some antibiotics act on mitochondrial translation initiation by the same mechanism as in bacteria. On the other hand, the authors showed that chloramphenicol, linezolid and telithromycin trap mitochondrial translation in a context-dependent manner. More interesting, during deep analysis of 5' end of ORF, the authors reported the alternative start codon for ND1 and ND5 proteins instead of previously known one. This is a novel finding in the field and it also provide another application of the technique to further study on mitochondrial translation.

      Strengths:

      This is the first study which applied mitoribosome profiling method to analyze mutiple antibiotics treatment cells. The mitoribosome profiling method had been optimized carefully and has been suggested to be a novel method to study translation events in mitochondria. The manuscript is constructive and well-written.

      Weaknesses:

      This is a novel and interesting study, however, most of conclusion comes from mitoribosome profiling analysis, as the result, the manuscript lacks the cellular biochemical data to provide more evidence and support the findings.

      Comments on revisions:

      The authors addressed most of my concerns and comments, although there is still no biochemical assay which should be performed to support mitoribsome profiling data.

      The author also carefully investigated the structure of complex I, however, I am surprised that the author chose to analyse a low-resolution structure (3.7 A). Recently, there are more high-resolution structures of mammalian complex I published (7R41, 7V2C, 7QSM, 9I4I). Furthermore, the authors should not only respond to the reviewers but also (somehow) discuss these points in the manuscript.

      We thank the reviewer for suggesting additional structural analyses. Of the suggested additional structures to look at, only 9I4I from Nguyen et al. 2026 was of human-derived complex I. Other structures were derived from species that utilized alternative codons at the 3’ end of these mRNAs, preventing comparison. However, for 9I4I, the authors were able to fit density for every amino acid of all mitochondrially-encoded proteins in complex I, except the first two residues of ND1 and ND5, which, as our manuscript suggest, are not translated.

      In addition to structures, we also attempted to identify additional publicly available mass spectrometry data for complex I. However, data which we identified either derived their proteins from bovine tissue, which utilizes different codons at the initiation sites, or did not capture the N-terminal region of the proteins. Therefore, we did not include this analysis.

    1. Author response:

      The following is the authors’ response to the previous reviews

      (1) Interpretation of LC3-II accumulation and phenocopying

      Therefore, LC3-II accumulation alone is insufficient to support phenocopying in my view.

      We agree with this assessment. Upon reconsideration, we concluded that LC3-II accumulation alone does not justify the use of the term "phenocopying." We have therefore removed this language from the manuscript.

      (2) Strength of conclusions regarding autophagosome biogenesis

      As presented, the findings support a correlative relationship rather than a defined role in autophagosome biogenesis.

      We agree that our original wording overstated the strength of the conclusions. To better reflect the data, we revised the text to state that our findings "expound upon" rather than "elucidate" the role of these membranes in autophagosome biogenesis.

      (3) Title wording

      The title states that ATG2A ‘engages’ Rab1A- and ARFGAP1-positive membranes during autophagosome formation... A more descriptive term, such as ‘associates,’ would more accurately reflect the data.

      We appreciate this suggestion. We revised the title to avoid implying a causal dependency. The title now states that ATG2A interacts with Rab1A- and ARFGAP1-positive membranes, emphasizing the membrane association observed in our study rather than a direct interaction with the proteins themselves.

      (4) ARFGAP1 knockdown phenotype

      The authors claim: ‘siRNA against ARFGAP1 had very little effect’ but the quantification and blots show actually no effect.

      We agree that the original wording was imprecise. The sentence has been revised to state:

      “siRNA against ARFGAP1 had no effect on flux.”

      This more accurately reflects both the data and our original interpretation.

      (5) Interpretation of ARFGAP knockdown experiments

      Conclusions drawn from KD experiments in Fig. S2 should be interpreted with caution, as knockdown efficiency is very low, particularly for ARFGAP1/3 in the triple knockdown.

      We agree with this caution and have revised the text accordingly. The manuscript now states:

      “Knockdown of ARFGAP2, ARFGAP3 and ARFGAP1-3 marginally increased autophagic flux (Fig. S2B,C), suggesting either no role or a minor role as negative regulators of autophagy.”

      This wording more appropriately reflects the limitations of the experiment.

      (6) Discussion of ERGIC/ERES remodeling literature

      It would strengthen the manuscript to discuss previous studies reporting ERES and ERGIC remodeling and formation of ERES-ERGIC contact sites (PMID: 34561617; PMID: 28754694).

      We appreciate the suggestion. These studies were already cited and discussed in the original submission, and therefore no additional changes were required.

      (7) Figure readability

      The font size in Figure 1A and Supplementary Figure S1G is too small for comfortable reading.

      We agree and have enlarged the labels in both figures to improve readability.

      (8) Clarification of starvation conditions in figure legends

      In Figures 2A-C and Figure 4, it is unclear how the cells were treated. Were they starved in EBSS?

      We have updated the corresponding figure legends to explicitly state the starvation conditions used in these experiments.

      (9) Interpretation of ARFGAP1 knockdown and LC3 lipidation

      In Figure 2A, ARFGAP1 knockdown appears to reduce LC3 lipidation without affecting Halo-LC3 cleavage.

      We do not observe a reproducible reduction in LC3 lipidation following ARFGAP1 knockdown and therefore do not believe this conclusion is supported by the data. No changes were made in response to this comment.

      (10) Clarification of protein-protein interaction statement

      The phrase ‘but protein-protein interactions appear to be limited to RAB1’ would benefit from clarification.

      We agree and have adopted the suggested wording. The manuscript now states:

      “but stable protein-protein interactions appear to be limited to RAB1.”

      We thank the reviewers again for their constructive feedback and for helping us improve the clarity and accuracy of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      King and colleagues generated a mouse with a point mutation in IL21R and investigated the influence on IL-21-mediated T and B cell activation and differentiation. They found that mutant mice show a reduced T and B cell response, with CD4 T cell differentiation into T follicular helper cells being primarily affected.

      Strengths:

      The authors combined in vitro and in vivo analysis, including bone-marrow chimeric mice.

      Weaknesses:

      The effect of the IL21R EINS mutant does not specifically affect STAT1, as clearly shown in Figure 1 H, I. Particularly at lower doses of IL21, which may be more relevant in vivo, the effects are very similar. A second key weakness is the very small Tfh response, a not very clear PD-1 and CXCR5 staining to identify Tfh, and a lack of a steady-state (prior to immunisation) comparison of Tfh numbers in the different mouse strains. The latter makes it impossible to know what fraction of the response is antigen-specific.

      Reviewer #2 (Public review):

      Summary:

      In the manuscript, "An IL-21R hypomorph circumvents functional redundancy to define STAT1 signaling in germinal center responses," Cecile King and colleagues identify a cytoplasmic site of the IL-21 receptor that differentially regulates STAT1 and STAT3 activation upon IL-21 stimulation. They further examine the immunological consequences of this site-specific alteration on Tfh differentiation and Tfh-dependent humoral immunity, raising important questions about how geneknockout models may obscure nuanced functional roles of signaling molecules.

      Strengths:

      The study convincingly highlights a non-redundant role for STAT1 downstream of IL-21-IL-21R signaling in the Tfh differentiation pathway. This conclusion is supported by in vitro analyses of STAT1 and STAT3 activation in CD4 T cells stimulated with IL-21 or IL-6; by in vivo assessments of Tfh and germinal center B cell responses in WT and IL21R-EINS mutant mice, including bonemarrow chimera systems; and by investigating the expression of Tfh-related molecules in WT versus IL21R-EINS CD4 T cells.

      Weaknesses:

      Although the experiments were carefully executed with appropriate controls, a key question remains unresolved: whether the Tfh differentiation defect in IL21R-EINS mice is directly attributable to reduced STAT1 activation. Rescue experiments that restore STAT1 signaling in IL21R-EINS TCR-transgenic CD4 T cells would provide strong evidence linking the mutation to impaired STAT1 activation and, consequently, defective Tfh differentiation. Without such evidence, it remains formally possible that additional, uncharacterized mutations introduced during ENU mutagenesis contribute to the phenotypes observed, particularly given the discrepancies between IL21R knockout and IL21R-EINS mutant mice.

      We agree that further experiments are needed to definitively show that the effect is attributable to reduced STAT1 activation alone. Rescue experiments will be a focus of future experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure1

      I would recommend changing the conclusion in Line 141 to 'potentially less affected' rather than unaffected as there is a clear and very consistent effect on pSTAT3 in every IL21 dose tested. Also, it seems that much more IL21 is required to induce STAT1 phosphorylation, which may explain the increased effect of the EINS mutant on this signalling pathway. Similarly, for the STAT5 data in Figure S1, there is very little phosphorylation beyond baseline phosphorylation (unstimulated), but there is a clear and consistent reduction in pSTAT5 in the EINS mutants.

      Neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells in either the percentage or MFI. We have edited line 141 to state “the levels of phosphorylated STAT3 in CD4+ T cells were significantly less affected by the Il21r<sup>EINS</sup> mutation (Fig. 1I).”

      A minor point: Why is the MFI in IL21R-/- mice at 300 in panel C and at 200 in panel E? How representative is the reduced baseline pSTAT5 in IL21r-/- mice? 

      This is likely due to machine voltage during acquisition in a different experiment.

      Collectively, I would suggest concluding from that data that the EINS mutation affects IL21R signaling, which results in reduced STAT1, 3, and 5 phosphorylation, with pSTAT1 being most strongly reduced, particularly at high IL21 concentration. 

      Please also see response to above comment. Since neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells, the data does not support that conclusion.

      All subsequent data therefore do not investigate the effect of the EINS mutation on STAT1, but on overall reduced IL21R signalling. This needs to be considered when interpreting the data. For example, the text in lines 171, 172, and 174 should be adjusted as the effect is neither only dependent on STAT1 nor is STAT3 signalling intact.

      Please also see responses above. It is possible that a different method for detection of phosphorylated STAT3 and STAT5 could have looked more closely into the effect at very low concentrations of IL-21. However, our findings using Westen Blot and flow cytometry only observed a significant difference in STAT1 activation. We have edited our sentence on line 333 to state “response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1”.

      A key signalling pathway downstream of IL21R is AKT and S6 phosphorylation. It would be important to also investigate the effect of the EINS mutation on these pathways.

      We agree and this will be a focus of future experiments.

      (2) Figure 2

      PNA or BCL-6 staining would be preferable to identify GC in Figure 1A, but the flow cytometry data in Figure 3 are convincing, so this is not absolutely necessary.

      Also, no conclusions can be made here about STAT1 specifically, and there could be other reasons why Tfh are slightly and temporarily reduced in EINS mice.

      (3) Figure 3

      Line 200: Please explain what is meant by 'despite an expansion of the IgG1 FAs B cell population on day5, the percentages of EINS IgG1* GC B cells were significantly lower (Fig. 3F). I cannot see any expansion of IgG1 FAS B cells, nor can I see a specific effect on day 7. Both total GC B cells and IgG1 GC B cells are similarly affected throughout the response. Some data points may not reach statistical significance, but the trend is very clear.

      Figure 3E shows the percentage of GC B cells increasing from day 3 to day 7 in WT and from day 3 to day 5 in IL21rEINS, with significant differences between WT and IL21rEINS on days 5 and 7. In Figure 3F he percentage of IgG1+ GC B cells increase from day 3 to day 14 in both genotypes, with a significant difference between IL21rEINS on day 7. We have edited to manuscript to state “Despite an increase in the IgG1<sup>+</sup> FAS<sup>+</sup> B cell population from day 3 in response to immunization, the percentages of Il21r<sup>EINS</sup> IgG1+ GC B cells were significantly lower relative to WT cells 7 days after SRBC immunisation (Fig. 3F).”

      (4) Figure 4

      How many days after SRBC immunization was the analysis done?

      The data show an intrinsic role of IL21R signalling to Tfh development, which may include a role for STAT1. The absence of any effect on GC B cells is somewhat surprising. Chimeric and irradiated mice sometimes mount poor immune responses, and GC B cell numbers are very low. What is the frequency of GC B cells in non-immunized mice? This would be important to know if the mice responded at all, and if they did, if the 'baseline' of GC activity differs.

      As stated in the figure legend for Figure 4 – on day 7 “Mixed BM chimaeras were reconstituted with equal ratios of WT CD45.1+ BM cells and Il21rEINS CD45.2+ BM cells. 8 weeks after transfer, the mice were immunized with SRBC and analysed 7days later.

      (5) Figure 7

      Please highlight that while only IL21R-/- mice showed a significant difference in the frequency of Tfh, a similar trend was observed in WT and IL21Reins mice. The data spread is smaller in the IL21R-deficient mice, facilitating statistical significance. As throughout the manuscript, this is not a STAT1 IL-21R mutant; it is a mutant with reduced IL21R signalling. In fact, the finding that IL-6 does not compensate for the EINS mutation may suggest that STAT1 plays a minor role in the biological effects observed.

      We can only report on the statistical significance of the data we have.

      (6) Other comments

      Figure 1 F/G. I think the Y axis should read pSTAT1 and pSTAT3, respectively.

      Thank you, we have corrected the graph accordingly.

      Figure S1A. Please change the order of WT, IL21EINS, and IL21R-/- to match the main figures (IL21R-/- last). Currently, A, B, and D have a different order, but C is like the main figures.

      The figure panels are aligned to show media, then either IL-2 or IL-6 and then IL21.

      Please provide complete flow cytometry gating strategies for all figures.

      Flow cytometry dating for P-STAT1 and p-STAT3 is shown in figure 1, for Tfh cells and Tfr cells in Figure 2, 4 and 5. Please also see supplementary figures for T cell gating and methods for detailed description of antibodies and dilutions used for immunostaining.

      Reviewer #2 (Recommendations for the authors):

      Line 332 requires revision.

      We have edited the final sentence to state” Taken together these findings demonstrate that, despite the strong ability of IL-6 to activate STAT1, IL-6 is ineffective at fully compensating for the germinal centre response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully considered all comments and have revised the manuscript to address the key points raised. We have also updated the author list to include Kinga Niedobecka, who performed the additional flow cytometric validation of the engineered THP1 cell lines included in the revised manuscript. In particular, we have strengthened the validation of the THP1-CD1c system, clarified and better signposted the characterisation of CD1c-autoreactive T-cells using existing data, and refined the explanation of the mechanisms underlying enhanced responses to Mtb-infected cells. Some of the suggestions represent significant additional experimental work beyond the scope of this manuscript, and in these instances we have amended the text to clarify interpretation and limitations.

      eLife Assessment

      The study investigates how CD1c-restricted T cells respond to Mtb-infected APCs, leading to increased cytokine production and cytotoxic activity that may help control Mtb infection. While the work is important and will interest researchers in the field, the supporting evidence is incomplete and could be strengthened by additional experiments. Experiments would: (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished, (ii) enhance confidence in the purity of the CD1c-specific T cell population isolated from blood, and (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished

      We thank the Editor for highlighting this important point. We agree that it is essential to establish whether conventional MHC-mediated antigen presentation could contribute to the observed T-cell responses. To address this directly, we repeated and extended our flow cytometric validation of the engineered THP1 system. These data are now presented in an expanded Fig. 1A and include assessment of CD1c, classical MHC class I, MHC class II, β2m, CD1b and HLA-E across WT THP1, THP1-KO and THP1-CD1c cells. Our THP1-KO system is based on CRISPR-mediated knockout of both β2microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. In this new analysis, WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c. Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, CD1b and HLA-E remained undetectable by flow cytometry.

      These extended validation data support the conclusion that residual MHC expression does not account for the observed T-cell responses, which are instead dependent on CD1c expression. We have revised the relevant section of the Results to incorporate these data and to clarify that the engineered THP1-CD1c APC system provides robust CD1c expression in the absence of detectable surface MHC-I or MHC-II (revised manuscript, page 5-6, lines 111-122; Fig. 1A and Fig. 1 legend).

      (ii) enhance confidence in the purity of the CD1c-specific T-cell population isolated from blood

      We agree that confidence in the specificity and purity of the CD1cautoreactive T-cell populations is essential. The relevant data were included in the original manuscript, but we recognise that they were not signposted clearly enough. We have therefore revised the Results to describe the enrichment, sorting, post-expansion validation and functional specificity of the T-cell lines more explicitly on page 8, lines 176-191.

      CD1c-autoreactive T-cell lines were generated from two independent donors using two complementary strategies. One line was generated by expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting and expansion. A second line was generated by direct enrichment using CD1c-endo streptamers, followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. Importantly, after expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. In the main figure, both donor-derived lines are shown to be strongly CD1c-endo tetramer-positive, with post-expansion tetramer positivity of 97100% (Fig. 3A and 3C). Both lines were αβTCR+CD4+ and lacked detectable γδTCR or CD8 expression (Fig. 3B and 3D).

      We also highlight the functional validation of specificity. Both T-cell lines were activated by THP1-CD1c APCs but not THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). Importantly, the specificity of these cells was further supported by TCR transfer experiments. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines were expressed in Jurkat reporter cells and conferred CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D, page 10, lines 225244). This provides independent confirmation that the enriched T-cell line contained CD1c-reactive TCRs capable of mediating CD1c-dependent recognition.

      Together, these data support that the T-cell populations used in the functional assays are highly enriched CD1c-specific T-cell lines rather than mixed or nonspecific populations.

      (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      We agree that identifying the additional signal provided by Mtb-treated THP1-CD1c cells is an important mechanistic question. We have now revised the Discussion to clarify our interpretation and to more explicitly outline the likely mechanisms (revised manuscript, page 16-17, lines 380-400).

      Our data suggest that the enhanced response to Mtb-infected THP1-CD1c cells is unlikely to be explained simply by increased CD1c expression, generic APC activation, or soluble cytokine release. CD1c expression was maintained but not increased on THP1-CD1c cells after Mtb infection, and stimulation with TLR2 or TLR4 agonists did not reproduce the enhanced cytotoxicity observed after Mtb infection. In addition, Mtb-treated THP1-CD1c cells alone produced IL-8 and RANTES, but not the broader cytokine profile observed in T-cell co-cultures. Together, these data suggest that Mtb exposure provides an additional CD1c-dependent activating signal.

      We now discuss that this signal is most likely an altered CD1c-presented lipid repertoire on Mtb-exposed APCs. Possible mechanisms include presentation of Mtb-derived lipids, infection-induced accumulation of host-derived stimulatory “stress lipids”, presentation of bacterial and mammalian shared lipids, or altered lipid processing and trafficking during infection. These possibilities are consistent with prior studies showing enhanced responses of autoreactive CD1-restricted T-cells to microbial stimulation and our TCR transfer experiments seemingly support a CD1c-TCR-dependent recognition mechanism. However, because we have not directly identified the lipid ligands presented by CD1c on Mtb-infected APCs, we now state this as a mechanistic hypothesis rather than a conclusion, and a key outstanding question.

      We have revised the Discussion (page 16-17, lines 380-410) to make this limitation explicit. Future studies will require isolation of CD1c molecules from Mtb-infected cells and then lipidomic analysis and mass spectrometry to define the CD1c-associated lipid species.

      Reviewer #1 (Public review):

      Strengths:

      (1) This study asks an important question. The single-cell transcription analysis suggests the inherent cytotoxic program of lipid-CD1c cells and provides insights into their phenotypic and potential functional profiles. Function experiments suggest that these autoreactive T-cells can react to Mtb infection, adding to the paradigm of infection control by these non-conventional T-cell populations.

      We thank the reviewer for this positive assessment of the importance of the study and for recognising the value of the single-cell transcriptional analysis and functional experiments. We are pleased that the reviewer agrees that our findings provide insight into the cytotoxic effector programme of CD1c-autoreactive T-cells and their potential contribution to immune responses during Mtb infection.

      Weaknesses:

      (2) The study lacks sufficient rigor; conclusions may be strengthened with the incorporation of more controls, and some deeper characterization of the THP1 system and the CD1c-specific T-cells isolated from blood. Crucial conclusions are drawn from the cell mixing experiments involving the engineered THP-1 system and CD1c-lipidspecific T-cells from blood. These cells need more in-depth characterization. The expression of MHC-I/II is clearly reduced in THP1-CD1c cells. However, it is important to ensure that it is completely abolished, since a residual expression can skew the result with activation of conventional T-cells in the blood or low levels of conventional T-cells that may be present in the CD1c-tetra/multimer sorted T-cells

      We agree that this is an important point and have addressed it by adding new experimental controls and by clarifying the validation of the CD1c-autoreactive T-cell lines.

      First, we repeated flow cytometric validation of the existing markers and extended the panel to assess additional surface molecules across WT THP1, THP1-KO and THP1CD1c cells. The revised Fig. 1A therefore includes repeat staining for β2m, classical MHC class I, MHC class II and CD1c, together with newly added staining for HLA-E and CD1b.

      The THP1-KO system is based on CRISPR-mediated knockout of both β2-microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. The repeated analyses confirmed the original staining pattern, while the additional HLA-E and CD1b stains further extended validation of the system. WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c.

      Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, HLA-E and CD1b remained undetectable by flow cytometry. We have revised the Results to describe these new validation experiments more clearly (revised manuscript, page 5-6, lines 111-122, Fig. 1A and Fig. 1 legend).

      Second, we have strengthened the description of the purity and specificity of the CD1c-autoreactive T-cell lines. These lines were generated using two complementary approaches, namely expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting, and direct enrichment using CD1c-endo streptamers followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. After expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. Both donor-derived lines were strongly CD1c-endo tetramer positive, with post-expansion tetramer positivity of 97 to 100%, and both were αβTCR+CD4+ with no detectable γδTCR or CD8 expression (Fig. 3A-D). Functionally, both lines responded to THP1-CD1c APCs but not parental THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). We have revised the Results to signpost these data more clearly (revised manuscript, page 8, lines 176–191).

      Finally, TCR transfer experiments provide independent confirmation of CD1c-specific recognition. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines conferred CD1c-endo tetramer binding and CD1c-dependent activation when expressed in Jurkat reporter cells (Fig. 5B-D, revised manuscript, page 10, lines 225244). Together, the absence of detectable MHC-I/MHC-II expression in the engineered APC system, the high CD1c-endo tetramer enrichment of the T-cell lines, the lack of activation against THP1-KO cells, and the TCR transfer experiments support the conclusion that the observed responses are driven by CD1c-dependent recognition rather than residual conventional MHC-mediated activation.

      (3) Figure 2: The immunohistochemistry appears to be shown only for one biopsy; it may be worth quantifying the immunohistochemistry of all five.

      We thank the reviewer for this helpful suggestion. We agree that quantitative analysis of CD1c immunohistochemistry across all biopsies would be valuable. We examined CD1c staining across all five TB lung biopsies and observed a consistent spatial pattern, with CD1c staining generally low or infrequent in central granulomatous regions and more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas.

      However, because these were diagnostic human biopsy samples with substantial variation in tissue size, architecture, granuloma representation and inflammatory composition, we do not think that simple bulk quantification of CD1c-positive area across biopsies would be robust or biologically interpretable. In particular, quantification would be strongly affected by whether a section captured granuloma centre, peripheral inflammatory regions, lymphoid aggregates, or uninvolved lung tissue. We have therefore retained the IHC as representative spatial evidence of CD1c expression in TB lung tissue, rather than presenting it as a quantitative comparison across anatomical compartments.

      We have revised the Results to make this clearer, stating that CD1c expression was observed across the biopsies analysed but was spatially heterogeneous, with staining most apparent away from the granuloma centre and in lymphoid/inflammatory regions (revised manuscript, page 7, lines 144-151). We have also tempered the interpretation in the Discussion to avoid overstatement and now highlight systematic quantitative spatial analysis of larger tissue cohorts as an important future direction (revised manuscript, page 18, lines 427-430).

      (4) The expression of CD1 molecules goes up during the differentiation of MoDC, and Mtb infection prevents or dampens the upregulation. Does Mtb infection downregulate the CD1 expression of mature DCs? Can the effect of Mtb on the expression of CD1a,b,c molecules be investigated using CD1c-expressing DCs from blood? What could be the reason THP-1 cells do not downregulate CD1 molecules upon Mtb infection, and how about the expression of CD1a and b?

      We agree that the distinction between impaired CD1 upregulation during MoDC differentiation and active downregulation of CD1 expression on already differentiated CD1-expressing DCs is important.

      In the revised manuscript, we have clarified that our primary cell data assess the effect of Mtb infection on differentiated MoDCs that already express CD1 molecules, rather than only examining failure of CD1 induction during differentiation. Specifically, we analysed a published RNA-sequencing dataset from differentiated human MoDCs infected with live Mtb and observed reduced expression of group 1 CD1 genes, including CD1A, CD1B and CD1C, at 48 hours after infection. We then validated this experimentally at the protein level by flow cytometry, showing reduced CD1c expression on primary MoDCs after live Mtb infection. These data support the conclusion that Mtb infection can reduce CD1 expression on CD1c-expressing primary DCs.

      We agree that analysis of freshly isolated blood CD1c+ DCs would be valuable. However, these cells are rare in peripheral blood and are technically challenging to isolate in sufficient numbers for live Mtb infection assays and downstream flow cytometric or functional analysis. For this reason, we used MoDCs as a tractable primary human DC model to assess infection-induced changes in CD1 expression. We now acknowledge in the revised Discussion that validation in primary blood-derived CD1c+ DCs would be an important future direction.

      We have also clarified why CD1c expression is not downregulated in the engineered THP1-CD1c system. In primary DCs, Mtb-mediated suppression of CD1c has been linked to host regulatory mechanisms, including post-transcriptional regulation by miRNAs such as miR-381-3p, which targets the 3′ UTR of endogenous CD1c transcripts. In contrast, CD1c expression in our THP1-CD1c cells is driven by a lentiviral CD1c-β2m fusion construct under a heterologous promoter and expressed from a cDNA lacking the native untranslated regions. Therefore, CD1c in this system is not expected to be regulated in the same way as endogenous CD1c in primary DCs.

      This is a deliberate feature of the model. It allows us to assess CD1c-dependent T-cell responses to Mtb-infected APCs without the confounding effect of infection-induced CD1c loss. THP1-CD1c cells do not express endogenous CD1a or CD1b because the parental THP1-KO cells lack β2m-dependent endogenous CD1 surface expression, and only CD1c is reintroduced through the CD1c-β2m fusion construct. We have revised the Results and Discussion to clarify these points (revised manuscript, pages 7- 8, lines 165174 and pages 17- 18, lines 411-430).

      (5) Figure 3: (F) What does the X-axis read for the no infection group? The value for MOI = 0 should be incorporated for the infected T-cell group.

      We agree that the original presentation could be clearer. The uninfected condition corresponds to MOI = 0, whereas the remaining points represent THP1-CD1c APCs exposed to increasing amounts of UV-killed Mtb. We have retained the figure layout but revised the figure legend to clarify that the MOI values on the x-axis apply only to the Mtb-treated conditions, and that the no-infection/no-treatment control represents MOI = 0 (Fig. 3 legend).

      (6) Figure 4: In the lysis assay, THP1-CD1c cells (uninfected and infected) incubated alone should be incorporated.

      We agree that APC-only controls are essential for interpreting the lysis assay, and we apologise that this was not sufficiently clear in the original manuscript. THP1-CD1c cells cultured alone, both uninfected and Mtb-infected, were included in all assays and used to define baseline target-cell viability for each matched condition.

      The data in Fig. 4 are presented as specific lysis to isolate the effect of T-cells on target cell viability. Specifically, THP1 viability in APC-only wells was used as the baseline and subtracted from the corresponding T-cell co-culture condition within the same experiment. Thus, lysis of uninfected THP1-CD1c cells was calculated relative to uninfected THP1-CD1c cells cultured alone, and lysis of Mtb-infected THP1-CD1c cells was calculated relative to Mtb-infected THP1-CD1c cells cultured alone. This presentation allows the T-cell-mediated effect to be visualised while accounting for baseline viability differences.

      We have revised the Methods and Fig. 4 legend to make this calculation more explicit (revised manuscript, page 25, lines 608-614; Fig. 4 legend).

      (7) A quantitative brief on the single cell TCR sequencing - including how many T-cells were sequenced and the frequency of different clone including EM1 and EM2 - should be shown.

      We agree that the single-cell TCR sequencing data required clearer quantitative description. We have expanded the Results and Fig. 5 legend to include the number of single cells analysed and the frequency of the dominant clonotypes. Single CD1c-endo tetramer-positive T-cells were sorted into individual wells for targeted TCR sequencing. After filtering and manual curation, 11 single cells yielded productive paired αβ TCR sequences. The repertoire was oligoclonal, with two dominant productive clonotypes accounting for 10 of 11 paired TCRs. EM1 was detected in 6 of 11 cells and EM2 was detected in 4 of 11 cells. These data support the selection of EM1 and EM2 for TCR-transfer experiments and clarify that they were dominant clonotypes within the CD1c-endo tetramer-positive T-cell line rather than arbitrarily selected TCRs. We have revised the Results and Fig. 5 legend accordingly (revised manuscript, page 10, lines 225-235, Fig. 5 legend).

      Reviewer #1 (Recommendations for the authors):

      (8) Perform an experiment to assess activation of T-cells expressing EM1 or EM2, upon mixing with CD1c-expressing dendritic cells isolated from human blood, with and without Mtb infection.

      We agree that testing EM1 and EM2 TCRs against primary dendritic cells is an important question. However, in the specific context of Mtb infection, the proposed experiment is difficult to interpret because Mtb downregulates CD1c expression on primary dendritic cells. This is supported by previous studies showing that Mtb and BCG suppress CD1c expression on DCs [1,2], and by our own data showing reduced CD1 group 1 transcript expression in Mtb-infected MoDCs and reduced CD1c protein expression on primary MoDCs following Mtb infection (Fig. 2B and 2C). Therefore, mixing EM1- or EM2-expressing Jurkat T-cells with Mtb-infected primary CD1c-expressing DCs would introduce a major confounder: reduced T-cell activation could reflect loss of CD1c expression rather than absence of a CD1c-dependent Mtb-induced activating signal. This is precisely why we used the engineered THP1-CD1c system, in which CD1c expression is preserved during Mtb infection (Fig. 2D). This model allowed us to test whether Mtb infection enhances CD1c-TCR-dependent activation without the confounding effect of infection-induced CD1c loss.

      Using this controlled system, we show that EM1 and EM2 TCRs confer CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, activation in response to THP1-CD1c but not THP1-KO APCs, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D). These data support the conclusion that the enhanced response to Mtb-infected APCs is mediated through CD1c recognition by the TCR.

      We have revised the Results and Discussion to clarify this rationale and to explain why the engineered THP1-CD1c system was necessary for these experiments (revised manuscript, page 10, lines 242-244; page 17-18, lines 411-430).

      (9) Conduct an experiment to assess T-cell cytotoxicity expressing EM1 or EM2, in the presence and absence of Mtb infection.

      We agree this would be a valuable experiment. As outlined in our response to 8, EM1 and EM2 were cloned into Jurkat T-cells to test TCR-dependent CD1c recognition and activation, not cytotoxic effector function. Jurkat T-cells are not cytotoxic effector cells<sup>3</sup>, so they are not suitable for target-cell killing assays.

      Cytotoxicity was instead assessed using the original CD1c-autoreactive T-cell lines (Figs. 3F-G and 4D-E). Testing EM1- or EM2-mediated killing would require engineering and validating primary human T-cells expressing these TCRs, which is a substantial additional workflow. We have clarified in the revised manuscript that the EM1/EM2 experiments demonstrate TCR-dependent recognition, while cytotoxicity was assessed using the CD1c-autoreactive T-cell lines (revised manuscript, page 10, lines 234–244).

      (10) A list of primers used for TCR sequencing should be provided.

      We have now provided the primer sequences used for targeted single-cell TCR sequencing in a new supplementary table (Table S1). We have also revised the Methods to provide additional detail on the single-cell TCR sequencing workflow, including CD1c-endo tetramer-guided single-cell sorting, oligo-dT reverse transcription, universal cDNA amplification, targeted amplification of TCR variable regions using TRAC-, TRBC-, TRGC- and TRDC-specific primers, well-specific 8-bp barcoding, size selection, library preparation and MiSeq sequencing. In addition, we now cite the SMART-seq2 protocol on which the approach was based (Picelli et al., 2014) (revised manuscript, page 20, lines 491-503; new Table S1).

      Reviewer #2 (Public review):

      Strengths:

      (1) The study is designed well and has developed many exciting tools to generate specific information.

      We thank the reviewer for this positive assessment of the study design and for recognising the value of the experimental tools developed in this work.

      Weaknesses:

      (2) The study has weaknesses in two important parameters - novelty and relevance in controlling TB. Further, the results could be better presented and discussed to allow easy understanding of the experimental design

      We accept that the novelty and relevance to TB control could be made clearer in the manuscript. However, we believe the study makes several important and previously unreported contributions, and we have revised the Introduction, Results and Discussion to improve the clarity of the experimental design and to state the conceptual advance more explicitly.

      First, to our knowledge, this is the first study to demonstrate that human CD1c-autoreactive T-cells respond more strongly to Mtb-infected CD1c+ APCs than to uninfected CD1c+ APCs. Previous work has shown that CD1c-autoreactive T-cells exist in human blood and can respond to CD1c-expressing cells in the absence of exogenous antigen. However, their role during infection has remained unclear. Our findings extend the field beyond the established steady-state, autoimmune and tumour contexts of CD1c autoreactivity by identifying Mtb-infected APCs as a biologically relevant setting in which these cells acquire enhanced effector activity. We propose that CD1cautoreactive T cells may not simply represent autoreactive bystanders that become pathogenic in disease, but instead form an evolutionarily conserved arm of lipid immune surveillance that can detect infection-associated changes in antigen presentation. Given the long-standing selective pressure imposed by microbial infection throughout human evolution, it is plausible that protection against infection represents a central physiological function of these cells, with their roles in autoimmunity and cancer reflecting the same capacity to sense altered self-lipid landscapes in other settings. Our data provide initial functional evidence supporting this model.

      Second, the study provides functional evidence that these cells are not simply activated by infected APCs, but can mediate effector functions relevant to antimicrobial immunity. CD1c-autoreactive T-cells showed enhanced activation, cytokine production and cytotoxicity in response to Mtb-infected APCs, and led to reduced Mtb burden under in vitro conditions. These findings are directly relevant to TB immunity because cytotoxic T-cell pathways and antimicrobial molecules such as granulysin have been implicated in control of intracellular Mtb.

      Third, the study links these functional observations to the ex vivo biology of human CD1c-autoreactive T-cells. Single-cell transcriptomic profiling demonstrates that these cells are enriched for cytotoxic effector-memory programmes and express molecules associated with target-cell killing and antimicrobial activity. This provides an independent, unbiased cellular basis for the functional assays and strengthens the conclusion that CD1c-autoreactive T-cells represent a plausible effector population in anti-mycobacterial immunity.

      We agree that the experimental design needed clearer presentation. In the revised manuscript, we have improved signposting of the stepwise logic of the study: (1) defining and validating the THP1-CD1c APC system, (2) demonstrating CD1cautoreactive T-cell enrichment and specificity, (3) testing responses to UV-killed and live Mtb, (4) confirming TCR-dependent CD1c recognition using EM1 and EM2 TCR transfer, and (5) integrating these functional data with single-cell transcriptomic profiling of ex vivo CD1c-autoreactive T-cells. These revisions aim to make the experimental design easier to follow and to clarify how each section supports the overall conclusion.

      We have revised the Introduction and Discussion accordingly to more clearly state the novelty and TB relevance of the work (revised manuscript, page 4-5, lines 83-104; page 14, lines 323-329).

      (3) At several places, UV-killed or live Mtb were used. What is the rationale behind that?

      We have now added a Methods statement explaining that UV-killed Mtb was used for controlled exposure to defined amounts of Mtb-derived antigen, particularly in dose-response cytotoxicity and cytokine-release assays, whereas live Mtb was used to assess T-cell activation, target-cell lysis and relative bacterial burden during APC infection with proliferating Mtb. We have also ensured that the figure legends clearly specify whether UV-killed or live Mtb was used in each experiment (revised manuscript, page 22, lines 534-540).

      (4) Why use irradiated THP1-CD1c cells for activating T-cells?

      Irradiated THP1-CD1c cells were used only during the T-cell expansion phase to provide sustained CD1c-mediated stimulation while preventing proliferation of the THP1 APCs. This was necessary because the expansion cultures lasted up to 12 days, during which non-irradiated THP1 cells would continue to divide and could overgrow the T-cell culture. Irradiation therefore allowed THP1-CD1c cells to function as APCs while maintaining controlled culture conditions and enabling selective expansion of CD1c-reactive T-cells. We have clarified this rationale in the Methods (revised manuscript, page 19-20, lines 472-474).

      (5) While functional assays identified only CD4+ cells as CD1c-restricted, scRNAseq shows that both CD4+ and CD8+ cells exhibit this phenotype

      We agree with the reviewer’s observation and have clarified this point in the revised Discussion. The functional assays were performed using CD1c-autoreactive Tcell lines generated from two donors. Both lines were CD4+αβTCR+, reflecting the outcome of the enrichment, sorting and expansion process used to generate sufficient T-cells for functional assays. These lines therefore provide mechanistic evidence that CD4+ CD1c-autoreactive T-cells can recognise CD1c+ APCs and respond more strongly to Mtb-infected APCs, but they are not intended to represent the full diversity of the CD1c-autoreactive T-cell compartment.

      By contrast, the single-cell RNA-seq analysis was designed to provide a broader ex vivo assessment of CD1c-endo-binding T-cells without relying on prolonged in vitro expansion. This revealed that CD1c-autoreactive T-cells include both CD4+ and CD8+ populations, with enrichment of cytotoxic effector-memory programmes. We therefore interpret the functional and single-cell datasets as complementary: the functional assays provide mechanistic validation using tractable CD1c-reactive T-cell lines, while the single-cell data demonstrate that the broader ex vivo CD1c-autoreactive compartment is phenotypically diverse and includes both CD4+ and CD8+ cytotoxic populations.

      We have revised the Discussion to clarify this point and to emphasise that combining in vitro functional assays with ex vivo single-cell profiling allowed us to capture both mechanistic activity and broader cellular diversity (revised manuscript, page 14-15, lines 341-349).

      (6) Identifying the specific lipid antigen presented by CD1c could add greater value to the study.

      We agree that identifying the specific CD1c-presented lipid antigen(s) would add important mechanistic insight. As noted in the response to the Editor, our data suggest that the enhanced response to Mtb-infected THP1-CD1c APCs is most likely due to altered CD1c-associated lipid presentation. However, defining these lipid species would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomics, which is a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight lipid identification as an important next step (revised manuscript, page 16-17, lines 380-410).

      (7) Since autoreactivity was independent of exogenous antigen, the cytotoxic activity should also be independent of exogeneous antigens? What additional signal a THP1-CD1c cells treated with UV-killed Mtb express that is absent from the untreated cells?

      CD1c-autoreactive T-cells likely recognise self-lipids presented by CD1c, but our data show that this response is enhanced after Mtb exposure. We interpret this as evidence that Mtb alters the quality or abundance of CD1c-associated lipid ligands, potentially through Mtb-derived lipids or infection-induced changes in host lipid metabolism. We have revised the Discussion to clarify that the precise lipid ligand(s) remain unidentified and will require future CD1c-lipidomic analysis (revised manuscript, page 16-17, lines 380-410).

      (8) The relative Mtb growth assay is confusing. CD1c cells with Mtb infection triggers massive lytic response, as shown in Figure 4. Under similar conditions, in Figure 6, the authors report a significant decline in Mtb growth in these cells. The problem is that with the kind of lytic response observed, a lot more Mtb could be present extracellularly and would evade killing. How do we reconcile the two observations?

      In the Mtb lux assay, extracellular bacteria were removed by washing after the initial infection step, so the starting bacterial population measured in the co-culture assay is expected to be predominantly cell-associated. We also recognise that the luminescence readout measures total viable lux-expressing Mtb under the assay conditions and does not distinguish intracellular from extracellular bacteria at later time points.

      The cytotoxicity observed in Fig. 4 reflects enhanced but incomplete lysis of infected APCs. Therefore, although T-cell-mediated lysis could release some bacteria from infected target T-cells, the reduced luminescence observed in Fig. 6C indicates a lower net viable Mtb burden under these co-culture conditions. We interpret this as the combined outcome of CD1c-autoreactive T-cell effector activity, including cytotoxicity and antimicrobial mediators such as granulysin and cytokines, rather than as direct evidence of selective intracellular bacterial killing.

      We have revised the Results and Methods to clarify that the Mtb lux assay measures relative viable bacterial burden/luminescence under in vitro co-culture conditions. We have changed the discussion to avoid over-interpreting this assay as distinguishing intracellular from extracellular Mtb killing (revised manuscript, page 11-12, lines 266274, page 26, lines 633-637).

      Reviewer #2 (Recommendations for the authors):

      (9) Nearly 40-50% of the samples did not respond to THP1-CD1 stimulation. What contributes to this diversity?

      We agree that there is clear donor-to-donor variability in the response to THP1-CD1c stimulation [4]. Approximately one-third of donors did not show detectable expansion under these assay conditions. This likely reflects differences in the precursor frequency and TCR repertoire composition of CD1c-autoreactive T-cells between donors, together with variation in activation state and responsiveness during short-term in vitro expansion. Apparent non-response may also reflect low-frequency CD1c-reactive populations that are present but fall below the detection threshold. We have revised the Discussion to acknowledge donor heterogeneity as an expected feature of primary human CD1c-autoreactive T-cell responses (revised manuscript, page 14-15, lines 341-349).

      (10) For lung biopsy staining, how is CD1 expression in healthy tissue or some unrelated inflammatory condition?

      The purpose of the lung biopsy staining was to determine whether CD1c-expressing cells are present in human TB lung tissue and to assess their spatial relationship to granulomatous inflammation, rather than to perform a formal comparison between healthy, non-TB inflammatory and TB lung tissue.

      Across the TB biopsies analysed, CD1c staining was spatially heterogeneous. CD1c expression was generally low or infrequent in central granuloma regions and in tissue regions remote from granulomatous inflammation, whereas staining was more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich regions. We have revised the Results to clarify that these data are presented as representative spatial observations within TB lung tissue, rather than as a quantitative comparison with healthy or unrelated inflammatory tissue.

      We agree that comparison with healthy lung and non-TB inflammatory lung tissue would provide useful additional context, particularly for distinguishing TB-associated changes from more general inflammatory induction of CD1c. We now acknowledge this as an important future direction (revised manuscript, page 18, lines 425-430). We have also revised the Results to clarify the spatial pattern of CD1c staining within TB lung tissue (revised manuscript, page 7, lines 144-151).

      (11) What was the rationale for using UV-killed or live Mtb for different experiments?

      This point is addressed in our response to 3 above.

      Reviewer #3 (Public review):

      Strengths:

      (1) The manuscript is well written, and the novelty, impact, and limitations of this study are precisely highlighted by the authors.

      We thank the reviewer for this positive assessment of the manuscript, particularly their recognition of the study’s novelty, impact and balanced discussion of its limitations.

      (2) Lipid antigen identification and direct lipid identification via lipidomics/MS of CD1c-bound lipids from Mtb-infected APCs would clarify whether the enhancement arises from altered self-lipids or subtle Mtb lipids

      We agree that direct identification of CD1c-bound lipids from Mtb-infected APCs would provide important mechanistic insight and help determine whether enhanced activation reflects altered self-lipids, Mtb-derived lipids, or shared lipid species. As noted above, this would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomic analysis, which represents a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight CD1c-lipidomic analysis as an important next step (revised manuscript, page 16-17, lines 380-410).

      Reviewer #3 (Recommendations for the authors):

      (3) Figure 2Ai-vi, lines 134-136. The authors should include the data from central granuloma staining to solidify their claim of the presence of CD1c expression remote from the centre of TB granulomas.

      Central granuloma regions are included in Fig. 2A, including panels showing staining within granulomatous tissue where CD1c expression is low or infrequent compared with distal inflammatory and lymphoid/B-cell follicle-rich regions. We agree that this spatial distinction was not sufficiently clear in the original text and figure legend.

      We have therefore revised the Results and Fig. 2 legend to more explicitly guide the reader through the central versus distal regions shown in Fig. 2A. The revised text now states that CD1c expression was observed across lung biopsies from all five TB patients, but was spatially heterogeneous, with staining most apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas, and generally low or infrequent in central granuloma regions (revised manuscript, page 7, lines 144-151; Fig. 2 legend).

      (4) Figure 2D, lines 149-151. The authors should clarify whether the CD1c resistance to downregulation is model-specific to THP1-CD1c-APCs or an overexpression artefact

      As described in our response to Reviewer 1, 5, we agree that this point required clearer explanation. We have clarified in the revised Results and Discussion that the preservation of CD1c expression in THP1-CD1c APCs likely reflects the engineered nature of this system, rather than a general feature of endogenous CD1c regulation during Mtb infection.

      Specifically, in primary MoDCs, Mtb infection reduces CD1c expression at both transcript and protein levels (Fig. 2B and 2C). In contrast, CD1c in THP1-CD1c APCs is expressed from a heterologous CD1c-β2m fusion construct rather than from the endogenous CD1C locus. Therefore, its resistance to downregulation is likely model-specific and related to the expression system. We now state this in the Results and Discussion, and explain that this feature allows CD1c-dependent T-cell responses to Mtb-infected APCs to be assessed without the confounding effect of infection-induced CD1c loss (revised manuscript, page 7-8, lines 165-174; page 17-18, lines 411-430).

      (5) Figure 6C. The relative Mtb burden is measured through luminescence. While this correlates closely with CFUs, confirmation with plating is better evidence.

      We have previously shown close correlation between luminescence in our system and CFUs<sup>5</sup> (Bielecka mBio 2017, PMID: 28174307). Perhaps controversially, we propose that luminescence is a better readout of total Mtb load. Luminescence captures all metabolically active Mtb, whilst CFUs may be confounded by clumping of bacteria, for example, giving an underestimate. However, we agree with the reviewer that luminescence is an indirect measure of bacterial burden and that CFU plating would provide additional confirmatory evidence. We have therefore revised the Results, Discussion and Methods to describe the assay more cautiously as a measure of relative viable Mtb burden/luminescence, and we now acknowledge the absence of CFU confirmation as a limitation of the study (revised manuscript, page 11-12, lines 266-274; page 17, lines 400-403; page 26, lines 633-637).

      (6) Figure 7. The authors show that CD1c-autoreactive T-cells exhibit cytotoxic effector memory phenotype. While the sc-RNAseq subsampling is robust, the number of sample donors being 2 might create a potential bias.

      We agree that the use of two donors for the single-cell RNA-seq analysis is a limitation and could introduce donor-specific bias. We have now stated this more explicitly in the Discussion. Importantly, the scRNA-seq data are not used alone to define function, but rather provide an ex vivo phenotypic framework that complements the functional assays showing CD1c-dependent activation, cytokine production, cytotoxicity and reduced relative Mtb burden.

      The subsampling analysis supports the robustness of the transcriptional patterns within this dataset, but we agree that larger donor cohorts will be required to determine how consistently these cytotoxic effector-memory programmes are represented across the broader human CD1c-autoreactive T-cell compartment. We have revised the Discussion accordingly (revised manuscript, page 17, lines 403-410).

      (7) The authors should on how it might compare with non-autoreactive CD1c-restricted T-cells.

      We agree that it is important to place these findings in the context of non-autoreactive, antigen-specific CD1c-restricted T-cells. Previous studies have shown that CD1c can present microbial lipid antigens, such as mycobacterial lipids, to T-cells with defined antigen specificity. In contrast, the CD1c-autoreactive T-cells studied here recognise endogenous ligands and appear to respond to infection through changes in the CD1c-presented lipid repertoire rather than through recognition of a single defined foreign antigen.

      Our findings suggest that autoreactive CD1c-restricted T-cells may provide a complementary mode of immune surveillance, capable of sensing infection-induced changes in lipid presentation, whereas non-autoreactive CD1c-restricted T-cells may respond more directly to specific microbial lipid antigens. A head-to-head comparison would clearly be very interesting, but an extensive new piece of work beyond the scope to the current study. We have expanded the Discussion to more clearly highlight this distinction, and that direct comparison is required (revised manuscript, page 16-17, lines 380-410).

      Concluding remarks

      In summary, we have addressed the key concerns raised by the reviewers by strengthening validation of the experimental system, improving characterisation of T cell populations, and clarifying mechanistic interpretation. We believe these revisions significantly improve the clarity and rigour of the manuscript. We accept that identification of the CD1-presented lipids is an important next step that will give significant mechanistic insight, but is beyond the scope of the current work.

      References:

      (1) Wen, Q. et al. MiR-381-3p Regulates the Antigen-Presenting Capability of Dendritic Cells and Represses Antituberculosis Cellular Immune Responses by Targeting CD1c. J Immunol 197, 580-589 (2016). https://doi.org/10.4049/jimmunol.1500481

      (2) Gagliardi, M. C. et al. Bacillus Calmette-Guerin shares with virulent Mycobacterium tuberculosis the capacity to subvert monocyte differentiation into dendritic cell: implications for its efficacy as a vaccine preventing tuberculosis. Vaccine 22, 3848-3857 (2004). https://doi.org/10.1016/j.vaccine.2004.07.009

      (3) Grailer, J. et al. A Novel Cell-based Luciferase Reporter Platform for the Development and Characterization of T-Cell Redirecting Therapies and Vaccine Development. J Immunother 46, 96-106 (2023). https://doi.org/10.1097/CJI.0000000000000453

      (4) Guo, T. et al. A Subset of Human Autoreactive CD1c-Restricted T Cells Preferentially Expresses TRBV4-1(+) TCRs. J Immunol 200, 500-511 (2018). https://doi.org/10.4049/jimmunol.1700677

      (5) Bielecka, M. K. et al. A Bioengineered Three-Dimensional Cell Culture Platform Integrated with Microfluidics To Address Antimicrobial Resistance in Tuberculosis. mBio 8 (2017). https://doi.org/10.1128/mBio.02073-16

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks whether synapses formed by the same broad neuronal class (excitatory pyramidal neurons, PN) adapt their presynaptic organization in a cortex-specific manner, comparing the prefrontal cortex (PFC) with the primary somatosensory cortex (S1). The authors combine sophisticated electrophysiology (paired recordings and extracellular minimal stimulation), pharmacological perturbations of presynaptic Ca<sup>2+</sup>-secretion coupling, bouton Ca<sup>2+</sup> imaging, and mechanistic modeling. Across two prominent excitatory connections (Layer 5 (L5) PN-L5PN and L2/3-L5PN), they provide convergent evidence that mature PFC synapses operate with looser Ca<sup>2+</sup> channel-release sensor coupling than their S1 counterparts.

      Overall, the study provides an appealing mechanistic link between synaptic nano/micro-architecture and cortical-area specialization. The idea that PFC synapses retain a more "plasticity-favoring" presynaptic state, while the primary sensory cortex emphasizes reliability and timing precision, is potentially impactful for how we think about circuit computation and plasticity across cortical hierarchies.

      Strengths:

      A major strength is the multi-pronged experimental strategy. The paper first establishes robust, area-dependent differences in synaptic efficacy, reliability, timing, and short-term plasticity (facilitation prevailing in PFC versus depression in S1), using both paired recordings and minimal extracellular stimulation paradigms. The coupling interpretation is then directly supported by differential sensitivity to EGTA (and appropriate positive-control effects of fast chelators). Finally, volume-averaged calcium signals are reported to be similar across areas, arguing against trivial explanations based on gross differences in calcium influx, and the modeling provides a quantitative framework for interpreting the observed chelator effects.

      Weaknesses:

      Limitations are minor and concern interpretation/clarity rather than core results. Some key inferences rely on indirect readouts (chelator sensitivity, fluctuation analysis-derived parameters, bouton-averaged calcium signals), each of which carries assumptions and potential confounds that should be discussed more explicitly. In particular, the repatching paradigm for the paired-recording EGTA experiment, though very impressive, and the limited number of extracellular calcium conditions used for fluctuation analysis (three concentrations), can influence quantitative estimates and the confidence intervals around them.

      We would like to thank the reviewer for his/her overall positive assessment of our manuscript and the constructive advice, which helped us to improve our manuscript. We discussed the limitations, assumptions and potential confounding factors in more detail. We addressed them pointwise in the recommendations for the authors.

      Reviewer #2 (Public review):

      Schwarze et al. investigated whether synaptic efficacy is brain-region specific. To this end, they compared synaptic connections established by layer 5 (L5) neocortical pyramidal cells and between L5 and L2/3 pyramidal cells. In order to identify the mechanism of this brain region specificity, the authors employed several experimental approaches, including paired electrophysiological recordings, extracellular stimulation, low- and high-affinity intracellular calcium chelators (EGTA and BAPTA), multiple probability fluctuation analysis (MPFA), and intracellular measurements of calcium transients as well as computational modelling. The findings of the present study indicate that synaptic connections in the primary somatosensory cortex (S1) are significantly stronger and more reliable than those in the prefrontal cortex (PFC).

      The study is timely, and the topic is of significant interest to the neuroscience community. Despite the extensive research that has been carried out on the neuroanatomy and receptor distribution of different brain regions, comparatively little attention has been paid to differences in synaptic physiology. The authors' approach is characterised by its elegance and comprehensive nature, and the conclusions drawn are compelling. Nevertheless, there are a number of unresolved issues.

      First, we would like to thank the reviewer for his/her detailed survey of our work, which was very helpful in improving our manuscript. We are happy about the overall positive evaluation and the constructive comments. To fully clarify all points, we performed new experiments and analyses, in particular we determined EGTA sensitivity in PFC and S1 from the same animal and we performed MPFA with an additional extracellular Ca<sup>2+</sup> concentration. We extended the discussion on the examined cell types. Overall, we carefully revised the manuscript to address all points. Please see below our point-wise response.

      Major points:

      (1) The authors state that data from the S1 cortex were obtained in a previous study. In the context of an explicitly comparative study (PFC vs. S1cortex), it would have been advantageous for the authors to perform a subset of experiments in which both cortices were obtained from a single animal. This is a feasible undertaking, given the spatial separation of the PFC and S1 cortex.

      This is only true for the paired recordings from L5PN-L5PN connections in S1, which were obtained in a previous study and partially reanalyzed. All recordings from L2/3-L5PN connections in S1 and PFC as well as the paired recordings on L5PN-L5PN synapses in PFC were obtained in the present study. To make this clearer, we have added Table 1. This lists which data and associated figures are from this study and which are from previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      Our experiments are lengthy and therefore it is challenging to achieve two successful recordings within the lifetime of acute brain slices. For this reason, the previous version of the manuscript did not include recordings from PFC and S1 of the same animal. We have now measured EGTA effects in L2/3-L5PNs from S1 and PFC of the same animal. Two new recordings were added to the EGTA-AM plots in Figure 2C-F and in the results section. Example recordings are shown in Figure S3A and B, as is the comparison of EGTA-AM effects in L2/3-L5PNs from PFC and S1 in Figure S3C, with data points from the same animal marked.

      “On the other hand, EGTA significantly reduced EPSC amplitudes only in PFC (0.54, 0.47-0.67; 55% of control) but not in S1 (0.85, 0.81-1.04; 100% of control). For a better comparison of EGTA effects some recordings were performed in PFC and S1 derived from the same animal to rule out interindividual effects (see example recording in Figure S3A-C).”

      (2) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We thank the reviewier for this helpful note. We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (3) PFC and S1 cortex in rodents differ markedly in their morphological organisation. For example, in all sensory cortices, layer 4 is very pronounced; however, in the PFC of rodent,s no clear layer 4 can be found. On the other hand, PFC shows a clear separation of layers 2 and 3, which is not visible inthe S1 cortex. Furthermore, PFC pyramidal cells in layers 2, 3, and 5 exhibit significant heterogeneity, diverging considerably from those found in layers 5a and 5b of S1 cortex. Thus, there is no clear correlation between L5 pyramidal cells in the PFC and the S1 cortex. In order to achieve a meaningful comparison of the data obtained in PFC and S1 cortex, it is necessary for the authors to determine whether the record is from similar pyramidal cell populations.

      (3) In addition, PFC pyramidal cells in layer 2, 3 and 5 are highly heterogeneous and differ markedly from those in layer 5a and 5b of S1 cortex. To achieve a meaningful comparison of the data obtained in the PFC and the S1 cortex, the authors need to determine whether the record from similar pyramidal cell populations.

      We apologize for having not been precise about the specific location and type of pyramidal neurons in the original manuscript. Extracellular stimulation in PFC and S1 was always performed in layer 2, where the first large cell bodies, relative to the pia mater, are located within a cortical column. Therefore, we assume that the same cell populations were stimulated in both brain regions (van Aerde & Feldmeyer, Cereb. Cortex 2015; Oberlaender et al., Cereb. Cortex 2012; Lefort et al, Neuron 2009). To stick with the standard terminology, we refer to it in the manuscript as upper layer 2/3. We kept stimulation intensity as low as possible to ensure that only a few presynaptic cells within the target region were activated.

      In S1, paired recordings were obtained from pyramidal neurons in layer 5A following the procedures and criteria described in detail in our previous work (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019; Bornschein et al., Science 2025). Briefly, these criteria are as follows: close proximity to layer 4 and the barrels as well as the PPR of 0.78 (0.69-0.90), which is consistent with depression dominating in L5A (Frick et al., Cereb. Cortex 2008; Bornschein et al., Front. Syn. Neurosci. 2019) and different from L5B with PPR ≥ 1 (Lefort & Petersen, Cereb. Cortex 2017). Within layer 5A we did not attempt to further differentiate between pyramidal neuron types. For recordings in S1 with extracellular stimulation we focused on the same locations as for the paired recordings. We extended the corresponding section in Materials and Methods of the revised manuscript.

      We agree that there is no clear layer 4 in PFC, making the distinction between layer 2/3 and layer 5 less clear. Layer 2/3 and layer 5 have approximately the same diameter (van Aerde & Feldmeyer, Cereb. Cortex 2015). Based on this, we performed recordings in upper layer 5 of PFC. We did neither morphologically nor electrophysiologically differentiate between pyramidal neuron cell types. It should be noted that within a cortical area (S1 or PFC), we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN-to-L5PN synapses, neither with regard to the EGTA-sensitivity of release nor with regard to the release probability. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would suggest stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighbouring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among pyramidal neurons appear to not be reflected on the level of their synapses.

      The Reviewer probably refers to such differences and heterogeneity in morphology and spiking patterns of pyramidal neurons. If he/she has more specific differences in mind, it would be helpful if references for the significant heterogeneity could be given.

      Please also note that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable for analyzing population differences among pyramidal neuron types.

      We refer to the problem of pyramidal neuron heterogeneity in the revised manuscript in the discussion.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).“

      “We did neither morphologically nor based on spiking patterns differentiate further between PN subtypes within a given layer. However, within a cortical area (S1 or PFC) we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN to L5PN synapses, neither with regard to the EGTA sensitivity of release nor with regard to p<sub>N</sub>. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would indicate stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighboring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among PNs appear to be not reflected on the level of their synapses.“

      (4) For the S1 cortex, in rats it has been found that L5 synaptic connection between pairs of L5a pyramidal cells and pairs of L5b pyramidal cells differ markedly with respect to mean EPSP amplitude, latency and coefficient of variation (cv, a surrogate measure for the synaptic release probability) (cf. Markram et al., 1997; Frick et al., 2008). It is therefore likely that PFC and S1 pre- and postsynaptic pyramidal cells are not only morphologically and electrophysiological distinct but also with respect to their synaptic properties. At least, the authors need to discuss these confounding issues and preferentially address them experimentally. For example, it would be helpful to demonstrate that paired recordings were made from the same pyramidal cell types, perhaps by documenting their morphology and/or firing patterns. In addition, they should discuss the marked difference in EPSP amplitude and putative release probability between their data and the earlier studies.

      We agree that Markram et al. (J. Physiol. 1997) and Frick et al. (Cereb. Cortex 2008) provided highly valuable insights into synaptic transmission between pyramidal neurons in S1. We referred to their work in detail in our previous work on developmental changes in the presynaptic organization of transmitter release in L5APN synapses in S1 (Bornschein et al., Cell Rep. 2019). Both studies were performed in young rats and EPSPs were measured, whereas we worked in mice and recorded EPSCs. This impedes a direct comparison of amplitudes.

      Markram et al. (J. Physiol. 1997) recorded in 2-week-old rats from thick tufted PNs, corresponding to L5BPNs. Given the longer lifespan and slower development of rats compared to mice, this likely reflects a maturation state that corresponds better to our previous measurements in 8 to 10-day-old mice. Markram et al. found small failure rate (median 7%), which is similar to what we found in our previous study for young L5APN synapses (low failure rates and high p<sub>N</sub>; Bornschein et al., Cell Rep. 2019).

      The study by Frick et al. (Cereb. Cortex 2008) is closer to our present study and to the mature age window in our previous study, although they also recorded from rats but from L5APN-L5APN pairs in almost 3-week-old animals in S1. Again, EPSPs rather than EPSCs were recorded, impeding a direct comparison of amplitudes.

      Both studies concluded, based on the synaptic failure rate and CV analysis of EPSP amplitude, that the synapses they investigated operate with high release probability. This is fully in line with our findings. Of note, we found no significant difference between L2/3-L5PN and L5PN-L5PN synapses within a given area, indicating that varibality on the synaptic level between PNs of a given area is not pronounced.

      In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1. We have discussed the results of these studies in relation to our own findings.

      “Two other previous studies on L5APN (Frick et al., 2008) and L5BPN (Markram et al., 1997) connections concluded that these synapses operate with high release probability, which nicely agrees with our previous (Bornschein et al., 2019b) and current results. It is remarkable that we did not even detect any differences between the L2/3-L5PN and L5PN-L5PN synapses within a given cortical area. Overall these results from different studies (Markram et al., 1997; Reyes and Sakmann, 1999; Frick et al., 2008; Bornschein et al., 2019b; Bornschein et al., 2019a) may indicate that variability on the synaptic level between PNs of a given area is not pronounced. In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1.

      (5) In order to perform multiple probability fluctuation analysis (MPFA), a parabolic fit with a mere three points is inadequate, particularly because 2 mM and 5 mM Ca<sup>2+</sup> are close to the peak of the variance-to-mean parabola, and only 1 mM Ca<sup>2+</sup> is on its initial linear part. A more meaningful result would have been obtained with an additional Ca<sup>2+</sup> concentration between 1.0 and 2.0 mM, as these are closer to the physiological range. In this context, the authors should have quoted the more recent and more detailed paper by the Silver group (Saviane and Silver, 2006; Lanore and Silver, 2016) and not just the Clements and Silver review paper.

      We used only three Ca<sup>2+</sup> concentrations for MPFA as these resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition, thereby clearly determining a parabola. Also Saviane and Silver (Nature 2006) performed MPFA with three extracellular Ca<sup>2+</sup> concentrations (1, 2, and 8 mM). We have now cited this paper, as well as the more recent work by Lanore and Silver (Neuromethods 2016), in relation to the MPFA method. To verify the reliability of MPFA, the determined parameters were compared with values estimated based on the EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001). The estimated values did in fact match those of the MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript.

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional Ca<sup>2+</sup> concentration of 1.5 mM did not improve the parabolic fit, as it yielded p<sub>N</sub> values very close to those determined with 2 mM Ca<sup>2+</sup>. Therefore, the additional experiments are shown in Author response image 1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Author response image 1.

      MPFA with four different extracellular Ca<sup>2+</sup> concentrations. (A) MPFA of EPSC amplitudes recorded at the indicated [Ca<sup>2+</sup>]<sub>e</sub> from L2/3-L5PNs in PFC. Top: Individual EPSCs (grey, average in black) recorded from L5PNs after extracellular stimulation in L2/3. Middle: Plot of EPSC amplitudes over time. Bottom: Corresponding mean-variance plot fitted with a parabola estimating the quantal parameters of release. p<sub>N</sub> is for 2 mM [Ca<sup>2+</sup>]<sub>e</sub>. Recordings were made in the presence of 10 µM Bicuculline, 0.25 mM Kynurenic acid and 50 µM 2-Amino-5-phosphonovaleriansäure. (B) As in (A), but for a L2/3-L5PN connection in S1. (C) Summary of determined p<sub>N</sub> values in PFC and S1 (dots represent individual experiments; P=0.485, Mann-Whitney U rank-sum test).

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown in Author response image 2.

      Author response image 2.

      Bootstrap analysis (A-C) Distribution of bootstrap 25% trimmed means of the quantal parameters p<sub>N</sub> (A), N (B) and q (C) in PFC (orange), S1 (blue) and S1 with gDGG (light blue). 10,000 bootstrap replicates were generated with replacement from the original data sample obtained by MPFA in L5PN-L5PN connections (cf. Figure 3D). (D) Same as in (A) but for MPFA in L2/3-L5PN connections.

      (6) Methods: The authors should clarify whether their paired recordings from L5 pyramidal cells involved whole-cell recordings from both pre- and postsynaptic neurons. From Figure 1B, it appears as if the presynaptic neurons were not recorded in whole cell mode but rather stimulated in cell-attached mode. This is also reflected in the artefact visible in the current trace recorded in the postsynaptic neuron. The authors should explicitly state their methodological approach and mention how reliable the timing of the presynaptic action potential was under these circumstances. The same holds true for the extracellular stimulation protocol. A significantly more detailed description of the experimental protocol is necessary here.

      In the paired recordings, presynaptic cells were stimulated in the cell-attached mode. For presynaptic EGTA application the whole-cell configuration was established after re-patching to allow buffer perfusion of the presynaptic L5PN. This is described in the methods section of the original version of the manuscript. We extended this description as follows:

      “In paired recordings, presynaptic L5PNs were stimulated in on-cell configuration (200-500 mV, 1-2 ms). In the chelator wash-in experiments, presynaptic neurons were repatched with a pipette solution supplemented with 10 mM EGTA (K-gluconate concentration was reduced to 135 mM to adjust osmolarity) and whole-cell configuration was established to allow EGTA perfusion of the presynaptic neuron.”

      The amplitudes were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (cf. Bornschein et al., J. Physiol. 2013). Synaptic delays were determined from the onset of stimulation to the fitted EPSC onset. We added this more detailed explanation to the methods section. in the timing of the presynaptic action potential was similar in recordings from PFC and S1. This applies to the paired-recordings with on-cell stimulation of presynaptic neurons as well as to the extracellular stimulation experiments.

      “Synaptic responses were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (Bornschein et al., 2013). Synaptic delays were determined from the onset of stimulation to the fitted onset of the EPSC. PPRs were calculated by dividing the second amplitude of two consecutive EPSCs by the first.”

      (7) Methods: The authors use Student's t-test for data comparison. The authors should verify that the data distribution was indeed normal, e.g. by using a Shapiro-Wilk test. If this is not the case, non-parametric tests should be used.

      We typically used non-parametric tests as stated in the figure legends of the corresponding figures. We have now explained the abbreviations for the Mann-Whitney U test (MWU) and the Wilcoxon signed-rank test (WSR) in the figure legends. A paired t-test was used only in Figure 5F after testing for normal distribution with the Shapiro-Wilk test. This is described in the methods section of the original version of the manuscript. Additionally, results of the Shapiro-Wilk test were now included in the figure legends.

      “Normality was tested using the Shapiro-Wilk test. Normally distributed data were compared with the t-test (two groups) or a one-way ANOVA (more than two groups). Non-normally distributed or small samples of data were compared with the Mann-Whitney U rank-sum test (MWU; two groups) or a Kruskal-Wallis ANOVA on ranks (more than two groups). (…). To compare pre- and post-treatment data the paired t-test or the Wilcoxon signed-rank test (WSR) was used, depending on the distribution of the data.”

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Max Schwarze and colleagues examined the coupling distance between presynaptic Ca<sup>2+</sup> channels and the vesicular release sensor at neocortical synapses in mice. They propose that Ca<sup>2+</sup> channel-release sensor coupling differs across cortical areas, with relatively loose (microdomain) coupling in prefrontal cortex (PFC) and tighter (nanodomain) coupling in primary somatosensory cortex (S1) for comparable pyramidal-neuron synapse types. To test this, they combine paired recordings and minimal stimulation with chelator manipulations (EGTA/BAPTA), mean-variance/MPFA-style analyses, presynaptic Ca<sup>2+</sup> imaging, and computational modeling. They conclude that presynaptic coupling organization is area-specific in the mature cortex and contributes to regional differences in synaptic timing, reliability, and short-term plasticity.

      Strengths:

      This study tackles an important question and is strengthened by a cohesive body of evidence assembled from multiple complementary approaches. A major asset is the inclusion of high-value datasets, particularly the paired recordings between L5 pyramidal neurons and the systematic assessment of EGTA sensitivity, which provide a solid functional foundation for the authors' central claims. The work is further distinguished by its genuinely multimodal design: combining electrophysiology with presynaptic calcium imaging (and integrating these observations with quantitative analyses and modeling) offers a more mechanistic view of neurotransmitter release than any single method could provide. Overall, the direct, within-framework comparison of presynaptic release-control mechanisms across cortical areas for comparable synapse types is compelling and gives the conclusions a level of robustness and interpretability that is often difficult to achieve in studies of cortical synaptic diversity.

      Weaknesses:

      Several aspects would benefit from clearer explanation, stronger integration with the existing literature, and a more explicit discussion of limitations and potential confounds. Without these additions, some conclusions remain speculative. Throughout the manuscript, the authors also often imply that different measurements reflect the same underlying synapse population. This is unlikely to be strictly true across all experiments and makes it difficult to integrate results from the various approaches into a single, unified set of functional synaptic properties. In addition, some statements-particularly those linking coupling mode to "higher-order neocortical functions"-appear broader than what is directly supported by the experiments and should be tempered or more precisely scoped.

      Below, I list several topics that could help better frame the main findings of the present study and clarify how it relates to previously published work.

      We would like to thank the reviewer for the comprehensive and detailed assessment of our manuscript and his/her overall positive evaluation. We have addressed all of the reviewer's points. We expanded the model description and discussion, and slightly toned down our conclusion.

      (1) The authors use EGTA sensitivity of EPSCs (together with additional metrics) to argue that S1 and PFC synapses differ in Ca<sup>2+</sup> channel-release sensor coupling. While this is a plausible interpretation, EGTA effects are not uniquely determined by coupling distance and can also reflect differences in Ca<sup>2+</sup> entry kinetics, action potential waveform, endogenous buffering/extrusion, or release-sensor/vesicle state. The authors use a constrained modeling approach, but the rationale for the different constraint sets is not fully clear from the current description. It would be helpful to expand and clarify the Methods section to explain how these constraints were defined, justified, and applied (and how alternative constraint choices would affect the results). In this context, the Abstract's broader claim that the study "reveals microdomain coupling as a presynaptic structure-function correlate of higher-order neocortical functions" appears overstated. Given the well-known diversity of cortical synapses even within a single region (e.g., synapses onto different interneuron subclasses or different PN cell types, extracortical sources like thalamus), the authors should clarify the intended scope: is the conclusion meant to apply broadly across synapse classes in S1 and PFC, or only to the specific connection type(s) examined here?

      We would like to thank the reviewer from pointing out that our description fell a bit short, in particular with respect to the interpretation of the EGTA effects. We addressed the points as follows in the revised manuscript: We discussed the interpretation of EGTA effects in more detail. We toned down the concluding statement in the last sentence of the Abstract.

      “Differences in the sensitivity of release to low to moderate concentrations of EGTA (≤ 30 mM) are a standard indicator of differences in the coupling distance (e.g. Adler et al., 1991; Bucurenciu et al., 2008; reviewed in Eggermann et al., 2012; Vyleta and Jonas, 2014; Kusch et al., 2018; Bornschein et al., 2019b). p<sub>N</sub> is determined by the size of the Ca<sup>2+</sup> signal at the release sensor and the binding kinetics and affinity of the sensor. The former in turn is determined by the details of the Ca<sup>2+</sup> influx and the diffusional coupling distance between the VGCCs and the sensor. The similarity of Ca<sup>2+</sup> signals between synapses in PFC and S1 (Figure 4) indicates that Ca<sup>2+</sup> influx is similar between boutons, although more subtle differences in the influx kinetics may have remained undetected in these volume-averaged signals. Regarding sensor affinity, results in a previous study indicate that differences in EGTA sensitivity show differences in coupling rather than sensor affinity even if k<sub>on</sub> of the sensor and its affinity should differ as much as ten-fold, which appears to be an unlikely scenario given that even the two major isoforms of Synaptotagmin that trigger synchronous release differ by less than a factor of three to four in their affinity (Bollmann et al., 2000; Schneggenburger and Neher, 2000; Bornschein et al., 2025). Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      “They suggest that microdomain coupling in pyramidal neuron synapses could be a presynaptic structure-function correlate of higher order neocortical functions.”

      (2) The chelator logic is sound in principle, but the Discussion should more explicitly acknowledge standard caveats and alternative explanations. The authors partly address this by including presynaptic Ca<sup>2+</sup> imaging and modeling, yet it would help to explain more clearly how the combination of (i) chelator sensitivity, (ii) presynaptic Ca<sup>2+</sup> signals, and (iii) model constraints rules out-or substantially reduces the likelihood of-changes in AP waveform, Ca<sup>2+</sup> influx kinetics, buffering/extrusion, or sensor/vesicle state as the primary drivers. In addition, recent hypotheses emphasizing vesicle priming and/or release-site occupancy as contributors to apparent EGTA sensitivity should be discussed as a complementary or alternative interpretation.

      Please see above the first part of the discussion to point one.

      (3) A substantial portion of the S1 comparison appears to rely on previously published datasets. This should be made unambiguous in the Results and Methods, and it would be helpful to summarize this clearly (e.g., in a table indicating which figures/analyses use new data versus reanalysis of published data). If this information is already present, it should be highlighted more prominently.

      Please excuse us for not having made it clearer which data had already been published. Only the paired recordings from L5PN-L5PN connections in S1 were obtained in previous studies and partially reanalyzed. Paired recordings on the same synapses in PFC as well as all recordings from L2/3-L5PN connections in PFC and S1 were obtained in the present study. At your suggestion, we have added Table 1 highlighting which data and associated figures are from this study and which were acquired in previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      (4) The modeling is informative, but the choice of a specific VGCC-release-site geometry and channel arrangement is not sufficiently justified. The manuscript adopts a particular spatial configuration, yet the rationale for selecting this geometry, rather than other plausible architectures discussed in the literature, is not clearly explained, nor is it meaningfully revisited in the Discussion. The authors should justify why the same organization is assumed across two distinct cortical areas and, ideally, include (or at a minimum discuss) a sensitivity analysis showing how key inferences (e.g., coupling distance and channel number) depend on the assumed geometry.

      We extended the discussion of why a ring-like structure of VGCCs was assumed in the model.

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015), and exclusion zones (Keller et al., 2015). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      (5) The calcium imaging data are valuable, but given the diversity of synapses within each cortical layer, it is not clear that imaged boutons can be confidently assigned to the specific connection types being interrogated electrophysiologically. A substantial fraction of boutons likely corresponds to different postsynaptic targets (including interneurons and distinct pyramidal-cell classes), and this heterogeneity could complicate interpretation. This limitation should be discussed explicitly

      Excitatory pyramidal cells make up 80-85% of cortical neurons, with the highest density in layer 5 (Keller et al., Front. Neuroanat. 2018). In the somatosensory cortex, inhibitory synapses account for only about 10% (Santuy et al., Brain Struct. Funct. 2018). We imaged a large number of presynaptic boutons within layer 5 (about 10 boutons per cell, in total 85 boutons in PFC and 100 boutons in S1, numbers of boutons were now included in Figure 4). In this respect, the impact of inhibitory synapses is minor. Since connectivity between neighboring PNs in layer 5A is high (Feldmeyer, Front. Neuroanat. 2012), we assume that a large proportion of the imaged boutons target neighbouring L5PNs. We added a sentence on potential postsynaptic targets in the results section.

      “The imaged presynaptic boutons most likely connect to neighboring pyramidal cells, as connectivity between L5PNs in layer 5A is high (Feldmeyer, 2012). Nevertheless, a small proportion of other postsynaptic targets, such as interneurons, cannot be ruled out.”

      (6) In unitary connections, the authors assess EGTA effects alongside other functional parameters (strength, delay, short-term plasticity), which is a major strength. However, for L2/3 to L5 connections, it appears that EGTA sensitivity was tested primarily using extracellular stimulation. Given anatomical and circuit differences between PFC and S1, extracellular stimulation may recruit different synapse populations across regions, potentially confounding regional comparisons of EGTA sensitivity. This limitation should be acknowledged explicitly. While I am not requesting technically demanding L2/3↔L5 paired recordings in S1, the possibility that different synapse identities are being sampled should be treated as a meaningful source of uncertainty. The Discussion would also benefit from placing the magnitude of EGTA effects in the context of prior "loose coupling" literature, where comparatively large EGTA effects have been reported in some systems. In addition, the reported difference between adult PFC EGTA effects and S1 inhibition appears small (on the order of <10%) and should be interpreted cautiously, especially given that PFC and S1 mature on different timelines and P21-P26 is unlikely to reflect a mature PFC circuit state. The adult cohort (P90-P100) is therefore important, but the age mismatch complicates PFC-S1 comparisons; ideally, S1 should be assessed at matched ages, or this limitation should be discussed explicitly. Finally, for statistical robustness, in panel D of Figure 2, were the comparisons corrected for multiple testing to control Type I error?

      EGTA sensitivity was examined in PFC and in S1, for two connections in each region - using paired recordings for L5PN-L5PN connections and using extracellular stimulation for L2/3-L5PN connections. The L5PN-L5PN data from S1 were collected in an earlier study (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019), have now been reanalyzed for the test period between 20 and 30 min, and included in Figure 2B for the sake of consistency (see Table 1). To emphasize this point, despite stimulating different input synapses with different stimulation methods, we obtained similar results in the respective brain regions. This suggests that the EGTA sensitivity observed in the investigated PFC connections is not a solely synapse-specific property.

      It is difficult to compare the absolute EGTA sensitivities from different synapses from different publications, since EGTA effects do not depend exclusively on the coupling distance, as the reviewer also noted in point 1. They are, among other factors, influenced by the Ca<sup>2+</sup> sensitivity of the release machinery, which differs between our Syt1-expressing cortical synapses and Syt2-expressing synapses in other brain regions (Schneggenburger et al., Nature, 2000; Bollmann et al., Science, 2000; Bornschein et al., Science 2025), such as the calyx of Held or the cerebellar basket to Purkinje cell synapse. Furthermore, direct patching and loading of the presynaptic bouton with EGTA - as feasible at the calyx of Held and other large synapses - results in higher effective EGTA concentrations compared to somatic loading of presynaptic terminals, despite identical pipette concentrations. The buffer-AM method introduces additional uncertainty regarding the effective intra-bouton EGTA concentration, since the loading efficacy has to be estimated. Thus, although differences in EGTA sensitivity primarily show differences in coupling distances, the comparison of absolute values between different publications is difficult. Consistently, data-constrained models are used to estimate the coupling topography and to compare these topographies rather than comparing the absolute EGTA effects (e.g. Buccurenciu et al., Neuron, 2008; Vyleta and Jonas, Science, 2014; Bornschein et al., Cell Rep., 2019; Chen et al., Neuron, 2024; Bornschein et al., Science, 2025).

      We include a note on this in the discussion.

      “Thus, although the absolute EGTA sensitivity is influenced by different factors, which necessitates data-constrained models for quantitative comparisons, the general sensitivity of release to EGTA indicates loose coupling.”

      In mouse neocortex postnatal maturation in S1 and PFC follows the same time course. Kroon et al. (Sci. Rep. 2019) reported that maturation of dendritic morphology and intrinsic properties of pyramidal neurons occurs within the first two weeks after birth, now cited in the discussion. Therefore, it is unlikely that the EGTA effect in PFC is due to a delayed maturation. The difference in EGTA sensitivity between PFC at P90-100 and S1 at P21-26 is indeed small but significant (P=0.009, Mann-Whitney-U rank sum test).

      “Since postnatal development follows the same time-course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, (…) (Kroon et al., 2019).”

      Thank you for the advice concerning statistical robustness. We replaced the Mann-Whitney-U rank sum test in Figure 2D by a one-way ANOVA and performed a Holm-Sidak post-hoc test correcting for multiple comparisons. Similarly, ANOVA was used to compare more than two groups in Figures 1K and 2F. We changed the corresponding P values and tests in the figure legends and added the performed post-hoc tests in the methods section.

      “For multiple comparisons post-hoc testing was performed with the Holm-Sidak (one-way ANOVA) or Dunn´s method (ANOVA on ranks).”

      (7) Alterations in initial release probability are often associated with changes in short-term plasticity. In the present manuscript, the authors report similar initial release probability at PFC and S1 synapses, yet observe differences in short-term plasticity profiles. The mechanistic basis for this apparent dissociation is not addressed and should be discussed explicitly, including potential explanations.

      Various other factors besides p<sub>N</sub> can influence short-term plasticity, that are the coupling distance, the number of occupied release sites (N<sub>occ</sub>), the replenishment of N<sub>occ</sub> or the recruitment of newly formed N<sub>occ</sub> as well as the expression of endogenous Ca<sup>2+</sup> buffers (Blatow et al., Neuron 2003; Felmy et al., Neuron 2003; Matveev et al., Biophys. J. 2004; Neher, Cell Calcium 1998; Regehr, CSH Perp. Biol. 2012) or fascilitation sensors (Turecek & Regehr, J. Neurosci. 2018; Shin et al., eLife 2025). Traditionally, p<sub>N</sub> had been assumed to have a major impact on short-term plasticity (STP) which is indeed the case at low replenishment rates (e.g. Feldmeyer and Radnikow, J. Physiol. 2009; Zucker and Regehr, Ann. Rev. Physiol. 2002). But at several synapses very fast replenishment rates have been described driving a progressive overfilling of the initial RRP and increasing N<sub>occ</sub> above baseline levels (Brachtendorf et al., Front. Cell. Neurosci. 2015; Doussau et al., eLife 2017; Miki et al., Neuron 2016; Valera et al., J. Neurosci. 2012) making replenishment the stronger determinant of STP.

      Additionally, the size and organization of sub-pools from which vesicle recruitment and release occurs affects the speed and reliability of vesicular release. In our previous study on L5PN-L5PN connections in S1 we found that developmental tightening of CDs was associated with an increase in PPR without altering p<sub>N</sub> (Bornschein et al., Cell Rep. 2019). We could show that the maturation of a replenishment pool during postnatal development increases vesicle recruitment and reliability thereby affecting STP (Bornschein et al., Front. Syn. Neurosci. 2019).

      We discussed this in the revised manuscript.

      “Classically, p<sub>N</sub> was considered as the major determinant of short-term plasticity (e.g. reviewed in Zucker and Regehr, 2002; Feldmeyer and Radnikow, 2009). More recently other factors, including the number of occupied release sites, their replenishment or an increase in their occupancy, or the expression of endogenous Ca<sup>2+</sup> buffers have been considered as more important determinants of short-term plasticity (Rozov et al., 2001; Blatow et al., 2003; Felmy et al., 2003; Matveev et al., 2004; Bornschein et al., 2013; Miki et al., 2016; Doussau et al., 2017; Jackman and Regehr, 2017; Neher and Brose, 2018). (…)”

      Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019).

      (8) There are multiple instances where the text appears to cite non-existent or misnumbered figure panels (e.g., references to "Figure 4G-I / 4J" when the relevant material appears elsewhere). These should be corrected throughout, as they currently reduce readability and confidence.

      We apologize for the misnumbering which originated from a previous version of this manuscript. Figure 4G-J is actually Figure 5A-D. We corrected the references to Figure 5 in the methods section.

      (9) The Methods describe P21-P26 animals, whereas the Results include older cohorts (e.g., P90-P100) and additional regions (e.g., mPFC). The Methods should be updated so that all cohorts and regions analyzed in the Results are fully described.

      Thank you for thoroughly reading the methods. We added the missing cohort (P90-100) and brain region (mPFC) to the method section.

      “C57BL/6J mice at P21-26 and P90-100 of either sex were decapitated under deep Isoflurane (Curamed) inhalation anaesthesia. (…) Coronal neocortical slices (150-250 μm thick) were cut from the lateral PFC, medial PFC (mPFC) or S1 region (Figure 1A) with a vibratome (HM 650 V, Microm).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      These are mostly points for discussion; there is no need for additional experiments.

      (1) Discuss potential effects of (re-)patching on presynaptic physiology and the "control" time course in Figure 2A.

      The repatching strategy is a strength, but it also introduces opportunities for physiological drift (dialysis effects, changes in access resistance, altered excitability, and the switch to presynaptic stimulation in whole-cell after repatching). Please discuss (and, if already quantified, briefly report) why EPSC amplitudes appear to increase in the control time course after repatching (as seen in Figure 2A). Even a short explanation (e.g., run-up after whole-cell access, recovery from on-cell stimulation, washout of endogenous buffering, improved spike waveform reliability, etc.) plus reassurance that baseline stationarity criteria were met would strengthen confidence in the repatching-based inference.

      Baseline recordings were usually performed with the presynaptic neuron in cell-attached mode. In this configuration intracellular ion concentration, second messenger systems as well as cell-specific resting membrane potential stay essentially unaffected. When switching to whole-cell mode after re-patching of the presynaptic cell, the pipette solution determines the intracellular environment, which may also affect second messenger systems and mobile endogenous buffers will be washed out. Although layer 5 pyramidal neurons do not express relevant concentrations of these mobile buffers (Helmchen et al., Biophys. J., 1996; Tran and Stricker, Biophys. J., 2018; Bornschein et al., Cell Rep., 2019), this might have contributed to the moderate and temporary run-up after whole-cell access to presynaptic cells.

      In postsynaptic neurons, we also routinely controlled the stability of R<sub>s</sub> and I<sub>leak</sub> during our prolonged measurements. In the controls, we found a median initial increase in relative EPSC amplitudes to 1.17 (1.02-1.38) between 0 and 10 min after the presynaptic whole-cell access was established, which correlated with a temporary decline in R<sub>s</sub> in 3 out of 6 recordings. EPSC amplitudes returned to their baseline values after 20 min at the latest (1.02, 0.75-1.20). We added potential reasons for the temporary amplitude increase in the controls in the results section.

      “In control recordings a temporary initial increase in EPSC amplitudes after whole-cell access to the presynaptic neuron was evident. Although L5PNs do not express large concentrations of mobile buffers (Helmchen et al., 1996; Tran and Stricker, 2018; Bornschein et al., 2019b), their wash-out might have contributed to the temporary run-up. On the other hand, run-up correlated with a temporary decline in R<sub>s</sub> in some recordings (3 out of 6). Run-up effects normalized after 20 min at the latest (1.02, 0.75-1.20).”

      (2) Clarify interpretation/robustness of MPFA-derived "N" given the large range, and discuss uncertainty from using only three [Ca<sup>2+</sup>]e conditions for L2/3-L5PN MPFA.

      The manuscript reports that "N" is markedly smaller in PFC than S1 (median ~2.1 vs 8), but the S1 range is very broad (3-19).

      (a) Briefly discuss whether/why such a wide N range is expected and how it should be interpreted (binomial "N" vs anatomical release sites; sensitivity to CV assumptions; potential dependence on connection geometry, bouton number, dendritic filtering, etc.).

      (b) Add a short statement on uncertainty/identifiability when fitting MPFA with only three conditions for L2/3-L5PN (e.g., whether confidence intervals/bootstraps were examined; how stable q and N are to small changes in the variance estimates). Even a qualitative note would help readers judge how much weight to put on the absolute N estimates versus the overall cross-area trend.

      (a) In our previous study on this connection we determined a wide range for N (8, 3-19) even though we used four extracellular Ca<sup>2+</sup> concentrations in MPFA (Bornschein et al., Cell Rep. 2019). The binomial parameter N can be considered to represent the number of release sites, including empty release sites (see Brachtendorf et al., Front. Cell. Neurosci. 2025). The range of release sites is likely to reflect the variability in the number of anatomical synaptic contacts ranging from 1 to 6 for these synapses (Frick et al., Cereb. Cortex 2008). Since 1 to 3 active zones/release sites per synaptic contact appear to be typical for small cortical synapses (e.g. Xu-Friedman et al., J. Neurosci. 2001), this results in a wide range of 1-18 release sites per connection.

      In our previous study (Bornschein et al., Cell Rep. 2019), we also investigated the effects of different values of CV1 and CV2 by repeating the MPFA fitting procedures for different combinations of CV1 and CV2 ranging from 0.1 to 1 each. We have quantified a deviation of ≤10% in the estimates of vesicular release probability from the typically used CV values of 0.3 across a wide range of CV value combinations (see also Schmidt et al., Curr. Biol. 2013). We have added a note regarding CV sensitivity in the methods section.

      “For CV assumptions that deviate from the standard value of 0.3, deviations in the calculated p<sub>N</sub> values of less than 10% are to be expected (Schmidt et al., 2013; Bornschein et al., 2019b).”

      (b) The three Ca<sup>2+</sup> concentrations we used for MPFA resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition. With this, a parabola is uniquely determined by three parameters. To further ensure the reliability of the parameters determined by MPFA, we compared them to values estimated from EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001), which yielded values similar to those from MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript).

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional experiments are shown for review purposes in Figure R1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown for review purposes in Author response image 2 (cf. point 5 of Reviewer#2).

      (3) Broaden the discussion, e.g. by linking to nanodomain/microdomain coupling as a general strategy for stimulus encoding, including sensory periphery examples.

      The work will resonate beyond the cortex if the authors explicitly connect their findings to broader principles: how the spatial coupling regime shapes the transfer function between Ca<sup>2+</sup> entry and vesicle fusion, thereby tuning reliability, timing, and dynamic range. Requested addition: Please consider adding a short subsection discussing analogous implementations in the sensory periphery, especially ribbon synapses of cochlear inner hair cells and rod photoreceptors, where nanodomain coupling has been discussed as a key determinant of encoding and release dynamics. Also, citing relevant work such as that by Scimemi and Diamond 2012 would further strengthen the paper.

      We agree that the work by Scimemi and Diamond (J.Neurosci. 2012) is important and we cited and discussed their work in several of our previous publications. We now also included the paper in the revised version of the present manuscript. As requested, we also included a discussion on findings from ribbon type synapses and also from the neuromuscular junction.

      “Nanodomain coupling was also found in the peripheral nervous system, in particular at retinal (Singer and Diamond, 2003; Jarsky et al., 2010) and auditory (Moser and Beutner, 2000; Brandt et al., 2005) ribbon-type synapses and at the neuromuscular junction (Harlow et al., 2001; Shahrezaei et al., 2006). These synapses have highly specialized properties and appear to be optimized for very reliable transmission and, in the case of ribbon synapses, also for high-frequency coding of sensory information (reviewed in Matthews and Fuchs, 2010; Eggermann et al., 2012). Thus, it appears that synapses in the sensory pathways, in particular those engaged in reliable high-frequency coding of sensory information, both in the periphery and in the lower processing stages of the CNS, up to primary sensory cortices, operate with nanodomain coupling. In the executing motor pathway, the neuromuscular junction uses nanodomain coupling and, as recent results from our group suggest, also PNs in the primary motor cortex (Yarim et al., in preparation). It is tempting to speculate that complete loops from or to the primary cortices to their peripheral target organs operate with nanodomains. Microdomain coupling, on the other hand, appears to come into play only if integration of information from multiple sources and plasticity are the main focus, as at certain synapses in PFC (this study) or hippocampus (Vyleta and Jonas, 2014).”

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015; Rebola et al., 2019), and exclusion zones (Keller et al., 2015; Rebola et al., 2019). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      “(…) Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      (4) Address limitations of basal/resting Ca<sup>2+</sup> estimates and make explicit that measured Ca<sup>2+</sup> signals are volume-averaged (not microdomain) readouts.

      (a) The reported basal [Ca<sup>2+</sup>]i values are in the ~tens of nM range. Given the stated in vitro KD for Fluo-5F in the authors' pipette solution (439 nM), the resting estimates are far below KD; this does not invalidate the approach, but it does warrant a brief discussion of sensitivity/uncertainty (influence of Rmin estimation, background subtraction, and how errors propagate into basal [Ca<sup>2+</sup>]i). Repeating experiments is not necessary-just clearer framing of limitations.

      (b) Please also emphasize more prominently (ideally in Results and/or Discussion) that the bouton signals are volume averaged and therefore do not directly report calcium microdomains at active zones or nanodomains at release sensors. The Methods already state this point; echoing it in the main text would prevent over-interpretation by readers.

      (a) We agree that Fluo5F is less suitable for determining absolute basal calcium levels. In a previous study (Bornschein et al., Science 2025) we determined the basal Ca<sup>2+</sup> concentration with OGB1 (K<sub>D</sub>=166 nM; basal [Ca<sup>2+</sup>]<sub>i</sub>=44 nM, 24-58 nM, n=43 boutons from 10 cells) and observed no significant difference to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with Fluo5F despite the K<sub>D</sub> of 439 nM (31 nM, 16-54 nM, 14 boutons from 3 cells; P=0.204, MWU; data not published). We added this limitation to the results section and swapped Figure panels 4E and F for confluence. The calibration curve of Fluo5F as well as the comparison to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with OGB1have been included in Figure S4.

      “The quantification of absolute basal [Ca<sup>2+</sup>]<sub>i</sub> was limited by the K<sub>D</sub> of Fluo5F (439 nM), which slightly underestimated basal [Ca<sup>2+</sup>]<sub>i</sub> values in comparison to quantification with OGB1 (K<sub>D</sub>=166 nM, Figure S4). Nevertheless, relative comparison of basal [Ca<sup>2+</sup>]<sub>i</sub> yielded no significant differences between PFC (30 nM, 21-34 nM) and S1 (22 nM, 13-38 nM; Figure 4F).”

      (b) In the results section, we have now emphasized that volume-averaged Ca<sup>2+</sup> signals were measured.

      “We performed dual-dye two-photon Ca<sup>2+</sup> imaging (Sabatini et al., 2002) to quantify volume-averaged Ca<sup>2+</sup> signals at presumed presynaptic boutons located on axon collaterals of L5PNs in PFC and in S1.”

      Reviewer #2 (Recommendations for the authors):

      (1) For a meaningful comparison, recordings from the PFC and the S1 cortex of the same animals should be undertaken. Additionally, I suggest performing additional experiments regarding the different cell types of L5 pyramidal cells in layer 5a.

      We performed new experiments to determine EGTA sensitivity in PFC and S1 from the same animal. The results from these experiments agree with the previous results. They are included in the results section , in Figure 2C-F and in Figure S3A-C.

      Additionally, we extended the discussion on the examined cell types. For S1 cortex we refer in more detail to our previous work, where we described in depth where and under consideration of which criteria our recordings were established and that based on these criteria we recorded from pyramidal neurons in layer 5A in S1 (Bornschein et al., Cell Rep. 2019; Bornschein et al. Front. Synapt. Neurosci. 2019; Bornschein et al., Science 2025). Within layer 5A, we did not attempt to further distinguish between types of pyramidal neurons. We include this in the methods section.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).”

      For the recordings in PFC and heterogeneity in pyramidal neuron types we refer to our detailed response to the point 3 of Reviewer 2. There we also discuss that the heterogeneity in morphology and spiking patterns is probably not reflected on the synaptic level. We would also like to emphasize that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable to differentiate between subpopulations of pyramidal neurons. This would require successful recordings form several tens of different pyramidal neurons, which is not feasible in our type of experiment. We discuss this limitation of the discussion.

      (2) The authors need to comment in depth on their MPFA data, and if feasibl,e perform additional experiments.

      Concerning the robustness of quantification of synaptic parameters by MPFA, we refer to our comments on point 2b of the recommendations for the authors to Reviewer 1. Additionally, we performed new MPFA experiments with four extracellular Ca<sup>2+</sup> concentrations that agree with our results with three Ca<sup>2+</sup> concentrations.

      (3) The statistical analysis should be revised and a test for the normality of data distribution should be implemented.

      A test for normal distribution (Shapiro-Wilk test) has already been described in the methods section in the previous version of this manuscript.

      (3) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (4) Is the relative variance of the mean EPSC amplitude and latency between connections larger in the PFC connections than in S1 cortex? This could indicate a variability in cell types.

      The relative variance of EPSC amplitudes calculated as median absolute deviation (MAD) was 0.46 in PFC and 0.50 in S1 arguing against differences in the variability in cell types. The larger variability in delays expressed as SD<sub>Delay</sub> is the result of the larger coupling distance in PFC compared to S1 (Bullmann et al., J. Neurosci. 2024). Consequently, also the relative MAD is larger (0.89) in PFC compared to S1 (0.14) and is therefore not able to detect differences in the variability of recorded cell types.

      (5) Reyes and Sakmann (1999) have previously described differences for L2/3-L5b and L5b-L5b synaptic connections in S1 cortex at different developmental stages. This paper needs to be cited as it is highly relevant to this study.

      Reyes and Sakmann (J. Neurosci. 1999) reported layer-specific differences in short-term plasticity in young sensorimotor cortex which disappeared as maturation progressed and short-term plasticity increased. In a previous study (Bornschein et al., Front. Syn. Neurosci. 2019) we also described a developmentally driven increase in short-term plasticity caused by the maturation of vesicle pools. In the present study we used mature animals and would therefore not expect layer-specific differences neither in S1 nor in PFC since the time course of postnatal maturation was described to be comparable in both neocortical circuits (Kroon et al., Sci. Rep. 2019).

      We discussed this paper in the context of developmental changes in short-term plasticity.

      “Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019a). Such pool maturation may also underlie the elimination of layer-specific differences in short-term plasticity between L2/3-L5B and L5B-L5B synaptic connections that were evident in young rats but eliminated during the first weeks of postnatal development (Reyes and Sakmann, 1999). Since postnatal development follows the same time course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, significant layer-specific differences in PN synapses are unlikely in both areas in our experimental time window (Kroon et al., 2019). Consistently, we found similar PPRs at L2/3-L5PN synapses and L5PN-L5PN synapses in both areas, with facilitation in PFC and depression in S1, irrespective of the presynaptic PN synapse type.

      (6) Regarding the point of loose or tight Ca<sup>2+</sup> channel coupling: Could some of the differences result from differences in the presynaptic Ca<sup>2+</sup> channel complement? Please comment.

      This can be excluded. Ca<sub>v</sub>2.1 and Ca<sub>v</sub>2.2 are the main channels gating release at PN synapses. The gating kinetics of these channels are very similar and they only differ somewhat in their peak current amplitude (Bornschein et al., Cell Rep. 2019, Figure 4). Since the number of open channels is a fit parameter in our simulations there would only be an effect on the estimate of the number of channels gating release but not for the estimate of the coupling distance. This is all the more true since the EGTA effect depends on the diffusion distance rather than on the gating kinetics. These considerations will also hold for Ca<sub>v</sub>2.3 channels, which have slower closing kinetics, but anyway play only a very minor role for triggering release.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study is a valuable contribution to the evidence base. However, the evidence provided is incomplete as the study results only partially support the study conclusions. Addressing the methodological and reporting issues raised by the peer reviewers and properly aligning the claim made for providing a tool for early warning with the study analysis/results would improve the study quality and usefulness of its findings.

      We are deeply encouraged by the editors’ recognition of this study as a valuable contribution to the evidence base. We fully concur with the eLife assessment that the manuscript, in its original form, required substantive methodological recalibration and rhetorical refinement to ensure the claims were strictly supported by the data. We accept these profound critiques unreservedly and have undertaken a comprehensive, ground-up revision of this study.

      This revision was guided by three overarching principles: (1) systematic decoupling of epidemiological confounders, (2) rigorous elimination of COVID-19-era surveillance bias, and (3) precise alignment of our claims with the actual scope of our predictive framework. Most importantly, we have fundamentally recalibrated our core claim: we have entirely removed all assertions of providing a “ready-to-use early warning tool.” Instead, we accurately reposition our contribution as developing a “robust, climate-informed predictive surveillance framework” that establishes an evidence-based foundation for understanding non-stationary climate-influenza dynamics in a subtropical urban setting.

      To ensure our analytical results definitively support this recalibrated conclusion, we executed a comprehensive remodeling of our entire dataset, rebuilding both the DLNM and LSTM networks. The major structural upgrades include:

      (1) Shift to Positivity Rate to Eliminate Testing Bias: To directly address the reviewers’ incisive concern regarding testing volume bias (i.e., raw case counts were artificially inflated or suppressed by massive fluctuations in PCR testing during the COVID-19 pandemic), we have fundamentally replaced “influenza positive case counts” with the “influenza test positivity rates” as our primary outcome metric. This relative metric mathematically standardizes the denominator, elegantly isolating intrinsic viral transmissibility from the severe artificial fluctuations of healthcare-seeking behaviors and diagnostic intensity.

      (2) Integration of New Epidemiological Covariates: We significantly enhanced our control for critical confounders (as suggested by the reviewers) by incorporating new, highly relevant covariates into our models. These include:

      - Weekly detection volumes: To capture residual testing capacity fluctuations.

      - Mask-wearing stringency indices (OWID): To computationally decouple the artificial suppression of cases caused by strict non-pharmaceutical interventions (NPIs).

      - Proportion of the non-local/transient population: To explicitly account for external viral importation risks.

      - Day of the Week (DOW, weekday vs. weekend): To reflect societal mobility and administrative patterns.

      Simultaneously, we removed redundant variables (e.g., raw COVID-19 case numbers) to minimize noise.

      (3) Comprehensive Re-optimization, Retraining, and Sensitivity Analyses: Following this massive feature engineering, we re-executed the Bayesian hyperparameter tuning process and completely retrained the LSTM networks. We also conducted rigorous sensitivity analyses (detailed in the new Table 2 and Supplementary files) by systematically ablating variables like testing volume and mask-wearing indices, definitively proving the necessity of these covariates in our framework. All corresponding codes, main figures (Figures 1– 6), supplementary figures, and statistical tables have been entirely updated.

      (4) Streamlining the Manuscript: We acknowledge the reviewers’ observation that the original methodology was overly meandering. In this revision, we have ruthlessly streamlined the text. We condensed routine laboratory protocols, reorganized the Methods to logically present the “Study area” prior to the “Study design,” and thoroughly rewrote the Discussion to weave our limitations directly into the interpretation of the results. This ensures conciseness, logical flow, and a sharp focus on the central narrative, tailored for the broad and rigorous readership of eLife.

      We believe these deep methodological revisions have profoundly elevated the scientific rigor, interpretability, and integrity of our study. Below, we provide detailed, point-by-point responses outlining how each specific comment was addressed.

      Reviewer #1 (Public review):

      A major concern is that the model is trained in the midst of the COVID-19 pandemic and its associated restrictions and validated on 2023 data. The situation before, during, and after COVID is fluid, and one may not be representative of the other. The situation in 2023 may also not have been normal and reflective of 2024 onward, both in terms of the amount of testing (and positives) and measures taken to prevent the spread of these types of infections. A further worry is that the retrospective prospective split occurred in October 2020, right in the first year of COVID, so it will be impossible to compare both cohorts to assess whether grouping them is sensible.

      We deeply appreciate this astute epidemiological critique. You have precisely identified the most formidable methodological challenge in pandemic-era time series modeling: the profound non-stationarity of influenza dynamics spanning the 2018–2023 timeline, driven by intensive Non-Pharmaceutical Interventions (NPIs), pandemic-related behavioral shifts, volatile testing volumes, and subsequent immunity debt. We fully agree that 2023 represented an atypical post-restriction “rebound” year and is unlikely to be straightforwardly representative of a stabilized post-2024 epidemiological steady state.

      First, we wish to clarify a minor but crucial methodological detail regarding the October 2020 split. You understandably raised concerns about comparing “both cohorts.” We must emphasize that there is no change in the cohort or the data collection methodology. The surveillance system, sentinel hospitals, and diagnostic protocols (managed by Putian CDC) remained identical and uninterrupted from 2018 to 2023. Although the original manuscript labelled the January 2018 – October 2020 and October 2020 – December 2023 phases as “retrospective” and “prospective” cohorts respectively, this distinction reflected only the administrative timing of ethical approval. It does not indicate any changes in sentinel hospital locations, ILI case definitions, specimen collection procedures, RT-PCR diagnostic protocols, or laboratory quality control. The patient population and clinical criteria are completely homogeneous. We have revised the “Ethics statement” subsection to eliminate this semantic confusion.

      Upon receiving this insightful feedback, our team conducted extensive mathematical evaluation on how best to address this timeline heterogeneity. We initially considered formal stratified temporal analyses (e.g., splitting models into strict Pre-COVID 2018-2019, COVID-disruption 2020-2022, and Post-COVID 2023 periods) and implementing a rolling-window validation scheme. However, after careful evaluation, we concluded that strict time-slicing of this particular dataset would introduce its own substantial methodological problems more severe than the heterogeneity it sought to resolve:

      (1) Statistical power constraints on for stratified Non-linear Lag analysis: The DLNM architecture requires continuous, robust longitudinal data to stably estimate two-dimensional exposure–lag–response surfaces with natural cubic splines, which typically demand approximately 16–25 effective degrees of freedom across the joint exposure × lag space. Stratifying our six-year dataset into the three intervals list above would yield:

      - Pre-COVID stratum (2018–2019): ~730 daily observations and only ~200 influenza B positive events, insufficient to stably identify cross-basis surfaces, with confidence intervals expected to widen to non-informative ranges.

      - COVID-disruption stratum (2020–2022): influenza circulation was substantially suppressed (though, importantly, not eliminated, a point we return to below), reducing signal density and risking that estimated surfaces reflect suppression dynamics rather than climate–transmission relationships.

      - Post-restriction stratum (2023): a single year cannot independently support a DLNM with meaningful lag structure given the typical 0–14-day lag window we examine.

      (2) The “training-domain contamination” problem in rolling-window validation: We carefully examined whether expanding-window rolling validation (e.g., train 2018–2020, validate 2021; train 2018–2021, validate 2022; …) would resolve the non-stationarity concern. We concluded it would not, for a subtle but consequential reason: each such window straddles the abrupt NPI transitions, meaning that within any single training window, the model is exposed to a nonstationary mixture of regimes without any explicit signal indicating which regime each observation belongs to. Implementing a rolling window in this specific context forces the algorithm to repeatedly train on fragmented, incomplete phases of this epidemiological cycle. The model is therefore likely to learn a confounded representation in which the climate signal is partially absorbed into the implicit regime-shift signal, a pathology that is in some respects more difficult to diagnose than that of unified modelling. By employing a single, continuous 88% training block (2018–2022), we structurally guarantee that the LSTM’s memory cell is exposed to the complete sequence of regime shifts, from pre-pandemic natural baseline, through extreme suppression, and ending right at the brink of the NPI relaxation. Reserving the entirely unseen 2023 “rebound year” as a chronological hold-out thereby serves as the ultimate extreme stress-test of the network’s capacity to dynamically synthesize NPI relaxation signals and climate variables to forecast a historically unprecedented surge.

      (3) Architectural mismatch between time-slicing and the LSTM’s design principle: A core motivation for adopting an LSTM rather than period-stratified statistical models was precisely to leverage its long-term dependency memory mechanism, which is designed to allow a unified architecture to learn how predictive relationships are modulated by time-varying contextual conditions. Presegmenting the data into NPI-defined strata forecloses this principal architectural advantage and would, in effect, reduce the analysis to a series of disconnected period-specific models, a design for which the LSTM’s complexity provides no benefit over simpler approaches.

      (4) Tension with the surveillance-bias correction: As we discuss in our response to Public review para 2 below, a separate and equally important critique from you concerns surveillance bias from temporally varying testing intensity. Stratified analysis would compound this problem, because each stratum carries a distinct testing-intensity profile (notably the surveillance surge during 2020–2022), and stratum-specific models cannot leverage cross-period testing-volume normalization. A unified model with explicit testing-volume covariates is protective against bias than period-stratified alternatives.

      Our Methodological Solution: Covariate-Driven Adaptive Learning

      Instead of artificially fracturing the timeline, we substantially restructured the analytical framework to teach the model how to contextualize the pandemic-era disruption while preserving longitudinal continuity:

      - First, we shifted the predictive target to Influenza Positivity Rates (laboratory confirmed cases ÷ ILI specimens tested): The positivity rate is the WHO recommended sentinel surveillance metric and intrinsically corrects for the substantial fluctuations in testing volumes, healthcare-seeking behaviour, and surveillance intensity that characterized the pre-pandemic, pandemic, and post-restriction periods, rendering the outcome metric far more comparable across the timeline.

      - Second, we integrated explicit Epidemiological Context Covariates: We structurally upgraded the LSTM network by feeding it vital time-varying covariates alongside meteorological data. Specifically, we incorporated the Our World in Data (OWID) mask-wearing stringency indices, weekly detection volumes, day of the week (DOW, distinguish between weekdays and weekends to capture administrative reporting patterns) and the proportion of non-local residents among tested patients (to account for population mobility and viral importation risks).

      - Third, we performed a covariate-ablation sensitivity analysis: To prove the model actively utilizes these contextual signals, we performed targeted ablation studies (new Table 2). Removing mask-wearing stringency indices increased the 2023 influenza A forecast MAE to 0.012 and the influenza B forecast MAE to 0.003 (compared with MAEs of 0.009 and 0.002, respectively, obtained when the covariate was retained), indicating a decline in the network’s predictive accuracy. Removing the weekly testing volume had an even more catastrophic impact on Influenza B, surging the MAE by 100% and SMAPE by 199.6%. This provides empirical evidence that the unified network does not passively average across regimes, but dynamically leverages NPI and surveillance signals to adapt to non-stationarity. We have added a paragraph to the revised Discussion interpreting these findings as suggestive (though not definitive) evidence that the unified-with covariates architecture is contextualizing rather than averaging across regimes.

      By explicitly providing the LSTM with these explicit contextual parameters, the network autonomously learned the “regime shifts.” It learned that when the mask-wearing index is high, transmission is dampened despite favorable meteorological conditions.

      Re-interpreting the 2023 Validation

      Guided by your critique, we fully agree that 2023 was an anomalous “rebound” year and not a steady-state reflection of a post-2024 “normal.” We have completely abandoned the framing that our model represents a steady-state tool for the future. Instead, we now explicitly frame the 2023 validation as an “extreme epidemiological stress test.” The revised framing rests on three explicit acknowledgements:

      - The 2023 validation tests forecasting performance during a transitional, post-restriction rebound year, and is best understood as an extreme stress test of the framework’s adaptive capacity, not as evidence of long-term predictive validity under stabilized future conditions.

      - Performance metrics observed in 2023 should not be naively extrapolated to 2024 and beyond.

      - Continuous prospective recalibration as 2024–2025 data accumulate, ideally combined with adaptive learning approaches capable of detecting regime shifts in real time, will be essential before any operational deployment.

      The fact that our updated LSTM network accurately forecasted the explosive, atypical 2023 viral rebound, despite being trained heavily on the suppressed 2020-2022 data, demonstrates its robust capacity to synthesize climate variables and NPI relaxation signals.

      We have substantially rewritten the Discussion section to honestly acknowledge this limitation and contextualize the 2023 results, stating that continuous recalibration will be essential for future forecasting.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018– 2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reliance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      We believe this integrated, covariate-based approach maintains the mathematical integrity of the time-series analysis while fully addressing your valid concerns regarding epidemiological non-stationarity.

      The outcome of interest is the number of confirmed influenza cases. This is not only a function of weather, but also of the amount of testing. The amount of testing is also a function of historical patterns. This poses the real risk that the model confirms historical opinions through increased testing in those higher-risk periods. Of course, the models could also be run to see how meteorological factors affect testing and the percentage of positive tests. The results only deal with the number of positive (only the overall number of tests is noted briefly), which means there is no way to assess how reasonable and/or variable these other measures are. This is especially concerning as there was massive testing for respiratory viruses during COVID in many places, possibly including China.

      We are exceptionally grateful for this incisive methodological observation. You have accurately identified a fundamental validity threat, “surveillance intensity bias”, that inherently constrains much of the existing climate–infectious disease literature. We fully concur with your assessment that raw case counts are jointly determined by underlying viral transmission and dynamic testing intensity. Furthermore, we recognize your highly valid concern regarding the risk of “circular reasoning,” whereby meteorological factors might simply trigger higher clinical suspicion and testing rates rather than genuine transmission events—a bias severely exacerbated by the massive respiratory testing surges during the COVID-19 pandemic. To systematically dismantle this threat and address your specific recommendations, we executed a ground-up restructuring of our analytical framework, implementing four complementary strategies:

      (1) Primary outcome redefined as Positivity Rates. Throughout the entire revised manuscript, both the DLNM and LSTM pipelines have been completely re-analyzed using influenza positivity rates (laboratory-confirmed cases ÷ total ILI specimens tested) as the primary outcome, rather than raw case counts. The positivity rate is the WHOrecommended metric for sentinel surveillance precisely because it mathematically standardizes the denominator, normalizing the raw testing volume variability. Its adoption effectively neutralizes the circular-reasoning concern raised by you. All predictive models, Figures 4, 5, and 6, and all primary metrics in Table 2 now reflect positivity-based estimates.

      (2) Weekly detection volume integrated as a dynamic LSTM covariate. Even after positivity-rate normalization, residual testing-intensity effects can persist (e.g., if testing patterns shift among demographic subgroups with systematically different positivity profiles). To capture this, our revised LSTM network incorporated weekly detection volumes as an explicit dynamic input feature. As detailed in our covariate-ablation sensitivity analysis (Table 2), removing the testing-volume covariate drastically degraded the forecasting accuracy for 2023, increasing the Mean Absolute Error (MAE) by 22.2% for influenza A and an astounding 100% for influenza B. This indicates that the model’s predictions are not driven solely by meteorological inputs but are appropriately and dynamically conditioned on the surveillance context.

      (3) SHAP quantification of surveillance bias. By incorporating testing volume into the LSTM, we made the surveillance-bias concern concretely visible and quantifiable. As shown in our new SHAP analysis (Figure 6E-H), weekly_detection emerged as the second most impactful predictor of the positivity rate for influenza B, and, when examining the influenza A results, weekly_detection likewise ranked fifth among the most impactful predictors. This proves that the deep learning algorithm autonomously recognized the profound impact of testing intensity and actively utilized it to dynamically adjust and calibrate its epidemiological forecasts.

      (4) Decoupling “Circular Reasoning” via DLNM Sensitivity Analysis. To definitively prove that our model does not merely “confirm historical opinions through increased testing,” we conducted a targeted DLNM sensitivity analysis (detailed in the new Additional file 4 and Figures S3–S4). We compared DLNM exposure-response curves predicting positivity rates with and without adjusting for weekly testing volumes. The results were striking: the non-linear exposure-response curves and extreme-weather lag patterns remained highly consistent across both models. This empirical stability proves that the identified meteorological drivers represent intrinsic biological/environmental triggers of viral transmission, independent of fluctuating surveillance intensity.

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version. We invite you to review the recently included sentences in the relevant subsections as specified above.

      “To rigorously address the inherent confounding effects of “surveillance intensity bias”, where fluctuations in raw case counts may merely reflect transient surges in clinical testing capacity rather than true community transmission, the primary outcome metric for all DLNM modeling was mathematically defined as the influenza positive rate, with meteorological factors and weekly detection volume serving as independent variables. The formula for calculating the daily influenza positivity rate is as follows:

      where R represents the daily influenza positivity rate, I represents the number of daily influenza positive cases, and N denotes the total number of daily influenza tests performed. By adopting this WHO-recommended surveillance metric, our our analytical framework explicitly standardizes the epidemiological denominator. This mathematical normalization effectively neutralizes the severe surveillance intensity bias caused by dramatic testing volume surges during the COVID-19 pandemic, ensuring that our models capture intrinsic viral transmissibility rather than artificial fluctuations in healthcare-seeking behavior or diagnostic capacity.” (Methods, page 16)

      “Furthermore, to evaluate the robustness of our findings against testing intensity, we conducted a targeted sensitivity analysis by reconstructing the DLNM models without the “weekly detection volumes” covariate. We then compared the non-linear cumulative risks and extreme weather lag effects between these ablated models and the original fully adjusted models. This allowed us to determine whether the identified climate-transmission associations were stable and biologically intrinsic, or merely artifacts of weather-correlated testing behaviors (Supplementary Information Additional file 4).” (Methods, page 17-18)

      “To assess the contribution of pandemic-related confounding variables, we performed targeted sensitivity analyses on the LSTM networks. Specifically, we sequentially removed covariates of mask-wearing stringency indices and weekly detection volumes from the input features while maintaining identical Bayesian-optimized hyperparameters. The performance of these ablated models was evaluated using MAE, RMSE, MAPE, and SMAPE, allowing us to quantify the exact necessity of incorporating testing and behavioral covariates in forecasting models during periods of epidemiological non-stationarity.” (Methods, page 22)

      “Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 40)

      “This study has several limitations that should be considered when interpreting the findings. Firstly, our primary analysis relies on influenza surveillance data collected from seven sentinel hospitals in Putian, which inherently captures only a fraction of all influenza cases occurring in the broader community. Although employing positivity rates as the primary outcome substantially mitigates the surveillance bias inherent in count-based analyses, residual selection effects may persist if testing patterns shift differentially across demographic subgroups with systematically divergent positivity profiles. Our inclusion of weekly detection volumes as an explicit LSTM covariate, complemented by parallel DLNM sensitivity analyses validating the independence of meteorological effects from testing volumes, were designed to characterize and partially account for this residual bias, but cannot fully eliminate it. Consequently, positivity rates likely underestimate the true burden of community influenza infection. Although the surveillance infrastructure in Putian remained uniform and uninterrupted throughout the 2018–2023 timeline, mitigating measurement bias and temporal confounding risks, the transferability of our findings to other global subtropical regions with different socioeconomic structures, healthcare systems, or population behaviors requires cautious, region-specific calibration. Accordingly, our framework should be strictly interpreted as a high-fidelity tool designed to forecast the observable public health surveillance signal, which is the most operationally relevant target for public health agencies, rather than for estimating unobserved, absolute community disease burden.” (Discussion, page 43-44)

      “This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNM-LSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version.

      (1) Although the authors note a correlation between influenza and the weather factors. The authors do not discuss some of the high correlations between weather factors (e.g., solar radiation and UV index). Because of the many weather factors, those plots are hard to parse.

      We sincerely appreciate your constructive feedback regarding the visual clarity of the correlation plots and the methodological implications of highly correlated meteorological variables. We agree that the original 11 × 11 scatterplot matrix was visually overwhelming and could obscure the statistical implications of highly correlated features (such as solar radiation and the UV index, Pearson’s r > 0.9).

      To address this, we have taken two specific revisions:

      (1) Improved Visualization (Revised Supplementary Figure S2): To make complex relationships easier to parse, we have completely redesigned Supplementary Figure S2. It is now logically partitioned into two distinct visual components:

      - Panel A features a high-contrast Pearson correlation heatmap, where color intensity and statistical significance asterisks allow for rapid, intuitive identification of highly correlated pairs (such as solar radiation and UV index).

      - Panel B retains the pairwise scatterplots (lower-left) and density distribution curves (diagonal) to facilitate the detailed visual inspection of non-linear trends, data skewness, and potential anomalies.

      (2) Methodological Defense on Potential Collinearity: We did not arbitrarily eliminate these highly correlated variables because our two-stage modeling approach inherently mitigates multicollinearity risks associated with potential collinearity:

      - For the DLNM analysis: To prevent coefficient instability caused by multicollinearity, all DLNM lagged analyses strictly employed univariate exposure-response models for meteorological factors (i.e., evaluating the relationship between a single meteorological factor and influenza positivity rates at a time, while controlling for long-term temporal trends and including other covariates as fixed terms). Consequently, the independent effect sizes and lag structures derived from the DLNMs are completely unaffected by inter-variable correlations.

      - For the LSTM network: Unlike traditional multiple linear regression (or ARIMA) where potential collinearity inflates standard errors and destabilizes coefficients, Deep Learning architectures (LSTM) are natively robust to redundant features. The network’s non-linear activation functions and gating mechanisms naturally weight overlapping signals during the optimization process, effectively using redundant variables as a form of algorithmic regularization without compromising predictive stability.

      We have expanded the Statistical Analysis and Discussion sections of the manuscript to explicitly articulate this rationale, ensuring maximum methodological transparency.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      “Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts.” (Discussion, page 40)

      We hope that our responses and revisions will meet your expectations and demonstrate our dedication to improving the scientific quality of this study.

      (2) The authors do not actually compare the results of both methods and what the LSTM adds.

      We sincerely appreciate your perceptive critique. We agree that our initial manuscript lacked a sufficiently rigorous, side-by-side comparison, and more importantly, it failed to adequately articulate why the LSTM architecture succeeds where classical statistical baselines fail.

      To rectify this, we have comprehensively overhauled the comparative analysis. We formulated our baseline as a multivariate ARIMA model, supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes), focusing on three dimensions:

      (1) Direct Quantitative Comparison: Following the restructuring of our outcome variable to the Influenza Positivity Rate, we re-evaluated both models using identical training (2018–2022) and validation (2023) datasets. The LSTM consistently and substantially outperformed the ARIMA model across all metrics. For Influenza A, the LSTM achieved an MAE of 0.009 and RMSE of 0.035 (vs. ARIMA: MAE 0.136, RMSE 0.238). For Influenza B, the LSTM yielded an MAE of 0.002 and RMSE of 0.011 (vs. ARIMA: MAE 0.049, RMSE 0.057). We have updated Table 2 and the corresponding Results section to explicitly present these side-by-side comparisons.

      (2) Visualizing “Where” the LSTM Outperforms: We have updated Supplementary Figure S3 to map the ARIMA predictions for the 2023 positivity rate, allowing for a direct visual comparison with the LSTM predictions in Figure 6 (A, B). The visualizations explicitly reveal the ARIMA model’s fundamental limitation: it tends to predict relatively flat or conservatively smoothed values, failing entirely to capture the extreme, explosive non-linear peaks of the 2023 viral rebound. Conversely, the LSTM network accurately tracks these sudden epidemic phase transitions.

      (3) Articulating “What the LSTM Adds” (Discussion Expansion): We have significantly expanded the Discussion section to intellectually articulate why the LSTM succeeds where ARIMA fails. To ensure a strictly fair methodological comparison, we formulated our baseline as a multivariate ARIMA model (ARIMAX), supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes). Therefore, the LSTM’s superior performance is not due to information asymmetry (i.e., it did not “see” more variables), but stems directly from its algorithmic architecture. The consistent superiority of the LSTM over both the linear sequential baseline (ARIMA) and the non-linear non-sequential baseline (XGBoost) isolates the recurrent gated architecture itself as the source of the predictive gain. This advantage arises from four architectural properties intrinsic to recurrent gated networks yet absent in both tree ensembles and linear autoregressive models:

      a) Sensitivity to temporal ordering, which decision-tree splits and linear regressors cannot natively encode;

      b) Gated propagation of long-range dependencies through forget–input–output mechanisms;

      c) Explicit accommodation of serial autocorrelation, violated by the i.i.d. assumptions underlying gradient-boosted trees;

      d) Paradigmatic comparability with sequential statistical models: the joint failure of ARIMA (linear, sequential) and XGBoost (non-linear, non-sequential) isolates recurrent sequential memory, rather than non-linearity per se, as the critical feature for forecasting under pandemic-era non-stationarity.

      We emphasize that this interpretation applies specifically to the present non-stationary epidemiological forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain competitive performance across many structured prediction domains.

      We invite you to review the recently revised table, figures and sentences in the relevant subsections as specified above.

      “The LSTM networks accurately captured both the timing and magnitude of these nonlinear epidemic surges, including the two outbreak peaks of influenza A during February-March and November-December of 2023, as well as the peak of influenza B in November-December, demonstrating good predictive performance. The predictive performance of the LSTM networks was quantitatively assessed using metrics such as MAE, RMSE, MAPE, and SMAPE. For influenza A, the MAE was 0.009, RMSE was 0.035, MAPE was 0.158, and SMAPE was 0.521; for influenza B, the MAE was 0.002, RMSE was 0.011, MAPE was 0.17, and SMAPE was 0.484 (Table 2).” (Results, page 31)

      “To benchmark the predictive value added by the LSTM architecture, we constructed two baseline models using identical training (2018–2022) and validation (2023) positivity-rate datasets, inclusive of all contextual covariates: the multivariate ARIMA model and the XGBoost gradient-boosting model. The LSTM model demonstrated decisive superiority over both baselines across all evaluated metrics (Supplementary Table S3). For influenza A, the LSTM achieved an MAE of 0.009 and SMAPE of 0.521, compared with the ARIMA’s MAE of 0.136 (SMAPE 1.212) and XGBoost’s MAE of 0.138 (SMAPE 1.081), representing approximately 15-fold error reductions relative to both benchmarks. For influenza B, the LSTM yielded an MAE of 0.002 (SMAPE 0.484), compared with 0.049 (SMAPE 0.810) for ARIMA and 0.070 (SMAPE 0.737) for XGBoost, approximately 25- to 35-fold reductions. Notably, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, suggesting that the LSTM’s advantage derives not from non-linear modeling capacity per se, but from architectural properties specific to recurrent sequential processing.

      Beyond global error metrics, visual comparison of forecasting trajectories (Figure 6; Supplementary Figures S5 and S6) reveals critical behavioral disparities. For influenza A, both the linear ARIMA model and the non-linear XGBoost model generated conservatively smoothed forecasts that entirely failed to capture the sudden, explosive peaks of the 2023 post-restriction viral rebound. For influenza B, ARIMA produced continuous spurious fluctuations during non-epidemic periods (when true positivity was near zero), likely overreacting to covariate variations, while XGBoost generated largely flat trajectories that missed the mid-year outbreak peak. In stark contrast, the LSTM’s recurrent gating mechanisms successfully filtered out covariate noise during low-transmission periods while accurately tracking extreme epidemiological phase transitions, demonstrating a qualitative advantage in handling non-stationary regime shifts.” (Results, page 33-34)

      “To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: the multivariate ARIMA model and the XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 41-42)

      Results

      “The MAE for the covariate-adjusted ARIMA model of influenza A was 0.136, RMSE was 0.238, MAPE was 1.115, and SMAPE was 1.212. For influenza B, the MAE was 0.049, RMSE was 0.057, MAPE was 1.426, and SMAPE was 0.810. Overall, the predictive performance of the ARIMA model was substantially inferior to that of the Bayesian optimized LSTM. As illustrated in Figure S5, the linear model struggled significantly with the non-stationary dynamics of the 2023 viral rebound. It either failed entirely to capture the extreme, explosive non-linear peaks (as seen in Influenza A) or generated continuous spurious predictions during zero-case periods due to mechanical linear reactions to covariate inputs (as seen in Influenza B).” (Supplementary information, page 16)”

      Results

      “The XGBoost model produced substantially higher forecasting errors than the LSTM across both influenza subtypes (Table S3; Figure S6). For Influenza A, XGBoost yielded MAE = 0.1379, RMSE = 0.2161, MAPE = 1.8195, and SMAPE = 1.0805, errors of a magnitude broadly comparable to the multivariate ARIMA baseline (MAE = 0.136) and approximately 15-fold higher than the LSTM (MAE = 0.009). For Influenza B, XGBoost achieved MAE = 0.0704, RMSE = 0.0739, MAPE = 1.9172, and SMAPE = 0.7371, exceeding both the ARIMA benchmark (MAE = 0.049) and the LSTM (MAE = 0.002) by approximately 35-fold relative to the latter.” (Supplementary information, page 19)

      We believe these additions provide a much more rigorous and academically satisfying comparative analysis.

      (3) The methods are long and meandering. They could be cleaned up and shortened. E.g., there is no need for 30 lines on PCR testing; the study area should come before the study design. The authors discuss similar elements in multiple places; this whole section can be shortened considerably without affecting the content.

      We sincerely appreciate the reviewer’s editorial guidance. We agree that the initial Methods section was somewhat disjointed and unnecessarily verbose, reflecting iterations from previous drafts. We have completely restructured and aggressively streamlined this section to ensure a logical and concise flow:

      - Structural Reorganization: As suggested, we have moved the “Study area” section to the very beginning of the Methods, providing the geographical and climatic context before detailing the study design and cohort.

      - Consolidation of Redundancies: We have merged the fragmented descriptions regarding ethical approvals, data anonymization, and sentinel hospital protocols into a single, cohesive “Study design and cohort” subsection.

      - Condensing Laboratory Protocols: We completely agree that 30 lines on standard PCR testing in the main text are unnecessary. We have condensed the “Sample collection and pathogen typing” section into a brief, 4-line summary explicitly stating the use of commercial assays (Da’an Gene Co., Ltd) and strict adherence to China CDC guidelines. The highly technical nuances (e.g., RNA extraction integrity, spectrophotometer ratios, and precise PCR amplification thresholds) have been relocated to the Supplementary Information (Additional file 1) for interested readers.

      We invite you to review the recent revisions in the Methods subsection as specified above.

      “Sample collection and pathogen typing

      Respiratory specimens (nasopharyngeal swabs) were collected from ILI patients during their initial visit, prior to treatment, and stored at 4°C in viral transport medium. All samples were delivered to the Putian CDC laboratory within 24 hours of collection. Nucleic acid extraction and one-step real-time fluorescent RT-PCR for influenza A/B subtyping (H1N1, H3N2, Victoria, and Yamagata lineages) were executed within 24 hours upon sample arrival. All assays utilized commercial diagnostic kits (supplied by Da’an Gene Co., Ltd., Guangzhou, China) and were processed in a Biosafety Level 2 (BSL-2) laboratory, strictly adhering to the manufacturer’s instructions and China CDC’s standardized protocols. Detailed laboratory procedures, including RNA integrity parameters and PCR amplification thresholds, are comprehensively documented in the Supplementary Information (Additional file 1).” (Methods, page 13)

      These revisions have significantly improved the readability of the manuscript without sacrificing methodological transparency.

      (4) How reliable is the "Our Word in Data" website for subnational coverage of restrictions? Some of the authors are from Putian and should be able to confirm the accuracy for both studied areas.

      We are very grateful for this insightful question. You astutely identify a common challenge in geospatial epidemiology in China: the systemic lack of publicly accessible, standardized, daily non-pharmaceutical intervention (NPI) datasets at the municipal (subnational) level. Local CDC policy records are typically maintained as internal administrative documents without standardized time-series data interfaces.

      Given this limitation, we utilized the national Our World in Data (OWID) stringency index as a proxy. To directly address the reviewer's excellent point regarding local verification, the co-authors of this study—who are frontline epidemiologists stationed at the Putian CDC and who directed the local pandemic response, including conducted a rigorous, retrospective cross-validation of the OWID index against local realities.

      Our local experts confirmed a high degree of fidelity between the OWID 0–4 scale and the actual policies enforced in Putian and Sanming:

      - 2020 (Initial Outbreak): Both cities enforced “mandatory face coverings outside the home at all times” (OWID Level 4), aligning perfectly with the dataset.

      - 2021–2022 (Normalized Control): Policies shifted to “required in all shared/public spaces” (OWID Level 3), which accurately reflects the local mandates required for public transit, schools, and commercial venues.

      - 2023 (Post-Pandemic Shift): Following the national policy pivot, mandates were downgraded to “recommended” (OWID Level 1), perfectly mirroring local ground truths.

      Therefore, while OWID provides a national-level index, our local CDC authors have empirically verified that its temporal variations accurately capture the intensity of behavioral restrictions experienced by the populations in our specific study areas. We have incorporated a concise statement regarding this expert validation into the revised Methods section.

      “While OWID stringency indices represent national-level policy, publicly accessible and standardized daily NPI datasets at the municipal level are currently unavailable in China. To ensure the spatial validity of these indices, co-authors from Putian CDC, who actively managed the local epidemic response, conducted a rigorous cross-validation. Our local public health experts confirmed that temporal fluctuations of OWID indices (ranging from Level 4 strict mandates in 2020 to Level 1 recommendations in 2023) exhibited high fidelity with the actual, on-the-ground enforcement of NPIs in both Putian and Sanming. Thus, OWID stringency indices serve as a highly reliable contextual proxy for local social contact restrictions.” (Methods, page 15)

      We hope that our responses and the additional statement will meet your expectations.

      (5) Figure 2A is hard to parse; it would make more sense to plot these as line plots (y=count, x=month).

      We completely agree with you. The original heatmap visualization obscured the temporal dynamics of the distinct influenza subtypes. We have entirely redrawn Figure 2A as multi-line plots mapping the monthly incidence trajectories of Influenza A (H1N1, H3N2) and Influenza B (Victoria, Yamagata) across the six-year study period. This new visualization (included in the revised manuscript) vastly improves readability and explicitly highlights the distinct phase shifts and interruptions caused by the pandemic.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3 is hard to parse. The flu counts are repeated in every figure, which makes them look very similar. I would recommend that the authors pick 1 or 2 as subpanels for the main paper and put the rest in the supplement.

      We sincerely appreciate your feedback. We completely agree that the original 3x3 square layout severely compressed the x-axis, making the daily temporal fluctuations difficult to parse and causing the flu curves to look visually redundant.

      To resolve this visual clutter without losing valuable environmental context in the main text, we adopted a highly effective structural solution simultaneously suggested by Reviewer #2. We have completely redrawn Figure 3 into an 8x1 vertically stacked format with a single, shared continuous x-axis across the entire study period. This extended aspect ratio vastly expands the timeline, dramatically revealing the highly distinct, daily microfluctuations of each meteorological factor alongside the epidemiological curves. We believe this new layout completely resolves the parsing difficulty you rightly pointed out. Given that a substantial portion of our readership relies heavily on the main text figures for immediate epidemiological context, retaining these cleanly formatted panels in the main manuscript maximizes the paper’s scientific impact. We hope you find this redesigned visualization satisfactory.

      (2) Using a thousand-separator throughout will make the manuscript more readable.

      We completely agree. We have meticulously applied thousand-separators to all relevant numerical values (e.g., 20,488; 17,333) throughout the revised manuscript to enhance readability.

      (3) Line 556 "quantified calculated", pick 1 word.

      We sincerely apologize for this typographical oversight resulting from the drafting process. However, the original sentence that led to the duplicated phrasing you highlighted has been removed, as we had already undertaken a comprehensive revision of the relevant material in that subsection in response to your earlier remarks. We invite you to review the newly substituted paragraph below.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used only 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year round. The authors further claimed that such an association could be able to provide an early warning in the community. In this direction, the current manuscript has several scopes of improvements and clarification of the claims, as I list here.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We sincerely thank you for the careful and constructive evaluation of our manuscript, and in particular for recognising the value of our prospective surveillance design and the importance of investigating influenza–meteorological associations in subtropical settings where year-round circulation patterns differ substantively from those in temperate regions. We are grateful that you have identified four specific dimensions in which the manuscript can be strengthened, rationale clarity, methodological/data-integration transparency, validation reporting, and the calibration of the “early warning” claim. We address each of these four points in detail below, and we have undertaken substantive revisions to the manuscript in response.

      Weaknesses:

      (1) The rationale of the study is not clearly stated.

      We sincerely thank you for this incisive observation. We agree that the original Introduction did not adequately articulate the study’s rationale, specifically, the causal chain linking public-health need, existing methodological limitations, and the incremental contribution of our integrated DLNM-plus-LSTM framework. The Introduction has been substantively rewritten to make this rationale explicit, structured around four logical pillars:

      - Disease burden grounding. We have added quantitative evidence on the global burden of seasonal influenza, such as annual mortality estimates, drawing on solid epidemiological sources, to establish the public-health magnitude that motivates the study.

      - Subtropical-specific knowledge gap. We articulated the distinctive epidemiological challenges of subtropical influenza transmission, including year-round circulation patterns, complex non-linear meteorological associations, and lag-structured exposure-response relationships, that fundamentally differentiate subtropical contexts from temperate epidemiological settings where most existing research has been conducted. This articulation directly motivates our adoption of distributed lag non-linear models (DLNM) as the appropriate analytical framework for capturing these complex non-linear and lag-structured associations.

      - Methodological gap and incremental contribution. We now position our integrated framework against three specific gaps in the existing literature: (a) studies using DLNM alone characterize lag-distributed exposure–response relationships but lack forecasting capability; (b) studies using LSTM alone provide forecasts but typically do not incorporate Bayesian hyperparameter optimization, do not stratify by influenza subtype, and do not account for COVID-19-era non-pharmaceutical interventions; (c) no existing study, to our knowledge, integrates DLNM-based mechanistic interpretation with Bayesian-optimized, subtype-specific LSTM forecasting in a subtropical Chinese setting under pandemic-perturbed surveillance conditions. Our study is positioned to fill this specific gap.

      - Adequacy of the six-year data window. We additionally address your implicit concern regarding study duration. While six years (2018–2023) is shorter than some long-horizon influenza time-series studies, this window was deliberately selected because it brackets a uniquely informative epidemiological transition: two prepandemic baseline years (2018–2019), three Non-Pharmaceutical Interventions (NPI)-suppressed years (2020–2022), and one post-suppression rebound year (2023). This structure allows the model to learn from a structural break that a longer but earlier-only series could not provide. We have made this argument explicit in the revised Introduction.

      - The dual-model rationale. We clarify that our core rationale is to bridge this gap. By utilizing DLNM to uncover the underlying environmental biological triggers and subsequently employing the Bayesian-optimized LSTM network, uniquely upgraded to ingest epidemiological context (mask-wearing stringency index, testing volumes, day and week, viral importation), we provide a comprehensive framework that achieves both mechanistic insight and operational forecasting agility. We have substantially rewritten the Introduction to reflect this explicit storyline.

      We close the revised Introduction by explicitly framing the study’s translational endpoint: providing a methodologically integrated framework (DLNM for mechanistic interpretation; Bayesian-optimised LSTM for forecasting) to support climate-informed influenza preparedness in subtropical settings. We have, however, calibrated the language used to describe this endpoint (see our response to Weakness #4) to avoid overstating the study’s operational readiness as an early-warning tool.

      We have substantially rewritten the Introduction to reflect this sharpened rationale. We invite you to read the whole section of the Introduction in the revised manuscript.

      (2) Several issues with methodological and data integration should be clarified.

      We sincerely appreciate your identification of methodological and data integration ambiguities in the original manuscript. We have substantively addressed this concern through three categories of clarifying revisions: (i) explicit articulation of the covariate framework, (ii) clarification of the analytical relationship between DLNM and LSTM components, and (iii) detailed specification of data sources and quality control procedures:

      - Clarification 1: Comprehensive covariate framework specification. We have explicitly articulated the complete covariate framework integrated into both DLNM and LSTM analyses. Beyond the meteorological factors (mean temperature, maximum temperature, minimum temperature, diurnal temperature range, relative humidity, atmospheric pressure, precipitation, sunshine duration), our analytical framework systematically incorporates: (a) influenza positivity rates as the primary outcome variable (replacing raw case counts to mitigate surveillance intensity bias, as detailed in our response to Reviewer 1’s Public Review Comment 2); (b) weekly testing volumes as an explicit covariate to control for residual surveillance-intensity variations; (c) mask-wearing stringency indices to capture pandemic-era public health intervention effects; (d) day-of-week indicators distinguishing weekdays from weekends to control for healthcare-seeking behavioral cycles; and (e) nonlocal population proportion to account for population mobility-related transmission dynamics.

      - Clarification 2: Articulation of the DLNM-LSTM analytical relationship. We have explicitly clarified the complementary analytical roles of DLNM and LSTM within our integrated framework, addressing potential confusion regarding whether these methods serve redundant or complementary functions.

      - Clarification 3: Data source specification and quality control documentation. We have substantively expanded the data source specification and quality control documentation to ensure full methodological transparency.

      We have revised the “Study design” subsection in the Methods to transparently outline how these multifaceted data streams were temporally aligned and fed into the dual-model architecture. We invite you to review these rewritten paragraphs.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records. To ensure consistency between meteorological measurements and influenza incidence records across all data sources, we applied rigorous quality control and pre-processing procedures, including: temporal alignment of all data streams to a unified daily resolution; missing value imputation using temporally adjacent observations for sporadic gaps (<5% of records); cross-validation of laboratory-confirmed cases against ILI consultation records to identify and resolve coding inconsistencies; and standardization of meteorological measurements against the regional monitoring network’s established calibration protocols. Final datasets underwent independent verification by two co-investigators to ensure analytical reliability.” (Methods, page 11)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the non-linear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather-influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics. Building upon this DLNM-derived foundation, the LSTM network constructs a time-series forecasting tool within an identical covariate framework, evaluating predictive capability for influenza transmission trends. Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency index (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 11-12)

      (3) Validation of the models is not presented clearly.

      We sincerely appreciate your identification of insufficient clarity in the validation framework presentation. We acknowledge that the original manuscript inadequately articulated the multi-tiered validation architecture underlying our analytical framework. We have substantively expanded the validation framework documentation through three categories of clarifications: (i) explicit articulation of the three-tier data partitioning architecture, (ii) detailed specification of validation procedures across each tier, and (iii) systematic enumeration of validation evidence supporting each analytical conclusion.

      - Clarification 1: Three-tier data partitioning architecture. Our validation framework employs a rigorously designed three-tier data partitioning architecture that addresses different validation objectives at each tier.

      Tier 1: Internal training and validation (Putian, 2018-2022). We allocated the 2018-2022 Putian surveillance data as the primary training set, within which 10% of samples were further randomly partitioned as an internal validation subset for hyperparameter tuning and overfitting monitoring during LSTM training. Early stopping mechanisms were implemented to terminate training when internal validation loss plateaued, preventing overfitting to training-specific patterns.

      Tier 2: Internal testing (Putian, 2023). We allocated the 2023 Putian surveillance data as the internal testing set, providing a temporally independent assessment of model predictive performance on data not utilized during training or hyperparameter optimization. The chronological partitioning preserves time series modeling validity by ensuring that all training data temporally precede testing data, avoiding data leakage that could artificially inflate performance estimates.

      Tier 3: External validation (Sanming, 2023). We obtained surveillance data from Sanming city for the period January 1, 2023, to December 31, 2023, matching the temporal coverage of the Putian internal testing set. This external validation set provides geographically independent assessment of model transferability across subtropical Chinese contexts, evaluating whether the Putian-derived model architecture generalizes to a different subtropical city sharing comparable climatic characteristics, influenza seasonality patterns, and public health intervention frameworks.

      - Clarification 2: Unified model framework across all validation tiers. A critical methodological feature of our validation framework is that the identical Bayesian-optimized LSTM architecture trained on Putian 2018-2022 data was applied without modification across all three validation tiers. This unified framework approach is methodologically essential because: (a) it tests genuine model transferability rather than evaluating differently-tuned models at each tier, which would conflate validation with re-optimization; (b) it enables direct performance comparison across internal testing and external validation, isolating the marginal performance degradation attributable to geographic transfer; and (c) it aligns with operational deployment scenarios where a trained model must be applied to new contexts without re-training.

      - Clarification 3: Validation evidence enumeration. The validation evidence supporting our analytical conclusions encompasses four complementary dimensions:

      Predictive performance metrics: Across both influenza A and B, the Bayesian-optimized LSTM achieved low error metrics on the internal testing set (influenza A: MAE = 0.009, RMSE = 0.035, MAPE = 0.158, SMAPE = 0.521; influenza B: MAE = 0.002, RMSE = 0.011, MAPE = 0.170, SMAPE = 0.484), substantially outperforming ARIMA benchmark models (Supplementary Figure S3).

      External validation: The Putian-derived LSTM successfully generalized to Sanming external validation data, with performance metrics maintaining comparable magnitudes to internal testing performance, substantiating model transferability across subtropical contexts.

      Sensitivity analyses: We conducted systematic sensitivity analyses across both DLNM and LSTM components. DLNM sensitivity analyses (Supplementary Figures S4-S5) demonstrate substantial concordance in cumulative risk patterns and lag-specific extreme condition responses across models with and without weekly testing volume adjustment, substantiating robustness of meteorological associations. LSTM sensitivity analyses (New Table 2) demonstrate that systematic covariate exclusion produces predictable and biologically plausible performance degradation patterns rather than artificially robust performance, confirming the absence of overfitting characteristics.

      Interpretability verification: SHAP interpretability analysis (Figure 6E-H) substantiates that the LSTM autonomously identified epidemiologically plausible feature importance hierarchies, providing independent verification that model predictions reflect genuine biological signal recognition rather than data artifacts.

      We invite you to review the corresponding manuscript clarifications.

      “The LSTM network was trained on data from Putian corresponding to the four categories described above, with the time period 2018-2022, and Putian’s 2023 data serving as the internal validation set. To rigorously evaluate model transferability beyond the training context, data from Sanming city, a mountainous subtropical city exhibiting comparable climatic characteristics, influenza seasonality, and public health intervention frameworks to Putian, were acquired for the period January 1, 2023, to December 31, 2023, temporally aligned with the Putian internal validation set. This Sanming dataset constituted our external validation set, facilitating a geographically independent assessment of model generalization within subtropical Chinese environments. The predictive performance of the LSTM algorithm for influenza A/B prevalence in 2023 Putian data was benchmarked against a parallel multivariate ARIMA model with exogenous variables. To ensure a fair methodological comparison, this baseline model was supplied with the exact same meteorological and epidemiological covariate matrix as the LSTM.” (Methods, page 12-13)

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Time-series data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale. First, we performed data normalization, a crucial step to ensure that training and test set data are compared on a unified scale. We normalized the training and test set data separately within the range [0, 1]. For the test set normalization, we used the maximum and minimum values from the training set as boundaries. This approach ensured consistency between the normalized test set data and the training set data. After the algorithm conducted predictions on the test set data, we performed denormalization to convert the predicted results back to the original data scale and rounded them to integers. These steps ensured that the final prediction results accurately and objectively reflected the LSTM’s performance in real-world scenarios and provided reliable data for subsequent calculation of evaluation metrics.” (Methods, page 18-19)

      (4) The claim for providing tools for 'early warning' was not validated by analysis and results.

      We are grateful for this incisive critique, which identifies a critical mismatch between our research achievements and the terminology employed in the original manuscript. Upon careful re-examination, we acknowledge unreservedly that the original manuscript’s use of “early warning” terminology overstated our actual research contribution. Our research has constructed and validated a methodologically rigorous LSTM-based influenza forecasting framework demonstrating strong predictive performance and external transferability; however, this constitutes a forecasting framework foundation rather than a fully-validated operational early warning tool ready for direct public health implementation.

      We recognize that genuine early warning tools require additional validation dimensions that our current research does not yet comprehensively address. In response to this important critique, we have implemented three categories of substantive corrections:

      - Correction 1: Comprehensive terminology revision throughout the manuscript. We have systematically revised “early warning” terminology throughout the manuscript, replacing it with more accurate descriptors that precisely characterize our actual research contribution. Specifically: “early warning system” has been revised to “forecasting framework” or “forecasting model”; “early warning tool” has been revised to “predictive modeling foundation”; and “early warning capability” has been revised to “predictive capability supporting future early warning system development”. These terminological refinements ensure that manuscript claims precisely correspond to demonstrated research achievements.

      - Correction 2: Manuscript title revision. We have correspondingly revised the manuscript title to remove “early warning” terminology and accurately reflect the study’s actual contributions: Revised title: “Meteorological Drivers of Influenza A and B Positivity in a Subtropical Chinese City: A Six-Year Surveillance Study Integrating Distributed Lag Non-Linear Models and Deep Learning”. This revised title precisely articulates the study’s actual scope: characterization of meteorological drivers (DLNM contribution), focus on positivity rates (methodological refinement addressing surveillance bias), specification of subtropical context (geographic scope), six-year temporal coverage (data scope), and integration of DLNM and deep learning (methodological framework).

      - Correction 3: Explicit articulation of forecasting framework versus operational early warning tool distinction. We have explicitly articulated the distinction between our current achievements and operational early warning tool requirements in both the Discussion and Conclusion sections, framing future research directions for operational early warning system development.

      We invite you to review the corresponding manuscript clarifications.

      “Importantly, however, this framework should be regarded as a methodological foundation for future operational developments rather than as a deployable early-warning system: routine use in public health practice would require prospective recalibration, integration with operational surveillance infrastructure, and additional validation beyond the scope of the present study.” (Discussion, page 43)

      “In conclusion, this study elucidates the distinct, non-linear meteorological drivers of influenza A and B transmission in a subtropical Chinese urban setting through an integrated dual-stage DLNM-LSTM framework. By adopting influenza positivity rates as the primary outcome and integrating socio-behavioral covariates, including mask-wearing stringency indices, weekly detection volumes, and non-local population proportion indicators, our approach mitigates surveillance-related biases and accommodates pandemic-era nonstationarity. It shows lower forecast error than a covariate-matched ARIMA baseline and provides preliminary evidence of portability within southeastern subtropical China, combining the interpretability of distributed lag modeling with the flexibility of deep learning. The framework offers an interpretable, climate-informed methodological foundation for future operational surveillance developments in subtropical settings. (Discussion, page 47)

      Reviewer #2 (Recommendations for the authors):

      (1) The title is not data-driven in different contexts, including 'early warning'; I was expecting substantial analyses in this direction to assess the 'early warning' in the manuscript. But I hardly found them in the text, merely utter as the implication of understanding the associations between meteorological drivers and influenza in advance. I suggest either revising the title or clarifying the claim by providing significant evidence and its impact on the epidemic onset and intensity. Further, revise 'Subtropical China' as 'a Subtropical Chinese city', as the former one is not accounted under this study.

      We completely agree with this constructive feedback. As detailed in our response to your Public Review weakness (4), we acknowledge that claiming an operational “early warning” system requires extensive real-world feasibility and threshold validations that exceed the scope of our current time-series analysis. We understand that achieving early warning capabilities necessitates the completion of at least three additional validation dimensions, as listed below.

      Dimension 1 - Threshold determination: Early warning systems require explicit thresholds defined by integrating predictions with established epidemiological thresholds, typically via ROC analysis to balance sensitivity and specificity for trigger activation; our framework provides predictions but does not define thresholds.

      Dimension 2 - Deployment validation: Validation of warning timeliness, false alarm control, and integration with existing CDC surveillance architectures is required; our work does not assess these operational dimensions.

      Dimension 3 - Robustness across heterogeneous scenarios: Early warning tools need systematic robustness testing across diverse social environments, public health policy contexts, extreme meteorological events, and co-circulation of emerging pathogens; our results cover 2018–2023 Putian-Sanming but not broader operational scenarios.

      We have modified the title to “a Subtropical Chinese city” and systematically revised “early warning” terminology to “forecasting” throughout the manuscript. These revisions ensure accurate representation of the study’s scope and contribution.

      (2) There are several studies establishing the potential association between influenza and the climatic drivers in several locations across the globe. I couldn't find sufficient text on establishing the rationale of this study from the perspective of existing literature. This should clearly be uttered in the introduction section itself.

      We are grateful for this critique, which echoes your Public Review weakness (1). We fully agree that the original Introduction lacked a cohesive narrative connecting the existing literature to our specific methodological innovations.

      To address this, we have comprehensively rewritten the Introduction section. The revised text now systematically establishes our rationale through a clear logical progression:

      - Acknowledging existing studies on climatic drivers but highlighting the unique challenge of non-linear, year-round influenza transmission in subtropical regions.

      - Identifying the methodological gap: existing models either use DLNM purely for retrospective explanation (lacking prediction) or employ deep learning (LSTM) purely for prediction (lacking epidemiological interpretability).

      - Highlighting the critical failure of current literature to mathematically adjust for the profound non-stationarity and surveillance intensity biases introduced by the COVID-19 pandemic (fluctuating testing volumes and NPIs).

      - Introducing our dual-stage solution: integrating DLNM and LSTM to forecast the Influenza Positivity Rate (rather than raw cases), structurally augmented with masking and testing volume covariates.

      We believe this robust literature review now unequivocally establishes the necessity and novelty of our study.

      (3) In connection with the above point, why the authors required the prospective cohort to assess a historical outcome should be highlighted clearly, which is one of the selling points of the study.

      You astutely highlight one of the core methodological strengths of our study design, and we appreciate the opportunity to emphasize this “selling point.”

      As noted in our response to Reviewer #1 regarding timeline splits, the phrase “prospective cohort to assess historical outcomes” reflects the administrative timeline of our study, but mathematically, the data stream is a continuous, longitudinal ecological surveillance.

      The critical selling point here, which we have now explicitly highlighted in the revised Methods section, is that our “historical” data (2018–2020) was not collected via traditional, unstructured retrospective chart reviews. Instead, it was derived from an already operational, highly standardized public health sentinel surveillance system. Because this system utilized identical clinical case definitions, swabbing protocols, and RT-PCR diagnostic assays continuously from 2018 through 2023, the historical data inherently possesses the high fidelity, standardized quality, and lack of recall bias typically reserved for strict prospective cohorts. We have modified the Ethics statement subsection to explicitly underscore this epidemiological advantage.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018–2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      (4) The authors retrieved the daily data on the cases, which is usually small in number for most of the time, can often be driven by the importation (by population mobility with risk of infections) for particularly in a small location like a city.

      We deeply appreciate this incisive epidemiological observation. We fully agree that in a municipal-scale study, relying solely on daily absolute case counts presents significant mathematical and epidemiological vulnerabilities: absolute numbers can be small, highly stochastic, and susceptible to sudden spikes driven by imported cases rather than indigenous climate-driven transmission.

      Driven directly by your comment, we have implemented two fundamental, structural upgrades to our study design:

      - Shift to Positivity Rates: As detailed in our previous responses, we have entirely abandoned daily absolute case counts. Our DLNM and LSTM models now strictly utilize daily influenza positivity rates (positive cases ÷ total daily tested samples) for influenza A and B as the primary outcome. Positivity rates inherently smooth out the stochastic noise of small daily counts and provide a robust, normalized metric of true transmission intensity.

      - Explicit Modeling of Importation Risk: To directly address the risk of importation via population mobility, we have integrated a novel covariate into our LSTM network: the daily proportion of non-local population/residents tested (labeled as outsiders_proportion in our SHAP analysis). By explicitly feeding this mobility proxy into the deep learning algorithm, the model is now mathematically equipped to contextualize and partial out the influence of imported infections when forecasting local transmission trends.

      (5) The authors have not considered this extrinsic factor in the account and not even discussed it.

      We apologize for previously neglecting this critical extrinsic factor. As outlined in our response to Recommendation 4, we have now explicitly operationalized this extrinsic factor by incorporating daily non-local population proportion (outsiders_proportion) as a dynamic input feature in our revised LSTM architecture.

      Furthermore, to empirically assess the magnitude of this importation risk, we conducted a retrospective analysis of the demographic data spanning our six-year study period. Our descriptive statistics reveal that days where the non-local population accounted for >50% of the daily tested cohort represented less than 1% of the total study days.

      This empirical finding allows us to draw two important conclusions: First, while importation undoubtedly occurs (and is now accounted for by our LSTM covariate), indigenous transmission remains the overwhelmingly dominant driver of the observed epidemic curves in Putian. Second, massive importation shocks are rare enough that they do not systematically skew the overarching climate-disease associations identified by our models. We have thoroughly integrated both the methodological adjustment and this empirical discussion into the revised Methods and Discussion sections.

      We invite you to review the corresponding manuscript clarifications.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records.” (Methods, page 11)

      “Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency indices (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 12)

      Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates.

      Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (6) Further, how could such a small number of cases (which can be sporadic) define the epidemic onset and its uncertainty?

      You are absolutely correct: defining an epidemic onset using a small, sporadic number of absolute daily cases introduces severe statistical uncertainty and false-positive onset triggers. This specific methodological vulnerability was a primary catalyst for our decision to fundamentally pivot our analytical framework from absolute cases to Influenza Positivity Rates.

      Unlike absolute counts, where a jump from 1 to 5 sporadic cases might artificially trigger an “onset” definition, positivity rates provide a continuous, normalized epidemiological signal. By assessing the proportion of positive tests against the total testing denominator, positivity rates mathematically stabilize the variance caused by sporadic daily testing. Consequently, an upward trajectory in positivity rates provides a highly reliable, low uncertainty signal of true epidemic onset and acceleration. The exceptional validation metrics of our revised LSTM model (e.g., MAE of 0.009 for Influenza A positivity rate) demonstrate that utilizing this normalized metric virtually eliminates the noise and uncertainty associated with sporadic small-number counts.

      (7) Figure 3 presents the time series of the cases. I wonder whether the data for these factors and outcomes are daily or aggregated by week/month? I suggest representing it in 9x1 format with a single x-axis to compare, instead of 3x3 format. Authors can refer similar plot in https://doi.org/10.1371/journal.pcbi.1012311 in Figure 1.

      We are extremely grateful for this specific and highly constructive visualization suggestion. To answer your query: the data plotted for both the meteorological factors and the influenza outcomes are indeed daily observations.

      We fully agree that the original 3x3 format severely compromised the readability of this daily data. Following your excellent advice and referencing the suggested literature, we have entirely redesigned Figure 3 into an 8x1 vertically stacked format with a single shared continuous x-axis. This structural upgrade has completely transformed the figure, eliminating the horizontal compression and elegantly exposing the fine-grained, daily temporal alignments between climatic extremes and viral surges. The newly rendered Figure 3 is now much more intuitive and analytically valuable. Due to space constraints and given that the revised Figure 3 has been included in our response to a similar comment from Reviewer #1, we will not reproduce the figure in this response. We invite you to review the revised manuscript or refer to our response to Reviewer #1’s recommendation 1 for Figure 3.

      (8) "Additionally, we plotted the loss function curves for the network on the training and validation sets to monitor LSTM convergence and the risk of overfitting." The authors validated the model with predefined training and validation sets. I suggest providing more details on the techniques and the length of the sets. How are these considerations safe for the assumptions and limitations of the models?

      We sincerely thank you for this crucial request for methodological transparency. You are absolutely correct that the techniques used for data partitioning are fundamental to the safety and validity of time-series modeling assumptions. To strictly respect the temporal dependencies of LSTM networks and avoid future-to-past data leakage, we avoided standard random train/test splitting. Instead, we implemented strict Chronological Out-of-Time (OOT) Three-Tier Validation Architecture.

      We have extensively expanded the Methods and Discussion sections to detail this partitioning scheme, our rationale, and its associated limitations:

      - The Three-Tier Validation Architecture and Length of Sets:

      Tier 1 — Training and Internal Cross-Validation (2018–2022, ~88%): Used for initial model fitting. Within this phase, 10% of the samples were held out during Bayesian optimization as an internal validation fold strictly for monitoring convergence, controlling overfitting (via early stopping), and guiding the hyperparameter search. This internal fold never contributed to the final reported performance metrics.

      Tier 2 — Internal Hold-out Test Set (2023, ~12%): A continuous 365-day block reserved as a strictly chronological hold-out, used solely for final performance evaluation. No information from this period influenced training or hyperparameter tuning.

      Tier 3 — External Independent Validation (Sanming 2023): The full surveillance time series from a geographically distinct subtropical city, providing the strongest evidence of cross-location generalizability.

      - Justification of the 88% / 12% Partition Ratio:

      This specific ratio was deliberately chosen to balance two competing epidemiological and computational considerations:

      Sufficient Training Memory: The model required a multi-year training continuum (2018–2022) to autonomously learn the structural breaks and complex non-stationarities introduced by the COVID-19 pandemic and strict NPIs.

      Epidemiological Gold Standard for Testing: Allocating exactly one year (2023) for testing is the epidemiological gold standard for seasonal infectious diseases. A full 365-day cycle ensures that model performance is evaluated across all seasonal phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Methodological Safeguards Against Information Leakage (Safe Assumptions):

      To ensure these partitions were safe for the model’s assumptions, we implemented strict safeguards:

      No Temporal Leakage: Tier 2 strictly follows Tier 1 in calendar time, preserving the sequential integrity assumed by LSTM architectures.

      No Distributional Leakage: All feature scaling and normalization parameters (means, standard deviations, min-max ranges) were derived exclusively from Tier 1 (Training) and applied unchanged to Tiers 2 and 3.

      - Explicit Acknowledgement of Limitations:

      We honestly acknowledge that while fixed chronological partitioning is the methodological standard for LSTM forecasting, a single fixed split point (Dec 31, 2022) fundamentally tests only one structural break. It does not exhaustively probe all potential future non-stationarities. Alternative strategies, such as expanding-window or rolling origin cross-validation, could offer additional robustness characterizations. We have integrated this crucial point into the Limitations section of the Discussion.

      We invite you to review the corresponding manuscript clarifications.

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Timeseries data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale.” (Methods, page 18)

      “Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reiance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      (9) The authors considered the DLNM analysis without considering the potential interactions among meteorological factors, which can't be avoided in real-world environmental contexts. I would suggest constructing such models by incorporating more reasonable interaction terms to reflect the complex relationships between these variables. Although the impact of COVID-19 was considered on the outcome of influenza directly.

      We sincerely thank you for raising this vital conceptual point. We completely agree that meteorological factors exhibit physically real interactions under real-world environmental conditions (e.g., the synergistic effect of extreme heat and high humidity on viral viability and aerosol dynamics). We welcome the opportunity to clarify how our analytical framework structurally addresses this precise complexity without compromising mathematical stability.

      Rather than forcing interaction terms into a single statistical model, we designed our dual-stage architecture specifically to create a functional division of labor between the DLNM and LSTM frameworks:

      - Methodological Constraints on Explicit DLNM Interaction Terms: While conceptually appealing, explicitly incorporating pairwise interaction terms across eight meteorological variables within the DLNM stage would generate 28 two-way interaction cross-bases, each with its own non-linear and lag-distributed spline structure. This parameter explosion leads to the “curse of dimensionality,” causing: (a) severe multicollinearity given the strong baseline correlations among weather variables; (b) profound instability of the cross-basis estimates and inflated standard errors; (c) a massive risk of overfitting; and (d) the complete loss of visual interpretability, which is the principal value proposition of DLNM. Consequently, as is standard practice in environmental epidemiology (Gasparrini et al., 2010), we restricted our DLNM stage to isolating interpretable, lag-distributed main-effect exposure–response surfaces.

      - The LSTM Stage as the Engine for Complex Interactions: This is precisely where the deep learning architecture provides its unique methodological value. The LSTM network does not require analysts to manually pre-specify rigid interaction terms. Instead, its multi-layer, non-linear gating architecture is intrinsically capable of autonomously extracting and representing arbitrary, high-dimensional interactions among all input variables simultaneously.

      Therefore, complex meteorological interactions are absolutely not ignored in our study; rather, they are absorbed into and resolved by the LSTM stage to maximize predictive accuracy, while the DLNM stage provides the lag-resolved, interpretable backbone for individual main effects. Our multi-stage variable integration ensures that the limitations of traditional statistical models do not artificially bottleneck the deep learning framework’s capacity to synthesize real-world complexities.

      We have now added explicit paragraphs to both the Methods and Discussion sections articulating this functional division of labor, ensuring maximum methodological transparency regarding how meteorological interactions are accommodated within our framework.

      We invite you to review the corresponding manuscript clarifications.

      “In the first stage, DLNMs characterize subtype-specific, non-linear, and lag-distributed associations between meteorological variables and influenza A and B positivity. In the second stage, a Bayesian-optimized LSTM network integrating meteorological, autoregressive, and socio-behavioral covariates is used for short-horizon forecasting, benchmarked against a covariate-matched multivariate ARIMA model and evaluated in an independent subtropical city (Sanming) as a preliminary test of model portability. Influenza positivity rate is used as the primary modeling endpoint to mitigate testing-related surveillance bias. Our aim is to provide an interpretable, climate-informed forecasting approach for subtropical influenza that can serve as a methodological foundation for future operational surveillance developments.” (Introduction, page 7-8)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the nonlinear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics.” (Methods, page 11)

      “Building on the significant non-linear and lagged effects through DLNM analysis of real world environmental exposures, this study further constructed multi-factor influenza A and B prediction LSTM networks. Leveraging a recurrent architecture with non-linear gating mechanisms, these networks are able to automatically capture and represent complex, high dimensional interactions among meteorological variables without the need for manual prespecification. This functional division of labor between the DLNM and LSTM models enhances predictive performance while preserving the interpretability and inferential stability established in the DLNM stage.” (Methods, page 18)

      “A central methodological feature of our framework is the deliberate division of labor between the DLNM and LSTM components. Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates. Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts (Zhu et al. 2022). Crucially, this LSTM stage explicitly incorporates weekly detection volumes, mask-wearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (10) In context with the above points, although COVID-19-related variables are included, important confounding factors such as population mobility, vaccination coverage, and school calendar (e.g., school openings/closings) are not adequately considered. Additionally, there is a potential risk of overfitting due to an imbalanced data split-too much data is allocated to the training set, while the validation set is relatively small on the other hand.

      We sincerely appreciate your comprehensive evaluation regarding confounding control and the risk of overfitting. These are highly pertinent methodological concerns, and we have implemented multiple refinements and empirical justifications to address each of them systematically.

      - Comprehensive Control of Confounding Factors

      To address the omitted confounders you rightfully identified, we have substantially expanded our covariate framework in the revised models:

      Population Mobility: We introduced the proportion of the migrant/non-local population as a new quantitative covariate (computed as the ratio of tested individuals reporting non-Putian residential addresses). This explicitly captures the extrinsic transmission pressure and viral importation risk exerted by mobile populations.

      Social/School Routines: To capture cyclical social contact patterns and surveillance reporting dynamics, we incorporated the Day-of-the-Week (DOW) indicator (distinguishing weekdays from weekends).

      Vaccination Coverage & School Calendar (Limitations): We honestly acknowledge that highly granular, municipal-level daily vaccination registry data and official macro-school holiday timelines were unavailable for integration into our daily time-series framework. However, for essential epidemiological context, the overall influenza vaccination coverage in mainland China during the 2018–2023 study window is historically estimated at a mere 2% to 3% of the general population, substantially lower than the 40–60% coverage typically observed in high-income temperate countries. Given this exceedingly low baseline, the population-level confounding contribution of vaccination on our predictive accuracy is expected to be minimal in absolute magnitude. We have now explicitly addressed this specific regional epidemiological context in the Discussion section.

      - Justification of the Data Split Ratio (88% vs. 12%)

      Regarding the perceived imbalance in the data partition, the chronological 88% / 12% split was not arbitrary; it was designed to balance two competing epidemiological necessities:

      Preserving Training Memory: The 88% training block (2018–2022) was strictly required for the LSTM to autonomously learn the multi-year seasonal cycles, the structural breaks induced by COVID-19 NPIs, and the suppressed-regime dynamics.

      The Epidemiological Gold Standard: The 12% testing block equates exactly to the 2023 calendar year (365 days). In seasonal infectious disease forecasting, evaluating performance across a complete, unbroken annual cycle is the epidemiological gold standard, ensuring the model is tested across all phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Safeguards Against Overfitting & Empirical Proof (Sensitivity Analysis)

      To definitively mitigate and disprove the risk of overfitting, we implemented a Chronological Three-Tier Validation Architecture:

      During the training phase, we randomly partitioned a 10% internal validation subset exclusively for hyperparameter tuning (via Bayesian Hyperopt) and implementing early stopping to terminate training the moment validation loss plateaued.

      We utilized the full 2023 surveillance data from an entirely distinct city (Sanming) as an external independent validation set. The fact that our model generalized excellently to Sanming is the strongest empirical proof against localized overfitting.

      Finally, to further substantiate model robustness, we conducted a controlled covariateablation sensitivity analysis (new Table 2). If a deep learning model is severely overfitted (i.e., memorizing noise), removing covariates often yields chaotic or random performance changes. However, when we explicitly removed the “mask-wearing stringency indices” or the “weekly detection volumes”, our model’s performance degraded in a predictable, biologically plausible manner (e.g., MAE increased by 33.3% to 100% across subtypes).

      This structurally proves that our LSTM architecture is not achieving artificially robust performance through overfitting, but exhibits appropriate sensitivity to the exact epidemiological features driving true viral transmission.

      We have extensively documented these justifications, safeguards, and limitations in the revised Methods, Results, and Discussion sections.

      We invite you to review the newly added limitation statement paragraph.

      “Thirdly, beyond the surveillance coverage limitation noted above, the aggregate-level nature of the available data further constrained our ability to adjust for individual-level confounders, including personal vaccination status, detailed comorbidities, healthcare seeking behavior, and socioeconomic status. This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNMLSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      (11) The authors should justify why the baseline model selection was made by comparing the LSTM model only with ARIMA? How the outcomes could be sensitive to other commonly used machine learning methods, such as Random Forest or XGBoost, etc, as a benchmark for their performance.

      We are deeply grateful for this methodologically incisive suggestion, which prompted us to substantially strengthen the manuscript’s benchmarking framework. Recognising the methodological importance of this recommendation, our team initiated the construction of an eXtreme Gradient Boosting (XGBoost) benchmark model in parallel with the Round 1 revision submission, and has now completed its full validation and interpretive analysis. XGBoost, one of the leading non-deep-learning machine-learning frameworks for structured predictive tasks, was implemented using the identical covariate matrix employed in the LSTM. The complete methodology, results, and interpretive analysis have been incorporated as Supplementary Additional File 6, with corresponding updates to the Methods (Study Design and Model Construction), Results, and Discussion sections of the main manuscript.

      Methodological design. The XGBoost model was implemented in R (xgboost package, version 3.2.1.1) and received a covariate set strictly matched to the LSTM: eight meteorological variables, weekly testing volumes, the mask-wearing stringency index, a day-of-week indicator, and the proportion of non-local (migrant) population. Critically, no additional lag features, sliding-window statistics, or autoregressive terms were introduced, ensuring that XGBoost’s information set was identical to the contemporaneous covariate stream consumed by the LSTM at each time step. This design deliberately isolates the predictive contribution of architectural inductive bias, the LSTM’s intrinsic sequential memory versus XGBoost’s memory-free tree-partitioning geometry, from any confounding due to hand-engineered temporal features. The training (2018 – 2022) and validation (2023) partitions and the evaluation metrics (MAE, RMSE, MAPE, SMAPE) followed the LSTM and ARIMA protocols precisely.

      Empirical findings.

      A complete performance matrix including MAE, RMSE, MAPE, and SMAPE for all three models is presented in Supplementary Table S3 (Additional File 6), which confirms the same rank-ordering (LSTM ≪ ARIMA ≈ XGBoost) across all four error metrics.

      Despite receiving identical inputs, XGBoost yielded errors approximately 15-fold higher than the LSTM for Influenza A and 35-fold higher for Influenza B, with errors broadly comparable in magnitude to the ARIMA baseline. Visual inspection of the 2023 forecasting trajectories (Supplementary Figure S6) revealed that XGBoost systematically under-predicted the explosive post-NPI Influenza A rebound in late 2023 and generated largely flat forecasts for Influenza B that failed to reproduce the mid-year peak.

      Interpretation. These findings are mechanistically informative and reinforce, rather than merely confirm, the manuscript’s central methodological claim. Four architectural properties of XGBoost account for the observed performance gap:

      Insensitivity to temporal ordering. Decision-tree splits operate on feature-dimensional thresholds and cannot structurally encode “the present depends on the past” in the manner natively accommodated by recurrent architectures.

      Absence of gated recurrence. XGBoost lacks any mechanism analogous to the LSTM’s forget-input-output gating for propagating, filtering, and dynamically re-weighting historical states, and therefore cannot represent long-range non-linear temporal dependencies.

      Violated independence assumption. Tree ensembles implicitly assume independent and identically distributed (i.i.d.) observations, an assumption fundamentally at odds with the strong serial autocorrelation of epidemiological time series.

      Non-comparable modelling paradigms. Both ARIMA and LSTM are explicit sequential models within a common paradigm, whereas XGBoost belongs to a fundamentally non-sequential family. The LSTM–ARIMA contrast therefore isolates the specific contribution of deep sequential learning within a paradigmatically comparable framework, while the LSTM–XGBoost contrast demonstrates that even a state-of-the-art memory-free non-linear learner cannot substitute for a genuinely sequential architecture under epidemiologically non-stationary conditions.

      The consistent superiority of the LSTM over both a linear statistical benchmark (ARIMA) and a non-linear memory-free machine-learning benchmark (XGBoost), all receiving identical inputs and evaluated on the identical validation window, isolates the recurrent gated architecture itself as the source of the predictive gain. We are indebted to the Reviewer for prompting this three-way benchmark analysis, which has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      This additional benchmark analysis, though completed subsequent to the Round 1 submission, has now been fully integrated into the Round 2 revision at the appropriate positions across the Abstract, Importance Statement, Methods, Results, Discussion, and Conclusion, together with the new Additional File 6. We are grateful to you for prompting this three-way benchmark, which we believe has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      We invite you to review the newly added Discussion paragraphs.

      “The optimized LSTM algorithm demonstrated strong predictive performance, with loss curves for both subtype-specific networks exhibiting favourable convergence, consistent with prior work (Du et al. 2023). To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: a multivariate ARIMA model and an XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 40-42)

      “Fourthly, our benchmarking strategy was deliberately structured as a paradigmatically layered three-way comparison (linear-sequential ARIMA, non-linear non-sequential XGBoost, non-linear sequential LSTM) rather than an exhaustive algorithmic survey. This design prioritized methodological clarity, isolating the contribution of recurrent sequential inductive bias, over horizontal coverage. Nonetheless, our evaluation does not extend to Transformer-based attention architectures, nor to hybrid ensemble strategies (e.g., LSTM–XGBoost stacking or multi-model Bayesian model averaging) that may offer complementary strengths. Systematic benchmarking against these emerging architectures, alongside prospective recalibration on additional subtropical surveillance streams and operational stress-testing under real-time data latency, will be essential next steps as the framework evolves toward operational deployment.”(Discussion, page 46)

      (12) I was expecting the statement of generalizability in terms of locations with the link to the results of the study. I suggest including such statements in the text along with limitations.

      We sincerely thank you for this highly constructive suggestion. We fully agree that an explicit, results-linked generalizability statement is crucial for appropriately interpreting and safely deploying our forecasting framework. To address this, we have substantially expanded the Limitations subsection within the Discussion to articulate a rigorous, three tier generalizability framework directly linked to our study findings:

      - Empirically Validated Generalization: The successful external validation of our framework in Sanming, a geographically distinct, inland prefecture-level city, constitutes direct empirical evidence of cross-location generalizability within the subtropical south eastern Chinese context. As detailed in our results, the framework retained high operational forecasting accuracy (e.g., Influenza A MAE of 0.015, SMAPE of 0.610) without any retraining on Sanming’s local data, proving that the meteorological exposure–response patterns and LSTM architecture are transferable across similar subtropical climates under analogous public health intervention frameworks.

      - Plausible Inferential Generalization: Beyond the directly validated Putian–Sanming pair, the framework’s architecture is plausibly generalizable to other subtropical cities across adjacent regions (e.g., Guangdong, eastern Guangxi) that share comparable monsoon climates, year-round multi-peak influenza circulation patterns, and similar sentinel surveillance infrastructures.

      - Boundaries of Transferability (Non-generalizability): We explicitly state the boundaries where our specific model parameters should not be directly extrapolated without rigorous local recalibration. These include: (a) regions located in subtropical latitudes but characterized by fundamentally distinct climate regimes (e.g., monsoonal systems with characteristics that are markedly incomparable), where the temperature–humidity coupling structures and seasonality with multiple winter peaks differ qualitatively from those in the study setting; and (b) regions with substantially different public-health intervention landscapes, such as settings with high population-level vaccination coverage or school-closure-based mitigation strategies.

      By explicitly defining these boundaries in the revised Discussion, we ensure transparent communication of the model’s appropriate application scope and methodological limitations.

      We invite you to review the newly added limitation statement paragraph.

      “Furthermore, the Bayesian-optimized LSTM architecture was directly applied, without re-training or hyperparameter adjustment, to both the Putian internal test set and the Sanming external validation set. This unified modeling framework ensures that performance differences between the two evaluation contexts primarily reflect geographic transferability rather than model re-optimization.” (Methods, page 21)

      Finally, it is imperative to explicitly delineate the generalizability of our forecasting framework within geographic and epidemiological contexts. The successful external validation in Sanming provides direct empirical evidence that our LSTM architecture and the identified meteorological thresholds are robustly transferable across the subtropical southeastern Chinese context, sharing comparable monsoon climates, year-round influenza circulation, and standardized public health intervention frameworks. However, we explicitly define the boundaries of this transferability. The specific meteorological coefficients, lag structure, and predictive parameters derived in this study should not be indiscriminately extrapolated to other regions, even within subtropical latitudes, where climatic regimes (e.g., monsoon regimes with markedly non-comparable characteristics) or public health contexts (e.g., higher baseline influenza vaccination coverage or distinct non-pharmaceutical intervention strategies) differ substantially. In such disparate settings, directly applying our pretrained model may yield substantial systematic biases. While the underlying modeling framework remains methodologically transferable, its parameterization requires rigorous local recalibration using region-specific surveillance data. Acknowledging these boundaries ensures that the framework can be deployed more safely and appropriately to realize targeted, climate-sensitive infectious disease forecasts.” (Discussion, page 44-47)”

      (13) The flow of the manuscript should be revised for general readers, for example author spent a lot of text on limitations in general without linking them to the outcomes and study design. The English language has a huge scope for improvement. I suggest paying attention to the presentation of the text for the English language use and continuity.

      We sincerely appreciate your candid feedback on the manuscript’s readability. We recognize that in our previous draft, integrating diverse critiques from multiple rounds of peer review inadvertently resulted in disjointed narrative flows, particularly in the Methods and Limitations sections.

      To address this, we have undertaken a massive structural overhaul and linguistic refinement of the entire manuscript:

      - Logical Flow: The Introduction has been sharply refocused; the Methods section has been streamlined (e.g., moving lengthy PCR protocols to the Supplement); and the Results have been reorganized with clear subheadings.

      - Contextualized Limitations: We completely rewrote the Discussion section. Rather than listing generic limitations, we have deeply anchored them to our specific study design and outcomes. For instance, we now explicitly discuss how our shift to predicting the “Positivity Rate” directly mitigates the “Surveillance Bias” limitation, and how the 2023 “rebound” limits steady-state generalizations.

      - Language Polish: The manuscript has undergone comprehensive editing by a native English-speaking academic expert to ensure grammatical precision, sophisticated vocabulary, and seamless continuity.

      We believe these revisions have dramatically elevated the clarity and scholarly tone of the text.

    1. Author response:

      Response to Reviewer 1:

      We thank the reviewer for their valuable and constructive suggestions.

      (1) The categorization of morphological traits was performed by researchers who are familiar with the taxonomy and morphology of this taxa. We acknowledge that this may introduce some degree of subjectivity, and we will provide photographs of other morphological traits for each species in the Supplementary data.

      (2) In our original analysis, we treated the PCAmix values derived from multiple discrete traits as continuous variables and used the lm () function to test the correlation between these values and net diversification fates (lines 352-357). We agree that a phylogenetic comparative approach would be more appropriate. We plan to re-analyze the data using either glm () or phylogenetic generalized least squares (PGLS) to properly account for phylogenetic relationships. In lines 331-333, we would clarify that this part refers to the HiSSE analysis based on discrete traits, which is independent of the linear regression analysis mentioned above. Nevertheless, we will ensure that both analyses are clearly distinguished and properly described in the revised Methods section. We will update the relevant sections accordingly.

      (3) We will carefully re-examine the entire manuscript and add appropriate measures of uncertainty.

      (4) We agree with the reviewer that two hypotheses are not mutually exclusive. In the revised manuscript, we will rephrase this paragraph to present both possibilities more neutrally.

      (5) We have observed mating behaviors in this group and found that ASE function occurs after the male has successfully grasped the female using its legs. However, we acknowledge that our study did not include direct experiments to quantitatively test the effect of ASE complexity on mating success. Therefore, our discussion in this section is indeed somewhat speculative.

      Response to Reviewer 2:

      We thank the reviewer for their encouraging and constructive comments. We agree that our study lacks direct experimental evidence to explicitly demonstrate the grasping and anti-grasping functions of the various male and female traits in Pseudovelia, and that our interpretations currently rely on comparisons with functional studies in more distantly related taxa. In the revised manuscript, we will explicitly state this limitation and refer to these traits as “putative” grasping or anti-grasping traits throughout the text where appropriate.

      We will also carefully address all minor editorial suggestions, including clarifying terminology, correcting typos, and improving figure legends.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Ramirez Carbo et al. use the powerful M. xanthus spore morphogenesis model to address fundamental mechanisms in coordinated peptidoglycan remodeling and degradation. As peptidoglycan is an essential macromolecule and difficult to study in vivo, the authors use indirect but important methodology. The authors first identify two lytic transglycosylase (Ltg) enzymes necessary for spore morphogenesis using mutant phenotypic studies. They characterize these mutants for their role in coordinating spore morphogenesis induced either in fruiting bodies (starvation-dependent) or in liquid-rich media conditions (chemical-dependent). They conclude from these phenotypic and epistatic analyses that LtgA is necessary for morphogenesis during chemical-induced sporulation, and LtgB appears to be necessary to coordinate LtgA activity by interfering with LtgA function. Under starvation-induced sporulation, the absence of LtgB interferes with the building of fruiting bodies. LtgA does not appear to play a primary role in promoting aggregation into fruiting bodies, nor in degradation of peptidoglycan as assayed by loss of signal in anti-PG immunofluorescence. The authors demonstrate that the purified periplasmic domain of LtgA is highly active in degrading purified PG sacculi in vitro, while that of LtgB is highly reduced (relative to LtgA or lysozyme). The authors use photoactivated mCherry Lyt fusions and PALM to track the fusion protein mobility, which they state correlates with activity as immobilization results from PG binding. They demonstrate that in vegetative cells, a greater proportion of LtgA-PAmCh is more immobile (more active) than LtgB-PAmCh, but that directly after chemical-induction of sporulation, LtgB-PAmCh becomes more immobile (active). These analyses in the partner mutant backgrounds suggest that LtgA-PAmCh is more immobile (less active) in the absence of LtgB, but the reverse is not observed. Finally, the authors demonstrate that overexpression of LtgA in vegetative conditions leads to cell rounding, likely because of uncontrolled PG degradation, while overexpression of LtgB displays no phenotype.

      Strengths:

      This paper capitalizes on a novel spore morphogenesis mechanism to define proteins and mechanisms involved in peptidoglycan reorganization. The authors use the powerful PALM microscopy technique to assess Ltg activity in vivo by assaying for immobility as a proxy for PG binding. The authors elucidate a novel mechanism by which two Ltg's function together- with one (LtgB) seeming to regulate the activity of the other (the primary Ltg).

      Despite some weaknesses, there is no question that this study provides important insight into mechanisms of peptidoglycan remodeling- a difficult but highly impactful area of study with implications for the development of novel therapeutics and the discovery of mechanisms of fundamental bacterial physiology.

      Weaknesses:

      In many places, the authors do not adequately justify interpretations of their assays, leading to some apparently unjustified conclusions. Many of these are minor and may just require citations to demonstrate that the interpretations are justified by previous studies (detailed in recommendations below), but two bigger concerns are as follows:

      (1) It is not clear how the muropeptides listed in Figure 1 were assigned, and it is missing in the methods. In the sporulating conditions, the spectra look like combinations of multiple peaks, and the data, as stated, is not convincing to the non-specialist eye.

      We thank the reviewer for raising this point. We've expanded the Methods section to give a fuller account of how muropeptides were identified. In particular, we now describe the chromatographic separation and the assignment process, which relies on comparison with published data, fragmentation patterns, retention times, and accurate mass values. We acknowledge that the chromatograms show several peaks; nevertheless, only those muropeptides that we could clearly identify by MS/MS analysis were annotated in Figure 1. This clarification is now made explicit in the figure legend in the revised version of the manuscript.

      (2) The observation that the lytB mutant prevents appropriate aggregation into fruiting bodies does not allow the interpretation that the absence of LtgB prevents PG morphogenesis in the starvation-induced sporulation pathway, per se. It is more likely that in the LtgB mutant, the morphogenesis program is not even triggered. This is because signaling proteins and regulators (specifically, C-signal accumulation/activated FruA), which are dependent on increased cell-cell signaling in the fruiting body, do not accumulate appropriately in shallow aggregates. C-signal/FruA are necessary to trigger the sporulation program in FBs. BTW: A hypothesis to explain the indirect effect of ltgB absence on aggregation could be that UDP-precursors are not regulated appropriately (unregulated LtyA (LtgA [sic])??), so polysaccharides necessary for motility are not properly produced.

      Along these lines, fruiting body formation does not equal sporulation, and even "darkened" fruiting bodies can be misleading, as some mutants form polysacchariderich fruiting bodies (that appear dark under certain light conditions in the stereomicroscope) but do not sporulate efficiently. The wording in the text suggests that the authors assume that sporulation levels are normal because fruiting bodies are produced (see specific comments for details).

      We deeply appreciate this question. Seeking the answer, we repeated the fruiting body assay and found that both the ΔltgA and ΔltgB mutants formed dark aggregates that were comparable to wild-type fruiting bodies. However, these “fruiting body-like” aggregates did not contain sonication-resistant spores. Thus, regardless of the signals, either glycerol or starvation, sporulation requires both LtgA and LtgB. We have corrected the mistakes in the first submission.

      (3) The authors repeatedly state that production of spore coat polysaccharides likely affects the PG IP staining (see below), but this is not well justified. A citation is needed if this has already been directly shown, or the language needs to be softened.

      We agree with the reviewer. We have softened our language as “However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.”

      (4) Better justification for the immobility of Ltg proteins in vivo as an assay for activity may be required. If this is well known in the field, it should be explicitly stated. The authors address this better in the discussion - but still state it is a correlation.

      We elaborated the justification, “Thus, when diffusive enzymes bind to PG, their mobility decreases (Lee et al., 2016; Zhang et al., 2023). For instance, DacB, another PG hydrolase, reduces its single-particle mobility in the conditions where its activity is activated (Zhang et al., 2023). By tracking single fluorescently-labeled enzyme particles, we can approximate their PG-binding in different physiological conditions and genetic backgrounds (Ramirez Carbo et al., 2024, Zhang et al., 2023, Ramírez Carbó & Nan, 2026).”

      We further discussed the correlation in discussion, “The simultaneous occurrence of reduced LtgA mobility and PG degradation during glycerol-induced sporulation indicates that the molecular dynamics of LtgA accurately mirrors its enzymatic activity. Such correlation between decreased particle mobility and increased enzymatic activity applies to many other PG-related enzymes, including multiple PG polymerases in E. coli and the endopeptidase DacB in M. xanthus (Lee et al., 2016, Zhang et al., 2023, Yang et al., 2021).”

      Reviewer #2 (Public review):

      Summary:

      The authors' initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LGT products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single-molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth, sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted from the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. How these small fragments were thus detected is unclear, and this suggests a more complex story concerning PG metabolism during sporulation. An anti-PG antibody is used to quantify PG in the spores, but it is not made clear what the specificity of this antibody is, and thus whether it would recognize the LTG -altered PG of the spore. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation".

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Details on places in the text that could be improved:

      (1) Line 77: more appropriate "homologs of the sporulation genes"

      Corrected.

      (2) Line 137-139. Where are the 14 KO mutant data? Only showing the ones with phenotypes? Needs a citation if published elsewhere.

      The phenotypes of other mutants are shown in the new Fig. S1.

      (3) Line 144 and later: why "ORF"? Are they not homologous to Ltg's??. Also, the orf designation is not in the figure, so it is hard to follow the data the authors are presenting.

      Following the reviewer’s recommendation, we deleted “ORF” and presented the ORF designation in the figure legend.

      (4) Figure 2B - add the y-axis legend to the figure (also S1B).

      Added

      (5) Line 156: What does "same below" mean?

      Deleted “same below”.

      (6) Line 161: emtA is not indicated on Figure 2D - can you reword or add to the legend??

      Changed to “MltE, an LTG encoded by Escherichia coli emtA”

      (7) Line 163: "opposite roles" could be a bit better defined... do you mean ltgA induces rounding and ltgB delays rounding???

      We agree with the reviewer. “opposite” was replaced by “different”.

      (8) Line 172: But did the ltgA mutant make the same number of viable spores (not just FBs)? Can't conclude "the slow sporulation pathway only requires ltgB" if ltgA mutant was not tested.

      We agree with the reviewer. In the revised manuscript, we quantified the starvation-induced spores and found that both LtgA and LtgB are required for forming mature spores. The data are shown in Figure 3 and supplement figures.

      (9) Line 202: might be appropriate to indicate the degree of homology (% identity over protein length).

      Added.

      (10) Line 178, 181, and thereafter: DIC microscopy.

      Corrected.

      (11) Line 182/183 and 190: What is the basis for the conclusion that "polysaccharides sustained unflattened structures"??

      We changed the description to “likely due to the deposition of spore coat polysaccharides that sustained unflattened cell structures (Wartel et al., 2013; Holkenbrink et al., 2014).”

      (12) Line 185/186: What is the basis for the conclusion that ltgB only contained a small amount of PG? If it is based on reduced intensity, it is not obvious from the images presented. Quantitative analysis would better support this statement.

      We changed the description to “Sacculi from the ∆ltgB pseudospores still contained PG but showed lower fluorescence intensity”. We also updated the figure to show the typical fluorescence intensities.

      (13) Figure legend 3 line 245/6: It is not totally clear what the difference is, in that the white arrows are pointing to in ltgA vs ltgB mutant. Is it the proportion of large spherical objects in ltgB? What is the significance of the puncta in A vs. B?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).” In the figure legend, we changed the description to White arrows point to the sacculi in irregular shapes.”

      (14) Line 190-192. Again, not really seeing what the authors are identifying to conclude "irregular shapes and ruptures" could authors pin-point more specifically and quantify such structures?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).”

      (15) Line 194: if the authors want to make such a strong conclusion, they really should demonstrate this. Anti-spore coat antibodies are available in the field. Or take out these statements.

      We took this statement out.

      (16) Line 210: either they lack PG, OR the antibody can't gain access? How to conclude both? Might be more accurate to say "although we can't rule out that the polysaccharide spore coat prevents access to PG"

      Following the reviewer’s recommendation, we changed the description to “These fruiting body spores lacked PG-specific fluorescence (Fig. 3A), consistent with their markedly reduced PG content (Fig. 1). However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.

      (17) Line 217-19: This is often stated in the literature, but full glycerol-induced spore maturation requires much more than 2 hours (note the authors are using O/N glycerol induction in Figure 1). And the long starvation-induced sporulation process is likely due to the differential start of sporulation, because the cells don't all enter the aggregate at the same time.... so they are not triggered to induce sporulation at the same time.

      We agree with the reviewer on the first statement and changed “two hours” to “four hours”. For the second statement, we do not completely agree. Much research showed that spore development in fruiting bodies is well synchronized, for example, in (Dworkin & Voelz, 1962), starvation-induced spores did not at 48 h.

      (18) Line 222: What is the evidence that it is really variation in production from the van promoter? Could it also just be due to differences in the cell cycle in the population? The authors may be correct, but should be less definitive about those conclusions since they haven't measured LtgA levels directly.

      We agree with the reviewer. We changed the description to, “Upon induction with 200 μM vanillate, cells exhibited heterogeneous morphology, likely resulting from variations in LtgA expression or differences in cell cycle stages within the population.” Following this suggestion, we also changed the description on murA overexpression, “Similar to the cells that overexpressed LtgA (Fig. 3B), this heterogeneity likely reflects variable murA induction or the unsynchronized growth stages within the population.”

      (19) Line 224: Where is the over-expressed LtgA data in Figure 2A? Is it that the authors are referring to over-expressed murA data, and the point is that overexpression of this kind of gene can lead to veg cell rounding? If the latter, this should be specifically stated in the text.

      We apologize for this mistake. The data were shown in Fig. 3B, rather than Fig. 2A, 3B.

      (20) Line 224: Did the authors test that LtgB is stably overproduced in veg cells? It could be that LtgB is turned over while LtgA is not. From the purified protein blot in Figure 3C, it does look like LtgB contains a degradation product (which is perhaps inhibiting the LtgB activity).

      (21) Line 229: AgmT? Do the authors mean Ltg?

      Corrected.

      (22) Line 238: Is "rate" the right word? Technically, kinetics haven't been measured. Could state "LtgB less active" or "less efficient"?

      Changed.

      (23) Line 255: "ltgB shows a slight increase in expression during slow sporulation". Do the authors mean over the entire dev time course or specifically during the sporulation phase (which will be different timing for DZ2 vs DK1622)?

      We clarified the description as “Consistent with its role in PG degradation, in a microarray-based transcriptome analysis, ltgA transcription was found to increase about twofold during rapid sporulation (4 h) but remain unchanged during slow sporulation (96 h). in the closely related DK1622 strain (Muller et al., 2010). Conversely, ltgB expression gradually rises during slow sporulation, reaching 1.8 times the vegetative level at 96 h, while remaining stable during rapid sporulation (Muller et al., 2010, Munoz-Dorado et al., 2019).”

      (24) Line 278: This needs to be corrected. The authors have shown that production of fruiting bodies is not affected by the fusions. (also in Fig. S1 legend). This is really not the same as sporulation efficiency.

      We changed the description to “the PAmCherry tags did not affect the formation of either glycerol-induced spores or starvation-induced fruiting bodies”.

      (25) Figure S1A: panel A: Is there a deg product partially cut off at the bottom of the gel? (or is this a non-specific cross-reactive band). What is the predicted molecular mass for both proteins with the PAmCherry fusion? What conditions were these lysates generated from: veg prior to glycerol induction? I understand there are probably no antibodies available to LtgA or B, but it is important to note that it is not possible to know if there is simultaneously wt LtgA or B produced (by cleavage and degradation of the mCh fusion). Panel B: Were the differences in l/w between wt and the fusion strains tested to see if there really were no significant differences? Please state in the text (it looks like the data variance is higher in the fusion strains relative to the wt at 1 hr).

      The bands at the bottom of the gel are the running front that appear in both lanes. They are not mCherry because there estimated molecular weight is much lower than that of mCherry (26.3 kDa). To avoid confusion, we cut these bands from the figure. The expression of both fusion proteins was detected from vegetative cells, which was clarified in the legend. The predicted molecular weights of them were provided in the legend too.

      (26) Figure S1C: The ability to make fb is not the same as the production of spores. The authors should test the number of viable spores produced under starvation conditions, if they want to state starvation-induced sporulation is not affected by the fusions.

      We agree with the reviewer. We quantified starvation-induced sporulation in the revised manuscript.

      (27) Line 399: Figure S2 looks at fb formation in the absence of vanillate, not sporulation.

      We changed the description to “cells grown without vanillate progressed normally through glycerol-induced sporulation and starvation-induced fruiting body formation”.

      We also quantified starvation-induced sporulation in the revised manuscript.

      (28) Line 281: How many fold is the ltgA transcript reduced compared to the ltgB? (i.e., If it is 1.2 fold reduced that may not be as worth mentioning as if it was 5-10 fold reduced).

      ltgA transcription in vegetative cells was detected in a microarray (Muller et al., 2010) but not reported in RNAseq (Munoz-Dorado et al., 2019), significantly different from that of ltgB. We pointed this out in the revised manuscript.

      (29) Figure 4B legend line 358. Define D (should it be italicized?).

      Corrected.

      (30) Line 363: Please define how significance was calculated. Is this the p-value?

      We deleted the word “significant”.

      (31) Line 298: How do the authors know immobility is from binding to PG? Provide a reference if this is well-known.

      The rationale and references have been mentioned at the beginning of this section.

      (32) Line 302 and thereafter: suggest "2.62 × 10-2 (plus minus) 2.0 × 10-3 μm2/s" is presented as "2.62 (plus minus) 0.20 × 10-2 μm2/s" for easier reading; switch to past tense (Were not are).

      Changed following the reviewer’s recommendation.

      (33) Line 310: "suggest" not "indicate", because binding of PG was directly tested.

      Corrected.

      (34) Line 341: "confirming it restricts access" seems very strong wording. Suggest: may compete with.

      Changed.

      (35) Line 392: SOME fb are larger- many are significantly smaller.

      Because fruiting body sizes do not reflect sporulation efficiency, we removed this description.

      (36) Figure S2 legend. Leaky expression; or on fruiting body formation.

      Corrected.

      (37) Line 436: stationary phase cells decrease PG synthesis- does this increase PG degradation? Perhaps it does lead to the death phase, which is striking in M. xanthus....

      How do cells die, either through death phase or under antibiotic stresses, is not well understood (Baquero & Levin, 2021) Very likely, cell death is due to the accumulation of oxidative damages (Kohanski et al., 2007) and cell lysis could be a byproduct of cell death, when cells lose control of the enzymes that break PG. While cell death is a great topic to investigate, it is beyond the scope of this study.

      (38) Line 507: washed.

      Corrected.

      (39) Line 524 (514 [sic]) and thereafter: sacculi.

      Corrected.

      (40) Line 587: reference for cell lysis procedure?

      The procedures of cell lysis, column loading and elusion were described in details “…cells were harvested by centrifugation at 6,000 × g for 20 min and lysed by sonication in buffer A (20 mM Tris-HCl pH 8.0, 200 mM NaCl), (Nan et al., 2010, Nan et al., 2006). Proteins were loaded to an NGC™ Chromatography System (BIO-RAD) and 5-ml HisTrap™ columns (Cytiva) and eluted by buffer B (20 mM Tris-HCl pH 8.0, 200 mM NaCl, 500 mM immidazole) (Pogue et al., 2018, Nan et al., 2010).”

      (41) Line 269: by microscopy (or "under THE microscope").

      Corrected.

      (42) Line 271: THE cell/PG.

      “The” added.

      (43) Line 603: OR not and.

      Corrected.

      (44) Line 497 (and elsewhere): Is it really CFU? If determined by OD, then not technically CFU because some cells will not grow into colonies.

      We agree with the reviewer. Cell concentrations were determined by OD. “CFU” was deleted.

      (45) Line 506: washed.

      Same as recommendation (38). Corrected.

      (46) Line 521: min.

      Corrected.

      (47) Line 532: How were peaks assigned?

      We described peak assignment in details in the revised manuscript, “Muropeptides were assigned based on: (i) accurate mass matching to theoretical monoisotopic masses of expected M. xanthus PG building blocks (Bui et al., 2009, White et al., 1968) and (ii) comparison of retention times with those reported in previous analysis with similar PG compositions.”

      Reviewer #2 (Recommendations for the authors):

      (1) If almost all the muropeptides detected in spores are anhydro products of LTGs, then it might be expected that these are all very small peptidoglycan fragments in the spores. If the anhydro units were at the ends of short PG chains, then muramidase digestion would release similar amounts of non-anhydro products, but none are detected. So, is muramidase digestion doing anything to the PG derived from the spores? Is muramidase digestion required to observe the spore muropeptide pattern? A control sample in which muramidase digestion is omitted would answer these questions.

      Muramidase digestion is essential for solubilizing PG into different muropeptides for UPLC analysis. In each sample, both anhydro and non-anhydro products were detected. In our original submission, we pointed out that “The two spore types showed similar profiles of a discernible presence of muropeptides that resembled those found in vegetative cells, albeit in significantly reduced quantities (Fig. 1).” We clarified the PG analysis. “The purification procedure yields only sedimentable PG, as all soluble fragments are removed during the washing steps. The resulting sacculi were then digested with muramidase, and the solubilized muropeptides were analyzed by UPLC (see Materials and Methods). The chromatograms in Fig. 1 reflect the muropeptides released specifically from the sedimented sacculus fraction.” Because muramidase release polysaccharides or disaccharides, so it’s digestion does not release nonanhydro GlcNAc species, which is why we did not see equal amounts of anhydro and non-anhydro products. In our case, over 90% of the polysaccharides and disaccharides contain Anhydro-MurNAc. The dominance of anhydro products indicates that LTGs cut very frequently on glycan chains.

      (2) This also raises the question of how these very small peptidoglycan fragments are even retained in the spores. They would be expected to be lost during spore purification or during PG purification prior to muramidase digestion. How do you even purify sacculi when the spores have no PG chains? One could theorize that the polysaccharide coats hold everything in, but then the PG would be protected from muramidase digestion.

      We thank the reviewer for this question. Our analysis suggests that M. xanthus spores do not lack PG entirely but retain a residual PG mesh that is still crosslinked. Such crosslinked material sediments during PG purification and remains accessible to muramidase, which hydrolyses internal glycosidic bonds and releases anhydro muropeptides. Non-crosslinked fragments would indeed be washed away during purification steps, so the detected muropeptides reflect the structure of this residual sacculus (Fig. 1).

      (3) What is the anti-PG antibody recognizing, the glycan backbone, the peptide side chain, or both? Can the antibody recognize the very short anhydro-containing disaccharides proposed to be predominant in spores? If not, then the PG quantification using the antibody is not accurate.

      The structures recognized by the anti-PG antibodies are unknown. Our samples do not contain small degradation products, which was clarified in the revised manuscript, “To answer this question, we purified cell sacculi and used immunofluorescence and an anti-PG serum (de Pedro et al., 1997) to visualize the remaining PG.” We used immunofluorescence to display the PG scaffolds remained in each sample and we did not perform any quantitative analysis based on the images.

      (4) Lines 325-327: This is a speculative conclusion and should be stated as such, i.e., "may control the pace.." This conclusion could be somewhat more strongly stated at the end of the next section, around lines 345-350.

      Moved following the reviewer’s recommendation.

      (5) Lines 403-404: This first sentence of the discussion seems completely dissociated from the topic of the paper; it should be deleted.

      Deleted.

      (6) Lines 404-405. I am not convinced that these findings "elucidate a new mechanism of sporulation." The study was undertaken because PG degradation was already tied to sporulation in a previous study. Furthermore, I am not sure that PG degradation is a "mechanism of sporulation." It is clearly an important step in sporulation of this species, but is it the driving "mechanism"?

      We tuned down our statement as “Our findings demonstrate that M. xanthus, a nonfirmicute bacterium, relies on PG degradation to change cell shape during sporulation”.

      (7) Lines 510-516 describe PG purification from vegetative cells. How was this process modified for spores?

      We apologize for the confusion. The process was clarified as “For PG analysis, samples were processed as previously described for Gram-negative bacteria (Alvarez et al., 2016; Desmarais et al., 2013). Vegetative cells were harvested at mid-stationary phase by centrifugation (30 min, 8,000 g). Vegetative cells and purified spores (as described in the previous section) were resuspended…”.

      Alvarez, L., Hernandez, S.B., de Pedro, M.A., and Cava, F. (2016) Ultra-Sensitive, High-Resolution Liquid Chromatography Methods for the High-Throughput Quantitative Analysis of Bacterial Cell Wall Chemistry and Structure. Methods Mol Biol 1440: 1127.

      Baquero, F., and Levin, B.R. (2021) Proximate and ultimate causes of the bactericidal action of antibiotics. Nat Rev Microbiol 19: 123-132.

      Bui, N.K., Gray, J., Schwarz, H., Schumann, P., Blanot, D., and Vollmer, W. (2009) The peptidoglycan sacculus of Myxococcus xanthus has unusual structural features and is degraded during glycerol-induced myxospore development. J Bacteriol 191: 494505.

      de Pedro, M.A., Quintela, J.C., Holtje, J.V., and Schwarz, H. (1997) Murein segregation in Escherichia coli. J Bacteriol 179: 2823-2834.

      Desmarais, S.M., De Pedro, M.A., Cava, F., and Huang, K.C. (2013) Peptidoglycan at its peaks: how chromatographic analyses can reveal bacterial cell wall structure and assembly. Mol Microbiol 89: 1-13.

      Dworkin, M., and Voelz, H. (1962) The formation and germination of microcysts in Myxococcus xanthus. J Gen Microbiol 28: 81-85.

      Holkenbrink, C., Hoiczyk, E., Kahnt, J., and Higgs, P.I. (2014) Synthesis and assembly of a novel glycan layer in Myxococcus xanthus spores. J Biol Chem 289: 32364-32378.

      Kohanski, M.A., Dwyer, D.J., Hayete, B., Lawrence, C.A., and Collins, J.J. (2007) A common mechanism of cellular death induced by bactericidal antibiotics. Cell 130: 797-810.

      Lee, T.K., Meng, K., Shi, H., and Huang, K.C. (2016) Single-molecule imaging reveals modulation of cell wall synthesis dynamics in live bacterial cells. Nature communications 7: 13170.

      Muller, F.D., Treuner-Lange, A., Heider, J., Huntley, S.M., and Higgs, P.I. (2010) Global transcriptome analysis of spore formation in Myxococcus xanthus reveals a locus necessary for cell diberentiation. BMC Genomics 11: 264.

      Munoz-Dorado, J., Moraleda-Munoz, A., Marcos-Torres, F.J., Contreras-Moreno, F.J., MartinCuadrado, A.B., Schrader, J.M., Higgs, P.I., and Perez, J. (2019) Transcriptome dynamics of the Myxococcus xanthus multicellular developmental program. Elife 8.

      Nan, B., Liu, X., Zhou, Y., Liu, J., Zhang, L., Wen, J., Zhang, X., Su, X.D., and Wang, Y.P. (2010) From signal perception to signal transduction: ligand-induced dimeric switch of DctB sensory domain in solution. Mol Microbiol 75: 1484-1494.

      Nan, B., Zhou, Y., Liang, Y.H., Wen, J., Ma, Q., Zhang, S., Wang, Y., and Su, X.D. (2006) Purification and preliminary X-ray crystallographic analysis of the ligand-binding domain of Sinorhizobium meliloti DctB. Biochim Biophys Acta 1764: 839-841.

      Pogue, C.B., Zhou, T., and Nan, B. (2018) PlpA, a PilZ-like protein, regulates directed motility of the bacterium Myxococcus xanthus. Mol Microbiol 107: 214-228.

      Ramirez Carbo, C.A., Faromiki, O.G., and Nan, B. (2024) A lytic transglycosylase connects bacterial focal adhesion complexes to the peptidoglycan cell wall. Elife 13.

      Ramírez Carbó, C.A., and Nan, B. (2026) Using Single-Particle Fluorescence Microscopy to Quantify Substrate Binding of Peptidoglycan-Modification Enzymes. Bio-protocol 16: e5696.

      Voelz, H., and Dworkin, M. (1962) Fine structure of Myxococcus xanthus during morphogenesis. J Bacteriol 84: 943-952.

      Wartel, M., Ducret, A., Thutupalli, S., Czerwinski, F., Le Gall, A.V., Mauriello, E.M., Bergam, P., Brun, Y.V., Shaevitz, J., and Mignot, T. (2013) A versatile class of cell surface directional motors gives rise to gliding motility and sporulation in Myxococcus xanthus. PLoS Biol 11: e1001728.

      White, D., Dworkin, M., and Tipper, D.J. (1968) Peptidoglycan of Myxococcus xanthus: structure and relation to morphogenesis. J Bacteriol 95: 2186-2197.

      Yang, X., McQuillen, R., Lyu, Z., Phillips-Mason, P., De La Cruz, A., McCausland, J.W., Liang, H., DeMeester, K.E., Santiago, C.C., Grimes, C.L., de Boer, P., and Xiao, J. (2021) A two-track model for the spatiotemporal coordination of bacterial septal cell wall synthesis revealed by single-molecule imaging of FtsW. Nat Microbiol 6: 584-593.

      Zhang, H., Venkatesan, S., Ng, E., and Nan, B. (2023) Coordinated peptidoglycan synthases and hydrolases stabilize the bacterial cell wall. Nature communications 14: 5357.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The revised manuscript includes several useful additions, and I appreciate the efforts to clarify parts of the analysis. The dataset remains valuable. However, several key issues raised previously are not yet fully resolved and continue to limit the clarity of the main conclusions.

      (1) I appreciate that the authors guide the reader to the relevant regions in the analysis of chromosome fusions (Fig. 2b). However, these subtelomeric regions are not clearly visualized, making it difficult to compare fused and unfused profiles, even though the conclusions rely largely on visual inspection of them. A more direct comparison between fused and unfused ends, together with quantitative summaries (e.g., binned Red1 enrichment and comparisons with internal regions), would make this experiment more convincing.

      Thank you for this suggestion. Figure 2 – figure supplement 1 now shows Red1 enrichment in 20-kb bins tiling in from the fusion points to more clearly show that there is no significant difference in Red1 enrichment between fused and unfused chromosomes. These data are consistent with the model that Red1 enrichment is not affected by the presence of telomeres and imply that Red1 under-enrichment near telomeres is primarily encoded in cis.

      (2) The SK1/S288c comparison (Fig. 2c) is an excellent approach, but is currently presented just as profiles, which again requires substantial effort from the reader to extract the relevant information. A systematic analysis across all informative chromosome ends-for example, comparing Red1 levels in syntenic regions using binned log2 fold-change-would more directly test the proposed in cis effect (L168) and clarify the contribution and range of Y'-associated effects. Other factors (e.g. distance from chromosome ends) could also be assessed within this framework.

      Thank you. Figure 2 – figure supplement 2 now shows the profiles placed in register using peak distribution. This analysis demonstrates that registered S288c and SK1 profiles have the same enrichment of Red1, indicating that there are no detectable long-range effects of Y’ elements or other telomere-associated sequences on the neighboring axis binding sites. Figure 2 – figure supplement 3a further quantifies the effect of registering profiles and separates the data based whether Y’ elements are present. These data are consistent with the interpretation that the presence of Y’ elements primarily affects the average axis protein enrichment profiles by displacing strong axis protein binding sites towards the chromosome interior. Our analyses also indicate that this effect is not limited to the Y’ elements as other telomere-associated sequences have a very similar effect on axis protein distribution near chromosome ends.

      Related to this, it is unclear if Y' elements themselves exhibit lower Red1 binding than the genome average. Providing the mean Red1 signal per Y' element would clarify this point and may also aid interpretation of the relationship between coding density and Red1 enrichment.

      Figure 1 – figure supplement 4a and Figure 2 – figure supplement 3b now show that the mean Red1 enrichment on Y’ elements is on average lower than in the rest of the genome. However, as shown in Figure 2 – figure supplement 3b, this effect is not unique to Y’ elements as other telomere-associated sequences show a very similar level of depletion.

      (3) The Dot1-Sir3 section is now simpler. However, I still find it difficult to follow the underlying rationale. In particular, it is unclear why a Dot1 function dependent on H3K79 methylation is introduced, given that the data in the previous section suggest H3K79 methylation is dispensable for subtelomeric Red1 depletion. A clearer statement of the authors' working model would be helpful.

      We apologize for this confusion. We restructured this section in an attempt to clarify the link between Dot1 activity and Sir3.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Raghavan and his colleagues sought to identify cis-acting elements and/or protein factors that limit meiotic crossover at chromosome ends. This limitation is important for avoiding chromosome rearrangements and preventing chromosome mis-segregation.

      By comparing protein axis recruitment in SK1 and S288C background, which differ in their number and distribution of Y' elements, the authors show that Y' element have a limited impact on axis protein enrichment. Genetic analyses coupled with ChIP experiments revealed that the differential binding of the Red1 protein in subtelomeric regions requires the methyltransferase Dot1. Interestingly, the lack of Red1 depletion in subtelomeric regions in this mutant does not impact DSB formation. Another surprising finding is that deleting DOT1 has no effect on Red1 loading in the absence of the silencing factor Sir3. Unlike Dot1, Sir3 directly impacts DSB formation, probably by limiting promoter access to Spo11. As now clearly stated in the abstract and the discussion, this explains only a small part of the low levels of DSBs forming in subtelomeric regions and the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      Strengths:

      This work provides intriguing observations, such as the impact of Dot1 and Sir3 on Red1 loading and the uncoupling of Red1 loading and DSB induction in subtelomeric regions.

      The separation of axis protein deposition and DSB induction observed in the absence of Dot1 is interesting because it rules out the possibility that the binding pattern of these proteins is sufficient to explain the low level of DSB in subtelomeric regions.

      The demonstration that Sir3 suppresses the induction of DSBs by limiting the openness of promoters in subtelomeric regions is convincing.

      Weaknesses:

      The section examining the impact of Dot1 and Sir3 remains complex, which is partly inherent to the intricate relationship between Dot1 and Sir3. However, the authors conclude that Dot1 acts independently of its catalytic activity based on the phenotype of the H3K79R mutant phenotype. Although this is possible it is not fully demonstrated as the H3K79R mutant may exhibit its own phenotype independently of Dot1. Unless the authors test the impact of the catalytic dead mutant Dot1-G401R on axis protein enrichment at subtelomeres they cannot claim that Dot1 act independently of its catalytic activity.

      Thank you. We softened the relevant statements and do not invoke Dot1 catalytic activity.

      Sir3's impact on DSB induction is compelling, yet it only accounts for a small proportion of DSB depletion in subtelomeric regions. Thus, the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      We explicitly state the fact that further regulation remains to be discovered in the abstract, results, and discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?

      We thank the reviewer for this important question regarding the lateral mobility of VDAC1 in supported lipid bilayers (SLBs).

      It is well established that membrane proteins can retain lateral mobility in SLBs formed on mica. This is enabled by the presence of a thin interstitial water layer (typically on the order of 10–20 Å) between the substrate and the bilayer, which reduces frictional coupling and preserves membrane fluidity. This property is a key advantage of SLB systems and has been extensively described [1].

      Furthermore, HS-AFM studies provide experimental evidence supporting such mobility. For example, Casuso et al. [2] showed that OmpF trimers—whose extracellular domain is larger than that of VDAC1— exhibit measurable lateral diffusion in supported membranes. This indicates that even relatively bulky membrane proteins are not immobilized by the mica support.

      In the specific case of VDAC1, our data provide direct evidence of mobility. As shown in Figure 4J of Ref. 17, individual VDAC1 pores display lateral displacements within the membrane, despite the presence of strong protein–protein interactions. This observation indicates that VDAC1 is not rigidly immobilized upon adsorption.

      Regarding the reviewer’s question about VBIT-4 treatment prior to membrane adsorption onto mica, we also performed experiments in which VDAC1-containing proteoliposomes were pre-incubated with VBIT4 before deposition onto mica and formation of supported lipid bilayers. These samples were not imaged at sufficiently high resolution to allow the same quantitative analysis of VDAC1 cluster organization as performed for the experiments shown in the main text. However, at the resolution obtained, we did not observe obvious large-scale changes in membrane organization compared with samples in which VBIT-4 was added after supported bilayer formation.

      Finally, in light of VDAC1 mobility observed under our experimental conditions and the well-established properties of SLBs, the mica support does not constitute a major limiting factor for lateral diffusion.

      Reviewer #2 (Public review):

      (1) The main limitation is that the conclusion that VBIT-4 does not affect VDAC1 oligomerization is strongest for the specific readouts used here: atomic force microscopy measurements of cluster compaction, VDAC1 channel properties, and simulated assembly behavior. These are direct and informative measurements, but they are not identical to the chemical cross-linking readouts used in much of the prior VBIT-4 literature. Readers should therefore distinguish between VDAC1 cluster organization in membranes, as measured here, and cross-linking-defined VDAC1 proximity.

      We agree with the reviewer that AFM-defined VDAC1 cluster organization and cross-linking-defined VDAC1 proximity are related but non-equivalent readouts, and we have revised the manuscript to make this distinction explicit. However, this distinction also highlights an important limitation in interpreting changes in cross-linking efficiency as direct evidence of altered VDAC1 oligomerization. Chemical cross-linking primarily reports the proximity and accessibility of reactive residues and does not directly provide information on the number, size, stability, or supramolecular organization of VDAC1 assemblies. Moreover, VDAC1 organization is strongly influenced by the lipid environment, as shown in giant proteoliposomes with reconstituted VDAC1 by fluorescence correlation spectroscopy [3] and AFM [4], and changes in protein spacing or orientation within dynamic lipid–protein clusters could alter cross-linking efficiency without disrupting the assemblies themselves.

      Cross-linking can provide valuable information when interpreted in the context of independently defined oligomeric structures or interfaces, as recently illustrated by Takeda et al. for yeast Por1 [5]. However, in most studies reporting inhibition of VDAC1 oligomerization by VBIT-4, the evidence relies primarily on SDSPAGE analysis of chemically cross-linked species, frequently quantified as changes in VDAC1 dimers. In contrast, our HS-AFM measurements directly assess the spatial organization and compaction of VDAC1 assemblies in lipid membranes, while our simulations independently assess assembly behavior. Although these approaches do not measure cross-linking efficiency, neither reveals measurable disruption of VDAC1 assemblies by VBIT-4 under the conditions tested. Our observations are therefore difficult to reconcile with the interpretation that VBIT-4 inhibits VDAC1 assembly formation. Instead, we propose that previously reported changes in cross-linking efficiency could reflect changes in protein proximity, orientation, or residue accessibility within dynamic lipid–protein clusters, particularly given the membrane-perturbing properties of VBIT-4 demonstrated here, rather than disruption of VDAC1 assemblies themselves.

      We have therefore revised the Discussion to acknowledge that these approaches probe distinct aspects of VDAC1 organization, while also clarifying that changes in cross-linking efficiency alone cannot be interpreted as direct evidence that VBIT-4 inhibits VDAC1 assembly formation.

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) A second limitation is the uncertainty around effective VBIT-4 concentration. Because VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive, nominal added concentration may differ substantially from the concentration of active compound available in each assay. This complicates comparisons across the different in vitro, simulation, cellular, and previously published assays.

      We thank the reviewer for this important comment. We agree that the nominal concentration of VBIT-4 does not necessarily reflect the effective concentration of active compound available in solution or within lipid membranes. We have therefore expanded the Discussion to explicitly distinguish nominal from effective concentration and to emphasize that, because of VBIT-4's poor solubility, aggregation, membrane partitioning, pH-dependent protonation, and limited stability during storage, the effective concentration cannot be readily determined or compared across experimental systems. “As a consequence, the nominal concentration of VBIT-4 added to an experiment is unlikely to correspond to the effective concentration of active compound available in solution or within lipid membranes. The effective concentration is expected to vary substantially with pH, storage conditions, formulation, and membrane composition, complicating direct comparisons between different assays and across studies. “

      Importantly, to facilitate comparison with the existing literature, we deliberately used the same nominal concentration range as previous studies investigating VBIT-4. Thus, although the effective membrane concentration is inherently uncertain, this limitation applies equally to previous studies using VBIT-4 and represents an intrinsic limitation of the compound rather than of our experimental approach. Measuring the effective membrane concentration is currently not feasible and is beyond the scope of the present study. We further note that this intrinsic uncertainty likely contributes to the variability observed between published studies.

      (3) The coarse-grained simulations provide a coherent mechanistic framework for membrane partitioning, aggregation, and defect formation. However, the VBIT-4 coarse-grained model is newly parameterized and is used to support a quantitative partitioning argument. The manuscript would be easier to interpret if the coarse-grained-derived partition coefficient were reported with uncertainty, convergence information, and protonation state, and compared with a matched all-atom octanol-water partition estimate from the same atomistic model used to build the coarse-grained mapping. This matters because the partitioning argument is used quantitatively to relate micromolar aqueous VBIT-4 to millimolar concentrations in the bilayer.

      Following the Reviewer’s comments, we have expanded the Methods section to add further detail and references on the transfer free-energy calculations and include the requested details. The protonation state used throughout is neutral VBIT, as alchemical free energy calculations of charged solutes require additional corrections, and reliable reference logP values for charged species are scarce given that most empirical predictors are parameterised for neutral molecules; this is the standard Martini pathway for nonbonded term validation.

      The calculated octanol/water logP for neutral VBIT, averaged over three independent replicates, is 3.52 ± 0.01 (replicate values: 3.54, 3.53, 3.51). Convergence and overlap diagnostics confirmed well-sampled simulations across all lambda windows; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supporting Information.

      Regarding the suggestion to compare against a matched atomistic free-energy calculation: we chose not to pursue this route, as atomistic MD-based logP estimates are not a more reliable reference than empirical predictors for this purpose. Benchmark studies have shown that empirical consensus methods generally outperform atomistic free-energy calculations when compared against experiment, while being substantially less computationally demanding (See [6,7]). We therefore benchmark against a consensus of five established empirical logP predictors (iLOGP, XLOGP3, WLOGP, MLOGP, SILICOS-IT via SwissADME), yielding a consensus logP of 3.42 ± 0.63 for neutral VBIT, in good agreement with our CG estimate of 3.52 ± 0.01.

      We replaced the following manuscript text in the Methods section:

      “These choices were validated by estimating CG octanol/water partitioning free energies, which were compared to predictors obtained via SwissADME[80] (iLOGP[81], XLOGP3[82], WLOGP[83], MLOGP[84], SILICOS-IT). The calculated partitioning free energies were obtained by thermodynamic integration as described elsewhere[77].”

      by the following:

      “These choices were validated by calculating CG octanol/water partitioning free energies for neutral VBIT and comparing them against a consensus of reference values obtained by theoretical predictors. Rather than comparing our Martini logP measurements against atomistic molecular dynamics calculations, we benchmark against established logP prediction methods. For equilibrium octanol/water partitioning, empirical predictors have generally demonstrated accuracy comparable to, or better than, atomistic free energy calculations when evaluated against experiment [6]. We further use a consensus of five empirical models, as consensus predictions have been shown to outperform individual predictors [7]. Reference logP values for neutral VBIT were obtained using the prediction methods available through SwissADME [8] (iLOGP [9], XLOGP3 [10], WLOGP [11], MLOGP [12], SILICOS-IT), yielding predicted logP values of 3.66, 3.85, 4.06, 2.25, and 3.28, respectively. The average of these predictions was used as a consensus estimate, giving a logP value of 3.42 ± 0.63.”

      “The calculated CG partitioning free energies were obtained for neutral VBIT by thermodynamic integration as described elsewhere [13]. In short, the solute is alchemically decoupled from each solvent environment independently across 12 lambda windows, gradually turning off all non-bonded interactions between the solute and its surroundings. The free energy change along this path corresponds to the solvation free energy in that solvent, and taking the difference between octanol and water (ΔG_octanol − ΔG_water) yields the transfer free energy, which is converted to a partition coefficient. Three independent replicates were run, yielding octanol-water logP values of 3.54, 3.53, and 3.51, with an average of 3.52 ± 0.01. Convergence and overlap analysis confirmed that all lambda windows were well-sampled and that the free energy estimates were statistically reliable as required per the guidelines for the analysis of free energy calculations [14]; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supplementary Figures 11 and 12.”

      (4) Finally, the cellular data strongly support VDAC1-independent cytotoxicity, but the lower-dose mitochondrial functional phenotypes were not directly compared between wild-type and VDAC1-knockout backgrounds. VDAC1 independence is therefore more directly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      We agree that our cellular data cannot confirm that the mitochondrial effect of VBIT-4 is independent of VDAC1. However, a previous study by Belosludtsev et al. shows a similar decrease in membrane potential upon VBIT-4 treatment due to inhibition of electron transport chain complexes [15]. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results and Discussion sections.

      Overall, this work provides a valuable and timely reassessment of VBIT-4, and its central conclusion will be useful for researchers interpreting studies that use this compound as a probe of VDAC1 function.

      Suggestions for authors:

      (1) Soften categorical statements such as "VBIT-4 does not alter VDAC1 oligomerization" by specifying the tested readouts: VDAC1 cluster compaction, channel properties, and simulated assembly behavior under the conditions used here. A matched cross-linking experiment under the authors' own VBIT-4 handling and concentration conditions could be useful, but is not essential; the essential point is to make clear that cross-linking-defined VDAC1 proximity and AFM/simulation-defined membrane cluster organization are related but non-equivalent readouts.

      We modified the manuscript to soften the tone. In the discussion, we now emphasize that cross-linking and HS-AFM/simulation probe related but non-equivalent aspects of VDAC1 organization, and that differences in cross-linking efficiency previously reported could be due to alteration of protein spacing or orientation.

      These findings demonstrate that VBIT-4 acts by perturbing lipid bilayers rather than through detectable direct modulation of VDAC1,

      Change title: VBIT-4 Does Not Alter VDAC1 Oligomerization

      To “VBIT-4 Does Not Measurably Alter VDAC1 Cluster Organization or Assembly Behaviour”

      This indicates that VBIT-4 neither prevents nor disrupts VDAC oligomerization.

      To “This indicates that VBIT-4 does not measurably prevent or disrupt VDAC1 assembly under the simulated conditions.”

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) More explicitly distinguish nominal added VBIT-4 concentration from effective available concentration, given the solubility, aggregation, pH-dependence, membrane partitioning, and storagesensitivity observations.

      We modified the Discussion (see public review)

      (3) Report the coarse-grained-derived octanol/water partition coefficient or transfer free energy numerically, with uncertainty, convergence information, and protonation state. Consider providing the corresponding all-atom octanol/water transfer free energy or partition coefficient for the same protonation state(s).

      We answered this comment and added two Supplemental figures 11 and 12. (see public review)

      (4) Either repeat the oxygen consumption rate, TMRM, and Rhod-2 assays in VDAC1-knockout cells, or narrow the wording so that only cytotoxicity is described as directly shown to be VDAC1-independent.

      Thank you for pointing this out. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results (suppression of “This demonstrates that the cytotoxicity is due to a loss of membrane integrity.”) and Discussion sections.

      We removed the direct link to VDAC1: At concentrations below 10 μM, VBIT-4 decreased mitochondrial calcium, respiration, and membrane potential in HeLa cells without affecting mitochondrial mass. They align with reports that VBIT-4 also accumulates in the mitochondrial inner membrane, where it inhibits respiratory complexes I, III, and IV and decreases mitochondrial membrane potential [15,16].

      (5) Consider moving the storage-stability observation into the main text, given its likely importance for interpreting variability in the broader VBIT-4 literature. It would also be useful to include clearer information on stock age, storage temperature, freeze-thaw history, solvent conditions, and whether precipitation or turbidity was observed.

      The Supplemental Figure 9C was moved the main text as new Figure 6, and additional information about storage conditions is added to the Methods section.

      (6) Minor correction: the parenthetical "10^3.5 = 3.2" should be corrected. Since 10^3.5 is approximately 3,162, the intended statement appears to be that a 1 µM aqueous concentration corresponds to approximately 3.2 mM in the bilayer.

      Thank you for pointing it out, it is corrected.

      References:

      (1) Castellana, E. T. & Cremer, P. S. Solid supported lipid bilayers: From biophysical studies to sensor design. Surf. Sci. Rep. 61, 429–444 (2006).

      (2) Casuso, I. et al. Characterization of the motion of membrane proteins using high-speed atomic force microscopy. Nat. Nanotechnol. 7, 525–529 (2012).

      (3) Betaneli, V., Petrov, E. P. & Schwille, P. The role of lipids in VDAC oligomerization. Biophys. J. 102, 523–531 (2012).

      (4) Lafargue, E. et al. Lipid composition of the membrane governs the oligomeric organization of VDAC1. 2024.06.26.597124 Preprint at https://doi.org/10.1101/2024.06.26.597124 (2024).

      (5) Takeda, H. et al. Oligomer-based functions of mitochondrial porin. Nat. Commun. 16, (2025).

      (6) Işık, M. et al. Assessing the accuracy of octanol–water partition coefficient predictions in the SAMPL6 Part II log P Challenge. J. Comput. Aided Mol. Des. 34, 335–370 (2020).

      (7) Calculation of molecular lipophilicity: State‐of‐the‐art and comparison of log P methods on more than 96,000 compounds - Mannhold - 2009 - Journal of Pharmaceutical Sciences - Wiley Online Library. https://onlinelibrary.wiley.com/doi/10.1002/jps.21494.

      (8) Daina, A., Michielin, O. & Zoete, V. SwissADME: a free web tool to evaluate pharmacokinetics, druglikeness and medicinal chemistry friendliness of small molecules. Sci. Rep. 7, 42717 (2017).

      (9) Daina, A., Michielin, O. & Zoete, V. iLOGP: a simple, robust, and efficient description of noctanol/water partition coefficient for drug design using the GB/SA approach. J. Chem. Inf. Model. 54, 3284–3301 (2014).

      (10) Cheng, T. et al. Computation of octanol-water partition coefficients by guiding an additive model with knowledge. J. Chem. Inf. Model. 47, 2140–2148 (2007).

      (11) Wildman, S. A. & Crippen, G. M. Prediction of Physicochemical Parameters by Atomic Contributions. J. Chem. Inf. Comput. Sci. 39, 868–873 (1999).

      (12) Lipinski, C. A., Lombardo, F., Dominy, B. W. & Feeney, P. J. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. Drug Deliv. Rev. 46, 3–26 (2001).

      (13) Souza, P. C. T. et al. Protein-ligand binding with the coarse-grained Martini model. Nat. Commun. 11, 3714 (2020).

      (14) Klimovich, P. V., Shirts, M. R. & Mobley, D. L. Guidelines for the analysis of free energy calculations. J. Comput. Aided Mol. Des. 29, 397–411 (2015).

      (15) Belosludtsev, K. N. et al. Effect of VBIT-4 on the functional activity of isolated mitochondria and cell viability. Biochim. Biophys. Acta Biomembr. 1866, 184329 (2024).

      (16) Belosludtsev, K. N. et al. Pharmacological and Genetic Suppression of VDAC1 Alleviates the Development of Mitochondrial Dysfunction in Endothelial and Fibroblast Cell Cultures upon Hyperglycemic Conditions. Antioxid. Basel Switz. 12, 1459 (2023).

    1. Author response:

      We sincerely thank the editors and the three reviewers for their thorough, highly constructive, and positive evaluation of our manuscript. We are gratified by the reviewers’ recognition of the genetic rigor of our study and the compelling nature of our findings regarding the dose-dependent role of Bcl11b in virtual memory CD8 T cell differentiation.

      We agree that the reviewers have raised fair and addressable points that will undoubtedly strengthen the final manuscript. Below, we outline our planned revisions to address the primary themes raised in the public reviews:

      (1) Genomic Targets and Bcl11b Occupancy (Reviewers #1 & #3)

      To address the request for direct genomic targets of Bcl11b, we will incorporate our existing Bcl11b ChIP-seq data from Bcl11b haploinsufficient and control peripheral naïve CD8+ T cells. We will provide comparative analyses demonstrating that Bcl11b ChIP-seq read density per region is highly concordant with population-matched ATAC-seq accessibility profiles, highlighting that direct Bcl11b occupancy closely mirrors the accessible chromatin landscape, which itself we have assayed thoroughly in thymic precursors. Furthermore, we will integrate this ChIP-seq analysis to cross-reference our bulk and pseudobulk differentially expressed gene lists to better define target overlaps.

      (2) Phenotypic Definitions, Cytokine Independence, and Quantification (Reviewers #2 & #3)

      We appreciate the reviewers’ suggestions to further solidify the phenotypic definitions of our populations. In our revision, we will:

      Perform targeted flow cytometry utilizing our Bcl11b haploinsufficient models to provide explicit quantification of CD5 expression across mature thymic (DP to CD8SP) and peripheral CD8 T cell populations.

      Include supplementary flow cytometry panels for CD122 alongside our standard CD44, CD62L, and CD49d gating strategies used throughout the manuscript in order to comprehensively lock down the T<sub>VM</sub> vs. T<sub>CM</sub> phenotypic definitions.

      Provide absolute cell counts (rather than just relative frequencies) for neonatal CD8 populations to explicitly confirm absolute expansion.

      Expand our evaluation of cytokine-independence by analyzing ImmGen-derived cytokine-response signatures against our scRNA-seq datasets.

      (3) Single-Cell Transcriptomic Alignments (Reviewer #3)

      To contextualize our findings within the broader literature, we will score our scRNA-seq datasets against the derived T<sub>VM</sub> transcriptional signatures recently published by Zhang et al. (2024). While exact cluster-to-cluster matching may be limited by differences in experimental models (e.g., steady-state ontogeny versus influenza infection), this alignment will allow us to demonstrate where our newly minted thymic T<sub>VM</sub>(/precursor) cells map along the established peripheral T<sub>VM</sub>-state continuum.

      Additionally, we will generate targeted split-violin visualizations of T<sub>VM</sub> module scores specifically within scRNA-seq Cluster 8. This will visually clarify the transcriptomic shifts driven by Bcl11b haploinsufficiency within this specific cluster, supplementing the DEG/GSEA tables currently provided.

      (4) Conceptual Clarifications and Discussion Expansions (Reviewers #2 & #3)

      We will expand our discussion to address several excellent conceptual points raised by the reviewers:

      Negative Selection vs. Fate Diversion: We will clarify why attenuated Bcl11b alters TCR-induced gene programs without triggering negative selection, emphasizing that Bcl11b dose reduction drives a portional failure of specific transcriptional repression rather than a general deregulation of global TCR-dependent signaling.

      Haploinsufficiency vs. Knockout: We will explicitly contrast our haploinsufficient (<2 fold reduction) virtual memory phenotype with the innate-like T (and ex-T) cell phenotypes previously reported in complete Bcl11b loss-of-function models.

      Mitochondrial Dynamics: We will refine our text regarding mitochondrial biology to more clearly distinguish between compensatory nuclear transcription (the mitochondrial gene module) and physical organelle performance (mitochondrial membrane potential).

      We look forward to submitting the fully revised manuscript and a detailed point-by-point response in the near future.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We appreciate the additional clarifications suggested by the reviewers, and we will include these in the final Version of Record.


      The following is the authors’ response to the original reviews.

      Reviewing Editor Comments:

      The reviewers are very enthusiastic about this study, but have pointed out a central issue: are "space-time attractors" really attractors?

      The reviewers would be willing to increase the assessment of significance if the comments are properly addressed, and in particular, the issue about space-time attractors.

      We thank the editors and reviewers for the feedback on our manuscript and have revised the paper to address their questions and concerns. This document includes (i) an overview of the major changes to the paper, and (ii) point-by-point responses to the reviewers. We have also attached a version of the revised paper that highlights substantial changes to the text.

      Briefly, the reviewers asked for improved intuition about the STA representation and dynamics, and its relationship to attractor networks. To address their questions, we have restructured the paper. It now starts by introducing the STA, which has a representation and connectivity that are both handcrafted. The revised manuscript characterises the resulting dynamics and fixed points in more detail, both empirically and analytically. We then introduce a new model that directly optimises the fixed points of a neural network to represent an explicit plan of the future. The optimal weights for inferring such representations resemble the STA connectivity empirically. Finally, we analyse our unconstrained recurrent neural network, which learns both optimal representations and connectivity. As also shown in the original paper, this network learns to implement an algorithm that closely resembles an STA. Together, these results show that attractor networks can infer PFC-like representations of the future, and this is an efficient solution to dynamic planning problems known to depend on PFC.

      RE1: Expanded theory of STA dynamics

      We have now formalised how the STA relates to a formulation of planning as an inference process over future trajectories, which has been previously proposed in cognitive science and reinforcement learning. We show in the revised paper that the STA dynamics resemble an algorithm for approximate inference in the corresponding probabilistic graphical model. This allows us to characterise the fixed points of the algorithm analytically and relate them directly to a well-established cognitive theory of planning. These analyses help bridge the gap between neural implementation and cognitive computation. They shine new light on previous results in the paper while also providing more intuition for the STA dynamics.

      We have also included a new model that directly optimises the fixed points of an attractor network to resemble a posterior distribution over future locations from planning-as-inference. This analysis complements the handcrafted STA, where we impose both the representation and connectivity, and the RNN, where both the representation and connectivity are learned. The new model imposes (i) an explicit spacetime representation, and (ii) the multiplicative structure of a message passing algorithm. We then train the weights associated with the forward and backward messages by gradient descent on the KL divergence between (i) the true posterior marginals and (ii) the approximate distribution over future locations implied by the network representation at the fixed point. Supplementary Figure S2 of the revised manuscript shows that the optimal weights reflect the transition structure of the environment, similar to the handcrafted STA model and the task-optimised RNN. This makes the connection between attractor dynamics and planning-as-inference more explicit by showing that the fixed points of an attractor network can be optimised directly for planning.

      RE2: Improved characterisation of fixed points

      We have clarified how and why the STA is an attractor network. Attractor networks are defined by the existence of stable fixed points. In ring and grid attractors, there is a continuum of such fixed points in the absence of structured inputs (but often with tonic excitation). In contrast, the STA has a discrete set of input-dependent fixed points. We show explicitly in the revised manuscript how these fixed points depend on the reward inputs to the network, and also how they relate to planning-as-inference.

      We are not claiming that the STA is exactly equivalent to continuous ring and grid attractors. Instead, we want to convey the intuition that the connectivity of the STA constrains the possible fixed points to be plausible trajectories through space and time. The reward inputs determine which of these possibilities is an actual fixed point in a given planning problem. This is not unlike ring attractors in the presence of strong visual inputs. The connectivity enforces a single bump of activity, and the visual input ‘yokes’ the bump to an appropriate orientation. These similarities and differences are highlighted in the revised paper.

      Finally, we have added a new Figure 3 to the main text that characterises the STA fixed points empirically. This figure:

      (a) Shows the evolution of the STA dynamics and convergence to different fixed points in different environments (panels A-B).

      (b) Shows that the network can converge to different fixed points on different trials in the same environment. This happens when there are multiple equally good paths to a goal (panels B-D).

      (c) Shows that other fixed points also exist that correspond to longer trajectories, but the dynamics of the network bias it towards representations of shorter paths. The STA reliably converges to fixed points representing longer trajectories if it is initialised within their basin of attraction (panel F).

      Updated main text:

      “Unlike ring and grid attractors, the fixed points of the spacetime attractor depend on tonic inputs. However, the connectivity constrains the fixed points to represent continuous trajectories for any combination of inputs. In this section, we show this empirically. Later, we will see that such connectivity is optimal for planning-as-inference.

      To compute a plan, it is necessary to know which states will be rewarding in the future. This reward information is provided as an input to the STA and enables fast adaptation without rewiring the synaptic connections. It alters the fixed points of the recurrent dynamics to only include trajectories that are also associated with high cumulative reward (Figure 3; Methods).”

      RE3: Ground truth rewards as an input to the network

      Both reviewers asked about the external input to the STA that specifies the reward available at different states in the future. In reinforcement learning and cognitive science, ‘planning’ is usually defined as the problem of computing a trajectory that maximises cumulative future reward, given an initial state and a reward function (e.g. Mattar & Lengyel, 2022). This is similar to many real-life situations, where we have a known but distant goal (win a game of chess, finish our paper before a deadline, …). When such a reward function is known, it remains challenging to determine the sequence of actions to get there. This has been the topic of much previous work in neuroscience, including (i) the successor representation, which combines a trial-specific reward function with stable transition statistics; and (ii) different types of sequential search, which use a known reward function to evaluate different possible future trajectories.

      To highlight the importance of planning, even when the reward function is known, Figure 4 of the revised manuscript shows that the STA performs better than a greedy baseline that acts according to the immediate reward input instead of planning to maximise cumulative reward. Planning is therefore distinct from learning or inferring a reward function, which is itself a major open question in cognitive science. While undoubtedly interesting, a solution to this problem is beyond the scope of our paper. That is why we decided to simply provide ground truth rewards as an input to the STA. We have clarified the distinction between planning and ‘reward learning’ in the revised paper, and the supplementary material now includes a discussion of where the reward input to the STA could come from.

      Author response image 1.

      All performance quantifications in Figure 4 now include an additional ‘greedy’ baseline (grey bars). This is an agent that acts according to the immediate future reward. The performance improvement of the STA over this baseline highlights the importance of planning to maximise cumulative reward.

      Updated main text:

      “There are several possible sources of reward input to a spacetime attractor (Supplementary Note). We focus on planning under a known reward function and therefore assume access to ground-truth rewards.”

      Reviewer #1 (Public review):

      Summary:

      This work builds a theory to implement planning trajectories towards a goal in a known environment, inspired by analyses of prefrontal neural recordings. Unlike standard neural architectures for this task, such as value-based learning and successor representations, their proposed theory is able to adapt to novel goal locations within a trial. The key to the theory is that future times are represented by orthogonal groups of neurons. The recurrent connectivity between groups of neurons selective to specific future times and locations reflects the learned knowledge of the task. Finally, the authors show that standard networks trained on the task approximate their proposed theory.

      Strengths

      The structure of the work is clear, and the presentation of the results is very well written, which is particularly noticeable given the consequential amount of results presented. The authors are able to link their theory with experimental findings in neural recordings. The reverse-engineering of trained recurrent neural networks is very thorough, by analyzing both dynamics and connectivity. The assumptions and predictions of their model are clearly stated.

      We appreciate the encouraging comments and hope our revised manuscript addresses the reviewer’s questions.

      Weaknesses

      (1.1) It is unclear whether their proposed theory, "space-time attractors", actually is an attractor network. The authors used recurrent neural networks with very few timesteps, and long single neuron time constants with respect to the task time scales. Attractor networks, as the ones the authors cite, refer to networks that generate nontrivial patterns of activity through recurrent interactions, after long periods of time.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, we show in the revised manuscript how the fixed points of the STA dynamics relate to planning-as-inference, and we clarify the similarities and differences between the STA and other attractor networks in the main text. We show in the new Figure 3 that (i) representations of future paths are stable over long periods of time, and (ii) multiple fixed points can exist when there are multiple paths to the goal. It is also worth noting that the RNN representation in Figure 6H remains stable for 75 time constants and recovers from perturbations. This is substantially longer than during training, where ‘planning’ lasted up to 14 network time constants, and it suggests that the network representation is a stable fixed point.

      (1.2) The authors gloss over how the reward inputs are calculated. Computing these reward inputs should be part of the planning process, and the authors are implicitly leaving this problem aside. How does the reward input, which includes future time and location, depend on the actions that have not yet been taken by the agent? It feels like most of the planning computation is already provided by these reward inputs at the beginning of the trial. It could be that the network is only learning to process the planned sequence of actions present in the inputs.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them. To make this point clearer, we show in Author response image 1 that the representations computed by the STA generate better behaviour than an agent acting greedily according to the reward function specified by the inputs. This highlights the importance of considering distant goals when choosing immediate actions.

      Reviewer #1 (Recommendations for the authors):

      The text is very nicely written, and I appreciated the way in which methods are presented, with a clear structure and a logical chaining of the different sections. My comments and suggestions refer mostly to the methods and the RNN implementation. Please find below a list of issues.

      Relatively major:

      (1.3) All the equations of the dynamics should be written in discrete time and not in continuous time. There is no notion of "iteration" in continuous time, so it is currently very hard to understand how the RNN works, and what the different epochs are ("the RNN performed 10 network iterations..."?).

      We have rewritten all equations in discrete time and clarified the notion of ‘iterations’.

      (1.4) This is pointed out in the public review, but the authors insist on making an analogy between the spacetime attractor implementation of planning and attractor networks. It seems to me that these two types of models are very different. What defines attractor networks (such as grid- or ring-attractor networks) is that recurrent connections internally generate stable states of activity for long periods of time, in the absence of inputs. Nothing like that is shown here. Robust input-driven trajectories are neither necessary nor sufficient for showing that an RNN is an attractor network. The fact that, given the inputs, networks are run for very short periods of time in this work seems to indicate that this is a very different type of network compared to the attractor networks mentioned previously.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, it is correct that the fixed points of the STA depend on the inputs, which is different from canonical ring and grid attractors.

      We show that the fixed points of the recurrent STA dynamics are reward-maximising paths when conditioned on those inputs. Briefly, we now (i) analytically characterise the input-dependence of the fixed points and show how they relate to planning-as-inference; and (ii) show empirically that STA representations remain stable for long periods of time (Figure 3). The revised paper also clarifies the similarities and differences to previous attractor models.

      Minor

      (1.5) It would be nice to show more clearly what the inputs are in a given trial, and how they change over time in the RNN (specifying the planning and execution phases). A supplementary figure may help.

      How is the information about walls provided to the network exactly? More generally, it would be nice to clearly indicate in the Methods what all the inputs "x" to the RNN are, and how they change over time (during planning, during execution, and how they change as the environment steps are updated).

      We have made a new Supplementary Figure S3 that illustrates the inputs to and outputs from the different models. Briefly, the information about the walls is provided to the RNN as a binary vector x<sub>w</sub> ∈ ℝ <sup>2𝑁</sup>. The elements of this vector indicate for each of the N states whether there is a wall (i) to the right of, and (ii) above it. In each trial, a subset of these is present, and a subset is absent. All ‘present’ walls are assigned a value of +1 in x<sub>w</sub> , and all absent walls are assigned a value of 0 in x<sub>w</sub> . We have also clarified this in the revised Methods.

      (1.6) It would help to clarify, at least in the methods, the shape of all the matrices and vectors that are trained.

      We have clarified the shapes of all matrices and vectors in the Methods.

      (1.7) N in the methods is not defined (I think N = 16, the total number of locations on the grid).

      N is indeed the total number of locations in the state space. This is 16 for almost all analyses in the paper, which involve planning on a 4x4 grid. The updated manuscript includes a few analyses in larger environments, where N is larger. We have clarified this in the Methods.

      Reviewer #2 (Public review):

      This well-written manuscript proposes to use attractors in space and time (STA) as a mechanistic explanation for planning in the prefrontal cortex. The main conceptual hypothesis is that planning is implemented as attractor dynamics in a representation that encodes states at each time step jointly. Depending on inputs, the network relaxes to a trajectory that already contains future states that will be visited at each time step, rather than computing a scalar value at each point in time and space like other classical approaches from RL. The authors compare this approach to implementations such as TD learning and successor representation, and further show that trained recurrent neural networks on specific tasks involving planning develop structured subspaces resembling the ones postulated in STA.

      The idea of treating attracting trajectories unfolding in time as the computational substrate for planning is very interesting and potentially important. The explicit construction of a state x time representational space and its implementation via recurrent dynamics are appealing and convincing in the idealized tasks considered. I found the manuscript to be refreshingly explicit regarding several of the assumptions and limitations of the models, for example, the fact that certain advantages can be viewed as properties of the state space itself and not necessarily of a fundamentally new planning mechanism.

      Overall, the manuscript presents a cool attractor model that extends in time and explores its performance in a subset of illustrative tasks involving planning. My doubts concern mostly the interpretation and scope of the claims made in the manuscript. Here are a few comments where I detail my questions/concerns:

      We appreciate the enthusiasm about the manuscript and its potential importance. We address the remaining questions and concerns below.

      (2.1) The authors nicely discuss that much of the difference between STA and classical TD or SR agents is "in some sense a property of the state space rather than the decision making algorithm," and that TD and SR could in principle be implemented in a comparable space x time representation. This is fair, but it also suggests that the central contribution of the manuscript lies primarily in the representational factorization (state x time tiling) and its dynamical implementation via attractors, rather than in a fundamentally new planning algorithm or theory, mechanistic or not. I think theory should be distinguished from mechanism, and it would therefore help the reader to describe the conceptual advancement more as a novel mechanism or implementation than a novel (mechanistic) theory for decision/planning.

      We respectfully disagree that ‘theory’ has to be distinguished from ‘mechanism’. We do agree that ‘computation’ and ‘mechanism’ can often be distinguished. However, we think theories can live at either of these (and other) levels of explanation. What we propose is indeed a potential mechanism for planning that combines recently characterised prefrontal spacetime representations with attractor dynamics to infer desirable ‘plans’. As we show in the revised manuscript, this mechanism resembles the computation of ‘planning-as-inference’, which has previously been proposed in cognitive science (e.g. Botvinick & Toussaint, 2012). Our theory is therefore not about the computation – it is about the mechanism. The title “A mechanistic theory of planning…” is meant to clarify what level of description our paper addresses.

      As an example of the importance of mechanistic theories, the computation of angular velocity integration can be implemented in many different ways. Seminal work by Skaggs et al. (1994) and others in the 1990s showed how it can be implemented in neural networks, inspired by experimental data. These theories paved the way for detailed experimental characterisations of the fruit fly head direction circuit more than two decades later (Turner-Evans et al., 2017; Kim et al, 2017; and others). Inspired by this and other success stories, we think an important role of theoretical neuroscience is to develop theories about neural mechanisms that can be tested in future experiments!

      (2.2) Related to my previous point, I think it would be helpful to position STA more explicitly relative to computational/theoretical literature in which attractor networks encode temporally ordered patterns (so effectively including future times). For example, classical extensions of Hopfield networks with asymmetric connectivity implement retrieval of sequences and ordered transitions between patterns (Sompolinsky & Kanter, 1986). More recently, sequential attractors and limit-cycle dynamics have been constructed in structured recurrent networks by the Morrison group (Parmelee et al., 2021). These works do not implement an explicit discretized state x future-time tiling as in STA and do not specifically discuss the usage for planning. However, they do provide concrete precedents for attractor dynamics over temporally structured trajectories in terms of mechanism. It would be useful to discuss this literature and clarify a little what's new mechanistically in the view of the authors.

      We agree that this is not the first use of attractor networks to represent or compute sequences. Instead, we show that a combination of spacetime representations with attractor dynamics is sufficient to compute plans in dynamic problems known to depend on prefrontal cortex. As the reviewer points out, the primary difference from most previous work lies in the fact that the entire sequence is encoded in a single fixed point of the STA dynamics. This differs from e.g. Sompolinsky & Kanter, where the population encodes one element at a time and generates sequences as limit cycles. The instantaneous encoding of an entire sequence in the STA is what enables planning through parallel message passing rather than sequential search. This is highlighted in the main text of the revised manuscript, which also includes a supplementary discussion of the similarities and differences between the STA and related work on sequences in attractor networks.

      Updated main text:

      “Entorhinal grid cells are also embed a world model in their connectivity (McNaughton et al., 2006), but they only encode a single location at a time (Vollan et al., 2026). Such networks can generate sequences, but the individual elements are represented one by one (Sompolinsky and Kanter, 1986; Kleinfeld, 1986; Widloski et al., 2025). The spacetime attractor suggests that circuit principles in prefrontal cortex resemble other cortical areas that use structural knowledge to infer features of the world. The major difference is that PFC instantaneously represents many points in time, which generalises known circuit principles to complex planning.”

      (2.3) A central claim of the manuscript is that space-time trajectories are attractors of the STA dynamics. The manuscript does provide empirical evidence consistent with attractor-like behavior. However, it is not explicitly shown whether trajectory representations persist in the absence of sustained external inputs. So it's not clear to me whether the trajectories should be interpreted as intrinsic attractors of the recurrent system, which can be selected by delivering transient inputs, or whether they must be stabilized by a specific continuous external drive. It would be useful if the author could clarify/discuss this point.

      We show in the revised paper that the fixed points of the STA dynamics take the form r <sub>δ</sub> = e<sup>R<sub>δ</sub></sup> ◦ (Ar<sub>δ−1</sub>) ◦ (A<sup>T</sup> r<sub>δ+ 1</sub>) (Methods). Here, r δ is the activity of neurons representing expected locations in δ actions; R δ is the reward function in δ actions; and A is the environment adjacency matrix. These fixed points depend on the reward inputs through the first term. In the absence of reward inputs, the fixed points are ‘diffusive’, while still respecting the transition structure of the environment. In the presence of reward inputs, they concentrate probability mass on trajectories with high expected reward. We have clarified these properties in the main text and introduced a new Figure 3 that characterises the fixed points of the STA in more detail. See also RE1 and RE2.

      (2.4) As far as I understand it, reward information is provided as input to specific populations encoding future time steps, and that's essential for rapid adaptation without rewiring connectivity. How such future-time-specific reward inputs would be generated and routed to distinct neural populations isn't entirely clear to me. Since this seems to be an essential component of the model, I think it would be important to discuss more deeply the source and plausibility of these reward signals related to different timesteps.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them (see Author response image 1). We agree that the challenge of estimating future reward is an interesting question, but it is beyond the scope of this paper. We have clarified this distinction in the main text and added a supplementary discussion that speculates about where reward information could originate in biological circuits.

      (2.5) The authors note that vanilla STA scales linearly with planning horizon, and discuss potentially hierarchical extensions for longer horizons. They acknowledge that learning abstractions remains an open challenge, yet the examples of planning in the manuscript are restricted to very short temporal horizons and limited branching complexity. It is not obvious to me in what cases the current implementation and interpretation of STA remains viable (for example, in terms of relaxation iterations) as the horizon and branching factor increase. Relatively simple planning can be managed by simpler, less costly models/algorithms, whereas complex planning is a lot harder to deal with, and it's something that a mechanistic "theory" should address. In the context of the claims of the paper in its present form, I think this is possibly the most important conceptual and practical limitation in the manuscript.

      It is correct that planning gets increasingly challenging with planning depth. In the absence of noise, the STA scales to sequences of up to 12-13 actions – and even longer if the minimum path length is known a priori. Performance gets progressively worse when recurrent activity and parameter noise increase. The revised paper includes a new Supplementary Figure S1A-C that shows how the STA planning ability depends on planning depth for different levels of noise.

      We do not consider planning depth to be a major limitation of the work, since humans are rarely thought to plan much more than 6 steps into the future at a single level of abstraction (e.g. van Opheusden et al., 2023). Instead, we believe that hierarchical planning is used to infer trajectories to distant goals (Eckstein & Collins, 2020). To illustrate this point, we have now implemented a proof-of-principle hierarchical STA in Supplementary Figure S1D-E. This simulation shows how an ‘abstract plan’ inferred by one STA can be treated as a goal to infer a more ‘detailed plan’ in a second STA. In principle, this enables the system to compute plans that are arbitrarily long, provided they can be broken down into chunks smaller than the limits imposed by the analyses in Supplementary Figure S1A-C.

      Finally, RNNs learn an STA-like algorithm when trained on dynamic planning problems with a planning depth of 6. It is therefore not clear to us whether simpler and less costly algorithms can be easily implemented in the dynamics of recurrent networks.

      (2.6) The RNN analyses show that trained networks develop structured subspaces aligned with future time indices and exhibit perturbation behavior consistent with attractor-like dynamics. The manuscript also explicitly notes differences between the trained RNN and the handcrafted STA (e.g., long-range couplings between subspaces and differences in behavior of lower-value trajectories under perturbation), which I much appreciated. My doubt is on the specificity of this result, as trained RNNs on fixed-horizon tasks can develop latent dimensions correlated with temporal progress within a trial or time-to-goal. I think it would help the reader to clarify whether the results demonstrate that STA-like computations emerge in RNNs trained on planning tasks, or that RNNs generally develop some kind of structured spacetime representations when tasks involve future timesteps and some degree of flexibility in the decisions.

      An important point to note is that the subspaces we identify do not encode time-to-goal, since they are all active at the very beginning of the trial. We also show that RNNs trained on simpler static tasks do not learn the same algorithm (Supplementary Figure S7) and do not generalise to dynamic problems (Figure 5F). Finally, other algorithms are capable of solving the dynamic problems we study (e.g. the ‘value agent’ in Figure 5B-E). We therefore do not think it is trivial that RNNs learn an STA-like algorithm.

      We do think that ‘structured spacetime representations’ generally emerge in RNNs trained on tasks that involve flexible behaviour in changing environments – in some sense that is the claim we are trying to make. It is known that spacetime representations are optimal for structured sequence memory tasks (e.g. Whittington et al., 2025; Dorrell et al., 2026), and we think this is for exactly the same reason. In sequence working memory, the reward function changes in time – for each action, the reward is only non-zero at the corresponding sequence element. However, the adjacency matrix is uniform for sequence memory – any sequence element can follow any other sequence element – so there is no need for planning. We are therefore not claiming that spacetime representations only emerge in the specific planning task we consider here. Instead, we expand the set of problems solvable by such representations to also include adaptive planning known to depend on prefrontal cortex. We have made this more explicit in the revised manuscript.

      Updated main text:

      “Together, our analyses show that RNNs trained on a dynamic planning task learn to approximate a spacetime attractor. This was also true across variations in model architecture (Methods; Figure S10; Figure S11). These results extend previous findings that explicit spacetime representations are optimal for sequence memory (Supplementary Note; Whittington et al., 2023; Dorrell et al., 2026; Wang et al., 2025). Additionally, RNNs with too few hidden units to learn a spacetime attractor failed to solve the task (Figure S12), suggesting that other solutions are not readily learned by gradient descent.”

      A few more minor points, mainly concerning clarity:

      (2.7) The main dynamical equation combines a log-domain recurrent term, a floor operation, and a log-sum-exp normalization step, followed by exponentiation. The intuition/logic behind this specific formulation could be clarified for the reader. For example it would be helpful to explain why the recurrent input appears inside a log, and also whether/how these operations relate to any multiplicative constraint.

      The specific form of these equations comes from the intuition that the STA approximates planning as an inference process over future trajectories. We have clarified this in the revised manuscript, which explicitly shows how these equations relate to planning-as-inference as formulated previously (e.g. Botvinick & Toussaint, 2012; Levine, 2017).

      (2.8) While the computational cost of successor representation in an expanded NT x NT representation is discussed, the corresponding scaling of STA in terms of number of units and connections (as a function, for example, of the planning horizon) isn't clear to me. Perhaps the authors could compare costs more explicitly.

      The memory cost of a spacetime-SR would be (NT)^2 and the computational cost (NT)^3 (it is possible that both of these could be reduced by taking advantage of the structured nature of the spacetime successor matrix, but that is beyond the scope of this work). The memory cost of the STA is NT, and the computational cost is (NT)^2 (each iteration of the network dynamics requires the calculation of T matrix-vector products of size NxN, and the number of steps to convergence is approximately linear in T). We have included this comparison in the Supplementary Discussion of the revised paper.

      (2.9) In the RNN analyses, structured subspaces aligned with future time indices are shown. I couldn't find a quantification of how much variance is captured by the subspaces, relative to other latent dimensions. Adding it would help get a feeling for the strength of the alignment.

      We have added a new Supplementary Figure S6 to the revised manuscript, which quantifies the variance explained by the future-coding subspaces over the course of a trial. The variance explained by the K dimensions encoded by these subspaces is substantially higher than a random baseline, and it approaches the upper bound given by the top K PCs. Interestingly, the future-coding subspaces all explain a lot of variance early in the execution period. During later stages of execution, only the ‘immediate future’ subspaces explain substantial variance. This suggests that the RNN only maintains information in subspaces that represent times before the end of the trial.

      References

      Botvinick, Matthew, and Marc Toussaint. "Planning as inference." Trends in cognitive sciences 16.10 (2012): 485-488.

      Dorrell, William, et al. "An Efficient Computing Theory of Prefrontal Structured Working Memory Representations." bioRxiv (2026): 2026-02.

      Eckstein, Maria K., and Anne GE Collins. "Computational evidence for hierarchically structured reinforcement learning in humans." Proceedings of the National Academy of Sciences 117.47 (2020): 29381-29389.

      Kim, Sung Soo, et al. "Ring attractor dynamics in the Drosophila central brain." Science 356.6340 (2017): 849-853.

      Levine, Sergey. "Reinforcement learning and control as probabilistic inference: Tutorial and review." arXiv preprint arXiv:1805.00909 (2018).

      Mattar, Marcelo G., and Máté Lengyel. "Planning in the brain." Neuron 110.6 (2022): 914-934.

      Skaggs, William, et al. "A model of the neural basis of the rat's sense of direction." Advances in neural information processing systems 7 (1994).

      Turner-Evans, Daniel, et al. "Angular velocity integration in a fly heading circuit." Elife 6 (2017): e23496.

      Van Opheusden, Bas, et al. "Expertise increases planning depth in human gameplay." Nature 618.7967 (2023): 1000-1005.

      Whittington, James CR, et al. "A tale of two algorithms: Structured slots explain prefrontal sequence memory and are unified with hippocampal cognitive maps." Neuron 113.2 (2025): 321-333.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this is an interesting and well-written manuscript on a fascinating question in a "charismatic" model system.

      Strengths:

      (1) The Introduction is concise, though it might be helpful to the non-specialist reader to learn a bit more about what is known about the social control of somatic growth across diverse species (including humans), which would help to make this work more generally interesting.

      (2) The experiment is well-designed.

      (3) The data collected are comprehensive.

      (4) The complementary analysis of both feeding and aggression/submission data with and without known social roles is a neat idea and compelling!

      Thank you for the positive feedback!

      Here, we investigate phenotypic plasticity associated with the adoption of social roles in the clown anemonefish, with strategic growth being just one aspect of that plasticity. Strategic growth, also known as social control of growth, is a fascinating form of adaptive phenotypic plasticity, whereby individuals modify their growth and size in response to fine-scale changes in social conditions (Buston & Clutton-Brock, 2022). In cooperative breeding systems with high reproductive skew, particularly fishes and mammals (possibly including humans), individuals have been shown to i) increase growth/size on the acquisition of dominant status (Dengler-Crish & Catania, 2007; Johnston et al., 2021; Thorley et al., 2018; Van Schaik & Van Hooff, 1996; Walker & McCormick, 2009), ii) increase growth/size when paired with size matched reproductive rivals (Huchard et al., 2016; Reed et al., 2019; this study), and iii) decrease growth/size to avoid conflict (Buston, 2003; Heg et al., 2004; Wong et al., 2007). While strategic growth is fascinating and clearly occurring in this study, we show coordinated changes of multiple aspects of the phenotype as fish adopt social roles. Therefore, we deliberately framed the Introduction broadly to avoid biasing the reader toward viewing growth as the sole or main driver.

      Weaknesses:

      (1) I was surprised that the HPA/stress axis was not considered here at all. Wouldn't we expect that subordinates have increased stress axis activation, which in turn could inhibit their growth and aggressive behavior?

      We also expected to see the HPA/stress axis activated in subordinates, which is why we carried out a targeted exploration of genes known to play a role in this axis. We did not find any genes that were significantly differentially expressed. We believe that there could be two explanations for this. First, from a methodological perspective, it could be due to our use of a whole-body RNA-seq, which may have masked this signal. Alternatively, the stress axis might play a more complex role than just acting as a simple on/off switch for reduced growth. Its activation may peak when competition over size is at its highest (during week one) or, conversely, it may peak later and help maintain reduced growth once hierarchies are firmly established (particularly after the dominant individual reaches its maximum size). To understand the role of the stress axis, future studies should observe how its activation varies over time. We acknowledge that the absence of a stress‑axis signal and its potential explanations were not clearly discussed in the original manuscript. In the revised version, we have addressed this in the Discussion at lines 564-567 and Methods at lines 795-797 and 802-805.

      Discussion lines 564-567 and 577-580:

      “These include appetite regulation (orexigenic and anorexigenic signaling), metabolic pathways (e.g., glycolysis, lactic fermentation, TCA cycle, fatty acid β-oxidation), growth-regulating pathways (GH/IGF, insulin/PI3K-AKT, mTOR, and Hippo), and potential molecular signatures to varying levels of social stress.”

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (2) To what extent are growth, food intake, agonistic behavior, and/or gene expression patterns coordinated across P1 vs P2 pairs? The lack of such an analysis seems like a missed opportunity.

      We had a similar thought. Specifically, we were interested in testing the hypothesis that the final size ratio of pairs, which is indicative of the amount of conflict remaining, would predict gene expression. We examined gene expression within pairs to test for coordinated changes and repeated the analysis, accounting for the pair size ratio. In both cases, we found no clear or consistent pattern within pairs. We have included these analyses in the revised manuscript, Methods lines 781-7946-804 and 807-814, and Supplementary Materials lines 144-155 and 164-172 along with three new Supplementary Figures: Fig. S8, S9, S10.

      Supplementary Materials lines 144-155 and 164-172:

      “We have identified genes associated with growth and ossification that showed strong positive correlations with body size. For this set of genes, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al., 2007), would predict variation observed in gene expression within social position (Fig. 5B). Visual inspection of the size ratio annotation in the heatmap of these genes (Supplementary Fig. S8), together with a PCA of paired individuals (P1 and P2) based on the same row Z-score values, was used to assess whether remaining conflict explained expression variation (PERMANOVA: p = 0.858; Supplementary Fig. S10A). These results suggest that size ratio between pairs was not associated with gene expression variation within social position.”

      “We have identified genes associated with appetite regulation and metabolism, which were significantly downregulated in P1 individuals compared to P2 and S fish. As above, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al. 2016), would predict variation observed in gene expression within social position (Fig. 6). We found no significant effect (PERMANOVA: p = 0.844; Supplementary Fig. S9 and S10B), indicating that size ratio does not explain gene expression differences within social positions.”

      Methods lines 781-794 and 807-814:

      “In the heatmap, samples (columns) were a priori grouped by social position (P1, P2, S) and subsequently ordered within each group by gene expression similarity, reflected by the column dendrogram.

      Additionally, we tested the hypothesis that size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of conflict remaining (Wong et al., 2016), would predict variation in gene expression. This was assessed by visual inspection of the same heatmap with size ratio annotation using Complex Heatmaps (Gu et al., 2016), together with a PCA of the expression of these genes of paired individuals (P1, P2) only. We also performed a PERMANOVA analysis using the same row Z-score values (adonis2 function in vegan package: Oksanen et al., 2017) accounting for social position and genetic background (clutch ID), using restricted permutations to account for the non-independence of P1 and P2 within pairs was carried out. Solitary individuals were excluded from this analysis because they lack size ratio data. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S8, S10A).”

      “In the heatmap, samples (columns) and genes (rows) are grouped a priori by social position (P1, P2, S) and pathway (APT, TCA, GLY), respectively, with dendrograms reflecting expression similarity within each group. For the candidate gene heatmap, we repeated the PCA and PERMANOVA exploration and associated heatmap visualization for pairs only (as described above), to test whether final size ratio (P2 SL / P1 SL) predicted gene expression variation within social positions. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S9, S10B).”

      (3) What was the rationale for using whole bodies for the transcriptome analysis? Given the hypotheses, the forebrain or hypothalamus and certain other organ systems (e.g.,liver, gonads, skin, etc.) would have been obvious candidate tissues here. I realize that cost is always a consideration, but maybe a focus on the fore-/midbrain could have been prioritized.

      We decided to use whole-body samples for this initial transcriptomic analysis to capture a broad view of gene-expression differences while keeping sequencing costs and sample requirements manageable. We agree with the reviewer that future work should explore specific tissues sampled from individuals at multiple time points to disentangle transcriptomic differences across tissue types. In our revised manuscript we explicitly state the limitations and outline future steps in the Discussion at lines 553-564, as well as in the Methods section we now explicitly state the rationale behind whole-body RNA-seq and acknowledge the limitations of this approach in lines 688-692.

      Discussion line 5536-564:

      “In this study, we used whole-body transcriptomics, which revealed overarching expression patterns, however, we acknowledge that this approach limited our ability to detect finer-scale signals. Obviously, this approach cannot resolve tissue-specific gene expression changes (Lu et al., 2020; Roux et al., 2023; Yin et al., 2023) which are critical for a complete understanding social role differentiation, such as adjustments of behavior, appetite, and growth (a list we consider non-exhaustive). To disentangle gene expression signatures associated with socially induced phenotypes such as the strategic up- and downregulation of growth, future studies should include sampling points aligned with the onset of these physiological shifts. Moreover, to understand underlying pathways involved, tissue-specific transcriptomics will be essential. Sampling multiple tissues across multiple time points would disentangle nuanced regulatory processes underlying coordinated functional shifts leading to different social phenotypes of strategic growth.”

      Methods lines 688-692:

      “To initially explore overall gene expression patterns associated with coordinated changes during the emergence of social roles, whole-body RNA-seq was performed to keep both sequencing cost and sample requirements manageable. We acknowledge the limitations of this approach and consider this study a hypothesis-generating tool that lays the foundation for future studies.”

      (4) Given the preceding point, why was a fold-change threshold used for assessing DEGs (supplementary Figure 3)? There is no biological justification to ever use a fold-change threshold, especially in bulk RNA-seq analysis. This is particularly true here, where wholebodies were used for RNA-seq analysis, which is a bit unusual. Relatively small cell populations (such as hypothalamic neurons that regulate growth or food intake) may show substantial gene expression variation across social types, yet will be masked by the masses of other cells in the whole body sample. However, gene expression may still vary significantly, albeit the fold-difference may be small. I therefore suggest a reanalysis that omits any fold-change threshold.

      We thank the reviewer for this important point, and agree that an arbitrary fold‑change cutoff is inappropriate/unnecessary. It should be noted that this fold-change cutoff was only used in this single figure, and all other analyses used p-values from the entire dataset. We have removed the fold‑change threshold cutoff and corrected the Figure (previously Supplementary Figure 3) now Supplementary Fig. S5, and all corresponding text.

      (5) Why is the analysis of color (hue, saturation) buried in the supplementary materials? Based on the hypotheses that motivated the study, color seems just as relevant as food intake, growth, and agonistic behavior, so even if the results are negative, they should be presented in the main paper.

      We agree that color can be an important social signal, so we included color measurements in our experimental design. However, after careful consideration of the color results, we decided that our experimental timing and husbandry changes introduced multiple confounding factors, preventing us from drawing confident conclusions. Specifically, our fish were ≈1 month old at the transfer from larval to experimental tanks and had already begun to deepen their orange hue, before our experiment. (In the wild, they would settle at one to two weeks of age, prior to the deepening of the orange hue). Once individuals attain a certain hue, it seems that color development can be halted, but not reversed. The transfer also involved changes in lighting, tank background, and diet, factors known to strongly affect coloration (Maytin et. al., 2018). Our results show a uniform shift in orange hue and saturation across social groups, suggesting that these confounding factors might have dominated changes in hue.

      For transparency, we report the color data in the Supplementary Materials, but we caution against drawing any strong conclusions. In the revised Supplementary Materials document, we added lines 231-236 recommending that future work should involve a targeted experiment to robustly test for the effect of the adoption of social roles on coloration or the effect of coloration on the adoption of social roles.

      Supplementary Materials lines 231-236:

      “Together, these considerations suggest that timing, environmental uniformity, and future reproductive potential may limit the expression or detection of socially mediated color plasticity under laboratory conditions. Future studies should carefully consider experimental design to minimize confounding factors and robustly test the effects of the adoption of social roles on coloration and the effects of coloration on the adoption of social roles.”

      (6) The Discussion is sometimes difficult to follow. The authors may want to consider including a conceptual graphic that integrates the different aspects of growth and satiety regulation, etc., into a work-in-progress model of sorts, which would also facilitate clearer hypotheses for future research.

      Thank you for flagging that parts of the Discussion are a bit difficult to follow. In the revised manuscript, we worked to improve readability of the Discussion. We also appreciate the suggestion of including a conceptual schematic. For this manuscript, we refrained from adding such a “work-in-progress” schematic, as we felt that our whole-body RNA-seq approach and single gene expression sampling time point substantially limited our ability to make predictions.

      Reviewer #2 (Public review):

      In this manuscript, the authors test growth, behavior, and gene expression in pairs of clownfish as they establish social dominance hierarchies, examining patterns of gene expression in these pairs after dominance has been established. The authors show solid evidence that emerging dominant clownfish show increased growth, aggression, and food consumption compared to their submissive or solitary counterparts, eventually adopting distinct gene expression profiles.

      Major Comments:

      (1) The Introduction is comprehensive, but it could be condensed. Likewise, the discussion could be condensed. There is considerable redundancy between the methods, the results,and the legend in Figure 1. The authors should consolidate and remove the redundancy.

      Thank you for flagging that parts of the manuscript could be condensed, we will work on this as we revise the manuscript.

      (2) For Figure 3, the authors are showing PC2 and PC3; why is PC1 not shown? There is so much overlap between the three groups in PC2 vs PC3; it seems unlikely that researchers could conclusively identify any individual as belonging to a group based on the expression profile. The ovals shown do not capture all the points within each of the groups, and particularly the grey S oval seems misaligned with the datapoints shown.

      We understand the concern raised by the reviewer about the overlap among points in the PCA. We have explored PC1-PC3 and found that PC2 and PC3 showed the clearest, statistically significant clustering by social position, while PC1 did not capture any variation due to social position. We have explored whether other factors might be masking differences, such as genetic relatedness, tank effects, total read count per sample, and found that none of these factors explained sample clustering. Regarding the ellipses shown around the points, they were not intended to capture all points, but rather they show the estimated 95% multivariate t-distribution for that given social group. We revised the figure legend (Fig. 3) to clearly reflect this. . In addition, for transparency we have revised the Results section (lines 279-285) and Methods section (lines 754-765) to clarify that we have performed PCAs on (1) all genes, (2) top 50%, (3) top 25%, (4) top 5% most variable genes, and all three pairwise PC comparisons (PC1 and PC2, and PC1 and PC3) which we show in the added Supplementary Fig. 3S.

      Results lines 279-285:

      “PCAs were performed on all genes, and the top 50%, top 10%, and top 5% of the most variable genes, all of which showed consistent clustering across all pairwise PC comparisons (PC1 vs PC2; PC2 vs PC3; PC1 vs PC3; Supplementary Fig. S3). Of all the examined PCAs, PC2 and PC3 with all genes showed the clearest clustering by social position (p = 0.003), and revealed overall no significant difference in gene expression between P2 and S individuals, with P1 individuals exhibiting more distinct clustering (Fig. 3A; replicate‑annotated version of Fig. 3A in Supplementary Fig. S4; for all other PCAs see Supplementary Fig. S3).”

      Methods lines 754-765:

      “To test for overall GE patterns between social positions and genotypes, data were rlog-normalized and the effect of genotype and social position were compared through a Principal Component Analysis (PCA), followed by PERMANOVA analysis using Euclidean distances in the vegan package (Oksanen et al., 2017). To assess whether the strength of clustering by social position differed depending on the genes included, additional PCAs were performed on four gene subsets: (1) all genes, (2) the top 50%, (3) the top 10%, and (4) the top 5% of the most variable genes. For each subset, the first three principal components (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3) were visualized to identify axes capturing variation associated with social position. To test for overall gene expression differences between social position and genotype, separate PERMANOVAs were performed for each PCA using the adonis2 in the vegan package (Oksanen et al., 2017).”

      (3) The authors indicate that the 15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling. Does this mean that each of the P1and P2 were pairs with each other? Have the authors tried examining the gene expression patterns in a paired manner? E.g., for the pairs that showed the greatest size differences,do they also show the greatest differences in gene expression? Do the P1s show the most extreme differences from P2s that also show the most extreme P2 differences? Perhaps lines on Figure 3A connecting datapoints from the P1 and P2 pairs would be informative.

      Yes, “15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling” refers to pairs of P1 and P2, we made sure this is clearly stated in the revised Results (lines 266-269) and Methods (lines 686-688). Yes, we have explored gene expression data considering the size difference between pairs, and found that it showed no clear differences in gene expression patterns (see our response and manuscript edits above under Reviewer #1 point 2). We also added a version of the main text Fig. 3A which clearly labels replicates within the PCA as Supplementary Fig. S4.

      Results lines 266-269:

      “To test the prediction that whole-body gene expression patterns will vary with social roles once clear social positions have emerged (P1, P2 and S), we performed gene expression profiling on 45 individual fish (experimental groups containing paired individuals and corresponding solitaries; N = 15 samples per social position; Fig. 1D).”

      Methods line 686-688:

      “Fish from these 15 replicates (n=45 samples; 15 paired individuals and their corresponding solitaries) were individually homogenized to allow equal RNA extraction from all tissues.”

      (4) For the specific target pathways that are up- and downregulated in the different backgrounds, I recommend that the authors include boxplots (or heatmaps) showing the actual expression values for these targets. Figure 6 shows a heatmap for appetite-related genes, and it would be great to see a similar graph for the metabolism and glycolytic genes; it would also be informative to see similar graphs for hormonal and sexual maturation pathways as well.

      We have explored genes across a broad set of metabolic pathways (glycolysis, TCA cycle, lactic fermentation, PDH complex, cholesterol biosynthesis, fatty-acid synthesis, and beta-oxidation) and show all metabolic genes that showed significant differential expression between P1, P2, and S in Figure 6. Overall, very few metabolism-associated genes were significantly differentially expressed, which is why we decided to combine appetite-regulation and metabolism-associated genes into a single figure (Figure 6). In our revised version of the manuscript, we have modified Fig. 6 to clearly indicate which pathways the genes belong to and modified Methods section lines 795-797 and 802-805.

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      We also examined hormonal pathways (glucocorticoid and thyroid signaling), but did not find genes in these pathways that were significantly differentially expressed. Finally, we would like to clarify that our samples consist of two-month-old juvenile individuals that are sexually immature —under ideal conditions, clown anemonefish can mature in one to two years, but they can also remain sexually immature for a decade or more (Buston 2004; Buston & García, 2007) — which is why we did not observe distinct molecular signatures of sexual maturation. We recognize that the sentence at line 520 was misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. We have revised the sentence in the Discussion (used to be line 520, now line 543-544) as well as in the Methods section we added lines 601-606.

      Discussion line 543-544:

      “In our current study, individuals within pairs were ultimately progressing towards rank 1 dominant female and rank 2 subordinate male roles.”

      Methods lines 601-606:

      “All fish used in this experiment were sexually immature juveniles (one-month-old at the beginning, two-month-old at the end), as clown anemonefish reach sexual maturity between the ages of one and two years old. Also, under ideal conditions, individuals can remain sexually immature for a decade or more (Buston 2004b; Buston & García, 2007). Therefore, in this study, the emergence of social roles and associated phenotypes do not reflect changes associated with sexual maturation.”

      (5) Particularly given that there is a relatively small number of genes enriched in the different rank conditions, I did not understand the need to do the WGCNA module analysis. I thought that an analysis of GO terms across the dataset would have been more meaningful than the GO term analysis shown in Figure 4, which considers only genes assigned to the "brown WGCNA module". This should be simplified or clarified.

      To clarify, GO enrichment analysis does not establish correlations with traits, it only describes which functions or pathways are over-represented in a given gene set. That is why we began by using WGCNA to define gene sets (modules) that are correlated to phenotypes. Our primary rationale for WGCNA was to identify modules of co-expressed genes that show significant statistical correlation with the phenotypes of interest (social role: P1, P2, S; growth; and food intake). Pairwise differential expression analysis (Figure 3B) identified a few hundred significantly differentially expressed genes, but those tests treat genes independently and are not able to help us link coordinated changes of co-expressed genes to phenotypes of interest. Because WGCNA is blind to traits, it first identifies groups of co-expressed genes, which can help resolve gene expression patterns.

      We therefore ran WGCNA on the rlog-transformed dataset to identify modules of co-expressed genes that show significant correlation with phenotypes of interest. For every module that showed such a correlation, we performed GO enrichment and carefully evaluated the resulting GO enrichment trees (see Supplementary Figs. S6, S7). The brown module was highlighted in the main text because it was one of the modules with a significant correlation to growth, and its associated GO enrichment showed clear growth-related signals that were not identified in the pairwise differential expression analysis results.

      In our revised manuscript we have clarified the rationale for the analytical approaches in the Results section at lines 297-303 and Methods section lines 767-772 and 775-779.

      Results lines 297-303:

      “...a weighted gene co-expression network analysis (WGCNA; Langfelder & Horvath, 2008) was conducted on the rlog-transformed gene expression (GE) dataset. Unlike pairwise DEG analysis, which treats genes independently, WGCNA identifies groups of co-expressed genes (modules) and tests whether modules of eigengene expression are correlated with variation in phenotypes of interest. For modules showing significant correlations with phenotypes, gene ontology (GO) analyses were performed using Fisher`s exact tests to identify over-represented pathways within each module.”

      Methods lines 767-772 and 775-779:

      “To test whether gene expression patterns are correlated with observed phenotypic changes in growth and appetite across social positions, a Weighted Gene Correlation Network Analysis (WGCNA) was conducted on the rlog-transformed GE dataset (Langfelder & Horvath, 2008). WGCNA identifies groups of co-expressed genes (modules) and correlates eigengene expression of each module with phenotypes of interest to determine modules of genes whose expression correlates with trait variation.”

      “For modules whose eigengene expression correlated with phenotypes of interest, gene ontology (GO) enrichment analysis was subsequently performed using Fisher's exact tests (presence/absence in a module) to identify over-represented pathways within each module across GO divisions of Biological Processes (BP), Molecular Functions (MF), and Cellular Components (CC) (Wright et al., 2015).”

      (6) The authors say that they have identified coordinated changes in behaviors and the"underlying gene expression, leading to the emergence" of social roles. This is a little bit misleading, since the gene expression analysis occurred well after the behavioral and phenotypic differences emerged. Presumably, the hormonal and genetic shifts that actually caused the behavioral and phenotypic difference occurred during the weeks during which the experiment was underway, and earlier capture of the transcriptome would presumably reveal different patterns, and ones that would be considered more causative.The authors acknowledge this in 434-435, but it could be emphasized further.

      We appreciate the reviewer raising this point. In the updated version of the manuscript, we have revised wording to convey that food intake, agonistic behavior, size and growth, and gene expression are all changing continuously, in response to each other and in response to social feedback (Introduction lines 133-136). An underappreciated aspect of this system (and likely many other systems) is that phenotype (including transcriptome) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including the transcriptome). Earlier capture of the transcriptome would reveal different levels of gene expression, reflecting the state of the system at that moment in time.

      Introduction lines 134-137:

      “Underlying this cascade is an underappreciated aspect of this system (and likely many other systems) where phenotype (including gene expression) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including gene expression).”

      (7) The authors have measured a number of differences between the different dominance classes of fish. All these differences were measured relative to the other classes, but in my view, the Solitary group was the closest to a baseline control. So I'm not sure that it is fair to say that "P2 and S individuals showed consistent downregulation of these genes and pathways" (line 401). I encourage the authors to emphasize the differences in gene expression from the "perspective" of the P1 individuals compared to the baseline of P2and S individuals. Line 474 says that "P2 fish showed significant upregulation" of a number of pathways. It should be very clear what that is compared to (compared to P1, presumably?)

      We agree with the reviewer that solitary individuals are the most intuitive baseline. Indeed, the experimental design included solitary fish because we expected they would serve as a useful control. Without social restraint, we anticipated they would show unrestricted growth, feeding, behavior, and associated gene‑expression patterns, similar to dominants.

      We initially ran analyses using solitaries as the baseline, but after examining the results, which showed subordinate‑like characteristics for the solitary individuals, we concluded that solitary individuals are not an ecologically appropriate control for this context. Removing juveniles from a social context and housing them in isolation may be stressful and can affect physiology and behavior in ways that do not reflect a natural baseline. From a life‑history standpoint, solitary living is not the typical state for A. percula.

      For these reasons, we reanalysed the dataset using the dominant (P1) as the reference to enable more ecologically meaningful comparisons (this choice was somewhat arbitrary, subordinates could also have been used as the reference). Given that gene expression is relative, we interpret results from both the dominant (P1) and subordinate (P2) perspectives in the Discussion to provide a complete view. We have clarified wording throughout the manuscript to make it clear that everything is relative (Introduction lines: 159-162 and 167-169; Methods lines: 622-625 and 726-728) as well as revised language throughout to make sure comparisons are clear.

      Introduction lines 159-162 and 167-169:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      (8) Along the same lines, the authors say in line 514 that subordinates and solitaries strategically downregulate their growth. I'm not convinced that this is the case: I would consider this growth trajectory to be the default and the baseline. I would interpret that under certain social conditions, a P1 dominant pattern of growth, behavior, and gene expression is allowed to emerge.

      We respectfully disagree with the idea that a single baseline/reference growth trajectory exists for any individual of this species. Growth of individuals is entirely social context-dependent: neither fast nor slow growth represents an inherent baseline. When two size‑matched juveniles meet and compete to establish dominance, accelerated growth is the expected trajectory. By contrast, juveniles joining an existing hierarchy are expected to exhibit reduced growth, which minimizes conflict and facilitates their social integration. Unlike species that show non socially mediated growth trajectories, clown anemonefish do not have a context‑independent growth rate, rather, individuals constantly readjust their growth according to their immediate social environment (Buston 2003).

      Therefore, growth trajectories must be considered from the perspective of all group members, because they emerge from interactions among individuals rather than reflecting an intrinsic baseline. In this study, we were interested in the establishment of dominance hierarchy and how individuals adjust their phenotypes during this process. By experimentally pairing size‑matched rivals, both individuals are initially expected to pursue the dominant trajectory, and thus neither individual represents a default state. Instead, the outcome reflects a social decision, after which both individuals reinforce their emerging social roles through coordinated changes. See our response and manuscript edits above under Reviewer #2, point 7 (public reviews).

      Reviewer #3 (Public review):

      Summary:

      The authors tested the hypothesis that interactions among size- and age-matched rivals will lead to the emergence of social roles, accompanied by divergence in four aspects of individual phenotypes: growth, feeding behavior, fighting behaviors, and gene expression in clownfish.

      Strengths:

      The data on growth, feeding rate, and fighting behaviors support the authors' claims.

      Thank you for the positive feedback!

      Weaknesses:

      Gene analysis conducted in this study is not sufficient to clarify how the relevant genes actually regulate growth and behavior.

      The information obtained from whole-body gene expression analysis is very limited.Various gene expression is associated with the regulation of fighting behaviors, food intake, growth, and metabolism, and these genes are regulated differently across tissues, even within a single individual. Gene expression analysis should be performed separately for each tissue.

      We understand the reviewer’s concern about whole‑body transcriptomes and agree that tissue‑specific sampling would provide greater resolution of the mechanisms linking gene expression to growth, agonistic behaviors, and food intake. For this initial study, however, we deliberately chose whole‑body samples to capture a broad, unbiased view of gene expression differences while keeping sequencing costs and sample requirements manageable. We explicitly acknowledge the resulting interpretational limits in the Discussion (lines 497; 553–5647), and suggest in the last paragraph that the patterns reported here should be used to build on in future studies exploring targeted, tissue‑specific hypotheses. See our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      Clownfish undergo sex change depending on social status and body size, as the authors mention in the manuscript. Numerous gene expressions are affected by sex change. It is unclear how this issue was addressed.

      We thank the reviewer for raising this point. Sex change and sexual maturation can indeed drive major transcriptional shifts in clown anemonefish, but our experiment did not encompass such a life‑history transition. All individuals in this experiment were juveniles (≈1 month old at the start, ≈2 months old at the end) and were sexually immature at these ages. Clown anemonefish reach sexual maturation around one to two years under ideal conditions, can delay sexual maturation for years under normal conditions (Buston 2004; Buston & García, 2007), and sex change in the genus Amphiprion is known to take over ~5 months (Moyer & Nakazono, 1978). Accordingly, individuals in this study were not sexually mature, and sex change was not biologically plausible over the five-week experimental period of our study. We recognize that the sentence at line 520 may be misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. During the revisions we made sure that it is clearly stated that the fish in this study were sexually immature. See our response and manuscript edits above under Reviewer #2, point 4 (public reviews).

      Recommendations for the authors:

      Reviewing Editor Comments:

      While appreciating the work presented in the manuscript, we note a few common concerns that the authors could prioritise in their revisions:

      (1) The transcriptomics

      The authors have used whole-body RNA-seq and so are constrained in the kind of mechanistic or tissue-specific insight it can really provide. There are also questions about how far the authors can go in linking these data to causation, given that sampling happened after the phenotypes had already diverged.

      We therefore recommend tempering the causal language, clarifying the rationale for their analytical choices (WGCNA, fold-change threshold...), and more explicitly discussing the limitations of the whole-body RNA-seq and what can (and cannot) be concluded from it.

      We thank the editors for these recommendations. We have addressed each point as follows:

      Tempering causal language: We have revised our manuscript and tempered with language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. We made changes to our Abstract (lines 39-41), Introduction (lines 115-118, 121-123), and Discussion (lines 431-434, 571-574), and Methods sections (lines 795-797). Whereas other parts of the manuscript already used non-causal language such as Discussion lines 333, 355, 358, 381, 437, etc.

      Abstract lines 39-41:

      “Here we identify associations between changes in gene expression, growth, and feeding behavior regulation that reinforce social role differentiation during dominance hierarchy formation in clownfish.”

      Introduction lines 115-118, 121-123:

      “Gene expression profiling offers a powerful approach to uncover coordinated changes in gene expression across multiple pathways, providing insight into associations between the development of social role-specific phenotypes and their underlying proximate mechanisms.”

      “Its well-annotated genome (Lehmann et al., 2019) enables transcriptomic analyses that can help characterize the molecular patterns associated with these processes.”

      Discussion lines 431-434, 571-574:

      “To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      “Together, these approaches provide a powerful framework for uncovering the dynamic, context-dependent mechanisms that regulate strategic growth and coordinated changes associated with the establishment and maintenance of dominance hierarchies in social vertebrates.”

      Methods lines 795-797:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      WGCNA rationale: We have revised our Results and Methods sections to explicitly define why we used WGCNA and GO enrichment analysis, and how these analyses provided more information than simple pairwise differential gene expression analysis. See our response and manuscript edits above under Reviewer #2, point 5, (public reviews).

      Fold-change threshold: In our revised manuscript, we have removed the arbitrary imposed fold-change threshold cutoff, which was only used in that single figure. See our response and manuscript edits above under to Reviewer #1, point 4, (public reviews).

      Whole-body RNA-seq limitations: We have revised the Methods and Discussion sections, and see our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      (2) Framing and interpretation

      Several of the reviewers' comments flag that the manuscript overstates coordination or "strategic" regulation, or where what's being treated as the baseline/derived isn't clear (for example, whether P2/S are actively downregulating or whether P1 represents the divergent trajectory). We recommend revisiting any wording that implies stronger mechanistic inferences than the data actually support (and defining more clearly what is meant by baseline and socially induced state).

      We thank the editors for these comments and feedback. We have revised the manuscript to clearly state that our system does not have the traditional baseline/socially induced states, rather it is all social context-dependent. In particular, we have adjusted the wording throughout the manuscript to make it clear that everything is relative, and we clarified that our choice of P1 as the reference group was arbitrary. We have made revisions in the Introduction (lines 159-162 and 167-169) and Methods (lines 622-625 and 726-728). See our response and manuscript edits above under Reviewer #2, points 7 and 8, (public reviews).

      We have also revised the manuscript to temper with wording that could be interpreted as implying stronger mechanistic or directional conclusions than our data allow. We adjusted language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. See our response and manuscript edits above under Reviewing Editor Comments, (1) the transcriptomics (recommendations to authors).

      Reviewer #1 (Recommendations for the authors):

      (1) The reader would benefit from a brief overview of the physiology and molecular basis of somatic growth, regulation of food intake, and aggressive behavior, especially in teleost fishes. This would also help the authors with formulating hypotheses that are a bit more explicit when it comes to the transcriptomic part of the study.

      We thank the reviewer for this suggestion, however, as another reviewer raised concerns about the length of the Introduction as is and we felt that adding detailed paragraphs on the physiology and molecular basis of somatic growth, food intake regulation, and aggressive behavior would significantly increase the length of our Introduction.

      As a compromise, we have revised the relevant sections of the Introduction (lines 109-115) to point readers to key review papers and primary literature where the molecular basis of these traits is covered in detail. We believe this approach balances the need for mechanistic context with manuscript length constraints, while also allowing readers with specific interests to follow up with the relevant literature.

      Introduction lines 109-115:

      “Most studies, often conducted without relevant social context of an individual or in species entirely lacking dominance hierarchies, have examined single pathways to uncover variation in coloration (Salis et al., 2019, 2021), appetite regulation (see review for teleosts: Volkoff, 2019), behavior (Bender et al., 2006; Renn et al., 2008; Santema et al., 2013; Solomon-Lane et al., 2022; see review for teleosts: St-Cyr & Aubin-Horth, 2009), and growth (Beckman, 2011; Lu et al., 2020; see reviews for teleosts: Reinecke, 2010; Zhou et al., 2024), leaving the broader molecular shifts associated with social role adoption unresolved.”

      (2) I certainly agree with that statement that "Understanding the proximate mechanisms that facilitate phenotypic adjustments is key to disentangling whether phenotypes are the cause or consequence of social rank" (line 90f.), though maybe the authors can elaborate a bit, as this relationship is quite dynamic and obviously goes both ways.

      We agree that this relationship is dynamic and bidirectional, and we have revised the Introduction to reflect this more explicitly. In particular, we added text noting that, in social vertebrates, phenotype, including gene expression, can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship (Introduction lines 93-96). We also revised the experimental framing in the Introduction and Methods (Introduction lines 159-162 and 167-16972; Methods lines 622-625 and 726-728) to emphasize that phenotypes are socially context-dependent in A. percula and that the statistical reference group was chosen for analytical clarity rather than as a true biological baseline.

      Introduction lines 93-96, 159-162 and 167-169:

      “This is further complicated by the fact that, in social vertebrates, phenotype (including gene expression) can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship.”

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members. This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role; however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      (3) Much of the information in the last paragraph of the Introduction (lines 151ff.) is best presented in the Methods section.

      We respectfully disagree with this suggestion. The eLife journal format does not include a standalone Methods section preceding the Results, so we expect most readers not to consult the Methods before reading the Results. We believe it is important to briefly orient the reader to the experimental approach and the ideas tested, thus providing some context before they encounter the Results and Discussion.

      (4) What count (TPM or similar) and abundance (above count threshold in fraction of samples) thresholds were used for the transcriptome analysis? Maybe I missed it, but how many genes were in the analysis?

      We thank the reviewer for pointing this out. We have updated our Methods section (lines 714-719) to explicitly include filtering steps used as well as included the total number of genes that were used in downstream analysis.

      Methods lines 714-719:

      “The read count file was then imported into R version 4.3.1 (R Core Team, 2021), size factors were estimated for each sample using the median ratio method in DESeq2 (Love et al., 2014) to account for differences in sequencing depth across samples. No outlier samples were identified, and all samples (n=45) were retained for downstream analyses. Genes with a mean raw count <10 were then removed, retaining 24,840 genes for downstream analyses.”

      (5) Were all these genes used in PCA, and if so, why? Would it not make more sense to only use the 50% or 25% most variable genes (which would likely enhance the separation of social types)? Also, did the authors inspect higher-order PCs to see whether any of them separate the social types or separate samples according to some other variable(e.g., size, hue, feeding, any technical factors, etc.)?

      We explored PCAs using four gene subsets: all genes, the top 50%, top 10%, and top 5% most variable genes, and all pairwise PC comparisons (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3; Supplementary Fig. S3). Of all these examined PCAs, PC2 and PC3, with all genes, showed the clearest, statistically significant clustering by social position, and more stringent filtering did not strengthen this signal. We therefore retained all genes in the primary analysis to avoid imposing arbitrary filtering thresholds that could exclude biologically relevant low-variance genes. See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      Regarding the inspection of higher-order PCs for potential confounding variables, we examined whether genetic background (clutch ID), tank identity, and sequencing depth explained clustering patterns across PC1–PC3 and found no evidence that any of these variables (PCA for sequencing depth not shown). See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      We did not explicitly explore whether body size at week 5 drove any of the observed separation among social positions. Rather, we investigated whether size ratio, an indicator of the amount of remaining conflict within a pair, could drive gene expression variation within social positions. See our response and manuscript edits above under Reviewer #1, point 2, (public review).

      Regarding orange hue, we do not believe that orange hue at week 5 (the time point at which whole-body samples were collected for gene expression profiling) represents a meaningful socially mediated result. See our response and manuscript edits above under Reviewer #1, point 5, (public reviews).

      Considering food intake, we chose not to explore this as a potential driver of PC variation because food intake could not be reliably assigned to all individuals at week 5. For a subset of pairs, individuals could not be confidently distinguished in videos, and we did not want to make assumptions that could introduce biases into the analysis.

      (6) Figures 5/6: Based on the gene expression shown in the heatmaps, I cannot see how the samples would cluster so cleanly by social type. Consider a bootstrapping analysis and provide bootstrap values and "confident" nodes. Also, what does "matrix" in the legend refer to? I assume some measure of gene expression level, maybe z-scored?

      We thank the reviewer for flagging this point. We acknowledge that the samples do not cluster freely by social position in the heatmaps and we have revised our Methods (see our response and manuscript edits above under Reviewer #1, point 2, public reviews) and Figure captions to clearly reflect this. To clarify, samples in Figures 5 and 6 were not ordered by unsupervised hierarchical clustering; rather, they were first grouped a priori by social position and then ordered within each group by similarity. We chose this presentation because it best illustrates the gene expression patterns associated with social position, which is the primary focus of the manuscript. We have updated the figure legends of Figures 5 and 6, to clarify that the color scale previously denoted as matrix in the heatmap represents row-scaled Z-scores of normalized gene expression values.

      To assess whether social position explains a gene expression variation in an unsupervised framework, we performed a PCA followed by PERMANOVA (using the adonis2 function in R, 999 permutations, Euclidean distance; see Author response image 1). Both analyses used the same row Z-score-scaled expression values as shown in Figures 5 and 6. Results showed a significant effect of social position on gene expression (PERMANOVA: A: social position p <0.001, B: social position p < 0.001; marginal clutch ID significance p=0.015), which was primarily driven by dominant individuals (P1) being significantly different from both subordinate (P2) and solitary (S) individuals. To include these into our manuscript.

      Author response image 1.

      Principal component analysis (PCA) of gene expression based on heatmaps of A) growth, B) appetite, and metabolism genes. PCA was performed on gene sets of the main text (growth: Fig 5B and appetite and metabolism: Fig. 6), using row Z-score scaled expression values, consistent with the heatmap scaling. Each point represents one individual. Ellipses represent 95% confidence intervals around each social position group. PERMANOVA results (adonis2; shown in the bottom right corners).

      (7) I applaud the authors for considering genetic/relatedness effects in their experimental design and analysis, but I am confused by the microsatellite vs. SNP analyses: why even use microsatellites in this day and age? Were only the SNP data used as the source of genetic information? This should be clarified.

      In our experiment, paired individuals were initially assigned a temporary rank based on their size at the start of the experiment; however, the initially larger individual (often only by a tenth of a mm) does not necessarily emerge as the dominant (P1). To avoid any assumptions regarding the identity of individuals in given social positions at the end of the experiment, we needed to verify that individuals within pairs were correctly identified. For the 45 individuals included in the gene expression dataset, this was done by calling SNPs directly from the TagSeq data. For the remaining 36 individuals not included in the gene expression dataset, we opted for microsatellite genotyping, as a validated panel with established markers was already available from a previous study (Rueger et al., 2025), making it a reliable and cost-effective solution. We have clarified the text in the Methods (lines 630-636) and Supplementary Materials (lines 42-49 and 65-67).

      Methods lines 630-636:

      “This step was necessary to avoid assumptions regarding social roles and fish identity. At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment using either microsatellite genotyping or SNPs called from TagSeq data, depending on whether individuals were included in the gene expression dataset (see Supplementary Materials).”

      Supplementary Materials lines 42-49 and 65-67:

      “This step was necessary to avoid assumptions regarding social roles and fish identity, and the microsatellite method allowed for a reliable and cost-effective solution as a panel with established markers was already available from a previous study (Rueger et al., 2025). At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1 mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment.”

      “Similarly, for the remaining 45 individuals, we applied a necessary correction step to avoid assumptions regarding social role and fish identity. For these 45 individuals, we called SNPs from our TagSeq data to assign clutch identity.”

      Reviewer #2 (Recommendations for the authors):

      (1) Line 520 indicates that individuals showed early gene expression signatures of sexual maturation, but I did not see where those results were presented.

      We have addressed this recommendation, the claim “individuals showed early gene expression signatures of sexual maturation” was incorrect and has been removed from the revised manuscript (Discussion lines 543-544). We have also updated the Methods section (lines 601-606) to explicitly clarify that all individuals were sexually immature juveniles throughout the experiment. See our response and manuscript edits above under Reviewer #2, point 4, (public reviews).

      (2) The paragraph starting at line 409 refers to GE profiles. I was confused about what that was.

      Do the authors mean GO profiles?

      We did not find any modules using GO enrichment analysis which showed strong enrichment of appetite- and metabolism-related GO terms. Therefore, we looked for differentially expressed genes in the rlog-normalized gene expression dataset that are known to be associated with appetite regulation and metabolism. We have revised the Discussion (lines 431-434) and Methods (lines 812-815) to explicitly state what we are referring to and avoid confusion.

      Discussion lines 431-434:

      “In our study, P1 individuals showed increased food intake compared to P2 individuals. To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      Method lines 802-805:

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (3) I saw several typos, e.g., Vulcano in Supplementary Figure 3.

      Thank you for flagging this. We have addressed typos such as Supplementary Fig. S4 (used to be Supplementary Fig S3) See our response and manuscript edits above under Reviewer #1, point 4 (public reviews).

      (4) The personal observations cited in 495 should be more explicit. Which author made these observations, over how long, in how many instances?

      We thank the reviewer for this comment. We have updated the text to explicitly name the authors who made these observations, to clarify that they were made during two independent long-term field studies, and to note that this pattern was observed in 8 or more instances across the two field studies (Discussion lines 516-520).

      Discussion lines 516-520:

      “Similar patterns have been anecdotally observed in the wild, where juvenile clownfish remained small for extended periods (over four months) following the loss of a dominant partner, only initiating changes in social role towards dominant characteristics upon the arrival of a new group member (personal observations of 8+ instances during two long-term independent studies by Pete Buston and Lili Vizer).”

      References:

      Buston, P. (2003). Forcible eviction and prevention of recruitment in the clown anemonefish. Behavioral Ecology, 14(4), 576–582. https://doi.org/10.1093/beheco/arg036

      Buston, Peter M. (2004). Territory inheritance in clownfish. Proceedings of the Royal Society B: Biological Sciences, 271(SUPPL. 4), 252–254.

      Buston, P. M., & García, M. B. (2007). An extraordinary life span estimate for the clown anemonefish Amphiprion percula. Journal of Fish Biology, 70(6), 1710–1719. https://doi.org/10.1111/j.1095-8649.2007.01445.x

      Buston, P., & Clutton-Brock, Tim. (2022). Strategic growth in social vertebrates (WITH REVIEWER COMMENTS). Trends in Ecology & Evolution, 37(8), 694–705. https://doi.org/10.1016/j.tree.2022.03.010

      Dengler-Crish, C. M., & Catania, K. C. (2007). Phenotypic plasticity in female naked mole-rats after removal from reproductive suppression. THE JOURNAL OF EXPERIMENTAL BIOLOGY.

      Heg, D, Bender, N, & Hamilton, I. (2004). Strategic growth decisions in helper cichlids. Proceedings of the Royal Society of London. Series B: Biological Sciences, 271(suppl_6). https://doi.org/10.1098/rsbl.2004.0232

      Huchard, E, English, S, Bell, M B. V., Thavarajah, N, & Clutton-Brock, T. (2016). Competitive growth in a cooperative mammal. Nature, 533(7604), 532–534. https://doi.org/10.1038/nature17986

      Johnston, R A., Vullioud, P, Thorley, J, Kirveslahti, H., Shen, L., Mukherjee, S., Karner, C. M., Clutton-Brock, T, & Tung, J (2021). Morphological and genomic shifts in mole-rat ‘queens’ increase fecundity but reduce skeletal integrity. eLife, 10, e65760. https://doi.org/10.7554/eLife.65760

      Maytin, Alexander K., Davies, Sarah W., Smith, Gabriella E., Mullen, Sean P., & Buston, Peter M. (2018). De novo transcriptome assembly of the clown anemonefish (Amphiprion percula): A new resource to study the evolution of fish color. Frontiers in Marine Science, 5(AUG), 1–11.

      Moyer, J. T., & Nakazono, A. (1978). Protandrous Hermaphroditism in Six Species of the Anemonefish Genus Amphiprion in Japan (No. 2). The Ichthyological Society of Japan. https://doi.org/10.11369/jji1950.25.101

      Reed, C., Branconi, R., Majoris, J., Johnson, C., & Buston, P. (2019). Competitive growth in a social fish. Biology Letters, 15(2), 20180737. https://doi.org/10.1098/rsbl.2018.0737

      Rueger, Theresa, Bhardwaj, Anjali Kristina, Turner, Emily, Barbasch, Tina Adria, Trumble, Isabela, Dent, Brianne, & Buston, Peter Michael. (2022). Vertebrate growth plasticity in response to variation in a mutualistic interaction. Scientific Reports, 12(1), 11238.

      Thorley, J, Katlein, N, Goddard, K, Zöttl, M, & Clutton-Brock, T. (2018). Reproduction triggers adaptive increases in body size in female mole-rats. Proceedings of the Royal Society B: Biological Sciences, 285(1880), 20180897. https://doi.org/10.1098/rspb.2018.0897

      Van Schaik, C P., & Van Hooff, J A. R. A. M. (1996). Toward an understanding of the orangutan’s social system. In Linda F. Marchant, Toshisada Nishida, & William C. McGrew (Eds.), Great Ape Societies (pp. 3–15). Cambridge University Press. https://doi.org/10.1017/CBO9780511752414.003

      Walker, S P. W., & McCormick, M I. (2009). Sexual selection explains sex-specific growth plasticity and positive allometry for sexual size dimorphism in a reef fish. Proceedings of the Royal Society B: Biological Sciences, 276(1671), 3335–3343. https://doi.org/10.1098/rspb.2009.0767

      Wong, M. Y. L., Buston, P. M., Munday, Philip L., & Jones, Geoffrey P. (2007). The threat of punishment enforces peaceful cooperation and stabilizes queues in a coral-reef fish. Proceedings of the Royal Society B: Biological Sciences, 274(1613), 1093–1099. https://doi.org/10.1098/rspb.2006.0284

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) This article purports to show that ML-SA8, a synthetic activator of the lysosomal TRPML1 channel, results in AMPK activation and glucose uptake in hepatocytes, and that this action has therapeutic potential for metabolic disease. The final figure shows that glucose levels are improved in db/db mice, although it is not entirely clear whether this is due to an effect on the liver, on other tissues, or on glucose production or uptake. The earlier figures try to make the case that SA8 causes activation and GLUT4 translocation and glucose uptake in liver cells; however, these data are not convincing. GLUT4 is expressed at such low levels in liver that it is likely not physiologically important. The authors use a fluorescent glucose analog to measure glucose uptake, and this molecule has been shown to enter cells largely by fluid phase endocytosis. Overall, this reviewer finds the premise misguided and the data unconvincing.

      Thank you for the critical comments and constructive suggestions. We have carefully considered all the concerns raised and provide our point-by-point responses below.

      (2) The initial figures show phosphorylation of AMPK on Thr172, but no downstream effects are shown. Usually, to convincingly show that AMPK activity is increased, it would be appropriate to immunoblot phospho-ACC or some other substrate. This is minor.

      Suggestion was taken! We will investigate the effect of SA8 on AMPK downstream effectors i.e. ACC activation by Western blotting. Ie p-ACC (Ser79) / total ACC.

      (3) Lines 135-148: GLUT4 is not expressed at levels that are significant for physiology in liver cells, and its function in liver is not particularly relevant. The authors cite references 38-40 to support that it may be expressed at low levels in liver, but no knockout studies have been done to show that this expression is physiologically important.

      We thank the reviewer for this critical comment. We agree that GLUT4 is not the predominant hepatic glucose transporter, but it is expressed at a relatively low level in the liver compared to other tissues.

      Regarding the physiological role of GLUT4, it has been shown to mediate glucose uptake in hepatic stellate and sinusoidal endothelial cells (Tang and Chen 2010, Karim, Liaskou et al. 2014). Furthermore, ischemia‑reperfusion (IR) significantly upregulated the expression of GLUT4 in the liver, rather than GLUT2, and the increased GLUT4 localized to the membrane peripheries of hepatocytes and enhanced glucose uptake, which in turn led to marked glycogen deposition (Kim, Jung et al. 2014, Kurabayashi, Furihata et al. 2022). These studies establish a clear physiological role for GLUT4 in the liver.

      Notably, Ranalletta, et. al. (2005) showed that GLUT4-null mice exhibit compensatory alterations in hepatic glucose and lipid metabolism, including increased hepatic glucose uptake and triglycerides conversion (Ranalletta, Jiang et al. 2005), indicating that GLUT4 ablation influences liver metabolism. Nevertheless, these existing evidences including the GLUT4 expression data and the functional changes observed in GLUT4 null mice—supports the relevance of GLUT4 in hepatic glucose metabolism. However, it is necessary to perform the liver-specific GLUT4 KO studies to clarify the role of GLUT4 in the liver. (We will incorporate these points and limitations into the revised Discussion section).

      (4) Figure 1e is not convincing. No controls are included to show the specificity of the antibody for immunofluorescent staining. No intracellular GLUT4 is visible in the unstimulated samples.

      We understood the reviewer’s concern about the specificity of GLUT4 antibody for immunofluorescent staining. The GLUT4 antibody (Abcam, ab33780) employed in our study has been extensively validated in previous studies for both Western blotting (Xie, Liu et al. 2024, Amanollahi, Holman et al. 2025, Ando, Takeda et al. 2025) and immunofluorescence (see also Johansson, Mannerås-Holm et al. 2013, Xiao, Zhang et al. 2025). in addition, we confirmed its specificity in our system by Western blot (Suppl. Fig.4), which showed a single band at ~45–55 kDa. Collectively, the combination of published validations, our own data supports the specificity of GLUT4 immunofluorescence detection.

      About the intracellular GLUT4 signal in Fig 1e. We apologize for the unclear GLUT4 signal in our original Fig. 1e. This was due to an inadvertently short exposure for the control condition. To improve this, we have now acquired new images with uniformly increased exposure time for all groups. As shown in the new Fig. 1e, GLUT4 is now clearly detected in control cells and mainly in cytosol, and ML-SA8 treatment obviously increases its accumulation at the plasma membrane.

      (5) In Figure 1f, again, the data are not convincing. The bands seem too sharp for GLUT4, which has 12 membrane-spanning domains as well as an N-linked glycosylation, so that it usually runs as a smear.

      We understood the reviewer’s concern. As an N-glycosylated membrane protein, GLUT4 may exhibit broader or diffuse migration patterns on immunoblotting. Nevertheless, the final band pattern is influenced by multiple factors, e.g. antibody specificity, sample preparation, electrophoresis conditions, and detection condition. Of note, multiple independent studies have shown endogenous GLUT4 as a relatively “sharp” immunoreactive band around 55–60 kDa (Gurley, Ilkayeva et al. 2016, Habtemichael, Li et al. 2021, Wu, Yu et al. 2024) (see also Ando et al., 2025; Amanollahi et al., 2025; Xie et al., 2024), which is very consistent with our results. Thus, we are confident that our GLUT4 band is specific and reliable.

      (6) Figure 1h. Data are not convincing. 2-NBDG is not a valid approach to measure glucose uptake. 2-NBDG enters cells largely via fluid phase endocytosis, and its accumulation is independent of known GLUT inhibitors such as cytochalasin B (Yazdani et al., MBoC 2022; PMID: 35921166; see also PMID: 42287154). The idea that such a bulky derivative of glucose could enter the transporter channel is not compatible with known structural data.

      We thank the reviewer for raising this point. We acknowledge that the uptake mechanism of 2-NBDG is controversial. As the reviewer noted, some studies have reported that it enters cells largely by endocytosis in certain cell types, and this remains a subject of ongoing discussion. Nevertheless, 2-NBDG continues to be widely employed as a glucose uptake tracer in this field, including several recent high-profile studies (Nobs, Kolodziejczyk et al. 2023, Xiong, Helm et al. 2023, Wu, Lv et al. 2025) as following.

      (1) Nobs, S.P., et al., Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes. Nature, 2023. 624(7992): p. 645-652.

      (2) Xiong, L., et al., Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis. Nat Immunol, 2023. 24(10): p. 1671-1684.

      (3) Wu, Y., et al., Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion. Phytomedicine, 2025. 149: p. 157508.

      In the current study, given the consistency of our results with parallel functional assays, the use of 2‑NBDG is justified in this context.

      (7) Supplementary Figure 5 uses 2-NBDG glucose uptake again. This reviewer is not convinced that the data reflect transporter-mediated glucose uptake, as suggested by the authors.

      Please see response to #6.

      (8) As well, although palmitate treatment of cells can cause an insulin-resistant-like phenotype in some cell types, this is not characterized in the present work.

      We understand the reviewer’s concern regarding the characterization of PA-induced insulin resistance model.

      First, this PA-induced insulin resistant hepatic model is well-established and validated in several literatures (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024), and it has been used for T2DM natural and synthetic drug screening (Faria, Calixto et al. 2025).

      Second, we have characterized this model in our system. As shown in Suppl. Fig. 5a, b, insulin (100 nM, 0.5 h) induced an increase of glucose uptake in HepG2 cells measured by 2-NBDG (Yamada, Nakata et al. 2000). In contrast, in PA-treated HepG2 cells, this effect was almost completely blocked, suggesting that PA-treated HepG2 cells are less sensitive to insulin. Overall, this model has been well validated and is suitable for the purposes of our study.

      (9) Finally, as noted, one would not expect hepatocytes to exhibit insulin-responsive glucose transport. Glycogen synthesis is the main insulin-regulated step that might be affected.

      As stated in the Responses#9, PA-induced hepatic insulin-resistance has become a widely accepted in vitro model for investigating therapeutic strategies (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024, Faria, Calixto et al. 2025).

      (10) The data in Figures 2b,c,f,g,k,l are not convincing. Again, 2-NBDG is used.

      Please see response to #6.

      (11) For the glucose consumption measurements in other panels of Figure 2, the methods section states that cells were cultured in 10 mM glucose. What volume was used? It is difficult to believe that a monolayer of cells would consume very much of the glucose that is present in the culture medium. Data are shown as a percent of controls, and look reasonable, but it would be helpful to include absolute as well as relative units.

      We appreciate the reviewer’s comments and apologize for any confusion about the method.

      First, we have revised the methods section to clarify the assay procedure “Following the manufacturer’s protocol, 2.5 μL of sample (medium/standard) was mixed with 250 μL of working solution a 96-well plate. The mixture was then incubated at 37 °C for 10 min and the absorbance was measured…” (line 421-423).

      Second, in response to the suggestion to include both absolute and relative units, we will provide the data with absolute value for reviewer’s reference. In the main figures, we have retained the normalized data as this format allows direct comparison of treatment effects across independent experiments.

      (12) In Figure 2, in experiments using the TRPML1 KO cells, no panel is shown to demonstrate knockout. The authors cite a previous paper for the construction of these cells, but the control immunoblot should still be shown here.

      We thank the reviewer for raising this point. The TRPML1 knockout cell line used in our study was originally generated and provided by Prof. Haoxing Xu’s laboratory, this cell line has been validated in several published literature (Wang, Gao et al. 2015, Zhang, Cheng et al. 2016). We understand the reviewer’s concern, so we will further validate this TRPML1 KO cell line.

      (13) In Figure 3, controls are missing in the BAPTA experiment in Figure 3a (only SA8-treated cells were treated with BAPTA and with EGTA). Again, it would be helpful to have p-ACC or some other readout of AMPK activity, and not just AMPK phosphorylation. 2NBDG is again used in this figure.

      Suggestion taken! We will add the controls including BAPTA-AM and EGTA only data. p-ACC/total ACC will also be measured. About the 2-NBDG, please see Responses#6.

      (14) Line 212-213 the text states "considering our finding that TRPML1-mediated Ca2+ release is essential for AMPK activation." This has not been shown. The work uses chelators and does not necessarily indicate a role for TRPML1. The drug may be specific, as suggested by the authors, but the way this phrase is worded is too strong. As well, AMPK was shown to be phosphorylated, but full activation towards its various substrates has not been shown.

      We thank the reviewer for raising this concern. We fully agree that the Ca<sup>2+</sup> chelators experiment alone could not specially attribute the effect to TRPML1.

      In fact, we have performed experiment to address the TRPML1-dependent mechanism in the original submission. As shown in Fig. 2i-l, In TRPML1 KO HAP1 cells (Qi, Xing et al. 2021), ML-SA8-induced AMPK phosphorylation and cellular glucose uptake were almost completely abolished compared to wild-type (WT) HAP1 cells. Moreover, pharmacological inhibition of TRPML1 with a TRPML1 specific synthetic inhibitor-ML-SI5 completely abolished ML-SA8-triggered AMPK phosphorylation (Fig. 2d, e). In addition, in IR-HepG2 model, ML-SI5 could substantially inhibited ML-SA8-induced cellular glucose uptake (Fig. 2f-h).

      Accordingly, we have also revised the statement to” considering our finding that TRPML1-mediated Ca<sup>2+</sup> release is necessary for AMPK activation” to make the sentence more rigorous.

      (15) Figure 4cd suggests that GLUT4 expression is increased by 2 or 3-fold in the liver of DB+SA8-treated mice, compared to controls. This may be the case, but its abundance is still likely ~1000-fold less in liver compared to skeletal muscle or adipose tissue. This reviewer is still not convinced that this is physiologically relevant. The images in Supplementary Figure 8 suggest a larger increase, but it remains uncertain whether the staining really represents GLUT4.

      We understood the reviewer’s concern. About the physiological importance of GLUT4, please see Responses #3. About the specificity of GLUT4 antibody, please see Responses#4.

      (16) Data showing that blood glucose and HbA1c are reduced in SA8-treated mice are reasonable, and GTTs and ITTs are shown. Unfortunately, there are no insulin concentrations, and it remains uncertain whether glucose production is reduced or uptake is increased (or if both effects are present).

      We will measure the insulin concentrations.

      (17) In the discussion, the authors again state that GLUT4 is present in the liver and that it regulates hepatic glucose homeostasis, and they cite reference 63. This review article does not argue that GLUT4 acts in the liver to regulate hepatic glucose homeostasis, but that its actions in muscle and fat have secondary effects on the liver.

      Sorry for the oversights. We have supplementary more precise references (Rossetti, Stenbit et al. 1997, Kurabayashi, Furihata et al. 2022, Fan, Jiao et al. 2023, Jiang, Luo et al. 2024).

      Reviewer #2 (Public review):

      The manuscript contains interesting studies suggesting that pharmacological activation of TRPML1 could be useful to treat T2D by increasing glucose uptake via activation of AMPK. Preclinical studies suggest the inhibitor improved blood glucose in Db/Db mice. Ex vivo studies in cell lines examine both pharmacologic and genetic manipulations, both to activate and to inactivate TRPML1, and the results consistently suggest that TRPML1 activates AMPK and increases glucose uptake.

      Strengths:

      The manuscript is well written, and the studies are carefully performed.

      (18) All mechanistic studies were performed in transformed cell lines; conclusions would be stronger if performed in primary cells. The in vivo studies were only performed in male mice. Performing metabolic studies in both sexes is standard practice now. Whether the findings would extend to females was not tested and remains uncertain. Some controls are missing, such as plasma membrane loading controls for fractionation studies. The GLUT4 staining was performed after fixation and permeabilization, yet control cells appear to be devoid of intracellular (and all) staining, a confusing result that doesn't reflect the expected biology.

      Thank you for your support and the constructive suggestions! We will answer these questions in the following point-to-point responses.

      Reviewer #3 (Public review):

      (19) Zhu et al. present a proof-of-concept for targeting the lysosomal calcium channel MCOLN1/TRPML1endolysosomal ion channels to restore type 2 diabetes mellitus (T2DM). Using synthetic TRPML1 agonists (ML-SA8) and genetic manipulation, the authors demonstrate that TRPML1 stimulation triggers localized lysosomal calcium release. This calcium efflux sequentially activates CaMKKβ and phosphorylates AMPK at Thr172 in various cell models, including palmitic acid-induced insulin-resistant HepG2 cells. This signaling pathway promotes GLUT4 translocation to the plasma membrane and increases intracellular glucose uptake. When administered daily to diabetic db/db mice over six weeks, ML-SA8 lowers fasting and random blood glucose, improves oral glucose and insulin tolerance tests, reduces hepatic steatosis, and lowers serum ALT and AST levels.

      Strengths:

      Based on the TFEB-independent pathway activated by TRPML1 and the experimental approaches described by Medina's group (PMID: 31822666), the authors use a combination of pharmacological and genetic tools to dissect such an intracellular signaling pathway. Additionally, the animal experiments show consistent phenotypic improvements across independent metabolic parameters. The ability of ML-SA8 to restore glycogen deposition and clear hepatic lipid accumulation in db/db mice without causing weight loss or overt toxicity provides a strong rationale for exploring lysosomal targets in metabolic disease.

      Thank you for the support!

      (21) The authors focus almost exclusively on hepatic GLUT4 to explain the observed glucose disposal. However, other glucose transporter isoforms such as GLUT2 dominate basal glucose transport. While the authors show increased AMPK phosphorylation in skeletal muscle and adipose tissue, they do not measure GLUT4 translocation or glucose uptake in these primary disposal organs. As a result, attributing systemic glycemic recovery primarily to hepatic GLUT4 translocation overlooks the major physiological roles of peripheral tissues.

      We thank the reviewer for this critical comment and we agree with the reviewer. Our initial focus on hepatic GLUT4 was driven by our primary interest in liver metabolism and the fact that we observed a consistent and robust effect of ML‑SA8 on hepatic GLUT4 translocation, AMPK activation and glucose uptake. We acknowledge that the possible contribution of GLUT2 to these effects cannot be excluded. Also, ML‑SA8 may exert similar effects in other tissues, such as skeletal muscle and adipose tissue as we found out that GLUT4 levels were upregulated in these tissues (Suppl.Fig.8b, c). We therefore agree that the systemic glycemic recovery should be attributed to the multi‑tissue effects of ML‑SA8, rather than a liver-restricted phenomenon. Accordingly, we will revise the Discussion section to incorporate this. Note that the precise mechanisms of ML-SAs on GLUT2 and in peripheral tissues warrant further investigation.

      (22) In both HepG2 cells and mouse liver tissues, ML-SA8 treatment increases total GLUT4 protein expression in addition to plasma membrane localization. Because total protein pools expand, the enrichment of GLUT4 in plasma membrane fractions cannot be cleanly attributed to acute vesicular translocation alone. The manuscript does not explain the timescale or mechanism behind this rapid total protein upregulation, leaving a mechanistic gap between acute ion channel gating and protein expression.

      Great suggestion. we will re-calculated the enrichment of GLUT4 in PM with total GLUT4 protein to clarify.

      (23) While the in vitro specificity of ML-SA8 is well-controlled, the systemic animal experiments lack a specific rescue or knockout control. Small-molecule agonists administered intraperitoneally over six weeks can exert off-target effects. Without demonstrating that co-administering the TRPML1 inhibitor ML-SI5 blunts the therapeutic effect in vivo, or showing that ML-SA8 lacks efficacy in TRPML1-null mice, the definitive link between in vivo glycemic recovery and TRPML1 activation remains incomplete.

      We understand the reviewer’s concern about the specificity of ML-SAs for TRPML1. In fact, the specificities of ML-SAs and ML-SIs have been rigorously validated in previous studies using TRPML1 knockout (KO) cells (Sahoo, Gu et al. 2017, Yu, Zhang et al. 2020) and further confirmed in the atomic-resolution co-structures (Schmiege, Fine et al. 2017, Schmiege, Fine et al. 2021).

      In the current study we provide additional evidence supporting this specificity. We show that ML-SA8-induced AMPK phosphorylation and glucose uptake were abolished in TRPML1 knockout cells (Fig. 2i-l) and blocked by specific inhibitor ML-SI5 (Fig. 2d, e, f, g). In addition, ML‑SA5 has been show to lack efficacy in TRPML1‑null mice in several studies (Yu, Zhang et al. 2020, Zhang, Wang et al. 2024, Xing, Wang et al. 2025), further reinforcing the target specificity of this class of compounds. Hence, multiple lines of evidence support the on-target effect of ML-SA8: (1) in vitro blockade by ML-SI5 and TRPML1 KO; (2) consistent effects across structurally distinct TRPML1 agonists; (3) dose‑dependent responses in vivo; and (4) absence of efficacy in TRPML1‑null mice.

      We agree with the reviewer that rescue experiment or in vivo knockout validation would provide the most definitive proof of target specificity. However, breeding TRPML1‑null mice or performing extensive dose-finding studies for ML-SI5 in vivo would require substantial time and resources which probably fall beyond the scope of the current study. We have therefore acknowledge this limitation in the Discussion section and noted that future validation with ML-SI5 co-administration or TRPML1-null mice is warranted. We hope the reviewer finds our response acceptable.

      Amanollahi, R., S. L. Holman, A. S. Meakin, M. Padhee, K. J. Botting-Lawford, S. Zhang, S. M. MacLaughlin, D. O. Kleemann, S. K. Walker, J. M. Kelly, S. R. Rudiger, I. C. McMillen, M. D. Wiese, M. C. Lock and J. L. Morrison (2025). "In Vitro Embryo Culture Impacts Heart Mitochondria in Male Adolescent Sheep." J Dev Biol 13(2).

      Ando, T., R. Takeda, R. Kano, T. Kusano, Y. Nonaka, Y. Kano and D. Hoshino (2025). "Effects of pyruvate administration on mRNA expression of inflammatory cytokines in adipose tissue and whole-body glucose metabolism in male mice." Physiol Rep 13(15): e70362.

      Casimiro, I., N. D. Stull, S. A. Tersey and R. G. Mirmira (2021). "Phenotypic sexual dimorphism in response to dietary fat manipulation in C57BL/6J mice." J Diabetes Complications 35(2): 107795.

      Fan, X., G. Jiao, T. Pang, T. Wen, Z. He, J. Han, F. Zhang and W. Chen (2023). "Ameliorative effects of mangiferin derivative TPX on insulin resistance via PI3K/AKT and AMPK signaling pathways in human HepG2 and HL-7702 hepatocytes." Phytomedicine 114: 154740.

      Faria, B. Q., P. S. Calixto, G. Picheth, L. M. Ferreira, F. G. M. Rego, J. F. C. Guerra and M. H. M. Sari (2025). "Palmitate-induced hepatic insulin resistance as an in vitro model for natural and synthetic drug screening: A scoping review of therapeutic candidates and mechanisms." Chem Biol Interact 420: 111717.

      Gurley, J. M., O. Ilkayeva, R. M. Jackson, B. A. Griesel, P. White, S. Matsuzaki, R. Qaisar, H. Van Remmen, K. M. Humphries, C. B. Newgard and A. L. Olson (2016). "Enhanced GLUT4-Dependent Glucose Transport Relieves Nutrient Stress in Obese Mice Through Changes in Lipid and Amino Acid Metabolism." Diabetes 65(12): 3585-3597.

      Habtemichael, E. N., D. T. Li, J. P. Camporez, X. O. Westergaard, C. I. Sales, X. Liu, F. López-Giráldez, S. G. DeVries, H. Li, D. M. Ruiz, K. Y. Wang, B. S. Sayal, S. González Zapata, P. Dann, S. N. Brown, S. Hirabara, D. F. Vatner, L. Goedeke, W. Philbrick, G. I. Shulman and J. S. Bogan (2021). "Insulin-stimulated endoproteolytic TUG cleavage links energy expenditure with glucose uptake." Nat Metab 3(3): 378-393.

      Jiang, Y., P. Luo, Y. Cao, D. Peng, S. Huo, J. Guo, M. Wang, W. Shi, C. Zhang, S. Li, L. Lin and J. Lv (2024). "The role of STAT3/VAV3 in glucolipid metabolism during the development of HFD-induced MAFLD." Int J Biol Sci 20(6): 2027-2043.

      Johansson, J., L. Mannerås-Holm, R. Shao, A. Olsson, M. Lönn, H. Billig and E. Stener-Victorin (2013). "Electrical vs manual acupuncture stimulation in a rat model of polycystic ovary syndrome: different effects on muscle and fat tissue insulin signaling." PLoS One 8(1): e54357.

      Karim, S., E. Liaskou, J. Fear, A. Garg, G. Reynolds, L. Claridge, D. H. Adams, P. N. Newsome and P. F. Lalor (2014). "Dysregulated hepatic expression of glucose transporters in chronic disease: contribution of semicarbazide-sensitive amine oxidase to hepatic glucose uptake." Am J Physiol Gastrointest Liver Physiol 307(12): G1180-1190.

      Kim, S., J. Jung, H. Kim, R. W. Heo, C. O. Yi, J. E. Lee, B. T. Jeon, W. H. Kim, J. R. Hahm and G. S. Roh (2014). "Exendin-4 Improves Nonalcoholic Fatty Liver Disease by Regulating Glucose Transporter 4 Expression in ob/ob Mice." Korean J Physiol Pharmacol 18(4): 333-339.

      Kim, S. J., A. Gajbhiye, A. R. Lyu, T. H. Kim, S. A. Shin, H. C. Kwon, Y. H. Park and M. J. Park (2023). "Sex differences in hearing impairment due to diet-induced obesity in CBA/Ca mice." Biol Sex Differ 14(1): 10.

      Kurabayashi, A., K. Furihata, W. Iwashita, C. Tanaka, H. Fukuhara, K. Inoue, M. Furihata and Y. Kakinuma (2022). "Murine remote ischemic preconditioning upregulates preferentially hepatic glucose transporter-4 via its plasma membrane translocation, leading to accumulating glycogen in the liver." Life Sci 290: 120261.

      Lee, J. Y., H. K. Cho and Y. H. Kwon (2010). "Palmitate induces insulin resistance without significant intracellular triglyceride accumulation in HepG2 cells." Metabolism 59(7): 927-934.

      Luo, J., H. Alkhalidy, Z. Jia and D. Liu (2024). "Sulforaphane Ameliorates High-Fat-Diet-Induced Metabolic Abnormalities in Young and Middle-Aged Obese Male Mice." Foods 13(7).

      Malik, S., S. Inamdar, J. Acharya, P. Goel and S. Ghaskadbi (2024). "Characterization of palmitic acid toxicity induced insulin resistance in HepG2 cells." Toxicol In Vitro 97: 105802.

      Nguyen-Phuong, T., S. Seo, B. K. Cho, J. H. Lee, J. Jang and C. G. Park (2023). "Determination of progressive stages of type 2 diabetes in a 45% high-fat diet-fed C57BL/6J mouse model is achieved by utilizing both fasting blood glucose levels and a 2-hour oral glucose tolerance test." PLoS One 18(11): e0293888.

      Nobs, S. P., A. A. Kolodziejczyk, L. Adler, N. Horesh, C. Botscharnikow, E. Herzog, G. Mohapatra, S. Hejndorf, R. J. Hodgetts, I. Spivak, L. Schorr, L. Fluhr, D. Kviatcovsky, A. Zacharia, S. Njuki, D. Barasch, N. Stettner, M. Dori-Bachash, A. Harmelin, A. Brandis, T. Mehlman, A. Erez, Y. He, S. Ferrini, J. Puschhof, H. Shapiro, M. Kopf, A. Moussaieff, S. K. Abdeen and E. Elinav (2023). "Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes." Nature 624(7992): 645-652.

      Qi, J., Y. Xing, Y. Liu, M. M. Wang, X. Wei, Z. Sui, L. Ding, Y. Zhang, C. Lu, Y. H. Fei, N. Liu, R. Chen, M. Wu, L. Wang, Z. Zhong, T. Wang, Y. Liu, Y. Wang, J. Liu, H. Xu, F. Guo and W. Wang (2021). "MCOLN1/TRPML1 finely controls oncogenic autophagy in cancer by mediating zinc influx." Autophagy 17(12): 4401-4422.

      Racine, K. C., L. Iglesias-Carres, J. A. Herring, K. L. Wieland, P. N. Ellsworth, J. S. Tessem, M. G. Ferruzzi, C. D. Kay and A. P. Neilson (2024). "The high-fat diet and low-dose streptozotocin type-2 diabetes model induces hyperinsulinemia and insulin resistance in male but not female C57BL/6J mice." Nutr Res 131: 135-146.

      Ranalletta, M., H. Jiang, J. Li, T. S. Tsao, A. E. Stenbit, M. Yokoyama, E. B. Katz and M. J. Charron (2005). "Altered hepatic and muscle substrate utilization provoked by GLUT4 ablation." Diabetes 54(4): 935-943.

      Rossetti, L., A. E. Stenbit, W. Chen, M. Hu, N. Barzilai, E. B. Katz and M. J. Charron (1997). "Peripheral but not hepatic insulin resistance in mice with one disrupted allele of the glucose transporter type 4 (GLUT4) gene." J Clin Invest 100(7): 1831-1839.

      Sahoo, N., M. Gu, X. Zhang, N. Raval, J. Yang, M. Bekier, R. Calvo, S. Patnaik, W. Wang, G. King, M. Samie, Q. Gao, S. Sahoo, S. Sundaresan, T. M. Keeley, Y. Wang, J. Marugan, M. Ferrer, L. C. Samuelson, J. L. Merchant and H. Xu (2017). "Gastric Acid Secretion from Parietal Cells Is Mediated by a Ca2+ Efflux Channel in the Tubulovesicle." Developmental Cell 41(3): 262-273.e266.

      Schmiege, P., M. Fine, G. Blobel and X. Li (2017). "Human TRPML1 channel structures in open and closed conformations." Nature 550(7676): 366-370.

      Schmiege, P., M. Fine and X. Li (2021). "Atomic insights into ML-SI3 mediated human TRPML1 inhibition." Structure 29(11): 1295-1302 e1293.

      Tang, Y. and A. Chen (2010). "Curcumin prevents leptin raising glucose levels in hepatic stellate cells by blocking translocation of glucose transporter-4 and increasing glucokinase." Br J Pharmacol 161(5): 1137-1149.

      Tukhovskaya, E. A., E. R. Shaykhutdinova, I. A. Pakhomova, G. A. Slashcheva, N. A. Goryacheva, E. S. Sadovnikova, E. A. Rasskazova, V. A. Kazakov, I. A. Dyachenko, A. A. Frolova, A. N. Brovkin, V. E. Kaluzhsky, M. Y. Beburov and A. N. Murashev (2022). "AICAR Improves Outcomes of Metabolic Syndrome and Type 2 Diabetes Induced by High-Fat Diet in C57Bl/6 Male Mice." Int J Mol Sci 23(24).

      Wang, W., Q. Gao, M. Yang, X. Zhang, L. Yu, M. Lawas, X. Li, M. Bryant-Genevier, N. T. Southall, J. Marugan, M. Ferrer and H. Xu (2015). "Up-regulation of lysosomal TRPML1 channels is essential for lysosomal adaptation to nutrient starvation." Proc Natl Acad Sci U S A 112(11): E1373-1381.

      Wu, D., H. C. Yu, H. N. Cha, S. Park, Y. Lee, S. J. Yoon, S. Y. Park, B. H. Park and E. J. Bae (2024). "PAK4 phosphorylates and inhibits AMPKα to control glucose uptake." Nat Commun 15(1): 6858.

      Wu, Y., W. Lv, S. Xiong, G. Cao, L. Fu, W. Liu, F. Shao, Y. Mei and Y. Lv (2025). "Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion." Phytomedicine 149: 157508.

      Xiao, B., W. Zhang, N. Ji and Q. Chen (2025). "Knockdown of CCNB1 alleviates high glucose-triggered trophoblast dysfunction during gestational diabetes via Wnt/β-catenin signaling pathway." Open Med (Wars) 20(1): 20241119.

      Xie, Y., X. Liu, W. Liu, L. R. Carr, L. P. Lee, N. Imai, E. A. Ortlund and D. E. Cohen (2024). "Activity and phosphatidylcholine transfer protein interactions of skeletal muscle thioesterase Them2 enable hepatic steatosis and insulin resistance." J Biol Chem 300(11): 107855.

      Xing, Y., M. M. Wang, F. Zhang, T. Xin, X. Wang, R. Chen, Z. Sui, Y. Dong, D. Xu, X. Qian, Q. Lu, Q. Li, W. Cai, M. Hu, Y. Wang, J. L. Cao, D. Cui, J. Qi and W. Wang (2025). "Lysosomes finely control macrophage inflammatory function via regulating the release of lysosomal Fe(2+) through TRPML1 channel." Nat Commun 16(1): 985.

      Xiong, L., E. Y. Helm, J. W. Dean, N. Sun, F. R. Jimenez-Rondan and L. Zhou (2023). "Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis." Nat Immunol 24(10): 1671-1684.

      Yamada, K., M. Nakata, N. Horimoto, M. Saito, H. Matsuoka and N. Inagaki (2000). "Measurement of glucose uptake and intracellular calcium concentration in single, living pancreatic beta-cells." J Biol Chem 275(29): 22278-22283.

      Yu, L., X. Zhang, Y. Yang, D. Li, K. Tang, Z. Zhao, W. He, C. Wang, N. Sahoo, K. Converso-Baran, C. S. Davis, S. V. Brooks, A. Bigot, R. Calvo, N. J. Martinez, N. Southall, X. Hu, J. Marugan, M. Ferrer and H. Xu (2020). "Small-molecule activation of lysosomal TRP channels ameliorates Duchenne muscular dystrophy in mouse models." Sci Adv 6(6): eaaz2736.

      Zhang, G., X. Cai, L. He, D. Qin, H. Li and X. Fan (2020). "Skimmin Improves Insulin Resistance via Regulating the Metabolism of Glucose: In Vitro and In Vivo Models." Front Pharmacol 11: 540.

      Zhang, H., Y. Wang, R. Wang, X. Zhang and H. Chen (2024). "TRPML1 agonist ML-SA5 mitigates uranium-induced nephrotoxicity via promoting lysosomal exocytosis." Biomed Pharmacother 181: 117728.

      Zhang, X., X. Cheng, L. Yu, J. Yang, R. Calvo, S. Patnaik, X. Hu, Q. Gao, M. Yang, M. Lawas, M. Delling, J. Marugan, M. Ferrer and H. Xu (2016). "MCOLN1 is a ROS sensor in lysosomes that regulates autophagy." Nat Commun 7: 12109.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful and constructive feedback. In response, we substantially revised the manuscript and clarified the rationale for using in vitro reconstitution assays and iPSC-derived neurons to determine how tau hyperphosphorylation alters its interaction with microtubules and its regulation of intracellular transport. We have also more clearly articulated how these findings relate to neurodegeneration and discussed the limitations of the model systems used.

      The primary concern of the reviewers was the justification for using COS7 cell lysates in reconstitution assays and iPSC-derived neurons as model systems. We have revised the manuscript to clarify that these experimental systems provided a means to isolate and examine how AD-related tau hyperphosphorylation alters tau-microtubule interactions and the regulation of intracellular transport. COS7 cells were selected because they are widely used for expression of mammalian proteins, including kinesins and tau. Human iPSC-derived neurons were chosen for their amenability to TIRF microscopy and transfection-based experiments, as well as the ability to compare CRISPR-generated tau knockout (MAPT-KO) neurons with their isogeneic control counterparts. Accordingly, we have revised the language throughout the manuscript to more clearly define the study’s objectives and emphasize that these systems were intentionally chosen as robust, well-controlled platforms for addressing specific mechanistic questions. We agree that they do not fully recapitulate AD pathology and that more representative models, such as mature, aged neurons or patient-derived neurons would be better suited to studying disease progression, and have included this limitation in the Discussion. However, because the central objective of this study was to dissect the mechanistic consequences of tau hyperphosphorylation on microtubule interactions and intracellular transport, we believe that these experimental approaches are well suited to address the questions asked.

      We also more explicitly addressed how background levels of phosphorylation may contribute to the effects observed with the pseudo-phosphorylation model of AD-related tau perturbations. We’ve addressed this by citing recent studies (Fan et al., 2025; Siahaan et al., 2026; Moretto et al., 2026) that quantitatively assess phosphorylation across expression systems and clarified how our experimental design, which directly compares WT, AP and E14 tau, effectively minimize uncertainty arising from background phosphorylation. While some degree of background phosphorylation is likely to be present, any resulting effects would be expected to occur consistently across all tau phospho-variants. We now discuss the limitation of our study that we did not directly quantify phosphorylation levels in cells.

      The reviewers also expressed concern about the potential influence of endogenous microtubule-associated proteins present in lysates and differences in tau occupancy on microtubules contributing to motility outcomes. To address this, we included additional analyses correlating tau intensity along microtubules with kinesin motility. We also expanded the Discussion to consider how tau competes with other MAPs for microtubule binding and how phosphorylation-dependent changes in tau–microtubule interactions may alter the MAP landscape. Consequently, the transport phenotypes observed with different tau phospho-variants may reflect both direct effects of tau and indirect effects arising from changes in MAP occupancy and competition on the microtubule lattice.

      We provide detailed, point-by-point responses to each reviewer comment below. We appreciate the thoughtful feedback from reviewers and are confident that the revisions, which include clearer language, strengthened justification of the experimental approaches, and additional supporting analyses, have substantially improved the clarity, rationale, and overall impact of the study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Beaudet and colleagues aims at exploring the effect of phosphorylation on the formation of tau envelopes and consequently on axonal transport, both in vitro on reconstituted microtubules and in human excitatory neurons derived from IPSCs.

      The authors found that a relatively widely used construct in which 14 serine or threonine residues, often hyperphosphorylated in Alzheimer's disease, are mutated to alanines (phosphodeficient), increases the density of tau envelopes compared to wildtype tau, whereas a phosphomimetic (same residues mutated to glutamic acid) reduces envelope density both in vitro and in human excitatory neurons derived from IPSCs.

      By analysing the trafficking of different kinesins (KIF1a and KIF5C), they observed different effects of tau phosphorylation status on the movement of these two motors.

      They then analyse transport of lysosomes by employing live imaging of lysotracker in human excitatory neurons derived from IPSCs transfected with wildtype, phosphodeficient or phosphomimetic tau, observing that phosphodeficient tau seems to reduce transport of lysosomes while phosphomimetic increases transport compared to wildtype tau.

      Strengths:

      (1) The work aims to study a novel and underexplored topic in the tau field, tau envelopes, and investigate their relevance to Alzheimer's disease pathology.

      (2) Experiments are well conducted and of high quality.

      Weaknesses:

      Relying only on in vitro reconstituted microtubules and human neurons derived from IPSCs leaves some doubts about the relevance of these results for Alzheimer's disease, considering the embryonic state of IPSCs-derived neurons.

      We agree with the reviewer that iPSC-derived neurons represent an immature state compared with the neurons most affected in Alzheimer’s disease. However, iPSC-derived neurons and in vitro reconstitution are robust experimental approaches that provide insight into (1) the effects of hyperphosphorylation on tau’s cooperative microtubules association and envelope formation, (2) how tau hyperphosphorylation affects the motility of kinesin motors that are sensitive to regulation by tau, and (3) how tau hyperphosphorylation alters the bi-directional transport of endogenous degradative organelles such as lysosomes. Our studies reveal the molecular effects of how hyperphosphorylation influences tau’s role in regulating intracellular transport and we believe that these findings will help to inform future studies examining how tau-related dysfunction first influences axonal transport, which would be expected to alter axonal health and homeostasis prior to the more severe pathological effects observed at later disease stages.

      We have included a paragraph under the subheading ‘Limitations of this study’ in the Discussion section to better contextualize our findings within the broader effort to understand tauopathies and Alzheimer’s disease. We clarify the limitations of using in vitro reconstitution and iPSC model systems on pages 20 and 21.

      Reviewer #2 (Public review):

      This manuscript examines how disease-associated hyperphosphorylation disrupts tau's role as a cooperative microtubule-binding regulator of intracellular transport. Using in vitro reconstitution assays and live-cell imaging in iPSC-derived neurons, the authors employ phosphomutant tau constructs (E14 to mimic hyperphosphorylation, AP to prevent phosphorylation) at 14 disease-associated residues to isolate phosphorylation effects independent of expression system-dependent PTM heterogeneity. The results show that hyperphosphorylated tau fails to form cooperative envelope-like structures on microtubules, instead binding diffusely and dissociating rapidly. In contrast, wild-type and phospho-resistant tau form cohesive envelopes that regulate motor protein access. At the single-molecule level, hyperphosphorylation reduces KIF5C inhibition while maintaining or enhancing KIF1A inhibition through altered processivity and detachment rates. In live neurons, hyperphosphorylated tau phenocopies tau knockout conditions, weakening tau-mediated inhibition of lysosome transport and increasing processive motility. The authors quantify tau binding using Gaussian mixture model-based image analysis and measure tau kinetics via FRAP, demonstrating that hyperphosphorylation-induced loss of cooperative binding correlates with dysregulated organelle transport. These findings establish a mechanism by which phosphorylation-driven disruption of tau's gatekeeper function on microtubules compromises axonal transport prior to aggregation in tauopathies. The paper provides interesting new knowledge for the field, but there are outstanding concerns that could be further addressed by the authors to strengthen and clarify the current manuscript:

      (1) Lack of Phosphatase-Treated Control and Explicit WT Phosphorylation Quantification

      Wild-type tau expressed in insect and mammalian cells is known to be phosphorylated by endogenous kinases (eg, GSK3, CDK5, MARK). The manuscript acknowledges this in the Discussion but provides no phosphatase-treated lysate control or quantification of endogenous phosphorylation on WT tau via phospho-specific Western blots. This leaves ambiguity about whether observed differences between WT and E14 reflect purely the introduced mutations or confounding baseline differences in phosphostate content.

      Tau contains ~85 putative phosphorylation sites and is modified by several kinases in cells. Studies by Siahaan et al. (2026) and Fan et al. (2025) provide detailed insight into tau phosphorylation heterogeneity, its role in protecting the microtubule lattice from severing enzymes, and the implications of phosphorylation patterns for aggregate formation. We reference these papers and include detailed description of these findings when initially establishing our justification for using pseudo-phosphorylation model.

      We used a pseudo-phosphorylation approach to test the effects of phosphorylation of specific residues in the proline-rich region and the pseudo-repeat domain in the C-terminus, which together with the microtubule-binding repeats, establish the minimal regions required for tau’s cooperative microtubule binding (Tan et al., 2019). This system enabled us to dissect the effects of tau phosphorylation without the added complexities of heterogeneity and multiple isoforms of tau that would otherwise be endogenously expressed. We’ve clarified these points in the revised manuscript (Pages 6, 7, 17, and 18).

      Background phosphorylation in the different phospho-variants used might contribute to the observed changes in tau’s MT interactions and regulation of transport. However, based on our results and the significance in the changes between the different phospho-variants, even if there is some basal level of phosphorylation, the results indicate that the effects of the pseudo-phosphorylation sites are strong enough to make observable changes above the basal levels of phosphorylation (see p. 6 of the revised manuscript).

      Disease-associated phosphorylation is likely more heterogeneous and dynamic than the pseudo-phosphorylation mutants used here, and phosphorylation at different sites may differentially regulate tau function (see p. 21 of the revised manuscript).

      (2) Limited Normalization of Motor Effects to Measured Tau Lattice Occupancy

      Although kinesin trajectories are classified inside vs. outside tau envelopes (inherently normalizing to local tau density), motor parameters are not systematically reported as functions of tau fluorescence intensity across all constructs. Co-purifying MAPs or microtubule-modifying enzymes in cell lysates is not quantified or excluded, leaving residual uncertainty about tau-specificity of observed motor inhibition. This should be at least acknowledged in the results section.

      As noted by the reviewer, it is challenging to compare conditions where the occupancy of tau on microtubules is dissimilar across conditions. To address this point, we performed a Spearman’s correlation analysis to compare how tau intensity affects kinesin dynamics along microtubules (Fig S3G). On page 12, our results show that kinesin dynamics are generally reduced in regions of high tau occupancy. However, in regions of comparable higher intensities, KIF5C is less inhibited by E14 tau, whereas KIF1A is less inhibited by AP tau.

      On pages 12 and 13, we acknowledge that while effects from other MAPs or motor proteins could potentially affect kinesin motility, we would expect that any effect from residual lysate components would be similar across tau phospho-variants.

      (3) Insufficient Citation of Prior Neuronal Tau Envelope Evidence

      In the Introduction, the authors state, "it was an open question if tau forms envelopes in neurons," but this understates existing evidence. Tan et al. (2019) report tau neuronal staining consistent with envelope formation, while Siahaan et al. (2021) provide more direct evidence in non-neuronal cells. The framing should acknowledge and integrate these prior findings.

      We agree with the reviewer that evidence from several studies using reconstitution systems, fixed neurons, and live cultured cells provides evidence of tau envelope formation in neurons. Specifically, tau envelopes have been observed along taxol-stabilized or GMPCPP-capped GDP microtubules in vitro (e.g., Dixit et al., 2008; Monroy et al., 2018; Tan et al., 2019; Siahaan et al., 2019), in 4% PFA-fixed and Triton X-100–extracted DIV7 mouse hippocampal neurons (Tan et al., 2019), and in live, non-neuronal U-2 OS cells following taxol treatment (Siahaan et al., 2022) or elevated pH (Siahaan et al., 2024). To our knowledge, our study is the first to demonstrate tau envelope formation in live neuronal cells under normal cell culture conditions. We revised the introduction (see pages 3 and 4) to more precisely position our findings within the context of prior studies.

      (4) Unclear Wording on Expression System-Dependent Phosphorylation

      The sentence "The phosphostate of tau is strongly dependent on the expression system" requires rewording. It is ambiguous whether this refers to the final phosphostate achieved after expression or the inherent phosphorylating capacity of each system. Clearer language would strengthen the methodological justification.

      On pages 6 and 7, we clarify the rationale for using COS7 cells to express GFP-tau and elaborate on recent papers demonstrating how different expression systems used to study tau (e.g., bacterial, insect, mammalian) produce tau with variable phosphorylation patterns (Siahaan et al., 2026; Fan et al., 2025).

      (5) Insufficient Quantification of Motor and Lysosome Transport Effect Magnitudes in Results Section

      The data on molecular motor motility and lysosome transport are densely described. The magnitude of effects (fold-changes, percentage differences) should be explicitly stated in the Results section when first presenting findings to orient readers to biological significance. For example, effect magnitudes for lysosome run lengths, velocities, and directional bias should be quantified in text, not left to figure inspection.

      We now incorporate the relevant quantifications in the text.

      (6) Incomplete Discussion of Projection Domain Necessity for Envelope Formation

      The Discussion states the projection domain is "a critical regulator of both tau-tau and tau-microtubule interactions," but does not engage with prior domain dissection work. Tan et al. (2019) found that the entire projection domain is not necessary for envelope formation in vitro. The authors should discuss which projection domain regions are specifically regulated by phosphorylation vs. required for cooperativity, providing a more nuanced interpretation than implied by their current framing.

      Tan et al. (2019) demonstrated that part of the proline-rich region (residues 198–244) within the N-terminal projection domain and the pseudo-repeat region within the C-terminus, together with the microtubule-binding repeats are the minimal region required to maintain tau’s ability to form cooperative envelopes along microtubules. We revised the text to better incorporate this previous work into the discussion and place our findings within this context. Our work demonstrates how phosphorylation within the proline-rich region and pseudo-repeats are important regulators of tau–tau cooperativity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is unclear how the method implemented by the authors to identify tau envelopes works exactly (GMM and BIC) and how appropriate it is. It does not appear similar to what others have done in the literature on tau envelopes. Moreover, when checking the intensity plots along microtubules, both in Figures 1C and 2A, one is left to wonder if the observed differences are not simply caused by the different thresholds. Indeed, the intensity profiles do not seem greatly different between conditions in Figure 1C, whereas it is evident that the threshold is quite different, with it being lower for AP tau and higher for E14 tau, which explains the differences in how many envelopes are detected. The authors could also try to quantify in a different way (e.g. even just a threshold based on Average+SD) to see if the results remain the same?

      We used a GMM/BIC approach to avoid biased comparisons of tau envelope formation on microtubules across phospho-conditions. In TIRF assays, intensity signals are inherently inconsistent, making it difficult to directly compare fluorescence intensity signals on different microtubules across regions within the same field of view. Additionally, tau distribution between microtubules and in solution varied between conditions (e.g., background tau signal is elevated in E14 conditions compared to WT or AP tau). Given these challenges, we quantified and compared tau intensities on a per-microtubule basis, which produced more robust results. While we initially attempted the reviewer’s suggested approach of using average + SD, per-microtubule variability in minimum and maximum signal, along with differences in local background, prevented the ability to set a threshold that reliably captured intensity differences along microtubules across and within replicates. We now more clearly explain why we chose this approach (see page 7).

      (2) The authors should discuss the possibility that the presence of a GFP tag at the N terminus of tau could affect the formation of envelopes, given the importance of this region. Also, they refer to tau GFP in some points of the text and other times to GFP tau. As it would seem they have always used tau tagged at the N terminus, they should refer to GFP-tau in order to avoid confusion in the position of the tag.

      We agree with the reviewer that the position of the N-terminal GFP could influence the projection domain, However, all tau constructs carry the GFP tag at the same position and differ only in their phospho-site mutations. The correct nomenclature for “GFP-tau” is now consistently used throughout.

      (3) Previous work (Tan et al., 2019) has shown that different isoforms of tau have different propensities to form tau envelopes. The authors should specify in each figure which isoform of tau they are expressing.

      The tau isoform used throughout this study is 4R0N. The tau isoform is clearly identified in the revised text.

      (4) Figure 1 C-E: It would be interesting to see the size of envelopes quantified, also.

      We now include a comparison of the mean envelope width for each phospho-variant (Fig 1F).

      (5) Figure 2F: As the FRAP experiment is not on tau envelopes but generally on axonal tau, this needs to be clearly stated to highlight how this limits the link between the different FRAP dynamics and the behaviour of tau envelopes.

      We changed the text to indicate that we perform FRAP on axonal tau and not specifically tau envelopes.

      (6) The authors should discuss whether they expect Kif1a and Kif5c to be responsible for transporting lysotracker-positive vesicles in neurons? This does not seem to be the case based on a quick literature search. If these are not the motors responsible for the transport of lysosomes, why do the authors decide to look at the transport of these organelles and not others? Also, what is the rationale for studying the transport of lysosomes, an organelle that is mainly transported retrogradely, after identifying defects in kinesin transport? The authors could either study in vitro the effect of tau phosphorylation on the movement of a kinesin more directly linked to lysosomes (e.g. KIF5B, KIF1B) or study the transport of some other organelle which is mediated by KIF1A and KIF5C.

      We revised the text to clarify this point. Several studies have shown that kinesin-1 and -3 are strongly inhibited by tau, whereas kinesin-2 and dynein are less sensitive (Hoeprich et al., 2017; Chaudhary et al., 2018; and others). Within this context, we asked how phosphorylation alters tau’s inhibitory effects on motors that are most sensitive to tau. The in vitro reconstitution assays were not intended to isolate the effects of tau on lysosome-specific motors. Rather, they were used to determine how tau phosphorylation affects representative kinesin-1 and -3 motors that drive a substantial fraction of anterograde axonal transport and are among the most sensitive to tau-mediated regulation.

      We next examined lysosome transport using LysoTracker to investigate how tau phosphorylation influences bidirectional cargo transport. Lysosomes are transported by teams of kinesin-1, -2, -3, and dynein, making them a useful model for assessing the consequences of tau regulation in a more physiological context. Current models of bidirectional transport proposed that cargo movement emerges from tug-of-war, which is a result of a balance of forces generated by opposing motors. Under this assumption, strong inhibition of kinesin by tau would be expected to reduce anterograde transport and/or enhance retrograde transport by shifting this balance towards dynein. We have clarified throughout the manuscript that our goal was to determine how tau phosphorylation affects bidirectional transport and to interpret these findings within this context. Because defects in degradative pathways are thought to contribute to neurodegeneration, these experiments may also provide insight into how tau hyperphosphorylation disrupts lysosome function during disease.

      Although KIF5C and KIF1A are not the primary kinesin homologs responsible for lysosome transport, we expect that other kinesin-1 and kinesin-3 motors respond similarly to tau. The in vitro findings provide mechanistic insight into how tau phosphorylation could alter motor function and ultimately contribute to changes in lysosome trafficking and distribution within axons. On page 19, we further clarify that the magnitude of tau-mediated regulation is likely to vary among kinesin family members due to differences in their intrinsic motor properties, and that the effects on lysosome transport are therefore expected to be more nuanced than those observed for individual motors in vitro.

      (7) In the trafficking experiments with lysotracker in human excitatory neurons, there seems to be a large fraction of anterogradely transported lysotracker-positive organelles. Based on a quick search, it would appear this occurs frequently in IPSC-derived neurons, but it's not the case in primary neurons (see, for example, Kulkarni et al., 2022). Given that IPSCs-derived neurons maintain an immature embryonic maturation status (as correctly stated by the authors when mentioning that they express mainly 3R tau) and that neuronal maturation influences transport in primary neurons (e.g. Moutaux et al., 2018), the authors should discuss these aspects highlighting the possible limitations, especially considering the claim of importance of their results for Alzheimer's disease, a pathology that hits neurons at full maturation stages. Alternatively, they could perform a similar experiment in murine neurons at mature stages.

      See response to comments from reviewer 1 under ‘weaknesses’.

      (8) In the context of the previous point, the immature phenotype of IPSCs could explain the apparent discrepancy between the results obtained by these authors and previously published work (Hallinan et al., 2019), which found that mature hippocampal neurons expressing E14 tau had reduced transport of lysosomes. Moreover, these authors also described patches of higher intensity of tau along the axons formed by E14 tau compared to WT tau, which are closely reminiscent of tau envelopes. The authors should discuss these discrepancies.

      We agree that there are discrepancies between our findings and those reported by Hallinan et al. (2019). In that study, E14 tau was shown to misfold in cultured mouse hippocampal neurons, forming MC1-positive axonal aggregates that impair lysosome transport. In contrast, we observe nearly the opposite effect: E14 tau remains diffusely distributed in axons and produces a phenotype resembling tau knockout conditions, with enhanced lysosome transport.

      These differences may stem from methodological factors, including the neuronal models used (murine hippocampal cultures versus human iPSC-derived neurons), fixation and immunolabeling compared with live-cell imaging, differences in neuronal maturity (DIV), and the presence of endogenous tau versus our knockout-and-rescue approach. Importantly, tau aggregation may reflect later stages of disease progression, where aggregates physically clog axons leading to obstructed axonal transport rather than tau acting as a regulatory “roadblock” to specific motor proteins. We now cite this paper and compare our results with this study and discuss these discrepancies and their potential implications in the ‘Limitations of this study’.

      (9) Figures 4 and 5 are quite hard to read. Perhaps the distinction between proximal, mid and distal axon, although valuable, could be moved to the supplementary, maintaining an overall average, or the most significant of the 3 in the main figures to improve readability?

      We made substantial revisions to figures 4 and 5 and the associated analysis. Because the effects of tau on lysosomal transport were largely consistent across proximal, mid, and distal axonal regions, we combined these datasets and report overall transport trends within the axon (from ~50 µm distal to the AIS to ~50 µm proximal to the growth cone). The region-specific analyses and figures showing lysosome motility in each axonal segment have been moved to Supplementary Figure S4.

      (10) In the discussion, the authors write an entire paragraph on how their results are important to stress the importance of the N terminus of tau in the formation of tau envelopes. This is based on the fact that most of the residues mutated in the phosphomimetic and phosphodeficient constructs are located in the N-terminal projection domain. However, some of these residues are located in the C terminus of tau, which also appears to have a role in tau envelope formation (Tan et al., 2019). The experiments presented do not discriminate the phosphorylation of which of the 14 residues is important to mediate the effects. Hence, I feel this paragraph needs to be toned down or removed entirely.

      See response to comment 6 from reviewer 1.

      (11) The authors make a point of using tau produced in mammalian cells in the experiments performed in vitro, stressing the advancement compared to previous work that used tau produced in bacteria or insect cells. Although this is certainly closer to physiological conditions, the production is done in cancerous kidney cells, so I feel the author should highlight that neurons might drive a distinct phosphorylation pattern. Could recombinant tau be produced in neuroblastoma cells?

      In the revised manuscript, we’ve addressed this comment in the “Limitations of this study” page 21. We used COS-7 cells, which are not cancerous but immortalized fibroblast-like cells derived from African green monkey kidney obtained from ATCC. These cells were chosen because of their widespread use for protein expression and their high transfection efficiency. We agree that the physiology of COS-7 lysates is not directly comparable to that of neurons. Our intention was to convey that proteins expressed in mammalian systems undergo post-translational modifications and are produced by cellular machinery that more closely resembles neuronal systems than bacterial or insect expression platforms. Although neuronal cell lines such as neuroblastoma cells may appear more physiologically relevant, they are often difficult to transfect (Alabdullah et al., 2019). This can create practical challenges in equalizing protein concentrations and obtaining sufficient amounts of overexpressed tau from lysates. While methods exist to improve transfection efficiency, there is no literature that we found stating that these cells would yield protein expression characteristics more comparable to neurons than COS-7 cells. A more comprehensive evaluation of alternative expression systems would require a substantially deeper literature search or systematic characterization of multiple cell lines, which falls beyond the scope of this manuscript. Therefore, we relied on the robust COS-7 expression system and will clarify this rationale in the revised manuscript.

      Alabdullah AA, Al-Abdulaziz B, Alsalem H, et al. Estimating transfection efficiency in differentiated and undifferentiated neural cells. BMC Res Notes. 2019;12(1):225

    1. Author response:

      The following is the authors’ response to the original reviews.

      As requested by all three reviewers we have added a new figure which applies our new end-to-end sorters on real openly available data. This demonstrates that our modular algorithms, optimized to work on simulated data, also performs well in real world cases. In addition, to demonstrate that our results are not due overfitting on simulated data on a specific probe geometry (Neuropixels 1.0), we added a Supplementary Figure demonstrating that our positive results for the end-to-end spike sorters are observed for numerous probe geometries (Neuropixels 2.0, SiNAPS, tetrode and Cambridge NeuroTech). We hope that the extended applications will convince the reviewers and readers of the robustness of our results.

      Reviewer #1 (Public review):

      Weaknesses:

      The reviewer identifies several weaknesses:

      (1) The main concern is the limited support for the claim that ’Lupin’ and individual modules’ outperform existing spike sorters.

      (2) Evidence is primarily from a single benchmark based on an intentionally simplified simulation. While the authors discuss the trade-offs between simulated and real data, the current evaluation does not provide enough diversity to justify claims of superiority.

      (3) While improving individual modules that run in a serial fashion could aid overall spike sorting performance, acknowledging that some end-to-end sorters work in an iterative fashion across multiple of these modules would be fair. Perhaps the optimal spike sorter is not a serial set of modules.

      (4) There is also a risk of benchmark overfitting. A modular approach makes it easy to select components that excel on specific benchmarks (or a specific project’s data characteristics) without generalizing.

      We would like to thank the reviewer for the comments and the valuable feedback. We revised our paper to answer the major concerns that were raised by the reviewer. Regarding the claims about Lupin (1,2), we modified the manuscript in two directions: (i) we attempted to stress that the goal of the paper is not the introduction of the new Lupin sorter per se, but rather presenting and highlighting the modularity of the sorting components framework. In doing so, we also toned down our claims of superiority; (ii) we added simulations on a diverse range of probes and included three real experimental datasets, obtained with different probes, in the results. Regarding the iterative aspect of some sorters – point (3) – we added a paragraph to the discussion highlighting that Kilosort4, unlike previous versions of KiloSort, [9] does not iterate over modules anymore, and thus it can be regarded as a serial algorithm. In addition, although all sorters we present are serial, our proposed framework does not prevent iterative schemes: for example, one could add a re-clustering step after a first template-matching pass. In fact, the modular framework makes this task even easier than before.

      Finally, related to point (4), we added some comments on the problem of overfitting with modular benchmarks in the Discussion, but we think that the risks have been mitigated with the addition of more end-to-end examples with various probe types and experimental data in the revised manuscript.

      The reviewer also points toward some possible ways to strengthen this work:

      (1) Evaluate on multiple simulation regimes, consider adding at least one biophysically detailed simulation, benchmark on multiple probe-geometries with neurons also clustered in different depth profiles (as this will affect drift solutions), and provide real-data validation. Even without full ground truth, real-data can be evaluated with expert curation, functional validation (e.g., refractory violations, quality metrics, unit waveform consistency), agreement across sorters, and consistency across time.

      (2) Related to real-data applicability, it is also important to acknowledge that modulatory approaches can enable overfitting to the needs of individual projects. Without real-data benchmarking (or benchmark diversity), it is unclear how the framework will guide users towards generalizable ’best practices’ rather than optimized configurations that work for their specific conditions.

      In response to these comments, the manuscript has been extended both with more simulated ground truth recordings and experimental data. For additional ground truth, we generated recordings with various probe geometries, to check that the results observed for our modular pipeline could generalize (see Supplementary Figure). Regarding real experiments, we chose three open dataset of various types (chronic and acute implantations, IMEC and Cambridge Neurotech devices) and compared the results of several sorters at a macroscopic level relying on high-level automatic curation tools [7, 3, 6] to quantify how many “good”, “oversplit” and “noise” units are found. This is now a new Figure 8 in our manuscript. We believe that the results from all these datasets demonstrate that Lupin is, as is claimed in the paper, on par with the most popular spike sorting algorithm, Kilosort4 [9].

      Regarding overfitting, we do not believe this is a direct consequence of the modular approach introduced in this article, but rather a general potential “risk” in the spike sorting field. Each spike sorter exposes a large array of parameters that users can tweak to attempt to optimize outcomes on specific datasets, but this end-to-end fine-tuning is hard to control and quantify. We believe that the modular benchmarks introduced in this paper may enable a finer, more controlled, and quantifiable parameter exploration. As an example, very high firing rates such as those observed in the cerebellum might require different parameters for peak detection than for neurons in the cortex. To test this, one could use our generation framework to mimic key macroscopic features of the system you are studying, and benchmark the peak detection step to find optimal parameters. Overall, we do not see this optimization strategy as a problem. Extracellular electrophysiology is so diverse based on brain regions, species, conditions, tasks, etc. that generalizable “best practices” might not exist.

      Reviewer #1 (Recommendations for the authors):

      (1) Tone down or further support the Lupin and specific modules’ superiority claims.

      This has been modified in the manuscript, and we added a final figure to discuss how Lupin is on par with Kilosort on real data, but with no claim of superiority.

      (2) Add benchmark diversity, this would test generalization and mitigate benchmark overfitting. Specifically: add more probe-geometries, and allow for different depth-profiles of clusters of neurons.

      We thank the reviewer for the suggestion, and indeed, we added some more benchmarks to convince the readers that the results observed can be generalized properly. More specifically, we extended the duration of the recordings to 30 min, and included additional benchmark datasets with four different probe geometries (Neuropixels 2.0, a Cambridge Neurotech probe, a SiNAPS probe and a tetrode).

      (3) Clarify how (x,y) positions of neurons are distributed around the shanks.

      This has been clarified in the Methods section. The (x,y) positions are generated uniformly within a rectangle covering the probe boundaries, plus a 20µm margin. Regarding the depth, z positions (distance from the probe) are drawn uniformly from the range [5,40] µm.

      (4) Add at least one real-data benchmark. Some suggestions for evaluation are: stability of firing rates over time, agreement across sorters, quality metrics, functional validation, expert curation.

      As suggested by almost all reviewers, we added some real-data benchmarks (see last figure). Since experimental data don’t have ground truth, we used automatic curation tools as a proxy for “goodness” of the results. We used automatic labels from Bombcell [3] and UnitRefine, which label units as good, multi-unit activity (MUA) and noise, and SLAy [7] for automatic merging, which correlates with the amount of putative oversplits. We felt it was fairer to use external curation tools instead of creating our own methods to assess quality.

      (5) Clarify recommendations for users facing drift. While your statement of ’not having drift is ideal’ is true, the reality is that many recordings have drift. Some practical solutions would be useful. For example, when to trust results and how to report drift sensitivity.

      The reviewer is right, and we added a sentence to clarify when motion correction methods should be used, in our opinions.

      (6) It would help to see what types of signals are most often missed, for example: low-SNR units, drifting units, bursty units, and show which modules affect which of these issues.

      This is already shown in Figures 4 of the manuscript, at the clustering level. These Figures show that cells with low firing rates and/or low SNR are most likely to be missed by all sorters. Our ground truth simulator does not include a bursting mechanism yet. We think that this would be an interesting aspect to simulate and we plan to include bursting units, with bursty spike trains and waveform modulation, in future releases. We thank the reviewer for the suggestion.

      (7) To overcome the issue of overfitting to specific datasets rather than generalization, it would help if a ’default’ or ’recommended starting point’ for users were described in more detail.

      We overcome the issue of overfitting by adding other artificial ground truth recordings (see Supplementary Figure S1), and also real world dataset (see Figure 8). In all these simulations, Lupin is used with default parameters, and this is, we believe, a good starting point. Of course, for very special needs (animal species, brain structures, ...), one might need to adapt parameters, but so as for any other sorters, and such an exploration of the parameter space is out of the scope of the current manuscript. This has been added in the Discussion.

      (8) A lot of the plots have ticks / labels too small to read in 100%, or show quite low-resolution. For example, Figure 3 and Figure 5. Please homogenize across all figures.

      The figures have been regenerated and homogenized as suggested.

      (9) Consider archiving the GitHub version used to generate the figures on Zenodo (DOI) for posterity.

      This has been done for the current state of the manuscript at https://zenodo.org/records/20407862 and the code is available at https://github.com/SpikeInterface/sorting_components_benchmark_paper

      (10) To make the manuscript more reader-friendly, I recommend adding graphics representing the different methods. For example, in Figure 3 one could add schematics of the two peak-detection methods.

      We thank the reviewer for this suggestion. However, this project and our manuscript is not introducing these methods, only re-implementing them in a modular framework. Hence we feel it is out of the scope of this manuscript to produce schematics of the many methods discussed in the paper readers should view the original sources to find out more information.

      (11) Discuss what is meant by KS-like clustering, as I was under the impression that KS4 is also iterative. This may be on the template-matching side, but it’s difficult to know where you draw the border between iterative clustering and (iterative) template-matching. Potentially, we would want to see these processes as one module together, as many sorters work in an iterative fashion across these steps.

      By KS-like clustering, we meant that the code is a direct port, in Python, of the clustering algorithm implemented in KiloSort4. However, we cannot guarantee that this is the exact same clustering method, because of the way the clustering of KiloSort is interleaved with some others steps part of the algorithm. To be more specific, Kilosort has its own special way of performing matched filtering, with a custom grid of templates generated internally with a higher resolution compared to the recording channels positions. These templates are then used to estimate a putative position of each spikes, and these positions are the ones that are used to start the clustering. Thus, the term KS-like comes from the fact that we do not reproduce this exact same mechanism. In SpikeInterface matched filtering is performed using an equivalent approach, but positions are not estimated in the exact same manner. Since the clustering code is equivalent, the results might differ. Regardless of these details, the clustering is not iterative. In former versions of KiloSort, the core algorithm used ideas from k-SVD algorithms, often used in the Machine Learning community. In such an algorithm, the goal is to learn a sparse dictionary of templates to reconstruct the signals, and indeed, there was some iterations between optimizing the templates and the spike times. However, this is no longer the case in KiloSort4. The algorithm works in a serial manner, following the global strategy mentioned in our paper.

      (12) Rather than showing results for one specific simulation, it would be more convincing if we saw the average result of multiple simulations. We don’t need to see individual neurons necessarily (e.g., Figure 3B and C, but this applies throughout the manuscript).

      While we tend to agree with the reviewer than averaging over multiple datasets might be more informative (as we did in the final figure, for end-to-end sorter comparison), we want to say that given the fact that we are using ground-truth recordings that are randomly generated, as long as we do not change the macroscopic properties of the recordings, results on various instances of the noise will be very similar, and averages might not be as informative as one could hope for. One option would be to vary the parameter space of these ground truth recordings, but then there are so many parameters (noise levels, firing rates, distributions of the cells, ...) than averaging everything, and/or even choosing what should be primarily studied is an open question on its own. However, to ease the redibility of the figures, individual neurons were removed from the plots.

      (13) Legends are often incomplete. For example, describe what individual dots are and what lines are in the different figures, even if it seems obvious.

      Legends have been updated.

      (14) Figure 7 would benefit from average thick curves for each model (optionally with individual thin lines for each instance, or error bars). Individual neurons can be left out. Figure 7E (and likewise Figure 5E) would benefit from having x labels to indicate precise labels, so the color attribution is reserved for specific models. It’s quite an intense ’search’ game to understand these figures.

      All the figures have been regenerated, and for the sake of readability, we removed the scatter plots for the individual neurons, focusing only on the averaged lines.

      (15) While I appreciate the effort is gigantic, it may be better for the reader to conclude that.

      This has been rephrased.

      Reviewer #2 (Public review):

      Major comments

      The model ground truth data used in the paper does not need to be a perfect match to experimental data to provide useful benchmarking. However, as with all measurements of spike sorting accuracy, extrapolation to experimental data can be complicated. Users of these tools will need to assess how well the simulated data matches their recordings.

      We agree with the reviewer that extrapolating our results to real data is difficult, and the same point was raised by the other referees. We have now extended the results to include three experimental datasets from different probes. Due to the lack of ground truth, we used automatic curation tools (Bombcell [3] and UnitRefine for labeling, SLAy for merging [7]) as a proxy for performance and showed that Lupin is on par with Kilosort4 on all datasets. We hope this gives users some idea of how well the new sorters will work on their data.

      Reviewer #2 (Recommendations for the authors):

      (1) Any comparison to experimental data would be welcome. Is it possible to add firing rate, amplitude, and template similarity distributions from measured recordings to the panels in Figure 2D?

      Experimental data is very diverse, depending on brain region, species, recording technology, etc. Instead of extending the comparison between simulated and experimental data, we rely on new Figure 8 to showcase the applicability and performance of the presented methods on real recordings.

      (2) Do yields of units passing basic quality metrics look different with the new sorter vs. established sorters? Another option that would help establish the range of applicability of the results would be a different model ground truth system, e.g., the hybrid ground truth data used by the authors in reference 7.

      As it can be seen on synthetic recordings (e.g Figure 7A-B), all sorters behave similarly with respect to firing rates, signal to noise ratios. To help establish the range of applicability, we added an extra Figure 8 in the manuscript, as described above.

      (3) Interpolation errors are likely larger for NP1.0 probes - which have 40 um vertical pitch - than NP2.0 probes, with 15 um pitch. If it’s possible to include even a small-scale comparison of the results from the 2.0 geometry, that would be very valuable for readers trying to decide what probe type to use.

      In the revised manuscript, we added new ground truth benchmarks with a NP2, and other, layouts to show that the key results do not depend on the geometry of the probe (see Supplementary Figure 1). But to address more specifically the point raised by the reviewer, we note that in the case of applying Lupin to NP2 layouts, there is only a loss of ≈ 7.5% of well-detected units between static and motion-corrected cases (see Supplementary Figure 1). Where when we apply Lupin to NP1 layouts (see Figure 7) there is a decrease of 20% of well-detected units. So clearly, a smaller pitch and the columnar arrangement of NP2 seems to help motion-correction methods, and thus reduce the failures due to motion. This has been added in the Discussion.

      About the text:

      (1) In panel E of Figures 5 and 7: Adding the name of the sorter algorithm under the bar charts would be helpful. It is encoded by the color of the outline of each box, but I found that cue rather subtle.

      The figures have been regenerated, with increased ticks and label fonts. We found that adding the names made the Figures too dense, and acronyms would not simplify the figures. In the end we decided not to add them.

      (2) Is the number of features (K) used for clustering already mentioned in the main text or methods? I couldn’t find it.

      We thanks the reviewer for pointing out this problem, and the answer (K = 5) has been added in the manuscript, in the methods section.

      (3) Equation 2 in Methods, defining accuracy, appears to be incorrect. I believe the correct equation is: accuracy = TPij/(Ni + Nj - TPij).

      The reviewer is right, and this has been corrected

      (4) In the abstract, line 18: Component based spike sorters => component based spike sorter.

      This has been corrected

      (5) Line 169 "we can artificially boost the signal-to-noise..." I don’t really see anything artificial about leveraging the extra information in neighboring sites. Maybe just remove that adjective?

      We removed “artificial” from the sentence.

      (6) Line 196-197, describing the result in Figure 3: Especially since 3A is a log plot, it would be helpful to add a percentage to the spikes missed on the high end. These are probably pretty unusual cases, under 0.1%?

      This has been commented in the text, but both because this number is only an approximation (the problem of pairing peaks between detected and ground truth is slightly ill-defined), and because it depends on the particular seeds and parameters of the artificial ground-truth, we avoided numerical values.

      (7) Line 567: "number of dimensions from M to K 5" => "number of dimensions from M to K[5]" That is, is the 5 meant to be a reference? Or is 5 the number of dimensions retained to describe the temporal waveform?

      5 is the value of K, and this has been corrected in the manuscript.

      (8) Line 585-590: Does this description of how to handle borders between groups of sites also apply to the KS-clustering method?

      For the KS-clustering methods, we used the original method implemented in KiloSort. In fact, KiloSort aggregates all the clusters found during the clustering steps, launched per bins, but duplicated templates (based on their shapes only) are removed before matching (and those are the cells at the borders found numerous times). After the template-matching step, once cells have been “populated” with theirs spikes, templates are once again removed and/or merged. However, as explained in the methods, currently this step has not been implemented in our framework. To be more explicit, during template matching, KiloSort stores the features of all discovered spikes, in a space where the effects of nearby spikes have been subtracted, to get a clearer picture of these features. In this “denoised” space, the clustering algorithms is launched again on all spikes (decimated), to assign labels.

      (9) Line 687: "In fact, a major main with Kilosort lies after the template matching step..." => "In fact, a major difference with Kilosort lies after the template matching step...".

      This has been corrected.

      Reviewer #3 (Public review):

      Major comments

      (1) The simulator itself has to be improved and extended. Right now, it simply generates, for every unit, a mother waveform from a sum of exponentials, scales that over channels, and then adds up multiple instantiations of every unit on every channel, along with noise. This is not a biophysical simulator: it is an ad hoc procedure, and the sentence "we firmly believe that.." (lines 482-483) does not make the procedure convincing. To make the simulator credible, the authors should: (1) use a set of biophysical equations, with multi-compartmental modeling of currents and return currents; (2) use noised data from extracellular recordings; or (3) some combination thereof.

      The reviewer is right when pointing out that the current ground truth generator is not “biophysical”, and this is why in the manuscript we used the terms “biophysically plausible”. However, we decided to changed this phrase to “phenomenological” in order to avoid confusion. We believe that our generation tool has the key ingredients to challenge (and also demonstrate limits of) modern spike sorters, which is the goal of the proposed simulator. Spike sorters performing well on such phenomenological simulated data should be a necessary, but not sufficient, condition to convince experimentalists that they will work well on real data. In previous papers [5, 4], we used MEArec [1], which relies on biophysical modeling of the reconstructed neurons to simulate extracellular potentials to generate ground-truth recordings. However, such biophysical simulations have two shortcomings: i) simulations are very slow and resource-hungry; ii) it is not guaranteed that the superior simulation environment (multicompartment modeling) translates to simulated data that are more similar to experimental data. In fact, the Kilosort4 paper [9] shows that the action potentials generated by MEArec [1] have an almost doubled duration compared to realistic data, which requires further ad-hoc parametrization. We believe this discrepancy can be due to the fact that virtually all multi-compartment models are built from in vitro slices, not in vivo recordings. Given these limitations, we decided to rely on a simpler but better controllable model for generating templates. Despite not being “biophysical”, the generation model can replicate, to some extent, the variability in waveforms (using different parameters for waveforms widths and spatial decays) and allows us to have a ground-truth model for drifting as well, that we can use to assess interpolation errors. Nevertheless, we agree with the reviewer that some aspects of the simulator can be improved, such as the structure of the noise in the data. We are currently working on the Spikeinterface side to improve the generation module so that it supports temporally correlated noise (spatial correlation is already supported and used in this manuscript).

      (2) The simulated dataset has to be extended in time. Maybe I missed something, but 500 units over 10 minutes, with some units having firing rates as low as 0.1 spikes/s, corresponds to some of the units firing an expected 60 spikes. This is clearly too short, and does not replicate the standard situation in extracellular experiments.

      We extended the simulated recording to an half an hour duration. However, as it can be seen in the Figures, this does not affect the main results of the paper.

      (3) The simulated dataset has to be extended in space. The choice of using NeuroPixels 1.0 geometry is a poor one. Many labs use other monolithic electrode arrays (MEAs, silicon probes, other rigid arrays); tetrodes remain a major tool, and flexible probes (polyimide, mesh) are evolving. Assessing algorithms over a single spatial architecture is likely to lead to local maxima in performance and potentially erroneous conclusions.

      We believe the that Neuropixels 1.0 geometry is a good choice: the NP1.0 paper it is the most highly cited paper about a high-density electrophysiology probe, suggested that it is currently the most widely used probe in the world. However, the reviewer is correct that there is a risk of overfitting. To demonstrate generalizability of our algorithm, we have added results obtained with NP2.0 layout, a Cambridge NeuroTech layout, a SINAPs probe layout and tetrodes. The results demonstrate that the Lupin sorter is not overfitted.

      (4) The existing spike sorters evaluated are not completely described. Some sorters (e.g., SpyKING Circus and KS4) were described in previous publications, but it is unclear whether the implementation that was used for the present tests is exactly the same as those previously published. More importantly, some of the sorters evaluated (e.g., TDC, TDC2, SpyKING Circus 2) were never described in a peerreviewed paper. This does not mean that they cannot be evaluated - but if they are, they must be described in full. Relying on the fact that the code is open source cannot replace a complete and accurate scientific description.

      The reviewer is rising a valid point, but we think it is beyond the scope of this paper to fully describe every component of every spike sorter mentioned in the paper. We made the deliberate choice to describe the sorters as chains of components in order to demonstrate the flexibility of our modular approach. Full details of every components can be found in the online documentation. In order to ensure the manuscript felt more complete and concrete, we revised the descriptions in our Methods section, in order to give some more details.

      (5) Related to the above, all relevant code should be made available online in permanent repositories, not only in author-controlled ones.

      We are not sure of what exactly is suggested by the reviewer. In order to clarify the situation, we pushed the notebooks and all the code needed to reproduce the Figures of the paper in a Zenodo archive https://zenodo.org/records/19695406

      (6) It is unclear why SpyKING Circus 2 and TDC2 are evaluated - these could potentially be described as straw men. I recommend reorganizing the manuscript so that after every module is evaluated separately based on a limited ground truth dataset, a single "best" sorter would be constructed, and then tested extensively (and compared to the de facto state of the art). Such reorganization would both demonstrate the utility of a modular approach and clarify the general usefulness of the outcome.

      Although we agree that these two sorters might be seen as straw men, we decided to keep them in the paper for various reasons. This has been clarified in the manuscript, but the primary reason is an historical one: the developers of these sorters decided to unite their efforts while designing new tools and algorithms, which led to the initial work on the modular framework described here. Because SpyKING CIRCUS and TriDesClous had to evolve for maintenance, it was decided to try to write a common “grammar” that would allow these two spike sorters to be described in the same framework. The reviewer is right in the fact that once the foundations were stabilized, most of the development efforts were put to Lupin, that was built as the best combination of all the expertise gained on the two aforementioned sorters. A second reason to keep them is that, once again, we decided that Lupin should not be the main focus of the paper. Of course, this is a nice illustration of what the modular framework can do, but we do not want to push it per se. What matters most is the methodology, especially since we can not claim, here, that Lupin would be the best spike sorter regardless of data types, probe geometries, .... We extended the paper with other probe geometries (see added Supplementary Figure 1) and real world data (see Figure 8), and we observed that Lupin was on par, and/or slightly better than KiloSort 4 with respect to number of False Positives for examples. But ultimately, the paper is really about the development of a common ecosystem such that all tools can be improved upon, at the community level.

      (7) The new algorithms developed, for example, clustering and template matching, have to be described in more detail, and demonstrated graphically on simple datasets. This can be done in supplementary material if the authors prefer not to extend the manuscript too much.

      As suggest by the reviewer, we tried to extend the methods of the clustering and the template matching steps, bearing in mind that some of them have already been published in detail elsewhere. We really want to underline that the central point of the paper is the modularity of the architecture, not so much the low-level details. To populate our framework and demonstrate its generalization, we implemented some key algorithms. But describing with schematics and in depth every individual methods is something that is not even done in papers focused on the spike sorters themselves. Later, the reviewer complains that the paper sounds like a technical report. Delving into more details would only amplify this problem.

      (8) This reviewer finds the description and interpretation of the results to be inadequate. As an example, focusing on Figure 5: The results in Figure 5A have to be supplemented and summarized as a scalar point estimate (e.g., median accuracy), an estimate of dispersion (e.g., using MAD, IQR, or SD), evaluated over multiple runs, and compared using statistical tests between tools and conditions (e.g., using a multi-dimensional analysis of variance, a mixed effect model, etc.). The results in Figure 5D must have an indication of dispersion. Any conclusions based on the numerical experiments must be based on these metrics and statistical evaluations.

      To try and simply the plots and their interpretation, we decided to remove the scatter plots of the individual neurons in all Figures, to really focus on the core trends. We believe that the results, such as the dependence on accuracy as a function of SNR, are difficult to capture with summary statistics, and that the results are best understood by looking at the plots we have made. If we wanted to make strong claims about one algorithm being more suitable for a specific task, then these summary statistics would be suitable. Instead, we are trying to demonstrate the general utility of the components framework, and we optimize Lupin simply to maximize the accuracy of each step.

      (9) The entire MS would benefit from expert proofreading; there are many language errors, mostly in indefinite articles and grammatical numbers.

      The manuscript has been intensively proofread by native english speakers.

      Reviewer #3 (Recommendations for the authors):

      (1) Lines 14-15: "...a... sorters...": either "...sorter..." or "...a... sorter...".

      This has been corrected

      (2) Line 30: "peak detection" or "event detection"?

      We prefer the term “peak detection", since this is exactly what the algorithm are looking for: spatio-temporal extrema in the signals

      (3) Line 30: the purely sequential structure of the modules is very limiting and generally incorrect. Template matching may replace peak detection, as is the case in many real-time hardware implementations. The feedforward process is limiting, and many sorters use feedback or multiple loops. It is unclear whether and how the proposed framework supports such structures.

      The reviewer raises a good point that we did not explain well in the manuscript. Since our framework is modular, with each component independent of the others, a developer has freedom to create “iterative" sorters. E.g. they could loop over a pair of steps until a criteria is met. In fact, this is one of the advantages of creating modular components. The examples we show in the manuscript are sequential feedforward structures, leading to this confusion. We have clarified the point in the text. We show sequential sorters because to our knowledge there are currently no iterative sorters in wide use. Older versions of Kilosort were iterative, but KiloSort4 is not. Modularity also allows us to isolate one component for a specific task. Hence, for a real-time implementation, we can pre-compute templates using the peak-detection and clustering components. Then these steps would be skipped for “online” sorting, which would only use the template matching component. The reviewer states that “template matching may replace peak detection”. Indeed, in Lupin, SC2 and TDC2, template matching does replace peak detection – the initial peak detection is only used to construct templates for downstream matching. We have clarified this point in the text.

      (4) Lines 62-63: Are all of these sorters supported by the spike Interface framework? Please include a table of which are and which are not. The same for every module.

      All the sorters listed are indeed supported by Spike Interface, and this has been added in the text.

      (5) Lines 84, 98, and elsewhere: TriDesClous 2 is mentioned. What about TriDesClous - is there such a sorter, and if yes, what is the scientific reference?

      TriDesClous (https://github.com/tridesclous/tridesclous) is a spike sorting pipeline that has not been properly published with a DOI, but that has been developed by the first author of the manuscript and has been used by many papers [8].

      (6) Line 84: Lupin - suggest reorganizing the manuscript around this sorter and evaluating it on multiple datasets.

      The paper is not about Lupin, and we tried to rewrite the manuscript in order to make this point more explicit. The paper is intended to be a proof of concept of the benefits that can be obtained thanks to a modular approach. This is why we do not want to reorganize the paper on Lupin itself.

      (7) Line 104, Figure 1, and elsewhere: What is the advantage of evaluating three sorters, if two are predicted to be worse than the third/state of the art? Suggest to reorganize the MS: (a) describe a proper simulator; (b) describe every individual modules: mention the existing algorithms and elaborate + demonstrate the new algorithms; (c) evaluate every individual module - on properly realistic dataset; (d) evaluate existing complete spike sorters + the proposed best combination - on the same ground truth dataset; (e) compare the best two sorters on multiple datasets. In addition, may identify the WEAKEST link in each sorter and demonstrate the improvement by replacing ("upgrading") that link alone.

      As clarified in the introduction and in the “End-to-end evaluation..." section, we decided to keep three sorters for various reasons. The first is historical: SpyKING CIRCUS 2 and TridesClous 2 were the first two sorters that motivated the creation of the sortingcomponents framework. Because both authors realized that they had so much code in common, they decided to unite their efforts while rewriting them and share some common building blocks. The second reason is that the paper is not about Lupin, per se, but more about the general philosophy of the modular architecture presented here. We want to push forward the idea, in the community, that we should share tools, ideas and algorithms in order to enhance the analysis pipelines. Keeping several sorters, even if sub-optimal, is a way to showcase the flexibility of the framework, and this is why we did not re-organize the manuscript as suggested by the reviewer.

      (8) Line 121: "can drastically cut the time.." - provide quantitative support. In general, avoid superlatives and unsupported statements.

      Spikeinterface has been primarily design to ease the comparison between spike sorting pipelines [2]. Thus all comparison metrics such as agreement matrices, false positives, false negatives, ... are available out of the box when using these Benchmark objects. It would be hard to provide a quantitative support for such a speedup since it will depend on the algorithm, but it allows developer to simply focus on the core implementation while benefiting for free of the whole ecosystem that will launch benchmarks and compare it, with appropriate metrics validated by a large community. Since this validation process, on its own, can be quite complex depending on the processing step, we truly believe the development gain is important, despite the fact that it might be hard to quantify. However, we rewrote the sentence in order to avoid superlatives.

      (9) Line 129: "powerful and fast way to generate artificial" - again, avoid statements with superlatives and lacking quantitative support.

      We rewrote the sentence to avoid superlatives.

      (10) Line 129: "powerful and fast way to generate artificial" - the assumptions made in constructing the artificial data critically and strongly affect the conclusions of the benchmark processes. For instance, it is well known that about 10% of the spikes in the cortex have a positive extrema, but the generator is limited to negative spikes. Also, the generator produces spatially-displaced and scaled versions of the waveform generated by the putative soma, but extracellular waveforms almost never behave that way. Even if the simulator is improved, it will always remain a simulator, and therefore the caveat should always be kept in mind.

      The reviewer is right: our simulator has some limitations compared to biological data. But the fact is that, even on synthetic data, there is still plenty of room for improvements of current spike sorting pipelines before even getting to real data, where ground truth are unknown. We believe that such ground truth simulator is a necessary, but not sufficient, condition to validate sorting algorithms. Other options such as hybrid recordings would also suffer from the same flaws, and might be even more questionable. In the revised version of the manuscript, we also added tests on real data to compare more qualitatively Lupin and KiloSort. The fact that results are in line with what is observed on synthetic data gives us confidence in our observations.

      (11) The authors should add a limitations section to the MS.

      Some limits have been more extensively discussed in the Discussion of the paper, with respect to the feedforward architecture, the validity of the ground truth data.

      (12) Line 129: Can the benchmark object be used with other data (e.g., existing)? The methods indicate that this is the case - please demonstrate.

      We think that a proper demonstration would be out of the scope of this paper, but indeed, as long as the user can provide a recording alongside with a sorting (exhaustive or not), then the Benchmark objects can be used. This allows the use of hybrid recordings, and/or manually curated datasets where users would be able to provide a ground truth. Of course, depending if the ground truth is exhaustive or not (if one knows the activity of all the neurons in the recordings), the metrics might not be the same. But everything is built-in in the object.

      (13) Line 144: 10 minutes is too short and nearly irrelevant. In particular, many units spike at rates much lower than 0.1 spikes/s, and in 10 minutes would emit less than a score of spikes (e.g., 0.01 spikes/s would accumulate, on average, 6 spikes..).

      We regenerated all the figures in the paper with 30 min long recordings, and the results remain similar to the 10 minute recordings.

      (14) Line 151: firing rates are not independent of the waveform as implicitly assumed; this should be accounted for, at least in the options in the simulator. Same for library waveforms - the firing rates should be a parameter that is optionally provided along with the waveform library.

      We are not sure what is meant by the reviewer. We believe that the point is that some cell types, with particular waveforms, have particular firing rates, such as fast-spiking interneurons. This could be dealt with in the current simulator, since users can provide, on a per cell basis, some particular values to generate the waveforms of the neurons. One could clearly imagine having some particular cell types with dedicated waveforms and firing rate parameters. However, for the sake of simplicity in the paper, we chose not to use such granularity. What the simulator can not do, at the moment, is to perform amplitude modulation of the templates as function of bursts for example. But this could easily be implemented, and should be part of future works to consolidate the generator.

      (15) Figure 2A, D: It is unclear why the specific zig-zag motion was simulated. Please rationalize, or use a motion from a real dataset. In particular, simulate (a) breathing-induced micro-motions, (b) gradual drift, and/or (c) a step jump.

      We used a zig-zag plus a Brownian motion to cover a rather broad range of continuous motion. Of course, as pointed out by the reviewer, the heterogeneity of real drifts is large, and again, we believe it is out of the scope of the manuscript to cover them all. In previous works, we explored how motion correction methods were working as function of the drifts [5], and the conclusion was that this simple continuous drift was already challenging enough to make the spike sorters fail. Adding discontinuities, as often encountered in experiment, would only make things worse. Finally, we added real world data (Figure 8) from three randomly picked dataset with heterogeneous probe geometries. They all come with some drift, that might be representative of what is typically dealt with.

      (16) Line 181: "more computationally demanding" - quantification?

      We forgot to add the reference to panel 3D, and this has been corrected in the manuscript.

      (17) Line 192: "(Figure 3A" - add ")"

      This has been corrected

      (18) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement should be quantified.

      We reformulated the sentence, but the point here is simply that template-matching based algorithms use the template-matching step exactly for this purpose: to label spikes that would have been ignored/missed by the peak detection and clustering steps. The fact that template-matching based pipelines such as KiloSort, SpyKING-CIRCUS, ... outperforms clustering-based solution when detecting spikes [4] is a clear support for our sentence.

      (19) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement, if correct, exposes a key weakness of the purely modular approach. For instance, assume that module 2.1 performance can be 0.5 and module 2.2 performance is 0.9, and that will have zero impact on the overall performance of a sorter that has module 4.1 as its fourth module, yielding an overall performance of 0.7. But when module 4.2 is used, module 2.1 performance of 0.5 results in an overall performance of 0.5, whereas module 2.2 performance of 0.9 translates to an overall performance of 0.9. The point should be clear now: evaluating every module in isolation cannot fully predict the behavior of the full system - even if a purely feedforward, single iteration (no loop) architecture is assumed. This requires algorithmic support in the proposed framework, and at the very least explicit discussion.

      The reviewer is right, benchmarking every module in isolation cannot fully predict the behavior of the full system. However, the fact that Lupin, built as an optimal combination of the components and can be on par, if not better in some situations, than KiloSort 4 on various artificial and real dataset makes a compelling argument in favor of assembling the best algorithmic pieces one after the other. In fully assembled spike sorting pipelines, the failures at one stage might be compensated by some algorithmic optimizations latter on. But still, being able to identify these failures, and eventually correct them at the appropriate level, i.e. as soon as they appear might be beneficial for the development of the tool, and to ease their readability/maintenance. This has been added in the discussion.

      (20) Line 203: "we hope that future efforts": This statement is strange for two reasons: (a) if the authors hope for it, why not simply do it? (b) if the authors cannot do it / it is out of scope, the proper place to mention their hopes is in the Discussion section.

      This has been moved to the Discussion section

      (21) Line 232: "motion is only partially compensated for" - so what is the utility of the compensation? This becomes clear later, but the statement is obscure at this point - please clarify.

      This has been clarified.

      (22) Figure 4A, right: The fact that the lines do not reach unity even for high FRs suggests that the FRs are NOT the key limiting factor. Please provide a scalar measure and a 2D analysis of the accuracy as a function of FRs and SNRs, separately for static and motion-corrected.

      We are not sure that we understand what is being requested by the reviewer. A scalar measure with a 2D analysis of the accuracy as a function of FRs and SNRs would mean, per case (static and motion corrected) at least 4 panels, thus a total of 8 panels. This seems like an overly dense figures. Further, a scalar measure is unlikely to provide more insight into the analysis.

      (23) Line 275: "kriging method" - describe.

      For a full description we refer readers to the original paper describing the method [9] and other work on motion correction [5]. We have added a brief description of the method in the text in the "Motion interpolation reduces spike sorting performance" section.

      (24) Lines 309-310: Where are the templates from in this case - true or estimated? If the latter, say it.

      This has been clarified.

      (25) Line 317: This is the point where I almost gave up - the MS up to this stage seemed very much like a technical report and is not organized in a clear, results-oriented manner. Even for a Methods-oriented paper, one expects to learn the key results, but those are lost in this MS.

      While we agree that this part of the manuscript is rather technical, this is something that we believe is at the core of some scientific questions that have not been yet properly addressed by most of the spike sorting pipelines. The idea of performing template-matching, i.e. to seek for spatio-temporal pattern in the signals while at the same time compensating only partially the motion has never been properly benchmarks. While template-matching is clearly a game-changer in static recordings, i.e. without motion, the extra-addition of motion might limit its performances. We firmly believe that our modular approach can be used to isolate such core questions from the whole integrated pipelines, and guide both users and developers.

      (26) Figure 6A, second and third panels: The fourth (rightmost) panel shows that there are differences between the "true" and "interpolated" templates, but this cannot be seen in the two middle panels - probably because too much information is overloaded on every channel. My suggestion is to show a reduced number of channels, but for each of those, show all five waveforms (corresponding to different spatial positions of the centers of mass) in a NON-overlapping display.

      The figure has been redrawn, as requested by the reviewer.

      (27) Figure 6 and the entire issue of the failure in motion correction + lines 331-332: It is not 100% convincing that the failure is not at the DREDGE level - please demonstrate that the motion is estimated perfectly.

      We checked that the DREDGE algorithm is correctly estimating the motion. To convince the reader this is the case, this has now been shown in Figure 2, panel D, where the simulated motion is displayed on top of the motion estimated by DREDGE. As shown, both are very similar.

      (28) Figure 6 and the entire issue of the failure in motion correction: If motion is estimated perfectly and interpolation is done properly, then is the failure an outcome of discrete spatial sampling? In other words, if the movement in space over time (in the simulator) were limited to discrete steps that correspond exactly to the positions of the electrodes, would the motion correction allow perfect performance? Stated differently: if the electrodes were not 15 or 20 micro-meters apart but rather 1 micro-meter apart, would the difference (e.g, between Figure 4A/4B, or between Figure 5A and Fig. 5B) be reduced? Please check and report.

      The reviewer is raising an interesting point, and indeed, this is something that we are planning to investigate. We felt it would have been too technical for the scope of this paper, which is aimed at showcasing the modular framework and its possibilities. But we are planning to explore such failures with ultra-dense probes as has been done in the DREDGE paper [10]. We also have the intuition that errors are originating from failures of discrete spatial sampling: hence the larger the sampling, the more pronounced the errors.

      (29) Figure 6B: It is impossible to discern which line is which - this underscores the general point of summarizing every CDF by a scalar with measures of dispersion (i.e., descriptive statistics) + statistical testing (i.e., quantitative statistics).

      While we agree with the reviewer that the plots are dense, they convey the global message of the paper and are the same that were used in [9]. To ease the comparison, we decided to stick to this representation.

      (30) Figure 6D: add error bars.

      There are no error bars, since there was only a single run in this figure, i.e. on a single dataset.

      (31) Line 318, line 328: So what is the result here - that the interpolation idea is not useful? If yes, what is the source of the failure? Is it discrete/poor spatial sampling as suggested above? Something else?

      The result is that interpolation, rather than motion estimation, is the bottleneck in recoding with motion. A careful analysis using data from dense probes, as already stated above, would be the focus of further work. But we suspect that errors results mostly from bad interpolation due to discrete spatial sampling.

      (32) Line 334: up to this point, there are no quantified results - and actually, no clear results at all.

      We rewrote the section accordingly, to make the key observations more striking. The goal of this section is not to propose quantified results and say which interpolation methods is the best (a full paper should be devoted to this question), but rather to show the possibilities offered by the modular framework, looking at questions that have been not yet properly addressed in spike sorting algorithms.

      (33) Line 345: "This is likely due" - provide support? Example?

      We rephrased the sentence accordingly.

      (34) Line 349: "this is mostly because both clustering..." - but feature detection was evaluated together with clustering, so a conclusion specifically about clustering is unwarranted. This is a general point - can the proposed software/modular framework evaluate feature extraction separately from clustering? If yes, please demonstrate.

      The reviewer is right about the fact that feature extraction was evaluated together with clustering, however, we want to stress that it is exactly the same method used by all the clustering algorithms developed in the paper. Thus, even if the feature detection method might not be optimal, it is not introducing any biases. Currently, almost all spike sorting algorithms these days are using PCA (or truncated SVD) as a feature detection method, on nearby channels where peaks are detected, and this is why we decided to only focus on this method.

      (35) Lines 358-374: If SpyKING-CIRCUS 2 and TriDesClous do not provide any advantage, why include them in the MS? Many other sorters could be included as well. In other words, what do we LEARN from the failure of these sorters to compete with KS4 and Lupin? If nothing, please remove them. If something, please state the conclusions clearly.

      Both these software have advantages, but we decided to keep them as they offer direct illustrations than implementation of fully integrated pipelines, other than Lupin, are possible within the proposed framework. TriDesClous2 is fast, and SpyKingCircus2 (in line with Spyking Circus [11]) is mostly tailored for in-vitro data, and thus can scale for more than thousands of channels while some other algorithms might not [9]. Of course, listing all the pros and cons of each software would be impossible within the scope of the paper, but we decided to keep them for illustrative purpose.

      (36) Line 369: "the main advantages of these sorters lie in their modularity" - modularity is an advantage only if it is useful for something - if the sorters are modular but perform poorly, the point is nullified.

      We agree with the reviewer, but we would like to point out that here, none of the modular sorters shown in the paper are performing poorly, and all have pros and cons as said in the previous point.

      We believe that this is important to show that modularity can offer some options both for users and developers with respect to probe geometries, animal species, data types, ...

      (37) Line 371: "on our dataset" - this is a very important limitation. See general comments, and discuss in the Limitations section of the Discussion.

      We extended the paper by adding more datasets, for different probe layouts, and also real datasets with meta comparison of state of the art spike sorters (see Supplementary Figure S1).

      (38) Figure 7C: the difference between the cyan (KS-like) and the orange (KS4) lines makes the KS-like irrelevant. Please improve or remove.

      The reviewer is right, and we removed the KS-like pipeline, since it at the stage of the manuscript it is not yet exactly like KiloSort 4.

      (39) Line 379: "each individual algorithmic steps" - should be "step".

      This has been corrected

      (40) Line 380: Is Lupin available for download and usage as a separate, standalone package - i.e., not as part of the spike interface framework? If not, please make it available. And if yes, indicate this clearly.

      Lupin is part of the SpikeInterface project, and thus can not be installed in a standalone mode, outside of the SpikeInterface framework

      (41) Line 385: "our... gigantic...effort" - avoid superlatives, especially to self.

      This has been removed

      References

      (1) A. P. Buccino and G. T. Einevoll. Mearec: a fast and customizable testbench simulator for ground-truth extracellular spiking activity. Neuroinformatics, pages 1–20, 2020.

      (2) A. P. Buccino, C. L. Hurwitz, S. Garcia, J. Magland, J. H. Siegle, R. Hurwitz, and M. H. Hennig. Spikeinterface, a unified framework for spike sorting. Elife, 9:e61834, 2020.

      (3) J. M. J. Fabre, E. H. v. Beest, A. J. Peters, M. Carandini, and K. D. Harris. Bombcell: automated curation and cell classification of spike-sorted electrophysiology data.

      (4) S. Garcia, A. P. Buccino, and P. Yger. How do spike collisions affect spike sorting performance? Eneuro, 9(5), 2022.

      (5) S. Garcia, C. Windolf, J. Boussard, B. Dichter, A. P. Buccino, and P. Yger. A Modular Implementation to Handle and Benchmark Drift Correction for High-Density Extracellular Recordings. eNeuro, 11(2):ENEURO.0229–23.2023, Feb. 2024.

      (6) A. Jain, R. Greene, C. Halcrow, J. A. Swann, A. Kleinjohann, F. Spurio, S. Graff, A. Pan-Vazquez, B. Kampa, J. Gall, S. Grün, O. Winter, A. Buccino, M. H. Hennig, and S. Musall. UnitRefine: A community toolbox for automated spike sorting curation.

      (7) S. Koukuntla, T. DeWeese, A. Cheng, R. Mildren, A. Lawrence, A. R. Graves, K. E. Cullen, J. Colonell, T. D. Harris, and A. S. Charles. SLAy-ing oversplitting errors in high-density electrophysiology spike sorting. bioRxiv: The Preprint Server for Biology, page 2025.06.20.660590, 2025.

      (8) J. Magland, J. J. Jun, E. Lovero, A. J. Morley, C. L. Hurwitz, A. P. Buccino, S. Garcia, and A. H. Barnett. Spikeforest, reproducible web-facing ground-truth validation of automated neural spike sorters. Elife, 9:e55167, 2020.

      (9) M. Pachitariu, S. Sridhar, J. Pennington, and C. Stringer. Spike sorting with Kilosort4. Nature Methods, 21(5):914–921, May 2024. Publisher: Nature Publishing Group.

      (10) C. Windolf, H. Yu, A. C. Paulk, D. Meszéna, W. Muñoz, J. Boussard, R. Hardstone, I. Caprara, M. Jamali, Y. Kfir, D. Xu, J. E. Chung, K. K. Sellers, Z. Ye, J. Shaker, A. Lebedeva, R. T. Raghavan, E. Trautmann, M. Melin, J. Couto, S. Garcia, B. Coughlin, M. Elmaleh, D. Christianson, J. D. W. Greenlee, C. Horváth, R. Fiáth, I. Ulbert, M. A. Long, J. A. Movshon, M. N. Shadlen, M. M. Churchland, A. K. Churchland, N. A. Steinmetz, E. F. Chang, J. S. Schweitzer, Z. M. Williams, S. S. Cash, L. Paninski, and E. Varol. DREDge: robust motion correction for high-density extracellular recordings across species. Nature Methods, 22(4):788–800, Apr. 2025.

      (11) P. Yger, G. L. Spampinato, E. Esposito, B. Lefebvre, S. Deny, C. Gardella, M. Stimberg, F. Jetter, G. Zeck, S. Picaud, et al. A spike sorting toolbox for up to thousands of electrodes validated with ground truth recordings in vitro and in vivo. Elife, 7:e34518, 2018.

    1. Author response:

      We thank the editors and reviewers for their thoughtful feedback on our manuscript. We are encouraged that they recognized the importance accounting for media consideration when modeling in vitro disease phenotypes, as well as the value of this dataset as a resource for the field. We appreciate the points raised regarding biological replicates and data normalization methods, and we plan to address these fully in our formal response and in revisions to the manuscript. Below, we provide preliminary responses to several comments and indicate how we anticipate addressing them in the revised manuscript.

      (1) Reviewer 1 & 3: Unclear definition of iPSC lines/clones used in data generation.

      We thank the reviewers for pointing out this ambiguity in the Methods. The project was completed with multiple donor iPSC lines, with at least 2 clones generated from each line, and each experiment was performed using at least three independent iPSC RPE lines. To minimize confounding variability, we took several precautions where possible: each experiment was performed within the same culture plate (coated with Matrigel from the same lot number), using RPE seeded at the same time to ensure comparable maturity, and RPE of the same passage number were used across multiple experiments to reduce de-differentiation or senescence effects. We agree that RPE derived from different iPSC differentiation batches can vary. For this reason, all iPSC RPE used in this study were generated from a single differentiation attempt. Because the goal of this project was to isolate the impact of nutrient composition on RPE phenotype, rather than to characterize variability arising from clonal or donor differences, we did not stratify our analysis by clone or donor. For imaging-based assays, multiple fields were selected at random from each well to ensure representative sampling. TEM analysis of sub-RPE deposits was performed with n=3 independent filters per medium condition, with three panoramic sections imaged per filter. We will add these details, including the number of clones used per iPSC line, to the revised Methods to clarify experimental unit and level of replication for each assay and will include sample number in legends.

      (2) Reviewer 2: In the Seahorse studies provided in Figure 3. basal readings for OCR are abnormally low compared to Oligomycin treatment and background, suggesting difficulties with the assay. Findings should be taken with caution.

      We thank the reviewer for this careful reading. The apparent discrepancy likely arises from comparing the raw OCR trace (left) versus the background-subtracted “Basal respiration” bar graph (right) in Figure 3A. In the raw trace, basal OCR (~45-65 pmol/min) is appropriately higher than both the oligomycin-treated (~35-45 pmol/min) and background (~30-45 pmol/min) rates, as expected. The bar-graph value is smaller (~7-25 pmol/min) only because the non-mitochondrial rate has been subtracted out, whereas the oligomycin plateau in the raw trace has not. These raw basal values fall within Agilent's recommended starting range for the XFe96 platform (~20-160 pmol/min), consistent with the modest basal energy demand of quiescent, differentiated RPE rather than an assay problem. A recent survey of 530 published Cell Mito Stress Tests [1] found that 17% report the implausible result of maximal OCR below basal OCR, and higher basal rate can compromise the FCCP-stimulated maximal rate. We titrated our cell numbers before this assay to ensure a clear FCCP response. The substantially increased maximal rates in all six media conditions indicate a technically sound assay. We will clarify this calculation in the revised Methods.

      (3) Reviewer 3: The metabolic analyses also require additional methodological clarification. For intracellular metabolomics, the culture format, cellular biomass, extraction volume, pooling strategy, and normalization method are not reported sufficiently. Normalization of extracellular measurements to unspent medium accounts for differences in starting metabolite abundance but not for differences in cell number or biomass. Similarly, normalization of intracellular signals to medium 1 does not correct for differences in the amount of cellular material extracted.

      We agree that these methodological details require clarification. Intracellular metabolomics was performed on RPE lysates collected from 12-well plates (n=3 independent wells/RPE lines per medium), each extracted separately without pooling. RPE were scraped directly into a fixed volume of 300 µL chilled 80% methanol per well, regardless of the medium condition; 10 µL of the resulting lysate was dried together with an internal standard (nicotinamide-D4), reconstituted in 100 µL of mobile phase, and 5 µL was injected for LC-MS/MS analysis. Media were processed in parallel using an identical workflow: 50 µL of conditioned media was collected at 24 and 48 hours, of which 10 µL was mixed with 40 µL cold methanol, and 10 µL of the resulting supernatant was dried with internal standard, reconstituted in 100 µL mobile phase, and 5 µL injected for analysis.

      Biomass data, including average nuclei count and total protein content reported in Supp. Fig. 2D, were obtained from RPE cultured in parallel under identical conditions. Because nuclei count and protein content did not consistently agree with one another across the six media, metabolite intensities were not normalized to either protein content or cell number. Instead, the intensity of each metabolite was instead normalized to Medium 1 to allow relative comparison across conditions, avoiding an additional, potentially skewed layer of correction from an imperfect biomass metric. We acknowledge this as a limitation of the method: since RPE size and biomass differ across media, our fold-changes reflect metabolite pool per well rather than per cell, which could over- or under-represent true per-cell differences in media that yield especially large or small RPE. We will clarify these details in the Methods and Discussion.

      (4) Reviewer 1: Since a major purpose of the manuscript is to highlight how cell culture conditions influence RPE biology and metabolism, it would be helpful to also report whether Mycoplasma testing was performed and confirmed to be negative across all cell lines.

      We thank the reviewer for the suggestion to include this information. To confirm, Mycoplasma testing was performed on all cell lines used in this study with negative results. We will add this information in the revised Methods.

      (5) Reviewer 3: public availability of the underlying metabolomics data would be important for a study intended to serve as a community resource.

      The metabolomics data has been deposited to UCSD Center for Computational Mass Spectrometry (CCMS) repository (Dataset: MSV000095024) and will be made publicly accessible upon publication.

      References:

      (1) Ransy C, Boissan M, Hammad N, Bouaboud A, Issad T, De Dieuleveult M, Miotto B, Ye M, Pasmant E, Bouillaud F. Extracellular flux analyses indicate low ATP yield and require refinement for accurate determination of maximal oxygen consumption rate. Sci Rep. 2026 Jun 10;16(1):18344.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Thach et al. report on the structure and function of trimethylamine N-oxide demethylase (TDM). They identify a novel complex assembly composed of multiple TDM monomers and obtain high-resolution structural information for the catalytic site, including an analysis of its metal composition, which leads them to propose a mechanism for the catalytic reaction.

      In addition, the authors describe a novel substrate channel within the TDM complex that connects the N-terminal Zn<sup>2+</sup>-dependent TMAO demethylation domain with the C-terminal tetrahydrofolate (THF)-binding domain. This continuous intramolecular tunnel appears highly optimized for shuttling formaldehyde (HCHO), based on its negative electrostatic properties and restricted width. The authors propose that this channel facilitates the safe transfer of HCHO, enabling its efficient conversion to methylenetetrahydrofolate (MTHF) at the C-terminal domain as a microbial detoxification strategy. Experimental data that shows an involvement of TDM in the reaction of HCHO with THF is less convincing.

      Strengths:

      The authors provide convincing high-resolution cryo-EM structural evidence (up to 2 Å) revealing an intriguing complex composed of two full monomers and two half-domains. They further present evidence for the metal ion bound at the active site and articulate a hypothesis for the catalytic cycle. Substantial effort is devoted to optimizing and characterizing enzyme activity, including detailed kinetic analyses across a range of pH values, temperatures, and substrate concentrations. Furthermore, the authors validate their structural insights through functional analysis of active-site point mutants.

      In addition, the authors identify a continuous channel for formaldehyde (HCHO) passage within the structure and support this interpretation through molecular dynamics simulations. These analyses suggest an exciting mechanism of specific, dynamic, and gated channelling of HCHO. This finding is particularly appealing, as it implies the existence of a unique, completely enclosed conduit that may be of broad interest, including potential applications in bioengineering.

      Weaknesses:

      Although the idea of an enclosed channel for HCHO is compelling, the experimental evidence supporting enzymatic assistance in the reaction of HCHO with THF is less convincing. The linear regression analysis shown in Figure 1C demonstrates a THF concentration-dependent decrease in HCHO; however, it is well established that HCHO and THF can react spontaneously in a non-enzymatic manner, raising the possibility that the observed effect does not require enzymatic involvement. I appreciate the authors' clarification that the data in Figure 1 were not intended to demonstrate enzymatic channelling or catalytic involvement in the HCHO-THF reaction, and that the assay does not distinguish between changes in HCHO production and downstream consumption. However, the statement "these findings show that TDM carries out two linked reactions: TMAO demethylation at one active site, and the HCHO produced can condense with THF at the C-terminal domain, connecting TMAO breakdown to one-carbon metabolism" (page 2) still implies a mechanistic and functional coupling that is not supported by the presented data and appears inconsistent with the authors' clarification. In light of this, I recommend revising this statement to avoid implying mechanistic or functional coupling between the two reactions unless additional experimental evidence is provided.

      We thank the reviewer for this clarification. We have revised as per recommendation (page 2).

      “Overall, these findings suggest that TDM-mediated TMAO demethylation generates HCHO, which can subsequently react with THF, potentially linking TMAO breakdown to one-carbon metabolism.”

      Overall, the authors were successful in advancing our structural and functional understanding of the TDM complex. They suggest an interesting oligomeric complex composition which should be investigated with additional biophysical techniques.

      Additionally, they provide an intriguing hypothesis for a new type of substrate channelling. Additional kinetic experiments focusing on HCHO and THF turnover by enzymatic proximity effects would strengthen this potentially fundamental finding. If this channelling mechanism can be supported by stronger experimental evidence, it would substantially advance our understanding and knowledge of biologic conduits and enable future efforts in the design of artificial cascade catalysis systems with high conversion rate and efficiency, as well as detoxification pathways.

      Reviewer #2 (Public review):

      Summary:

      The manuscript reports a cryo-EM structure of TMAO demethylase from Paracoccus sp. This is an important enzyme in the metabolism of trimethylamine oxide (TMAO) and trimethylamine (TMA) in human gut microbiota, so new information about this enzyme would certainly be of interest.

      Strengths:

      The cryo-EM structure for this enzyme is new and provides new insights into the function of the different protein domains, and a channel for formaldehyde between the two domains.

      Weaknesses:

      (1) The proposed catalytic mechanism in this manuscript does not make sense. Previous mechanistic studies on the Methylocella silvestris TMAO demethylase (FEBS Journal 2016, 283, 3979-3993, reference 7) reported that, as well as a Zn2+ cofactor, there was a dependence upon non-heme Fe2+, and proposed a catalytic mechanism involving deoxygenation to form TMA and an iron(IV)-oxo species, followed by oxidative demethylation to form DMA and formaldehyde.

      In this work, the authors do not mention the previously proposed mechanism, but instead just say that elemental analysis "excluded iron". This is alarming, since the previous work has a key role for non-heme iron in the mechanism. The elemental analysis here gives a Zn content of about 0.5 mol/mol protein (and no Fe), whereas the Methylocella TMAO demethylase was reported to contain 0.97 mol Zn/mol protein, and 0.35-0.38 mol Fe/mol protein. It does, therefore, appear that their enzyme is depleted in Zn, and the absence of Fe impacts on the mechanism, as explained below.

      The proposed catalytic mechanism in this manuscript, I am sorry to say, does not make sense, for several reasons:

      (i) Demethylation to form formaldehyde is not a hydrolytic process; it is an oxidative process (normally accomplished by either cytochrome P450 or non-heme iron-dependent oxygenase). The authors propose that a zinc (II) hydroxide attacks the methyl group, which (a) is unprecedented, (b) even if it were possible, would generate methanol, not formaldehyde.

      (ii) The amine oxide is proposed to deoxygenate, with hydroxide appearing on the Zn - unfortunately, amine oxide deoxygenation is a reductive process, for which a reducing agent is needed, and Zn2+ is not a redox-active metal ion;

      (iii) The authors say "forming a tetrahedral intermediate, as described for metalloprotease," but zinc metalloproteases attack an amide carbonyl to form an oxyanion intermediate, whereas in this mechanism, there is no carbonyl to attack, so this statement is just wrong.

      So on several counts the proposed mechanism cannot be correct. Some redox cofactor is needed in order to carry out amine oxide deoxygenation, and Zn2+ cannot fulfil that role. Fe2+ could do, which is why the previously proposed mechanism involving an iron(IV)-oxo intermediate is feasible. But the authors claim that their enzyme has no Fe. If so then there must be some other redox cofactor present. Therefore, the authors need to re-analyse their enzyme carefully and look either for Fe or for some other redox-active metal ion, and then provide convincing experimental evidence for a feasible catalytic mechanism. As it stands the proposed catalytic mechanism is unacceptable.

      Revised version. The authors have essentially not changed the proposed mechanism. They have removed the reference to zinc metalloproteases, but still propose a mechanism mediated only by Zn2+. As explained above, attack by zinc (II) hydroxide is unprecedented and would generate methanol, not formaldehyde, and amine deoxygenation is a reductive process that cannot be fulfilled by Zn2+. So the proposed mechanism is still not feasible at all. The authors now say that "oxidative chemistry....remains unresolved", I'm sorry, but that is not acceptable.

      I have urged the authors to re-examine the metal content of their enzyme, In the Supporting Information (Figure S5) they give ICPMS data that indicates a Zn stoichiometry of 0.5 mol Zn/mol protein, and Fe is not detected. Have the authors analysed for other redox active metals? The authors say that there is no evidence for any other metal binding site, but there is only 50% occupancy of Zn in their protein, so could there be a different metal ion present in place of Zn in the other 50% of the protein, that accounts for the observed activity?

      Since there is clearly a major discrepancy here, the onus is on the authors to explain the discrepancy, rather than just returning with the same data. For example, they could treat the enzyme with EDTA to remove all metals (and check the treated enzyme by ICPMS), and then add different metal ions to test activity with different metals (could even titrate with different molar equivalents of metal ions). They could then test a range of different redox-active metal ions.

      We have re-examined our data and repeated experiments on the reviewer's opinion. We have repeated the IC-PMS several times with different preps, including full scans (data presented). Our enzyme is active, but no iron signal is detected. Moreover, the experimentally determined structure does not support the presence of a non-heme iron-binding site (Bugg TDH, Ramaswamy S., doi:10.1016/j.cbpa.2007.12.007, and other papers). More detailed response in the recommendations to authors.

      (2) Given the metal content reported here, it is important to be able to compare the specific activity of the enzyme reported here with earlier preparations. The authors have now done this in the revised version.

      (3) The consumption of formaldehyde to form methylene-THF is potentially interesting, but the authors say "HCHO levels decreased in the presence of THF", which could potentially be due to enzyme inhibition by THF. Is there evidence that this is a time-dependent and protein-dependent reaction? Not yet addressed.

      We thank the reviewer for this important point. At present, we have not performed detailed time-dependent or protein-dependent analyses to determine whether the observed decrease in HCHO levels in the presence of THF reflects enhanced downstream consumption or indirect effects, such as inhibition of TDM activity by THF. We acknowledge that further kinetic and protein-dependence studies will be important directions for future work.

      Also in Figure 1C, HCHO reduction (%) is not very helpful, because we don't know what concentration of formaldehyde is formed under these conditions; it would be better to quote in units of concentration, rather than %. This point has been addressed by the authors in the revised version.

      (4) Has this particular TMAO demethylase been reported before? It's not clear which Paracoccus strain the enzyme is from; the Experimental Section just says "Paracoccus sp.", which is not very precise. There has been published work on the Paracoccus PS1 enzyme, is that the strain used? Details about the strain are needed, and the accession for the protein sequence. Addressed in the revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      As noted above, there is still a major problem with the proposed mechanism not being feasible, and remaining questions about the presence or absence of a redox-active metal ion in their enzyme. They should:

      (1) Re-examine for other metal ions (apart from Zn and Fe) using ICPMS. The redox metal ion could, in theory, be some other transition metal ion.

      (2) Seek evidence for the role of metal ions in the activity of this enzyme, for example, by treating with EDTA to remove metal ions, and then adding different metal ions, to correlate activity with a particular metal ion.

      We thank the reviewer for these valuable suggestions regarding the metal identity and its functional role in TDM activity. We also acknowledge the reviewer’s concerns regarding the catalytic mechanism. In response, we have substantially revised Scheme 1 and the associated Discussion text to focus on the observed interactions of the substrate (TMAO) and products (DMA and HCHO) within the Zn<sup>2+</sup>-containing active site, rather than proposing a detailed catalytic mechanism that is not fully supported by the current data.

      (1) We repeated the ICP–MS analysis using a wide-range full-scan survey to examine the presence of additional metal-associated isotopes beyond Zn and Fe. The corresponding experimental details have been added to the revised ICP–MS Methods section (Page 9). Full-scan ICP–MS profiling of purified TDM detected Zn as the predominant associated metal species (Author response image 1A). In contrast, signals corresponding to Fe and other transition metals were either undetectable or present only at trace levels comparable to, or lower than, those observed in the digested HNO<sub>3</sub> solution control. To further validate this observation, we performed targeted ICP–MS quantification for both Zn and Fe on the same purified samples. These measurements confirmed that Fe was below the detection threshold, whereas Zn was consistently detected at an approximate ratio of 0.5 Zn<sup>2+</sup> per protein monomer (Author response image 1B, C, Figure S5).

      The observed 0.5 Zn<sup>2+</sup>-to-protein stoichiometry is consistent with the previously discussed 2 full-length + 2 half-domain (2+2½) assembly. In this complex, only the intact core domains retain the complete metal-binding motif, whereas the truncated half-domains lack the Zn<sup>2+</sup>-binding region. Consequently, only two metal-binding sites are expected per assembled complex, in agreement with the ICP–MS measurements. We additionally note that Zn<sup>2+</sup> was not intentionally supplemented during purification. Based on the current cryo-EM and biochemical data, both metal-binding sites in the full-length subunits appear similarly occupied, with no evidence for asymmetric metal loading.

      Importantly, the previously published FEBS Journal model proposed an Fe<sup>2+</sup>-binding site based on metal analysis and homology modeling rather than direct experimental determination. The residues implicated in Fe<sup>2+</sup> binding are well resolved in our experimental maps and do not define a metal-coordination environment compatible with a second mononuclear metal-binding site. Consistent with the ICP–MS results, the experimental structure provides no evidence of a second metal-binding site. Nevertheless, the enzyme remains catalytically active under these conditions.

      (2) We attempted metal depletion experiments using EDTA treatment to evaluate the functional role of the bound metal ion. However, removal of metal ions resulted in rapid protein aggregation, preventing subsequent activity measurements. These observations suggest that the bound Zn<sup>2+</sup> ion plays an important role in maintaining the structural integrity and stability of the TDM complex. While metal reconstitution experiments would be informative, the aggregation observed following metal depletion precluded a meaningful assessment of alternative metal ions in the current study.

      Author response image 1.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      They use cultures of insulinoma MIN6 cells that form spheroids in a micro-patterned PEG-hydrogel to measure Ca <sup>2+</sup> oscillations in multiple cells simultaneously.

      Strengths:

      They demonstrate that insulinoma spheroids are formed in multi-well plates and that Ca <sup>2+</sup> imaging can be performed on them.

      Weaknesses:

      The type of equipment and multi-wells used for the experiments are very specialized to be used as a common tool. Insulinoma cells are tumoral cell lines that divide, unlike primary beta cells. Pancreatic islets are very different from this preparation, as they are highly heterogeneous, whereas these cells all respond equally. It would be good to see the same technique applied to primary cells.

      MIN6 cells do not respond to glucose and other secretagogues in the same way as primary cells, and they cycle, depending on the phase of the cycle to which they are exposed.

      The authors should report the number of cells per spheroid and the number of cells that are alive and dead.

      I would like to examine the effects of calcium channel blockers on calcium transients, and the use of pregnenolone is already described in the literature, but remains less well known.

      MIN6 cells secrete much insulin, because detecting the hormone in ELISAs requires too many primary cells. The authors should discuss the model in greater detail and compare it with primary beta cells. Also, they take 3 mM glucose as the basal concentration, which is low.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript contains numerous typos, including the combination of numbers and units without a space.

      All spacing inconsistencies between numerical values and unit symbols (e.g. mM, μM, µL, and Hz) have been corrected throughout the text and images. In addition, the following typographical and grammatical errors have been addressed:

      - Abstract: "the frequency of Ca <sup>2+</sup> oscillations correlate" → correlates

      - Figure 2A caption: "200uL" → 200 µL

      - Figure 5A caption: "glimepirde" → glimepiride

      - Figure 5 title: "KATP-antagonists induces" → induce

      - Figure 7D caption: "concentrations" → concentration

      - Figure 7 caption: "Tukey’s post-hoc testm" → Tukey’s post-hoc test

      - Results section header: "increases insulin secretions" → insulin secretion

      - Discussion: "we obtained an EC50 values 7.4 ± 0.4 mM" → "we obtained EC50 values of 7.4 ± 0.4 mM"

      - Discussion: "a EC50 value" → an EC50 value

      - Discussion: "concentration dependence profiles that matches" → match

      - Discussion: "PS-induced activation TRPM3" → activation of TRPM3

      - Conclusion: "found the Islets of Langerhans" → found in the Islets of Langerhans

      - Acknowledgements: "grant agreement No. 955643" → agreement No. 955643

      (2) I suggest showing the experiments in primary beta cells because of the many differences from insulinoma cells, as the most important result is a better culture technique for calcium imaging.

      We thank the reviewer for this valuable suggestion. We did indeed attempt to perform comparable experiments using the Cellartis<sup>®</sup> hiPS Beta Cell Media Kit (Takara, cat. no. Y10108) as a more physiologically relevant cell model. To minimize cellular stress during the transition, we adjusted our protocol by seeding the differentiated cells into the micropatterned plates 24 hours prior to imaging, thereby maintaining the recommended culture conditions for as long as feasible. However, during calcium imaging, the cells failed to respond to either the elevated glucose stimulus or the positive control, suggesting that the cells did not survive the transfer to our plate format with sufficient viability to mount a functional response.

      We acknowledge that the use of primary beta cells or hiPS-derived beta cells would strengthen the physiological relevance of the platform. Nevertheless, we would like to emphasize that the primary objective of this study is to demonstrate the feasibility of high-throughput calcium imaging in 3D cell culture using our optimized protocol: a proof-of-concept that is, by design, independent of the specific cell model employed. Adapting the protocol to accommodate more sensitive or terminally differentiated cell types is a meaningful avenue for future work, but falls outside the scope of the current manuscript. We have added a brief note to the Discussion section to explicitly acknowledge this limitation and to identify hiPS-derived beta cell compatibility as a priority for subsequent optimization.

      (3) These kinds of cultures in three dimensions are interesting, but it has been shown that it is even better to have the liquid flow, simulating blood flow; this can at least be discussed.

      We thank the reviewer for this insightful comment. We agree that the introduction of perfusion-based flow represents a meaningful improvement over static 3D culture systems. We have incorporated this point into the Discussion section, where we now explicitly acknowledge the potential benefits of dynamic culture conditions, including improved viability, insulin secretory function, and morphological integrity of 3D β-cell tissues, and identify the integration of perfusion-induced flow as a potentially meaningful path to explore for future development of the platform.

      Reviewer #2 (Public review):

      Summary:

      The study by Robben et al., show 3D beta-cell spheroid platform, a valuable tool allowing high-throughput monitoring of cytoplasmic Ca concentrations and insulin secretion, with Ca signals comparable to those recorded in primary islets. The authors demonstrate a solid method to culturing MIN6 cells in a 3D culture system, recording Ca signals in a high-throughput format and characterizing these Ca signals using pharmacological tools, including TRPM3 channel and K-ATP channel modulators. This highlights the utility of the 3D beta-cell spheroid for screening new ion channel modulators in beta-cells of the pancreas.

      Strengths:

      - The study shows that the MIN-6-based 3D beta-cell model is better to study Ca-signaling and insulin secretion compared to 2D culture of single MIN-6 cells.

      - The method allows imaging of Ca signaling in many spheroids in parallel followed by collecting medium to measure insulin release and correlate both effects.

      - The authors demonstrate that this system is suitable for screening new pharmacological modulators and used as an agonist of the ATP-sensitive potassium channel (diazoxide) and the agonist and antagonist of the TRPM3 channel.

      Weaknesses:

      - The study is based on only one cell line, the MIN6 insulinoma cells, which may not fully mimic the pancreatic beta-cells within the islet.

      - The authors show only spheroids cultured overnight. A long-term culture is missing to assess beta-cell viability long term function.

      - The authors tested their platform using only two compounds. Testing a larger compound library is necessary to make a clear conclusion about the suitability of the platform for high-throughput screening.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #2 (Recommendations for the authors):

      Major Points

      (1) In this study, only 2 pharmacological compounds (for TRPM3, and K channel) were tested. Testing a larger compound library would be necessary to fully demonstrate its suitability for high-throughput screening applications. If this is not possible at the moment, including data on additional pharmacological compounds, e.g., modulators of voltage-gated Ca channels, which are key regulators of Ca signaling in beta cells of the pancreas would strengthen the study. (e.g., use voltage gated Ca channel blocker such as verapamil or nimodipine).

      We thank the reviewer for this constructive suggestion. We would like to clarify that our pharmacological characterization was not limited to two compounds. In total, six compounds spanning two distinct ion channel targets were evaluated: the K-ATP channel modulators diazoxide, glimepiride, tolbutamide and nateglinide, and the TRPM3 modulators pregnenolone sulphate and isosakuranetin. This panel includes both agonists and antagonists across two mechanistically distinct targets, and we believe this is sufficient to demonstrate the platform's suitability for high-throughput compound screening in the context of a proof-of-concept study.

      We nonetheless agree with the reviewer that extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers like verapamil or nimodipine, would further demonstrate the versatility of the platform. We have added a statement to the Discussion explicitly identifying this as a valuable direction for future work.

      (2) Testing another beta-cell line (e.g., INS-1 cells) would strengthen the manuscript.

      We thank the reviewer for this suggestion. We agree that validating the platform using an additional β-cell line, such as INS-1 cells, would further broaden the applicability of the approach. We did indeed attempt experiments with INS-1 cells; however, the results were inconclusive due to cell quality issues at the time of testing, and we were unable to generate reliable data suitable for inclusion in the manuscript.

      We would also like to emphasize that the primary aim of this study was to demonstrate the methodology of the high-throughput screening platform, rather than to provide a comprehensive cross-cell-line validation. In this context, the use of the well-established MIN6 β-cell line is sufficient to serve as a proof-of-principle demonstration of the platform's capabilities.

      Nonetheless, we consider a systematic evaluation of INS-1 cells on this platform as an important and natural next step and have included this explicitly as a future perspective in the Discussion.

      (3) I find the presentation of the results and analysis of Ca oscillation frequency and area under the curve excellent. Could you please provide more details on the analysis method used to quantify the frequency of glucose-induced Ca oscillation. If a custom script was used, sharing this information with the scientific community would be great.

      We thank the reviewer for this positive feedback. We confirm that the Ca <sup>2+</sup> oscillation analysis was performed using a custom Python script (v3.10.11), the key steps of which are described in the Data Analysis section of the experimental procedures.

      In line with our commitment to open and reproducible science, the script will be made publicly available upon acceptance of the manuscript, allowing the broader scientific community to apply, adapt, and build upon the analysis pipeline.

      (4) Please move the Supplementary Figure to the main Figure 3. This will allow a direct comparison between Ca signals in spheroids and in single cell (2D cultures) under identical conditions.

      We thank the reviewer for this suggestion. We have partially incorporated the supplementary figure into the main manuscript. The mean Ca <sup>2+</sup> response of 2D-cultured MIN6 cells at 20 mM glucose has been added as Figure 3D, enabling direct visual comparison with the spheroid data under identical stimulation conditions. The remaining panels of the supplementary figure — showing representative Ca <sup>2+</sup> traces across multiple glucose concentrations and the corresponding dose-response curves for peak frequency and area under the peaks in 2D monolayers — have been retained in the supplementary information, as their inclusion in the main figure would substantially increase its complexity. The figure legend has been updated accordingly.

      (5) A direct comparison of insulin secretion between 3D cultured spheroids and 2D cultures should also be shown.

      We thank the reviewer for this suggestion. We attempted to include a direct comparison of insulin secretion between 3D spheroids and 2D MIN6 monolayers; however, the 2D measurements proved unreliable for quantitative comparison. Insulin values in the 2D condition consistently exceeded the upper detection limit of the ELISA, and inter-well variability was too high to draw meaningful conclusions. We therefore chose not to include this comparison and instead present the Ca <sup>2+</sup> imaging data in Figure 3D as a functional readout enabling direct comparison between the two culture formats under identical stimulation conditions.

      (6) The authors should further discuss the remaining effects of pregnenolone sulphate on insulin secretion.

      We thank the reviewer for this comment. We would like to clarify that in our experimental setup, pregnenolone sulphate and isosakuranetin were applied simultaneously rather than sequentially. As a result, the incomplete inhibition of PS-induced Ca <sup>2+</sup> oscillations and insulin secretion observed in the presence of isosakuranetin may in part reflect a kinetic offset between the two compounds, whereby PS-induced TRPM3 activation and downstream signalling may have been initiated prior to the establishment of effective TRPM3 blockade by isosakuranetin. In addition, as noted in the Discussion, TRPM3 may not be the only molecular target of PS in these spheroids, and alternative signalling pathways may contribute to the residual insulin secretion observed in the presence of the antagonist. We have added a brief clarification to the Discussion to explicitly acknowledge the potential influence of this kinetic limitation on the interpretation of these results.

      (7) Is it possible to collect 3D cultured spheroids after each experiment to measure for example intracellular insulin content or protein levels by Western blot.

      Physical recovery of spheroids from the PEG hydrogel plates for downstream biochemical analysis, such as intracellular insulin content measurements or Western blot, would indeed be a meaningful addition to the platform's capabilities. We would like to note that the firm attachment of spheroids to the glass substrate, while essential for maintaining spheroid positioning during the extensive washing and liquid handling steps, does present a practical challenge for post-experimental recovery. Although we have successfully extracted spheroids of other cell types from comparable plate formats, reliable recovery of the MIN6 β-cell spheroids without compromising their structural integrity has not yet been achieved. We therefore identify the optimization of spheroid recovery as a valuable direction for future development of the platform.

      (8) Ca signals in response to the application of glucose appears more robust in spheroids compared to single MIN6 cells. What are the possible mechanisms underlying this difference. Whole RNA-seq experiments would be one approach to identify differentially expressed genes in 2D versus 3D culture (this is maybe a whole project by itself). An alternative is to look by RT-qPCR analysis for key β-cell markers and genes encoding ion channel and ion channel subunits.

      We thank the reviewer for this thoughtful comment. The more robust Ca <sup>2+</sup> signals observed in 3D spheroids compared to 2D monolayer cultures likely reflect several interconnected factors. First, and importantly, it should be noted that Ca <sup>2+</sup> measurements in 2D monolayer cultures typically represent an averaged signal across a large population of cells, which tends to obscure individual oscillatory events and reduce the apparent amplitude and regularity of Ca <sup>2+</sup> responses. In contrast, our 3D spheroid platform enables Ca <sup>2+</sup> measurements at the level of individual spheroids, allowing discrete oscillatory peaks to be resolved with much greater fidelity. Beyond this methodological distinction, the 3D architecture also promotes enhanced cell-to-cell communication, better recapitulation of in vivo β-cell coupling, and a more physiologically relevant microenvironment, all of which are likely to contribute to the improved oscillatory Ca <sup>2+</sup> dynamics observed.

      We agree with the reviewer that elucidating the transcriptional underpinnings of these differences, through whole RNA-seq or targeted RT-qPCR analysis of key β-cell markers and genes encoding ion channel subunits, would be highly informative and represents an elegant approach to understanding the molecular basis of the observed functional improvements. As the reviewer rightly acknowledges, however, such experiments constitute a substantial research effort in their own right. We have added a statement to the Discussion identifying this as a valuable direction for future investigation, alongside the other platform development priorities already outlined.

      (9) A more detailed discussion of the limitations of the platform and potential strategies to further improve this system would strengthen the manuscript.

      We thank the reviewer for this constructive suggestion. In response, we have expanded the Discussion to provide a more comprehensive overview of the current limitations of the platform and the strategies we envision for future development. Specifically, the revised Discussion now addresses the following points:

      First, the platform is currently optimized for the MIN6 insulinoma cell line, which differs from primary pancreatic beta cells in several important respects, including glucose sensitivity and secretory capacity. Initial attempts to adapt the protocol to hiPS-derived beta cells were unsuccessful, likely due to insufficient cellular viability following transfer to the micropatterned plate format. Optimizing the platform for use with primary beta cells or hiPS-derived beta cells is therefore identified as a priority for future development.

      Second, while the current study demonstrates proof-of-concept pharmacological characterization using six compounds across two mechanistically distinct ion channel targets, extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers, as well as validation using alternative insulinoma cell lines such as INS-1, would further demonstrate the versatility and generalizability of the platform.

      Third, the introduction of perfusion-induced flow, which has been shown to improve viability, insulin secretory function, and morphological integrity of 3D beta-cell tissues under dynamic culture conditions, is identified as an additional avenue for optimization.

      Fourth, while the current platform operates in a 96-well plate format, which already provides a substantial throughput of up to 1824 individual spheroid measurements per plate, adaptation to higher density plate formats such as 384-well or 1536-well plates would be a necessary step towards true high-throughput screening compatible with industrial drug discovery pipelines. Miniaturization of the hydrogel design and adaptation of the molding procedure to accommodate these formats therefore represents an important direction for future development.

      Minor points:

      (10) Please clarify the glucose concentration at which the MIN6 cells the cultured 3D beta cell spheroids were maintained overnight prior to the experiments.

      Prior to the experiments, both 2D MIN6 cells and 3D MIN6 β-cell spheroids were maintained overnight in standard high-glucose DMEM (25 mM glucose), consistent with widely used MIN6 culture protocols. This information has been added to the Methods section.

      (11) In Figure 3A, is the presented Ca trace derived from one single spheroid? Showing representative Ca traces (e.g., 5 i traces per condition) would show the reproducibility of the recordings.

      We thank the reviewer for this suggestion. The trace shown in Figure 3A is indeed derived from a single representative spheroid. To address the concern regarding reproducibility, we refer the reviewer to the updated Figure 3C, which now displays the individual Ca <sup>2+</sup> response traces of all spheroids within a single well in response to 20 mM glucose stimulation. Grey lines represent individual spheroids, while the black line denotes the mean response. This panel illustrates not only the reproducibility of the oscillatory response across the imaged population but also highlights an important consequence of inter-spheroid variability: because individual spheroids oscillate asynchronously, their peaks cancel out when averaged, resulting in a mean trace that appears relatively flat and lacks the oscillatory features visible in individual recordings. We note that this same cancellation effect likely underlies the comparably flat mean response observed for 2D-cultured MIN6 cells in Figure 3D. Additionally, the difference in the initial response profile between 2D and 3D cultures may in part reflect the geometry of the hydrogel microenvironment. In 2D cultures, the glucose stimulus equilibrates rapidly and near-uniformly across the culture plane, potentially driving a more synchronized initial response and the early peak visible in Figure 3D. In contrast, the PEG hydrogel surrounding the 3D spheroids may act as a diffusion barrier, causing the stimulus to reach individual spheroids with variable delay and thereby further desynchronizing response onsets across the population. The figure legend has been updated accordingly.

      (12) In Figure 3A, the potassium application bar looks a bit shifted. Please check and correct if necessary.

      We thank the reviewer for carefully examining the figure. The apparent shift in the potassium application bar is not an error. In these experiments, glucose was added after an initial 10-minute baseline period, and the observed delay in the Ca <sup>2+</sup> response reflects two contributing factors. First, the glucose solution was pipetted at the top of the well, and diffusion to the level of the spheroids introduces a short lag before the stimulus reaches the cells. Second, the spheroid shown in Figure 3A was located at the outer edge of the well, where mixing is slower and the stimulus arrives with additional delay compared to centrally positioned spheroids. Together, these factors account for the offset between the start of the glucose application bar and the onset of the visible Ca <sup>2+</sup> response.

      (13) Is the system also compatible with the ratiometric Ca imaging dye Fura-2?

      The platform is indeed compatible with ratiometric Ca <sup>2+</sup> imaging using Fura-2. The glass-bottom plate format is a prerequisite for Fura-2 imaging due to the requirement for UV excitation at 340/380 nm, and the transparency of the PEG-based hydrogel ensures that the optical properties of the platform are fully compatible with this approach. Furthermore, the µCELL FDSS fluorescence plate imager used in this study supports dual-excitation ratiometric imaging, making it instrumentally compatible with Fura-2 without any additional hardware modifications.

      (14) In Figure 4 legend: 100 µm diazoxide should be corrected to 100 µM diazoxide.

      This has been addressed in the revised manuscript

      (15) Please comment on the cost of spheroid generation compared with conventional 2D MIN6 cultures.

      We thank the reviewer for this relevant question. The cost of spheroid generation using our platform is largely comparable to conventional 2D MIN6 cultures, with the primary additional expense being the specialized micropatterned PEG-based hydrogel plates. All other aspects of the workflow, including cell culture reagents, imaging consumables, and instrumentation, remain identical. It should be noted that providing a precise cost comparison is difficult, as the price of the hydrogel plates represents the dominant variable cost and is subject to change depending on production scale and supplier agreements. Nevertheless, we consider the platform to be cost-effective relative to alternative 3D culture systems, which often require more complex fabrication procedures or proprietary consumables.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study is to develop high-throughput screening assays utilizing homogeneous 3D cell cultures that more accurately replicate the intricate architecture and cellular communication found in tissues. The authors have chosen pancreatic islet β-cells as a model system to evaluate agents that modulate insulin release, which is particularly relevant given the increasing prevalence of diabetes mellitus-a significant global health concern. Moreover, the incorporation of human-based 3D spheroids, organoids, or organ-on-chip technologies into drug discovery protocols is essential for enhancing clinical translation, as candidate compounds identified using animal models have often demonstrated limited success in clinical settings.

      Strengths:

      This study was thoughtfully planned and skillfully carried out. The use of micropatterned hydrogels to observe 19 spheroids at once is an ingenious aspect, which has been effectively validated with Ca microfluorography. Overall, I found this investigation to be exceptionally well-executed and free from notable flaws, as the results clearly back up the conclusions. Additionally, the developed method achieved the proposed aims, providing a high-throughput format with 3D cultures. I believe this study deserves publication.

      Weaknesses:

      For an HTS assay, authors should incorporate the Z-factor.

      We thank the reviewer for their valuable comment, which we have directly addressed in the revised version, as outlined below.

      Reviewer #3 (Recommendations for the authors):

      (1) The study is very well performed, but to support the claim of suitability for HTS, the Z-factor of the method should be reported.

      We thank the reviewer for this helpful suggestion. In response, we have now included the Z′-factor in the manuscript to explicitly quantify assay performance and robustness. Specifically, the Z′-factor (0.6) has been added and discussed in the Results section, incorporated into the Discussion to contextualize assay suitability for high-throughput screening, and included in the Methods under “Statistical Analysis,” where its calculation is described.

    1. Author response:

      We thank the Reviewing Editor and the reviewers for their highly constructive feedback and their positive assessment of our study. We plan to submit a revised manuscript that addresses these critiques through textual revisions, contextualization of our data, and explicit discussion of the study's limitations.

      To address the comments from Reviewers 1 and 3 regarding the sensing mechanism and cue specificity, we agree that the precise sensor for diacetyl remains an open question. While we cannot specifically pinpoint the mechanism with our current data, we hypothesize that this response relies on a non-canonical sensing mechanism, given our multiple negative results for canonical diacetyl receptors (odr-10, sri-14), signalling pathways, and cilia-defective mutants (daf-19; daf-12). We will revise the text to suggest the mechanism could be cilia-independent or cell-autonomous. Additionally, we will explicitly frame the investigation of AWA-ablated worms, diacetyl derivatives such as acetoin, and additional volatile food cues as important future directions to establish the generalizability of this response.

      Regarding the contrasting survival phenotypes observed by Park et al.24 raised by Reviewer 1, our manuscript currently discusses how chronic odour exposure represses the longevity benefits of dietary restriction48, hypothesizing that this may stem from age-dependent olfactory decline49 and the confounding variable of olfactory learning8. To make this connection clearer, we will explicitly cite Park et al. in this section to directly link their findings with our hypothesis that repeated odour exposure without a nutritional reward extinguishes its efficacy as an anticipatory cue.

      In response to Reviewer 2’s feedback on our metabolic profiling, we will ensure it is clear in the text that we highlighted the lipid species most prominently affected by diacetyl, explicitly noting that triglycerides were not significantly altered. We will also better direct readers to Supplementary Table 1, which contains the comprehensive lists of all metabolites and lipids detected in our metabolomics and lipidomics, and quantifies how they are affected by the treatments in our study. As noted in our manuscript, both DHAP and Gro3P were successfully detected in our metabolomics platform but were not significantly altered by diacetyl exposure. Glycerol, however, was measured via a commercial enzymatic assay because it was not detected by our specific LC-MS platform. We will acknowledge that independently measuring the effects on cellular redox states and utilizing GC-MS for broader metabolic profiling are valuable future directions.

      Finally, to address Reviewer 3’s queries regarding experimental readouts and nhr-49, we will clarify our rationale for using thrashing as our primary readout for hyperosmotic stress. Because diacetyl exposure triggers an acute induction of the DHAP-glycerol shunt, we specifically chose an acute behavioural readout to match this rapid timeline. Regarding nhr-49, we will add a new point to the discussion proposing that the distinct phenotypes may come down to expression thresholds. We hypothesize that while nhr-49 RNAi partially reduces gpdh-1 expression, this residual level of expression might still be sufficient to allow development and survival during sustained hyperosmotic stress."

      References:

      (8) Choi, J. I., Yoon, K., Kalichamy, S. S., Yoon, S.-S. & Lee, J. I. A natural odor attraction between lactic acid bacteria and the nematode Caenorhabditis elegans. ISME J. 10, 558–567 (2016).

      (24) Park, S. et al. Diacetyl odor shortens longevity conferred by food deprivation in C. elegans via downregulation of DAF‐16/FOXO. Aging Cell 20, e13300 (2021).

      (48) Zhang, B., Jun, H., Wu, J., Liu, J. & Xu, X. Z. S. Olfactory perception of food abundance regulates dietary restriction-mediated longevity via a brain-to-gut signal. Nat. Aging 1, 255–268 (2021).

      (49) Suryawinata, N. et al. Dietary E. coli promotes age-dependent chemotaxis decline in C. elegans. Sci. Rep. 14, 5529 (2024).

    1. Author response:

      We thank the editors for the eLife Assessment and the reviewers for their thorough and constructive evaluation of our manuscript.

      We are glad that the evidence for reduced metacognitive sensitivity in relation to the compulsive hypersensitivity dimension was considered solid. In hindsight, we agree that the evidence for the abstraction findings is currently incomplete, and we plan to address this with additional analyses as outlined below (to clarify the robustness of the abstraction metric and its association with symptom dimensions).

      Below, we outline how we plan to address the main points raised by the reviewers, grouped thematically, given the overlap across the three reviews.

      A full, detailed point-by-point response accompanied by the corresponding analyses will follow in our formal revision response.

      (1) Exclusion rate and lack of relevant sensitivity analysis

      In the revision, we will report a detailed comparison of included versus excluded participants on demographic and symptom variables, and we will conduct the originally preregistered sensitivity analysis to estimate abstraction and metacognition metrics, including the excluded sample, and compare these metrics with psychopathological variables, rather than omitting this analysis.

      (2) Abstraction metric and its deviation from the preregistered definition

      We recognise that our primary measure of abstraction differs from the preregistered metric (the proportion of blocks better fit by the Abstract RL model), and that this change appears to affect the pattern of results. In the revision, we will report split-half reliability for the abstraction metric, present the bootstrapped analysis of the preregistered metric, and provide a clearer justification for the switch, making the source of the discrepancy (metric versus inference method) transparent.

      (3) Generalisability of the transdiagnostic factor structure

      We acknowledge that our claim of consistency between samples of the transdiagnostic dimensions requires more support and more careful framing. In the revision, we will provide full methodological detail on both factor analyses (extraction method, criteria for the number of factors, etc.), and we will moderate our interpretation, especially in the discussion, to reflect the actual strength of correspondence rather than describing the structure as broadly consistent.

      (4) Parameter and model recovery

      We will extend our recovery analyses to include a model recovery/confusion analysis between the Abstract RL and Feature RL models, and we will investigate the source of the relatively low learning-rate recovery (e.g., by testing recovery stratified by block length and without injected noise), reporting these results in the revision.

      (5) Documentation of the attention-check procedure

      We will provide all details on the attention-check items and criteria used to identify inattentive participants, including a reference where applicable, and we will standardise terminology (e.g., infrequency vs. inattention items) throughout the manuscript and supplementary information.

      (6) Other methodological and presentational clarifications

      We will review the manuscript, methods, and supplementary materials to resolve the remaining methodological ambiguities and presentational issues raised by the reviewers. This includes, among other points, clarifying whether the reported regression coefficients are standardised and correcting labelling inconsistencies, missing references/DOIs, and other minor textual issues throughout.

    1. Author response:

      We thank the editors and all three reviewers for their careful and constructive evaluation of our manuscript. We recognize that a single concern, the possibility that cells classified as CD56<sup>dim</sup>CD16<sup>dim</sup> after co-culture represent activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16 rather than a pre-existing subset, underlies the majority of the comments. We therefore address this concern first, in a central response, and then respond to each reviewer and editor comment in turn. Where a comment relates to this shared concern, we point to the central response rather than repeating the argument.

      The analytical and presentational revisions described below are complete: the statistical analyses have been re-run with corrections for multiple comparisons, the Discussion has been rewritten, the figures and supplemental tables have been renumbered and corrected, and existing data on pre-stimulation receptor expression and NKp30 have been incorporated. These will appear in the revised manuscript. The new experiments described will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Central are CD56<sup>dim</sup>CD16<sup>dim</sup> cells a pre-existing subset, or activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16?

      We agree with the premise that CD16 is rapidly shed by ADAM17 upon activation, and that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion. For this reason, our conclusion does not rest on post-assay classification. The evidence below, from experiments already in the manuscript, argues against activation-induced conversion, and we will strengthen it with the expanded sorted-subset experiments described at the end.

      Cells sorted before target exposure establish the advantage independently of any during-assay shedding (Fig. 2, unchanged in the revised manuscript).

      The most direct evidence comes from subsets purified before the assay. In Fig. 2, NK cells were sorted into CD56<sup>dim</sup>CD16<sup>dim</sup> and CD56<sup>dim</sup>CD16<sup>bright</sup> populations before any exposure to target cells, and their cytolytic function was measured as specific lysis of autologous HIV-infected T cells. Purified CD16<sup>dim</sup> cells lysed infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, reaching 77.58% versus 39.18% at 1:1. The difference was significant at 1:4 (p = 0.0008), 1:2 and 1:1 (both p < 0.0001); at the lowest ratio tested, 1:8, specific lysis was low in both subsets and the difference did not reach significance (p = 0.2121; Supplemental Table 4). The additional donors described below will allow this comparison to be made across a larger data set, including at the lowest ratios where specific lysis is low in both subsets.

      Because the subsets are defined by sorting before target contact, and because the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from activation-induced CD16 shedding during the assay. Public reviewer 2 and the peer reviewer both identified this experiment as the strongest evidence in the manuscript. Its interpretation is secure; what it requires is additional donors for statistical robustness, which we provide in the planned expansion below.

      The two subsets respond to ADAM17 inhibition in opposite directions (Figs. 7 and 9, now Figures 6 and 8).

      If CD56<sup>dim</sup>CD16<sup>dim</sup> cells were simply CD56<sup>dim</sup>CD16<sup>bright</sup> cells that had shed CD16, the two would be one population sampled at different points along a shedding continuum, and inhibiting ADAM17 would move them in the same direction. Instead, ADAM17 inhibition moves them in opposite directions. In the antibody-dependent degranulation assay (Fig. 7B, now Figure 6B), ADAM17 inhibition increased CD56<sup>dim</sup>CD16<sup>bright</sup> degranulation but decreased CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation across VRC01 concentrations. The same opposition is seen when serial degranulation is resolved by the number of degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased multiple degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells, with cells undergoing three events rising from 0.79% to 5.12%, while in CD56<sup>dim</sup>CD16<sup>dim</sup> cells it reduced them, with three events falling from 15.16% to 2.24% and the non-degranulating fraction rising from 61.56% to 92.13%.

      A single population would not be expected to respond to the same perturbation in opposite directions, and these observations are difficult to reconcile with the CD16<sup>dim</sup> cells being activated CD16<sup>bright</sup> cells; rather, they point to two functionally distinct subsets with opposite dependence on ADAM17 activity. This is consistent with our model, in which CD16<sup>dim</sup> cells use ADAM17-mediated shedding to detach and serially re-engage, whereas CD16<sup>bright</sup> cells are hindered by the loss of CD16.

      The CD16<sup>dim</sup> degranulation advantage is driven by NKG2D through a mechanism separable from ADAM17 (Fig. 8, now Figure 7).

      The change in subset frequency on exposure to VRC01-treated infected cells (Fig. 8B, now Figure 7B) is abolished by anti-NKG2D even though VRC01 and ADAM17 remain present, indicating that this frequency shift is driven by NKG2D-dependent activation rather than by antibody-CD16 engagement alone.

      The degranulation data in Fig. 8A (now Figure 7A) show that the two perturbations act differently in the two subsets. In CD56<sup>dim</sup>CD16<sup>dim</sup> cells, both reduce the response, and the combination reduces it further than either alone: from 12.60% under vehicle to 5.24% with anti-NKG2D (p < 0.0001), 5.05% with ADAM17 inhibition (p < 0.0001), and 2.03% with both (p < 0.0001 versus vehicle; p < 0.0001 versus anti-NKG2D alone; p = 0.0001 versus ADAM17 inhibition alone). If anti-NKG2D acted only by removing the activation trigger for ADAM17, that is, if NKG2D and ADAM17 lay on a single linear pathway, blocking the pathway at two points would not be expected to add to the effect of either alone. The further reduction therefore indicates that NKG2D and ADAM17 contribute through separable mechanisms.

      In CD56<sup>dim</sup>CD16<sup>bright</sup> cells the two perturbations act in opposite directions. Anti-NKG2D reduced degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells, where ADAM17 activity is required. Two populations differing only in the extent to which they have shed CD16 would not be expected to respond to the same two perturbations in opposite ways.

      Supporting evidence: pre-sorted subsets are stable and differ before stimulation.

      Two further observations support a pre-existing subset. First, we have directly tracked the fate of each subset sorted before target exposure. NK cells were sorted into CD16<sup>bright</sup> and CD16<sup>dim</sup> subsets, exposed to HIV-infected cells for one hour, and reanalyzed for CD16 expression. One hour is the point at which we observe the highest frequency of degranulating cells in both subsets (Figure 3—Figure Supplement 1A in the revised manuscript), and therefore the point at which activation-induced shedding would be most likely to be detected. Sorted CD16<sup>bright</sup> cells remained predominantly CD16<sup>bright</sup> (approximately 62%); of those that lost CD16, most became CD16<sup>negative</sup> (approximately 35%) rather than CD16<sup>dim</sup> (approximately 4%). Sorted CD16<sup>dim</sup> cells likewise shifted predominantly to a CD16<sup>negative</sup> phenotype (approximately 75%). Activation-induced CD16 shedding therefore directs cells of both subsets toward the CD16<sup>negative</sup> gate rather than generating the CD16<sup>dim</sup> population from CD16<sup>bright</sup> cells. These data are presented in Author response image 1.

      Author response image 1.

      Phenotype of sorted CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> NK cells after exposure to HIV-infected T-cells. NK cells were sorted into CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> subsets, exposed to purified autologous productively HIV-1<sup>SHM-1</sup>-infected T cells for 1 hour at a 1:1 effector-to-target cell ratio, and reanalyzed for CD16 expression. Bars show the percentage of each sorted NK cell population (CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup>) falling into the CD16<sup>bright</sup>, CD16<sup>dim</sup>, and CD16<sup>negative</sup> gates after exposure, as the mean ± standard deviation (SD) of three replicates. One hour is when the highest frequency of degranulating cells is observed in both subsets.

      Second, the subsets differ before any stimulation. CD56<sup>dim</sup>CD16<sup>dim</sup> cells express higher NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), a difference of approximately 1.6-fold; across all conditions tested the subset means were 2521.5 versus 1772.2 gMFI, or 1.42-fold (two-way ANOVA: subset F(1, 20) = 1897, p < 0.0001, 70.97% of the total variation; Fig. 6, now Figure 5; Supplemental Table 18). This difference is specific to NKG2D: NKp46, measured on the same cells in the same wells, did not differ between the subsets in the no-target condition (mean difference 45.67 gMFI, p = 0.0668), and was higher on CD56<sup>dim</sup>CD16<sup>bright</sup> cells when targets were present. The subsets therefore differ in NKG2D density before activation, and not in activating receptor density generally.

      Planned strengthening.

      To place this beyond doubt, we will expand the sorted-subset experiments, performing the specific-lysis assay (Fig. 2) and the antibody-dependent degranulation assay on subsets purified before target exposure across additional donors, together with uninfected-target controls. If cell yields from the sort permit, we will also perform the serial degranulation assay on sorted subsets; because the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells and the serial degranulation assay requires four sequential labelling and washing steps, we cannot commit to this in advance of the sort. These experiments require sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      We note that performing every functional assay in this study on sorted subsets is not feasible within the scope of this revision. Sorting the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, which constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, from a sufficient number of donors to repeat the full panel of assays would require resources beyond those currently available to us. We have therefore prioritized the specific-lysis and antibody-dependent degranulation assays, which bear most directly on the concern raised by the reviewers and the editor, and will extend the approach to the remaining assays as resources allow.

      Public Reviews:

      Reviewer #1 (Public review):

      Overall organization

      Overall, the manuscript includes many data in nine figures plus supplemental figures, and would benefit from some focusing of the results.

      We agree, and we have reduced the main figures from nine to eight. The NKG2D ligand histograms, previously Figure 5A, have been removed for the reason given in our response to the peer reviewer, Recommendation 4. The degranulation against wild-type and ΔVpr-infected targets, previously Figure 5B, and the degranulation and killing frequency measurements in mixed populations, previously Figure 3, have been moved to the supplementary material. The inhibitory receptor analyses have been reduced from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. Each Results section now opens with a statement of the principal finding before the supporting data, and the Discussion synthesizes what the findings mean rather than restating them.

      Figure 1 and Figure 1—figure supplement 2

      The observation that CD16dim NK cells responded more strongly by degranulation to K562 cells and HIV-1-infected cells could be due to the shedding of CD16 following activation. In other words, more strongly activated NK cells express higher levels of CD107 but also shed CD16, resulting in higher CD107 expression in CD16low NK cells. The authors should investigate this, for example by performing the degranulation assays shown in Figure 1 in the presence and absence of an ADAM17 inhibitor.

      We thank the reviewer for raising this important point, which we recognize is shared by all three reviewers and the editors, and which we address in full in the central response above. We note for clarity that the K562 data are not in Figure 1; they are presented in Figure 1—figure supplement 2, both in the reviewed preprint and in the revised manuscript. They were included to reproduce a previously established finding (Amand et al., Front Immunol 2017;8:699), and were not intended as a central experimental claim; the mechanistic focus of this study is the response to HIV-infected cells. Notably, that same study addressed the question by sorting the CD56<sup>dim</sup> subsets before stimulation rather than gating after, and our study applies the same approach and extends it to the HIV-infected setting. Four lines of evidence argue against the interpretation that CD16<sup>dim</sup> degranulation reflects activation-induced CD16 shedding of CD16<sup>bright</sup> cells:

      (1) In cells sorted before target exposure, purified CD16<sup>dim</sup> cells lyse HIV-infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, significantly so at 1:4 and above (Fig. 2; Supplemental Table 4); because the subsets are defined before any activation and the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from shedding during the assay.

      (2) The two subsets respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased degranulation in CD16<sup>bright</sup> cells but decreased it in CD16<sup>dim</sup> cells.

      (3) Blocking NKG2D together with ADAM17 reduced CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A), indicating that the CD16<sup>dim</sup> advantage is driven by NKG2D through a mechanism separable from ADAM17-mediated shedding.

      (4) The subsets also differ before any stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells in the no-target condition (p < 0.0001) and no corresponding difference in NKp46 under the same condition (Fig. 6, now Figure 5).

      We will further strengthen these findings by expanding the sorted-subset experiments across additional donors, as described in the central response.

      Figure 2

      The authors sorted CD16dim and bright NK cells for these experiments and observed higher lysis of HIV-1-infected CD4+ T cells. Important controls should be included in these experiments - how strong was the lysis of HIV-1-uninfected CD4+ T cells by these different NK cell subsets? It also appears that the results shown were derived using NK cells from one donor, and "representative of two independent sort experiments performed with separate donors, each yielding similar results". Why are the authors now showing the respective data? One or two experiments appear too few to come to these conclusions. To support the broad conclusions drawn by the reviewers, the experiments should be performed in a larger number of individuals.

      We thank the reviewer for these constructive points.

      (1) Uninfected-target control. We agree this is an important control and will include lysis of uninfected autologous CD4 T cells by the sorted CD16<sup>dim</sup> and CD16<sup>bright</sup> subsets in Figure 2, confirming that the observed lysis is specific to HIV-infected targets. We note that the corresponding CD107a degranulation controls against uninfected targets are presented in Figure 1—Figure Supplement 4C of the revised manuscript.

      (2) Number of donors and presentation of data. We agree that the conclusions require more than the representative donor shown. As described in the central response, we will expand these sorted-subset experiments to a larger number of individuals and will present the data from all donors rather than a single representative experiment. Because they require cell sorting and primary-cell work, these experiments will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Figures 3 and 4

      It appears that experiments were performed again using bulk NK cell populations, and superior degranulation and killing frequencies by CD16dim NK cells might reflect different levels of activation again, as described above for Figure 1. The same applies to Figure 4 - lower degranulation events in CD16bright NK cells are consistent with lower activation of these cells, resulting in less CD16 downregulation. Also, it is not clear to the reviewer why CD107a expression and killing frequencies decrease with higher effector-to-target ratios (Figure 3).

      (1) Activation-induced shedding in bulk experiments (Figs. 3 and 4, now Figure 2—figure supplement 1 and Figure 3). We agree that these figures use bulk NK cell populations gated by CD16, and we address the underlying shedding concern in full in the central response. The concern that the lower serial degranulation of CD16<sup>bright</sup> cells in the direct-killing assay simply reflects lower activation and therefore less shedding is addressed directly by our ADAM17-inhibition data. At 0 µg/mL VRC01, that is, in the absence of antibody, ADAM17 inhibition already affects the two subsets differently rather than in the same direction (Figs. 7B and 7C, now Figures 6B and 6C), as would be expected if they were one population differing only in activation level. This differential response is also seen across the antibody-dependent conditions in Figs. 7B and 9 (now Figures 6B and 8). As described in the central response, we will additionally repeat the specific-lysis and antibody-dependent degranulation measurements on subsets purified before target exposure across additional donors, and the serial degranulation assay as well if cell yields from the sort permit.

      (2) Decrease in CD107a and killing frequency at higher effector-to-target ratios (Fig. 3, now Figure 2—figure supplement 1). This reflects the nature of the readout. CD107a mobilization is measured per effector cell, as the percentage of NK cells that degranulate, and is therefore maximized when targets are in excess. At low effector-to-target ratios, nearly every NK cell can encounter and engage a target, yielding a high percentage of CD107a-positive cells; at high ratios, targets become limiting, so a large fraction of NK cells never contact a target and remain unstimulated, and the rapid destruction of the limited target pool further reduces the stimulus available to the remaining cells. This lowers the measured per-effector degranulation frequency even as the absolute number of targets killed is maintained, and it is distinct from a lysis assay, which measures the fate of the target population and accordingly rises with increasing effector-to-target ratio. The same per-effector readout behavior applies to the degranulation data shown in Figs. 5C and 6C (now Figures 4A and 4B, and Figure 5C). This explanation has been added to the revised Discussion.

      Pages 19-25

      It would be helpful if the authors could provide some conclusions regarding their findings - it is very difficult for the reader to follow the many reported frequencies and p-values. What does this actually mean? Overall, the results appear to follow prior observations that licensed (KIR3DL+) NK cells respond more strongly than unlicensed (KIR3DL1neg) NK cells. The consistent observation within these different subanalyses that CD16dim NK cells degranulate more than CD16bright NK cells is probably the result of activation-induced CD16 downregulation in these assays, as mentioned above. Providing two-way ANOVA analysis results for these very many observations would furthermore require, in the opinion of the reviewer, adjustments for multiple comparisons.

      We thank the reviewer, and we have addressed this in three ways.

      (1) Readability. We agree that these sections were difficult to follow as presented. We have rewritten them, opening each with a statement of the principal finding before the supporting statistics, and reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. The Discussion now synthesizes what the findings mean.

      (2) CD16<sup>dim</sup> degranulation in these subanalyses. The consistent observation that CD16<sup>dim</sup> cells degranulate more than CD16<sup>bright</sup> cells across these subanalyses is addressed in full in the central response, where several lines of evidence, including subsets sorted before target exposure (Fig. 2) and the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8), argue against activation-induced CD16 downregulation as the explanation.

      (3) Multiple comparisons. We agree, and we have re-analyzed these comparisons, applying the post-hoc test matched to each comparison structure: Dunnett's where every subset is compared against a single designated subset, Tukey's where all pairwise comparisons are of interest, and Šidák’s where a prespecified subset of comparisons is of interest. Adjusted p-values are reported throughout, and the design and post-hoc test used for each figure and panel are given in a new supplemental table. The streamlining described above has also reduced the number of comparisons reported in the main text.

      Figures 5 and 6

      These figures demonstrate that NK cell-mediated activation by HIV-1-infected cells depends on NKG2D ligands and can be inhibited by blocking this interaction - this is consistent with data presented by the Barker group and others previously, and does not provide new information.

      We agree that the dependence of NK-cell recognition of HIV-infected cells on NKG2D and its ligands is established, including in our own earlier work (Ward et al., PLoS Pathog 2009;5(10):e1000613) and by others, and we do not present that dependence as a novel finding.

      On review, the histograms in Figure 5A were reproduced from that earlier study and should not have been included without attribution. We have removed that panel and cited the original finding in its place. Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, establishing that the degranulation response in this system depends on Vpr, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      We would also distinguish Figure 6 (now Figure 5), which we consider a substantive finding rather than a restatement of the established NKG2D-ligand dependence. That figure shows that NKG2D expression differs at the level of the individual subsets, and that the difference is present before stimulation and is specific to NKG2D. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), approximately 1.6-fold, while NKp46 measured on the same cells in the same wells did not differ (mean difference 45.67 gMFI, p = 0.0668). This provides a candidate mechanism for the superior effector function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, in addition to their serial-degranulation capacity, and it bears directly on the central question of whether the two subsets differ intrinsically rather than as a consequence of activation. The revised text presents the established NKG2D-ligand dependence as context while making the subset-level NKG2D difference, and its mechanistic significance, more prominent.

      Figure 7

      The authors extended their functional analyses of NK cells to ADCC function. It is very well established that CD16 is downregulated in the context of ADCC following activation of NK cells. Consistent with this, higher degranulation is observed by CD16dim NK cells.

      We agree that CD16 is downregulated during antibody-dependent responses, and this is precisely why we included the ADAM17-inhibition experiments within Figure 7 (now Figure 6), to determine whether the higher degranulation of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is a consequence of that shedding or a property of a distinct subset. As detailed in the central response, these experiments argue against the shedding interpretation. In Fig. 7B (now Figure 6B), inhibiting ADAM17 affects the two subsets in opposite directions: it increases the degranulation of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing that of CD56<sup>dim</sup>CD16<sup>dim</sup> cells. If the CD16<sup>dim</sup> cells were simply CD16<sup>bright</sup> cells that had shed CD16, blocking shedding would be expected to move the two in the same direction; the opposite responses instead indicate two distinct populations with opposite functional dependence on ADAM17 activity. The same opposition is seen when serial degranulation is resolved by the number of events per cell (Fig. 9, now Figure 8). Thus, while CD16 downregulation during antibody-dependent responses is well established, these data indicate that the superior response of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is not explained by it. This interpretation is now explicit in the revised Discussion.

      ADAM17-inhibition data (final figures)

      These data are of interest, but should be presented in a more structured way. First of all, does the addition of ADAM17 inhibitors change the overall proportion of CD16bright and dim NK cells following activation, independent of whether these cells degranulate or not? Overall, the proportion of CD16dim NK cells that degranulate appears to be reduced in the presence of the ADAM inhibitor, which is consistent with reduced CD16 shedding and maintenance of CD16 expression on activated NK cells - and this is supported by the increase in CD107a-positive NK cells that express CD16 (Figure 8a). Overall, the differences between CD16bright and dim NK cells in their level of activation appear to disappear in the presence of an ADAM17 inhibitor, based on the data shown in Figure 8b, suggesting that CD16 downregulation is occurring in response to activation of NK cells as a consequence of CD16 shedding, and can be inhibited by an ADAM17 inhibitor.

      We thank the reviewer for these suggestions, which we have used to present the ADAM17-inhibition data more clearly.

      (1) Effect on subset proportions, independent of degranulation. This is shown in Fig. 8B (now Figure 7B). Because the two subsets differ greatly in baseline frequency, with CD56<sup>dim</sup>CD16<sup>bright</sup> cells constituting the large majority of CD56<sup>dim</sup> NK cells before stimulation, a change in raw bulk proportion is small and difficult to interpret, for example a shift from roughly 95% to 92.5% of the bright population. To place the two subsets on comparable footing, the panel reports, for each subset, the frequency following target exposure minus its frequency in the matched unstimulated condition. Presented this way, ADAM17 inhibition clearly reduces the activation-associated change in subset proportions, consistent with reduced CD16 shedding. This normalization is stated in the legend and is now described in the Results text so that the analysis is not overlooked.

      (2) Interpretation of the ADAM17-inhibition data. We agree that CD16 downregulation occurs as a consequence of activation-induced shedding and is prevented by ADAM17 inhibition; this is not in dispute. We would, however, offer an additional observation that bears on whether the between-subset functional difference is itself a product of that shedding. In Fig. 8A (now Figure 7A), combining ADAM17 inhibition with NKG2D blockade reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation to 2.03%, below both anti-NKG2D alone at 5.24% and ADAM17 inhibition alone at 5.05% (p < 0.0001 and p = 0.0001 respectively). If the CD16<sup>dim</sup> advantage were solely a consequence of CD16 shedding, and if NKG2D blockade acted only by reducing that shedding, the combination could not reduce degranulation further than ADAM17 inhibition alone.

      The same figure also shows that the two perturbations act in opposite directions within the CD56<sup>dim</sup>CD16<sup>bright</sup> subset: anti-NKG2D reduced their degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells. Together with the opposite responses of the two subsets to ADAM17 inhibition described in the central response, this indicates that NKG2D and ADAM17 contribute through separable mechanisms and that the two subsets are not one population at different stages of shedding. These data are now presented in a more structured form and the interpretation is explicit in the revised text.

      Summary statement

      Taken together, many of the data presented in the manuscript are consistent with the very well-established downregulation of CD16 expression on activated NK cells, suggesting that the observed association between reduced CD16 expression on CD56dim NK cells and enhanced effector functions is a consequence of higher activation of these NK cells.

      We appreciate the reviewer articulating the central concern so clearly. We agree that CD16 downregulation on activated NK cells is well established and occurs in our assays; where we reach a different conclusion is on whether the enhanced function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is a consequence of that downregulation. As set out in the central response, three observations argue that it is not: the advantage is present in cells sorted into subsets before any target contact, where post-assay CD16 changes cannot apply (Fig. 2); the two subsets respond to ADAM17 inhibition in opposite directions, both in magnitude (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8), which is difficult to reconcile with their being one population at different activation levels; and blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone (Fig. 8, now Figure 7), indicating that the advantage is driven by NKG2D through a mechanism separable from shedding. We therefore interpret the association between low CD16 and enhanced function not as activation-induced downregulation of a single population, but as a property of a distinct, pre-existing subset. This interpretation is stated and defended explicitly in the revised Discussion, and will be strengthened with the expanded pre-sorted experiments.

      Reviewer #2 (Public review):

      (1) The central conclusion is weakened by the use of CD16 as a stable phenotypic marker. CD16 is well established to be rapidly downregulated following NK-cell activation and target cell (K562 or infected cells) engagement through ADAM17-mediated shedding. NK cell shedding regulates NK cell effector functions by promoting target cell detachment, boosting serial killing capacity, and preventing overstimulation. Therefore, NK cells displaying a CD56dimCD16dim phenotype after co-culture cannot be assumed to represent a pre-existing subset with intrinsically superior cytotoxic activity, but may instead correspond to activated CD56dimCD16bright NK cells that have downregulated CD16 during the assay. Because the vast majority of the functional experiments classified NK cell subsets based on post-assay CD16 expression, it is difficult to distinguish intrinsic functional differences between NK cell subsets from activation-induced phenotypic conversion. This limitation affects the interpretation of most of the study's principal findings.

      We thank the reviewer for this careful and well-articulated concern, which we recognize as the central issue of the review, and which we address in full in the central response above. We agree with the reviewer's premises: CD16 is rapidly shed by ADAM17 upon activation, and this shedding is itself functionally important, promoting target detachment, supporting serial engagement, and limiting overstimulation. Indeed, ADAM17-mediated shedding is integral to the serial degranulation mechanism we propose. We also agree that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion.

      For this reason, our conclusion does not rest on post-assay classification. As detailed in the central response, the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is demonstrated in cells sorted into subsets before any target contact, where the readout is direct lysis rather than post-assay gating and where activation-induced shedding therefore cannot account for the difference (Fig. 2), an experiment both this reviewer and the peer reviewer identify as the strongest in the manuscript. This is reinforced by evidence that the two subsets are functionally distinct rather than one population caught at different stages of shedding: they respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8); blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone, indicating a mechanism separable from shedding (Fig. 8, now Figure 7); and the subsets differ before stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We will strengthen this further by expanding the sorted-subset experiments across additional donors, and the revised text rests the manuscript's conclusions explicitly on the pre-sorted data.

      (2) The "killing frequency" analysis presented in Figure 3 is based on a mathematical estimate rather than a direct experimental measurement. Since total target cell killing is measured in mixed NK cell populations, it cannot be attributed to individual NK cell subsets. This experiment must be repeated using purified NK cell subsets.

      We agree that the killing frequency in Fig. 3 (now Figure 2—figure supplement 1C) is a mathematical estimate rather than a direct measurement. We would add that this is intrinsic to the metric: killing frequency is a derived quantity whether calculated from mixed or purified populations, so repeating it on purified subsets would not convert it into a direct measurement. This analysis has been moved to the supplementary material, and its limitations are stated in the Discussion, namely that killing frequency is an estimate and should be interpreted as such. Direct, subset-resolved killing is instead provided by Figure 2, in which NK cells sorted before target exposure show that purified CD56<sup>dim</sup>CD16<sup>dim</sup> cells lyse HIV-infected targets more efficiently than purified CD56<sup>dim</sup>CD16<sup>bright</sup> cells; this is the measurement on which our conclusion regarding direct killing rests, and it is the experiment we will expand across additional donors.

      (3) The serial degranulation assay presented in Figure 4 does not directly measure serial target cell killing and therefore does not support the conclusion that CD56dimCD16dim NK cells possess superior serial killing capacity. Furthermore, the increased serial degranulation observed in the CD16dim population could simply reflect activation-induced CD16 downregulation rather than an intrinsic property of this subset. This experiment should therefore be repeated using purified NK cell subsets.

      We agree that Fig. 4 (now Figure 3) measures serial degranulation, not serial killing directly. The text has been revised throughout to describe this as serial degranulation rather than serial killing, so that our conclusions match what was measured. Regarding the concern that increased serial degranulation in CD16<sup>dim</sup> cells reflects activation-induced CD16 downregulation, we address this in the central response; the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8) argue against that interpretation. As the reviewer suggests, we will repeat the serial degranulation assay on subsets purified before target exposure if cell yields from the sort permit. We note that this assay requires four sequential labelling and washing steps and that the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, so the number of sorted cells recovered may be limiting; we will report the outcome either way.

      (4) The finding that CD56dimCD16dim NK cells exhibit greater ADCC activity is somewhat counterintuitive given the central role of CD16 in mediating ADCC. Moreover, these experiments are likely confounded by activation-induced CD16 downregulation, which is expected to be even more pronounced during ADCC. Thus, the apparent superiority of the CD56dimCD16dim subset may simply reflect the conversion of activated CD56dimCD16bright NK cells into the CD16dim gate rather than intrinsically greater ADCC activity. To directly compare the intrinsic ADCC capacity of each subset, these experiments should be repeated using purified NK cell populations prior to target-cell stimulation.

      We agree that the greater antibody-dependent response of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is counterintuitive given the central role of CD16, and we regard it as an informative finding rather than an artifact. As set out in our revised Discussion, the surface density of gp120 on HIV-infected primary T-cells is approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), two orders of magnitude below high-density antigens such as CD20 on Raji cells at approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Antibody-dependent responses against HIV-infected cells therefore proceed under conditions of limiting antigen, and both CD56<sup>dim</sup> subsets face the same constraint. What differs between them is not the constraint but the NKG2D available to meet it: CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry higher NKG2D before target contact and respond more strongly, despite their lower CD16. Consistent with a requirement for a second signal under these conditions, ADAM17 inhibition reduced CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation even at 0 µg/mL VRC01 (Figs. 7B and 7C, now Figures 6B and 6C), where no antibody is present to engage CD16.

      Regarding the concern that this superiority reflects conversion of CD16<sup>bright</sup> cells into the CD16<sup>dim</sup> gate, we address this in full in the central response; the opposite responses of the two subsets to ADAM17 inhibition, in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), argue against it. As the reviewer recommends, we will directly compare the intrinsic antibody-dependent capacity of each subset using cells purified before target-cell stimulation, extending the pre-sorted approach of Figure 2 to the antibody-dependent setting across additional donors.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As discussed in the public review, the majority of the functional assays should be repeated using purified NK cell subsets. This approach would eliminate the confounding effect of activation-induced CD16 downregulation and allow the intrinsic functional properties of each subset to be directly compared.

      We agree, and this is the central experimental commitment of our revision. As described in the central response, we will repeat the specific-lysis and antibody-dependent degranulation assays using NK cell subsets purified before target-cell exposure, so that the intrinsic functional properties of each subset are compared directly and are not subject to activation-induced changes in CD16 expression. We will also perform the serial degranulation assay on sorted subsets if cell yields permit; that assay requires four sequential labelling and washing steps, and the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56 <sup>dim</sup> NK cells, so we cannot commit to it in advance of the sort. This extends the pre-sorted approach already used in Figure 2, which the reviewer identifies as the strongest evidence in the manuscript, across additional donors and across the functional readouts. These experiments require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (2) The experiments performed with purified NK cell subsets in Figure 2 provide the strongest evidence supporting the authors' conclusion that CD56dimCD16dim NK cells exhibit greater direct cytotoxicity against HIV-infected target cells. These data are the most convincing in the manuscript because they are not confounded by post-assay changes in CD16 expression. However, unlike the other functional assays, no representative gating strategy or raw flow cytometry plots are provided, and the results appear to be based on a single representative experiment. Given the importance of these data to the manuscript's central conclusion, this experiment should be expanded to include biological replicates from additional donors, representative flow cytometry plots, and validation using additional HIV-1 infectious molecular clones.

      We appreciate the reviewer identifying the sorted-subset experiments in Figure 2 as the strongest evidence for our conclusion, and we agree these data warrant expansion. In the revised manuscript we will:

      (1) Expand the experiment to include biological replicates from additional donors, with all donors shown rather than a single representative experiment.

      (2) Provide the representative gating strategy and flow cytometry plots for the sorted subsets. The reviewer is correct that these should be included, and we will add them, including for the expanded experiments.

      (3) Validate the finding using additional HIV-1 strains. We note, for clarity, that the virus used throughout this study is a primary patient isolate (HIV-1<sup>SHM-1</sup>), as stated in the Materials and Methods, rather than an infectious molecular clone. We have now compared NK cell degranulation against autologous CD4<sup>positive</sup> T-cells productively infected with HIV-1<sup>SHM-1</sup>, with the X4-tropic infectious molecular clone HIV-1<sup>NL4-3</sup>, and with the R5-tropic laboratory-adapted strain HIV-1<sup>BaL</sup>, at three effector cell to target cell ratios. CD56 <sup>dim</sup> CD16 <sup>dim</sup> cells degranulated more than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells against every virus at every ratio, in all nine comparisons at p < 0.0001. Both subsets responded less to HIV-1<sup>NL4-3</sup> and HIV-1<sup>BaL</sup> than to HIV-1<sup>SHM-1</sup>, and did so in proportion: the ratio of CD56 <sup>dim</sup>CD16 <sup>dim</sup> to CD56 <sup>dim</sup> CD16<sup>bright</sup> degranulation ranged from 2.7 to 3.8 across all nine conditions. The magnitude of the response therefore varies with the virus, whereas the relationship between the two subsets does not. These data are included in the revised manuscript as Figure 1—figure supplement 3, with the statistical analysis in a supplemental table.

      The experiments described in points 1 and 2 require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (3) The mechanism underlying the enhanced effector function of CD56dimCD16dim NK cells remains unclear. Although the phenotypic characterization presented in Figure 5 (and related supplement figures) is informative, NK cell receptor expression was assessed after target cell stimulation, when it may already have been altered by activation and CD16 downregulation. Receptor expression should therefore be evaluated prior to stimulation. In addition to NKG2D, the authors should also consider assessing additional activating receptors, notably NKp30, which has recently been implicated in the elimination of autologous HIV-1-infected cells (PMID: 41079618).

      We agree that receptor expression should be assessed before stimulation, and in fact it is. In Fig. 6 (now Figure 5), the receptor gMFI data include the no-target condition, showing that CD56<sup>dim</sup>CD16<sup>dim</sup> cells express 1182 gMFI more NKG2D than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact (p < 0.0001), approximately 1.6-fold, and therefore before any activation-induced change in receptor expression. NKp46, measured on the same cells in the same wells, did not differ between the subsets under the same condition (mean difference 45.67 gMFI, p = 0.0668), indicating that the difference is specific to NKG2D rather than a general difference in activating receptor density. The baseline condition is now labelled explicitly as 1:0 in the revised figure and described as such in the Results text.

      Regarding NKp30, we have assessed this receptor. Degranulation did not differ between NKp30 positive and NKp30 negative cells within either CD56<sup>dim</sup> subset, whereas both CD56<sup>dim</sup>CD16<sup>dim</sup> groups exceeded both CD56<sup>dim</sup>CD16<sup>bright</sup> groups regardless of NKp30 status, indicating that the enhanced degranulation of the subset is not attributable to NKp30, paralleling our finding for NKp46. These NKp30 data are included in the revised manuscript as Figure 5—figure supplement 1, with the corresponding statistical analysis in a supplemental table. We note that this analysis addresses whether NKp30 accounts for the difference between the subsets; it does not exclude a role for NKp30 in NK-cell recognition of HIV-infected cells more generally, consistent with the study the reviewer cites, which we now discuss.

      (4) In Figure 5, the histograms corresponding to the uninfected and ΔVpr conditions appear to be identical. If this is indeed the case, this represents a serious concern, as these are two distinct experimental conditions and should not be represented by the same flow cytometry plot. This raises the possibility of an inadvertent panel duplication. The authors should carefully verify the figure and replace the duplicated panel if necessary.

      We thank the reviewer for this careful observation. On review, the histograms in Figure 5A were reproduced from our earlier study (Ward et al., PLoS Pathog 2009;5(10):e1000613) and should not have been included without attribution. We have removed that panel and cite the original finding in its place.

      Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      (5) The ADCC experiments and calculation require additional methodological clarification, particularly the analyses presented in Figure 7C. Although NK cell degranulation is commonly used as a surrogate marker of ADCC, the data presented in Figure 7 do not appear to isolate the antibody-dependent component of the response. To specifically quantify ADCC-mediated degranulation, the degranulation induced by HIV-infected target cells alone (i.e., in the absence of VRC01) should be subtracted from that measured in the presence of VRC01. Notably, in Figure 7C (DMSO), the CD56dimCD16dim population appears to exhibit similar levels of degranulation in the absence and presence of VRC01, suggesting that antibody-dependent degranulation may be limited in this subset, which does not support the author's conclusions.

      We thank the reviewer for raising this, and we agree that the antibody-dependent and antibody-independent components of the response should be distinguished. We would, however, respectfully argue against the subtraction as a means of doing so, and we believe the experiments already in the manuscript address the underlying question more directly.

      The condition without VRC01 is not a background to be removed. It is the NKG2D-driven response of the same cells to the same infected targets, measured through the same degranulation machinery, and it is one of the principal findings of the study. Subtracting it treats the two components as though they were independent and additive, when both converge on a single immunological synapse and a single degranulation event per cell. The difference between the two conditions is therefore not the antibody-dependent response; it is the increment in total degranulation produced by adding antibody, which is a different quantity and one that carries no clean interpretation at the level of the individual cell.

      The question the reviewer raises, whether the antibody-dependent component differs between the subsets, is answered directly by the two-way ANOVA of these data. Across the VRC01 titration, the effect of NK cell subset accounts for 85.91% of the total variation (F(1, 16) = 424.0, p < 0.0001) and the effect of VRC01 concentration for 10.18% (F(3, 16) = 16.74, p < 0.0001), while the subset × VRC01 interaction is not significant (F(3, 16) = 1.112, p = 0.3732) and accounts for 0.68% (Supplemental Table 20 in the revised manuscript). The absence of an interaction means that adding antibody raises the response of both subsets by a comparable amount, and that the difference between the subsets is the same at every VRC01 concentration tested. This is a statistical statement about the antibody-dependent component, obtained without subtracting one condition from another.

      We agree with the implication the reviewer draws from this, and we state it plainly in the revised Discussion: the antibody-dependent increment is modest in both subsets. We attribute this to the very low surface density of gp120 on HIV-infected primary T-cells, approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), which is two orders of magnitude below high-density antigens such as CD20 on Raji cells, approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Under these conditions the antibody-dependent signal available to any NK cell is limited, and this applies equally to both subsets.

      Where we differ from the reviewer is on the conclusion this supports. That the antibody-dependent increment is modest in both subsets does not weaken our central claim, which is comparative: at every VRC01 concentration tested, including in the presence of antibody, CD56<sup>dim</sup>CD16<sup>dim</sup> cells degranulate more than CD56<sup>dim</sup>CD16<sup>bright</sup> cells against antibody-coated HIV-infected targets. That comparison is what the manuscript reports, and it is unaffected by how the response is partitioned between its antibody-dependent and antibody-independent components.

      We also note that the experiments in Figs. 7B and 7C (now Figures 6B and 6C) do isolate a component of the response experimentally rather than arithmetically. Inhibiting ADAM17 removes the contribution that depends on CD16 turnover, and it does so in opposite directions in the two subsets, reducing CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation and increasing that of CD56<sup>dim</sup>CD16<sup>bright</sup> cells at every VRC01 concentration. These are direct experimental manipulations of the antibody-dependent pathway, and they are more informative than the arithmetic difference between two conditions.

      Finally, we take the reviewer's point that the analyses in Fig. 7C require clearer explanation. In the revised manuscript we state explicitly what is plotted, namely the percentage of each CD16 subset among CD107a positive CD56<sup>dim</sup> NK cells, we describe the background subtraction that applies to all CD107a data in this study, and we report the statistical analysis of each panel in full.

      (6) It is also unclear how the authors interpret the effects of ADAM17 inhibition. While ADAM17 inhibition increases the ADCC activity of the CD56dimCD16bright population, it simultaneously decreases that of the CD56dimCD16dim population. An alternative explanation is that inhibition of CD16 shedding prevents activated CD56dimCD16bright NK cells from transitioning into CD56dimCD16dim during the assay. This possibility should be discussed and experimentally addressed, as it provides a plausible alternative interpretation of the observed phenotype.

      We thank the reviewer for articulating this alternative, which we address in full in the central response. We agree that the opposite effects of ADAM17 inhibition on the two subsets are central to interpreting these experiments, and we interpret them as evidence that the two are distinct populations rather than one transitioning into the other.

      The reviewer's alternative, that ADAM17 inhibition prevents CD56<sup>dim</sup>CD16<sup>bright</sup> cells from transitioning into the CD56<sup>dim</sup>CD16<sup>dim</sup> gate, predicts that blocking shedding should reduce the CD56<sup>dim</sup>CD16<sup>dim</sup> population by cutting off its supply from CD16<sup>bright</sup> cells. Three observations argue against this being the explanation for the functional difference. First, the effect is not merely a change in population size but a change in per-cell function in opposite directions: ADAM17 inhibition increases the number of serial degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing them in CD56<sup>dim</sup>CD16<sup>dim</sup> cells (Fig. 9, now Figure 8), which is difficult to explain if the dim cells were simply bright cells prevented from converting. Second, blocking NKG2D together with ADAM17 reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A); if NKG2D blockade acted only by reducing the shedding that drives the putative transition, the combination could not exceed the effect of ADAM17 inhibition alone. Third, we have tracked the fate of each subset sorted before target exposure: at one hour, when degranulation is maximal, only approximately 4% of sorted CD16<sup>bright</sup> cells were found in the CD16<sup>dim</sup> gate, while approximately 35% had moved to the CD16<sup>negative</sup> gate (Author response image 1 accompanying this response). Shedding therefore directs CD16<sup>bright</sup> cells past the CD16<sup>dim</sup> gate rather than into it.

      We discuss this alternative explicitly in the revised Discussion and will address it further experimentally by repeating these assays on subsets purified before target exposure, where no transition can occur during the assay.

      Reviewing Editor Comments:

      The conclusion that CD56dimCD16dim NK cells are intrinsically superior effectors against HIV-infected target cells requires additional evidence because CD16 is rapidly downregulated following NK-cell activation. Throughout most of the study, NK-cell subsets are classified after target-cell encounter, making it difficult to distinguish pre-existing CD56dimCD16dim cells from activated CD56dimCD16bright cells that have undergone ADAM17-mediated CD16 shedding. The authors should repeat functional experiments using NK-cell subsets purified before target-cell exposure and determine the extent to which ADAM17 inhibition alters subset frequencies and functional readouts. These experiments are essential to establish whether the observed functional differences reflect intrinsic biology rather than activation-induced phenotypic conversion.

      We thank the editor for this clear synthesis of the central concern, which we address in full in the central response above. In brief, our conclusion does not rest on post-encounter classification: the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is established in cells sorted into subsets before any target contact, using direct lysis as the readout (Fig. 2), and is reinforced by the opposite responses of the two subsets to ADAM17 inhibition in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), by the separable contributions of NKG2D and ADAM17 (Fig. 8, now Figure 7), and by pre-stimulation differences between the subsets, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We agree these questions are central and will repeat the functional experiments on subsets purified before target exposure, and the effect of ADAM17 inhibition on subset frequencies is presented explicitly in Figs. 7C and 8B (now Figures 6C and 7B).

      Major conclusions should be supported by more rigorous experimental validation. In particular, the sorted NK-cell experiments should be expanded using multiple independent donors, include killing of uninfected target cells as controls and provide representative gating strategies and flow cytometry plots. Likewise, the current analyses of killing frequency, serial killing, and ADCC should be strengthened by direct measurements using purified NK-cell subsets rather than mathematical estimates or analyses performed in mixed NK-cell populations.

      We agree and will strengthen the validation as follows. The sorted-subset experiments will be expanded across multiple independent donors, with all donors shown. Uninfected-target controls will be included for the sorted-cell lysis experiments (Fig. 2); the corresponding CD107a controls against uninfected targets are already presented in Figure 1—Figure Supplement 4C of the revised manuscript. Representative gating strategies and flow cytometry plots will be provided for the sorted-cell experiments, as for our other assays. Regarding direct measurement: the killing-frequency metric (Fig. 3, now Figure 2—figure supplement 1C) is a mathematical estimate whether derived from mixed or purified populations, and it has been moved to the supplementary material with this limitation noted in the Discussion, while direct, subset-resolved killing is provided by the pre-sorted lysis experiment (Fig. 2), which we will expand; the serial degranulation assay (Fig. 4, now Figure 3) is now described as serial degranulation rather than serial killing, and will be repeated on purified subsets if cell yields from the sort permit; and the antibody-dependent comparison will be performed on subsets purified before stimulation.

      Some aspects of the data analysis and presentation require clarification. The authors should evaluate receptor expression before target-cell stimulation, clarify the ADCC analyses and interpretation of ADAM17 inhibition, verify the apparent duplicated flow-cytometry panel, apply appropriate statistical corrections for multiple comparisons where necessary, and streamline the presentation by emphasizing the principal conclusions rather than extensive descriptive analyses.

      We have addressed each of these. Receptor expression before stimulation is shown in the gMFI data of Fig. 6 (now Figure 5) at the no-target condition, where CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry 1182 gMFI more NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells (p < 0.0001) with no corresponding difference in NKp46 (p = 0.0668); this condition is now labelled explicitly as 1:0 in the figure and described as such in the Results text. The antibody-dependent analyses and the interpretation of ADAM17 inhibition are clarified in the revised text, as detailed in our responses to the three reviewers and the central response.

      On the duplicated panel: the histograms in Figure 5A were reproduced from Ward et al. (2009) and should not have been included without attribution. That panel has been removed and the original finding is cited in its place. Figures 5B and 5C are both new results and are retained, becoming Figure 4—figure supplement 1 and Figures 4A and 4B respectively.

      We have applied appropriate multiple-comparison corrections, using the post-hoc test matched to each comparison structure and reporting adjusted p-values throughout; the design and post-hoc test used for each figure and panel are given in a new supplemental table. Finally, we have streamlined the presentation by opening each Results section with a statement of the principal finding before the supporting data, and by reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detailed data retained in the supplemental tables and their significance synthesized in the Discussion.

      References

      Amand M, Iserentant G, Poli A, Sleiman M, Fievez V, Sanchez IP, Sauvageot N, Michel T, Aouali N, Janji B, Trujillo-Vargas CM, Seguin-Devaux C, Zimmer J. 2017. Human CD56<sup>dim</sup>CD16<sup>dim</sup> cells as an individualized natural killer cell subset. Frontiers in Immunology 8:699. doi:10.3389/fimmu.2017.00699.

      Lallemand C, Liang F, Staub F, Simansour M, Vallette B, Huang L, Ferrando-Miguel R, Tovey MG. 2017. A novel system for the quantification of the ADCC activity of therapeutic antibodies. Journal of Immunology Research 2017:3908289. doi:10.1155/2017/3908289.

      Vasiliver-Shamis G, Tuen M, Wu TW, Starr T, Cameron TO, Thomson R, Kaur G, Liu J, Visciano ML, Li H, Kumar R, Ansari R, Han DP, Cho MW, Dustin ML, Hioe CE. 2008. Human immunodeficiency virus type 1 envelope gp120 induces a stop signal and virological synapse formation in noninfected CD4+ T cells. Journal of Virology 82:9445-9457. doi:10.1128/JVI.00835-08.

      Ward J, Davis Z, DeHart J, Zimmerman E, Bosque A, Brunetta E, Mavilio D, Planelles V, Barker E. 2009. HIV-1 Vpr triggers natural killer cell-mediated lysis of infected cells through activation of the ATR-mediated DNA damage response. PLoS Pathogens 5(10):e1000613. doi:10.1371/journal.ppat.1000613.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We have made major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We have changed the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions have greatly improved our work and presentation and we are grateful to the reviewers and our editors. We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have now included an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS strengthens our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results are reported in the new Supplemental Table 6.

      We have now included a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results are now reported in the manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We have addressed other recommendations from this reviewer as follows. We cite and discuss previous SCD rodent model work observing decreased butyrate in disease, and we have updated our results and discussion sections regarding associations between the microbiome and virome and clinical and molecular features.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself. (REVISION POINT 3)

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have tempered our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We have strengthened our control of clinical confounders, and added clearer statistical correction, as described below, with corresponding updates to the manuscript. We look forward to conducting future studies to test causality and understand mechanism.

      We have now done sensitivity analysis within SCD patients to determine whether our microbiome and virome results associate with treatment and clinical severity. To evaluate whether microbiome and virome features were explained by clinical or demographic heterogeneity within the SCD cohort, we restricted analyses to SCD participants. We fit a separate multivariable regression model for each feature. Each model included age, sex, hydroxyurea use, transfusions in the past year, and acute care utilization in the past year as predictors.

      feature_z <sup>~</sup> age_z + sex_F + HU + log1p(TxPastYr)_z + log1p(AcuteCarePastYr)_z

      Non-negative abundance, ratio, pathway, viral, and count-like variables were log-transformed to reduce skew, using feature-specific pseudocounts for zero-containing microbiome/virome variables and ln (1 + x) transformation for count covariates. Diversity and indicator scores were not log-transformed. Continuous variables were then standardized to Z-scores before modeling. Models were fit using ordinary least squares with HC3 robust standard errors. FDR correction was applied separately for each model term across the tested microbiome and virome features. No microbiome or virome feature showed an FDR-significant association with hydroxyurea use, transfusions in the past year, or acute care utilization in the past year. These results are reported in the manuscript and in the new Supplemental Table 7.

      We have reported multiple-testing correction results for all associations between microbiome and virome features and clinical and molecular features and updated manuscript figures accordingly.

      To evaluate relationships between microbiome/virome features and clinical or immune markers within the SCD cohort, we performed Spearman correlation analyses. Microbiome and virome features were organized into four prespecified feature groups: community metrics, taxa, functional pathways, and viral features. Clinical and immune markers were grouped into marker sets for visualization and multiple-testing correction, including hematologic clinical markers, hemolysis markers, creatinine, acute care burden, and cytokines/chemokines. Spearman correlation coefficients were calculated for each feature–marker pair. Benjamini-Hochberg FDR correction was applied within each prespecified feature group by marker group block. Nominal associations were defined as p < 0.05, FDR-significant associations as q < 0.05, and trends as q < 0.10. In the heatmap figure, boxes now indicate nominal p < 0.05, asterisks indicate q < 0.05, and daggers indicate q < 0.10.

      We have addressed other recommendations from this reviewer as follows. We have modified our language describing prophage/immune associations. We have revised the Methods to clarify how abundance data were processed before MaAsLin2 modeling. Taxonomic profiles were analyzed using MaAsLin2 with total-sum scaling normalization and log transformation, while pathway profiles were analyzed without additional normalization because the input pathway table had already been normalized prior to MaAsLin2 analysis; MaAsLin2 log transformation was then applied. We agree that relative abundance metagenomic data are compositional, and we have revised the text to clarify that these analyses identify covariate-adjusted associations with transformed relative abundance rather than absolute abundance. We also now note this as a limitation of the study. We also note the limitations of F:B as a metric. We have added text describing the need for further analysis of the prophages to understand their patterns of host range and transmission and their associations with features such as shared geography, health care exposure, and diet. Finally, we have updated Figure 5 to reflect our updated analysis with significance indicated.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We have now noted in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single-centre, cross-sectional. We have further made modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review)

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revised version.

      With the last round of revisions, the authors have now addressed my concerns.

      Thank you to re-review this revision, and we all appreciate you kindly contributing to substantially improve the conceptualization, statistics and statements for this manuscript.

      Reviewer #4 (Public review):

      Summary:

      The current study tested the effects of repeated sessions of tDCS targeting the DLPFC on procrastination behavior. The main outcome is that anodal versus sham DLPFC tDCS reduces procrastination behavior on both a short-term and a long-term scale up to six months after the stimulation sessions.

      Strengths:

      The current study tests competing models of procrastination with state-of-the-art high-definition transcranial electric stimulation. The study assesses stimulation effects on procrastination on both a short-term and a long-term scale, suggesting that repeated stimulation of the prefrontal cortex reduces procrastination on a time scale of up to six months.

      Weaknesses:

      The manuscript has already been reviewed and revised before, and it seems that the quality of the manuscript has substantially improved as a result of this revision process. I agree with the other reviewers that one must be cautious with drawing conclusions regarding the cognitive mechanisms underlying this effect, as many different cognitive functions are implemented by the DLPFC.

      We do appreciate you to take valuable time offering those insightful and helpful comments on this revised manuscript. As you kindly raised, this revision has redrawn conclusions and statements on domain-specific mechanistic roles of DLPFC in interpreting why this neuromodulation treatments are effective.

      One aspect of the current results that puzzles me is the strength of the current stimulation effects. Meta-analyses suggest that tDCS shows only small-to-moderate effect sizes (with Cohen's d around 0.5). While the authors report no effect sizes for their statistical models, the small p values, in combination with the unusually small sample size of 18 participants per group, suggests that the effect size must be rather large. Can the authors provide an estimate of the effect size of their stimulation effects? If they are considerably larger than to be expected, could the authors give an explanation for why their stimulation setup is showing much stronger effects than comparable high-definition tDCS studies on cognition or decision making?

      Thank you for raising this very crucial question in effect size determination. We fully understand that this large effect size makes you puzzled, and that the limited sample size indeed attenuates detectable power in the statistics. We completely agree that reporting the actual effect sizes is essential for interpreting the magnitude of our findings, and we appreciate the opportunity to clarify why the observed effects in our study appear substantially larger than the small-to-moderate effect sizes typically reported in meta-analyses of tDCS studies on cognition and decision-making.

      Following your suggestion, we have calculated the effect sizes for our primary outcomes. Based on the simple effect analyses (pre- vs. post-neuromodulation within the active neuromodulation group), we computed Cohen’s d for the within-group changes: for task-execution willingness, d = 2.37 (95% CI [1.49, 3.25]); for the actual procrastination rate, d = 1.52 (95% CI [0.87, 2.16]).

      We acknowledge that these effect sizes are considerably larger than the typical d ≈ 0.5 reported in the tDCS literature. After careful consideration, we attribute this discrepancy to three methodological and conceptual differences between our study and typical cognitive/decision-making tDCS studies. First, the small-to-moderate effect sizes are observed in studies utilizing a single-session tDCS protocol, yet our study employed an intensive 7-session HD-tDCS protocol over 15 days. Therefore, multi-session tDCS that induces cumulative, activity-dependent long-term potentiation (LTP)-like plasticity may substantially amplify and consolidates behavioral effects compared to single-session stimulation (Ke et al., 2023; Zhong et al., 2021). Second, most tDCS studies on cognition recruit healthy young adults who often perform near ceiling on laboratory tasks, leaving little "room for improvement" and thereby constraining the observable effect size. In our study, we strictly screened for severe chronic procrastinators. Because our participants had severe baseline deficits in task execution, the "ceiling space" for behavioral improvement was much larger, naturally inflating the observable effect size of the intervention. Lastly, given all the procrastinators completed tasks in the last session (0% procrastination rate in the active group, without within-group variance), mathematically, this boundary variable (0% vs 100%) artificially inflates the effect size estimate when calculating Cohen’s d with a near-zero post-test standard deviation.

      Nevertheless, as a sensitivity analysis, those findings are confirmed by Beta regression model addressing the risks of boundary variables, indicating that the potential inflation of effect sizes is statistically acceptable:

      In summary, while the observed Cohen's d values are unusually large, they are contextually justified by the cumulative nature of our multi-session protocol, the targeted clinical-like population, and the mathematical properties of bounded behavioral metrics.

      Results Section (Page 9, Line 441-444)

      “... For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001, Cohen d = 2.37, 95% CI: 1.49-3.25; Fig. 3A and Fig. S2a).”

      Results Section (Page 9, Line 454-458)

      “... Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004, Cohen d = 1.52, 95% CI: 0.87-2.16), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group.”

      Regarding the strengths of the stimulation effects, I moreover found remarkable that the post-test procrastination rate was 100% in all (!) participants in the DLPFC group (figure 3F). I admit that it is hard to trust results that have no individual variation at all. This means that all participants are perfect responders to tDCS, which is again at variance what one typically expects for tDCS (where one usually has many non-responders). Do the authors have an explanation for this?

      Thank you for raising this highly important and helpful comment. Indeed, we fully understand that this result (a 100% task complete rate among all participants in the DLPFC group) is somewhat extraordinary. This pattern was equally striking to us when unblinded the data. After carefully scrutinizing the data and statistics, we are thrilled to confirm that this pattern is true. In support of this observation, we were gratified to receive numerous thank-you letters from participants who engaged in active neuromodulation. They expressed gratitude to us, and reported that they have substantially ameliorated procrastination behavior in real-life activities after completing the trial. While this does not constitute formal scientific evidence, we are also glad to see the benefits of this neuromodulation for those procrastinators.

      Two reasons could account for this pattern herein. One interpretation is to attribute this pattern to “floor effect”. In the present study, the procrastination rate was calculated as 1 minus the task-completion rate (e.g., 80%, 60%, 40%) by the deadline. At last stimulation sessions (#6 and #7), all the participants completed their real-life tasks before the deadline, yielding a 0% (1 minus 100% completion rate) procrastination rate, without any between-individual variation. Thus, rather than there being no individual variation in procrastination, this scalar – the procrastination rate - is too insensitive to capture subtle differences per se. For instance, although participants #1 and #2 both showed a 0% procrastination rate - meaning that both completed their tasks before the deadline - Participant #1 might have completed it 3 hours before the deadline, whereas Participant #2 might have completed it only 10 minutes before. In this case, the “scalar inflation” emerges to let us perceive that both participants have equivalent procrastination rates, although participant #2 may have a higher procrastination level than #1. As conceptually defined in the field, procrastination is contextualized as “not completing a task before the deadline”. Thus, if this task is completed before the deadline, regardless of whether it was finished close to or far in advance of the deadline, this case is defined as “no procrastination”. In the present study, the primary outcome is whether a participant procrastinated on a real-life task before the deadline in real-world settings, irrespective of when she/he completed this task. Thus, this scalar - procrastination rate - fits our conceptualization of procrastination.

      Another reason is the potential accumulative effects from sequential multi-session tDCS stimulation, as we explained above. As shown in Mann-Kendall trend tests, the procrastination rates show a significant linear downtrend in the active neuromodulation group across sessions, even after removing sessions #6 and #7. This indicates that the improvements of going against procrastination may be sequentially accumulative along with the increase in sessions, implying a potential “dose-dependent effect”. Despite a speculative interpretation, this “dose-dependent effect” in neuromodulation has been well-documented in previous studies, showing the robustly linear association between the number of sessions and effectiveness (c.f., Cole et al., 2020; Hutton et al., 2023; Sabé et al., 2024; Schulze et al., 2018). Therefore, although this extreme pattern is somewhat extraordinary compared to previous observations, it makes sense.

      We also conducted robustness check by removing sessions #6, #7, and both, to validate whether this results were biased by “scalar inflation”. We do believe that this analysis could support statistical robustness to go against potential biases from extreme cells. By doing so, we found that all the group*treatment_day interaction effects remained significant when removing either session #6 or session #7 (or even both, all p-values < .05), indicating high statistical robustness. Please see Table S3 and Table S4.

      Taken together, in spite of their being extraordinary, we confirm that those findings are statistically robust to extreme outliers. As you kindly suggested, we have added those findings of the robustness check into the revised Supplemental Materials section.

      In any case, I am surprised by the rather small sample size. Due to the small effect sizes for tDCS, it is common to have a minimum of 30 subjects per group in between-subject designs. According to G*Power, a between-subject design with 17 subjects per group could detect only relatively large effect sizes of Cohen's d = 0.99 (alpha = 5%, power = 80%, independent-samples t-test). As explained above, this is far above the effect size that can be expected for tDCS. In addition, small samples bear the risk that results strongly depend on outliers in the data, which might explain the strong effect size observed in the current study. The small sample size should be discussed as a major limitation of the current study and that the results need to be replicated by studies with larger sample sizes. Moreover, to rule out that the results are driven by outlier in the data, the authors should show individual data points in all plots showing empirical data.

      We sincerely thank you for this highly constructive and methodologically sound critique. We completely agree that sample size is a critical consideration in tDCS research, and that visualizing individual data points is essential to rule out the possibility that our findings are driven by outliers.

      We acknowledge that our sample size is smaller than the ~30 per group often recommended for detecting small-to-moderate effects in general cognitive tDCS meta-analyses. We have determined this a priori effect size based on the existing work we published previously (Xu et al., 2023, J Exp Psychol Gen;152(4):1122-1133). In our pilot study (Xu et al., 2023), we identified a significant interaction effect between the single-session tDCS stimulation (active vs sham) and time (pre-test vs post-test) (t = 2.38, p = .02, n = 27; 95% CI [0.14, 1.49]) for changing procrastination willingness in the laboratory settings, indicating a medium effect size. Based on this specific empirical foundation, GPower indicated that a total sample size of 34 (17 per group) was sufficient to achieve 80% power (please see GPower output below). To account for potential attrition, we aimed to recruit 36 participants (18 per group), ultimately retaining 46 participants (23 per group) after exclusions. While we stand by this a priori justification, we fully agree with your overarching point that this remains a constraint.

      We completely agree with your observation regarding the plots. The apparent absence of data points in the previous versions of Figures 3B and 3F was not due to data exclusion, but rather to severe overplotting. Because multiple participants in the active neuromodulation group achieved identical scores (e.g., 0% procrastination rate or 100% task-execution willingness in later sessions), their data points perfectly overlapped, making it appear as though only ~10 points were present. As you helpfully suggested, we now employ jittered scatter plots with adjusted transparency, ensuring that all 23 individual data points per group are clearly visible, even when values are identical. As these revised figures demonstrate, the significant group differences reflect a consistent, cohort-wide shift rather than the influence of isolated outliers. Those

      As you rightly suggested, we have explicitly framed the small sample size as a major limitation and emphasized the necessity for large-scale replication. We have strengthened the wording in the Limitations section to explicitly mention the risk of outlier dependency and the need for larger cohorts.

      Legend Section (Page 28, Line 1161-1163)

      “… To ensure transparency and rule out outlier-driven effects, individual data points for all participants (N=23 per group) are overlaid on the bars using a jittered distribution to prevent overplotting of identical values.”

      Discussion Section (Page 13, Line 691-696)

      “… a major limitation of the current study is the relatively small sample size (total N = 46). While this was determined a priori based on our specific pilot study, small samples inherently bear a higher risk of being influenced by outliers and may overestimate effect sizes compared to large-scale meta-analytic expectations for tDCS. Therefore, these findings warrant caution in generalization and necessitate rigorous replication in larger, adequately powered cohorts.”

      Related to this, in the figure showing individual data points (3B/F), I count only around 10 data points per tDCS group for the 18 participants per group. I ask the authors to modify the plot that the data points from all participants can be seen (for example, by adding some noise on the x-axis for participants with the same value on the y axis).

      Thank you for this kind reminder. As we replied above, those plots have been redrawn by adding the jitters, which favor the readability as you kindly suggested.

      Another surprising aspect of the data is that repeated sessions of tDCS change procrastination behavior up to six months after stimulation. Do the authors think that their tDCS setup leads to such long-lasting neuroplastic changes, and if yes, can they cite prior work where similar dosages of tDCS also showed such long-lasting effects? Or could the results be explained by learning effects, for example because participants in the DLPFC group learned during the repeated tDCS sessions that it feels internally rewarding to finish one's tasks instead of procrastinating them, and they still benefit from this kind of "learned industriousness" 6 months later? In any case, in my view it is important to be more specific about how seven sessions of tDCS can affect behavior half a year later.

      We sincerely thank the reviewer for this highly insightful and thought-provoking comment. The concept of "learned industriousness" is particularly apt and captures a crucial alternative mechanism that we must address. We agree that explaining how seven sessions of tDCS can affect behavior half a year later requires a nuanced discussion of both neurobiological and behavioral learning mechanisms.

      Regarding the first point, we do believe that our multi-session protocol can induce long-lasting neuroplastic changes. While single-session tDCS effects are typically transient, cumulative neurobiological evidence demonstrates that repeated, multi-session protocols (typically ranging from 5 to 10 sessions) can induce activity-dependent, long-term potentiation (LTP)-like plasticity that consolidates over time (Agboada et al., 2020; Au et al., 2017; Jannati et al., 2023). Our 7-session protocol falls squarely within this range of "intensified dosing" designed to promote such consolidation. Meta-analyses and empirical studies on multi-session tDCS have shown that such protocols can produce behavioral and neurophysiological effects lasting weeks to months, particularly when targeting prefrontal regions involved in value-based decision-making and cognitive control (e.g., Brunoni et al., 2013; Sabé et al., 2024; Woodham et al., 2025).

      Furthermore, we completely agree with you for this alternative explanation regarding learning effects. It is highly plausible that participants in the active group, experiencing reduced task aversiveness and increased outcome value during the intervention, learned that completing tasks is internally rewarding. This aligns perfectly with the psychological concept of "learned industriousness" (Eisenberger, 1992), where the reinforcement of effortful behavior makes future engagement more likely. We explicitly acknowledge that repeated exposure to the experience-sampling protocol and the positive feedback of task completion could facilitate this kind of behavioral learning. More importantly, we argue that the learning effects are not bad things in this neuromodulation, and the learning effect and the neuroplasticity may be synergistic. The sham control group underwent the exact same experience-sampling protocol, reported real-life tasks, and had the identical opportunity for "learned industriousness" through feedback. However, as identified in the half-year follow-up, the sham group did not exhibit the same progressive improvement during the intervention, nor did they sustain a significant reduction in procrastination at the 6-month follow-up (their rates returned to near-baseline levels). This divergence suggests that while learning may play a role, the active neuromodulation likely provided the necessary neuroplastic "boost" (e.g., by enhancing prefrontal value-encoding circuits) that facilitated, accelerated, and consolidated this learning, making the behavioral change durable. Without the neuromodulatory enhancement, the mere exposure to the protocol was insufficient to produce long-term change.

      As you kindly suggested, we have explicitly incorporated this nuanced discussion into the revised manuscript, by citing relevant literature on multi-session tDCS plasticity, explicitly acknowledging the "learned industriousness" hypothesis, and reiterating the limitation of having only a single follow-up point.

      Discussion Section (Page 12, Line 627-634)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Participants in the active group may have learned during the repeated sessions that completing tasks feels internally rewarding, thereby benefiting from a form of “learned industriousness” (Eisenberger, 1992) that persists months later.”

      Discussion Section (Page 14, Line 719-723)

      “... we explicitly note that a single 6-month follow-up timepoint cannot definitively establish the stability or trajectory of these effects. Future studies incorporating multiple longitudinal assessments (e.g., 1-month, 3-month, 6-month, 12-month) are required to substantiate claims about long-term retention and to disentangle the precise contributions of neuroplasticity versus behavioral learning.”

      Lastly, the link to the data repository works, but I could not inspect the data because I was asked to request access to the data, which I did not do in order to remain anonymous.

      Thank you a lot to take invaluable to review our data and code in this repository. As we reported previously, all the data and code to support those findings have been deposited in the eLife online submission system for your reviews and scrutiny before this manuscript is formally published. As the editorial policy of eLife on VOR (Version of Record) instructed, to prevent from mixture of codes and data across multiple round of revisions, those data and codes in the final version would be released once this paper is formally published. Please do not worry for the anonymity policy. This is a public peer review, and it thus enables those helpful comments that you kindly suggested to be public when this manuscript is formally published. Again, thank you to substantially contribute on this revised manuscript by sharing those helpful suggestions.

      Recommendations for the authors:

      Editors note: We encourage the authors to consider the remaining reviewer concerns and revise the manuscript accordingly.

      Thank you so much for this warm and kind reminder. We have addressed all of those concerns that remained by the new Reviewer #4, point-by-point. All the co-authors do appreciate you for handling our manuscript, and for contributing those fruitful and helpful comments. We do believe that the quality of this manuscript has been substantially improved, benefiting from this editorial process.

      References

      Agboada, D., Mosayebi-Samani, M., Kuo, M. F., & Nitsche, M. A. (2020). Induction of long-term potentiation-like plasticity in the primary motor cortex with repeated anodal transcranial direct current stimulation - Better effects with intensified protocols? Brain Stimulation, 13(4), 987–997. https://doi.org/10.1016/j.brs.2020.04.009

      Au, J., Karsten, C., Buschkuehl, M., & Jaeggi, S. M. (2017). Optimizing transcranial direct current stimulation protocols to promote long-term learning. Journal of Cognitive Enhancement, 1(1), 65–72. https://doi.org/10.1007/s41465-017-0007-6

      Brunoni, A. R., Boggio, P. S., Ferrucci, R., Priori, A., & Fregni, F. (2013). Transcranial direct current stimulation: challenges, opportunities, and impact on psychiatry and neurorehabilitation. Frontiers in Psychiatry, 4, 19. https://doi.org/10.3389/fpsyt.2013.00019

      Cole, E. J., Stimpson, K. H., Bentzley, B. S., Gulser, M., Cherian, K., Tischler, C., Nejad, R., Pankow, H., Choi, E., Aaron, H., Espil, F. M., Pannu, J., Xiao, X., Duvio, D., Solvason, H. B., Hawkins, J., Guerra, A., Jo, B., Raj, K. S., Phillips, A. L., … Williams, N. R. (2020). Stanford accelerated intelligent neuromodulation therapy for treatment-resistant depression. The American Journal of Psychiatry, 177(8), 716–726. https://doi.org/10.1176/appi.ajp.2019.19070720

      Eisenberger, R. (1992). Learned industriousness. Psychological Review, 99(2), 248–267. https://doi.org/10.1037/0033-295X.99.2.248

      Hutton, T. M., Aaronson, S. T., Carpenter, L. L., Pages, K., Krantz, D., Lucas, L., Chen, B., & Sackeim, H. A. (2023). Dosing transcranial magnetic stimulation in major depressive disorder: Relations between number of treatment sessions and effectiveness in a large patient registry. Brain Stimulation, 16(5), 1510–1521. https://doi.org/10.1016/j.brs.2023.10.001

      Jannati, A., Oberman, L. M., Rotenberg, A., & Pascual-Leone, A. (2023). Assessing the mechanisms of brain plasticity by transcranial magnetic stimulation. Neuropsychopharmacology, 48(1), 191–208. https://doi.org/10.1038/s41386-022-01453-8

      Ke, Y., Liu, S., Chen, L., et al. (2023). Lasting enhancements in neural efficiency by multi-session transcranial direct current stimulation during working memory training. npj Science of Learning, 8(1), Article 23. https://doi.org/10.1038/s41539-023-00200-y

      Sabé, M., Hyde, J., Cramer, C., Eberhard, A., Crippa, A., Brunoni, A. R., Aleman, A., Kaiser, S., Baldwin, D. S., Garner, M., Sentissi, O., Fiedorowicz, J. G., Brandt, V., Cortese, S., & Solmi, M. (2024). Transcranial magnetic stimulation and transcranial direct current stimulation across mental disorders: A systematic review and dose-response meta-analysis. JAMA Network Open, 7(5), e2412616. https://doi.org/10.1001/jamanetworkopen.2024.12616

      Schulze, L., Feffer, K., Lozano, C., Giacobbe, P., Daskalakis, Z. J., Blumberger, D. M., & Downar, J. (2018). Number of pulses or number of sessions? An open-label study of trajectories of improvement for once- vs. twice-daily dorsomedial prefrontal rTMS in major depression. Brain Stimulation, 11(2), 327–336. https://doi.org/10.1016/j.brs.2017.11.002

      Woodham, R. D., Selvaraj, S., Lajmi, N., Hobday, H., Sheehan, G., Ghazi-Noori, A.-R., Lagerberg, P. J., Rizvi, M., Kwon, S. S., Orhii, P., Maislin, D., Hernandez, L., Machado-Vieira, R., Soares, J. C., Young, A. H., & Fu, C. H. Y. (2025). Home-based transcranial direct current stimulation treatment for major depressive disorder: a fully remote phase 2 randomized sham-controlled trial. Nature Medicine, 31(1), 87–95. https://doi.org/10.1038/s41591-024-03305-y

      Zhong, M., Cywiak, C., Metto, A. C., Liu, X., Qian, C., et al. (2021). Multi-session delivery of synchronous rTMS and sensory stimulation induces long-term plasticity. Brain Stimulation, 14(4), 884–894. https://doi.org/10.1016/j.brs.2021.05.003

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This interesting paper demonstrates that transgenic over-expression of sphingosine 1-phosphate receptor 1 (S1PR1) on neutrophils alters their phenotype, resulting in (1) accumulation of neutrophils in blood, spleen, lung, and liver; (2) a shift in homing receptor expression with reduced CXCR2 and elevated CXCR4; (3) altered transcriptional profile with an increase in "G5c" neutrophils and reduced "module scores" for apoptosis and inflammatory response; (4) reduced ROS production upon fLMP stimulation; and (5) altered responses to bacterial and viral infections of the lung. It raises many interesting questions about how S1P signaling regulates neutrophil biology, and hence will be the basis of future studies. These include: (1) What is the physiological role of S1PR1 signaling in neutrophils? Although there is no dramatic effect on numbers upon S1PR1 loss, is there an effect on any of the other parameters measured? (2) What is unique about the lung that S1PR1 over-expression is particularly impactful there? (3) What distinguishes the bacterial context in which S1PR1 over-expression is maladaptive from the viral context in which S1PR1 over-expression is protective? and (4) Can treatment with an S1PR1 agonist mimic S1PR1 over-expression? As a possibly related question, when in neutrophil development does S1PR1 signaling function to shift the phenotype?

      Strengths:

      (1) A comprehensive characterization of S1PR1-transgenic neutrophils.

      (2) Opens many interesting areas of investigation.

      We thank the reviewer for a very positive assessment of our work and for raising very interesting questions, which will be useful to extend this work in the future.

      Weaknesses:

      Although some characterization of the neutrophil-specific Mrp8-Cre is done, most of the experiments use the more widely expressed LysM-Cre. The redistribution phenotype is much stronger with LysM-Cre than with Mrp8-Cre, so it is unclear what effects are attributable to a cell-intrinsic role of S1PR1, even in studies of neutrophils analyzed ex vivo.

      We acknowledge that the data from Mrp8-Cre mouse strain are more limited than the LysM-Cre counterparts. This is because we initially characterized the LysM-Cre S1pr1 KO and TG strains and confirmed key findings relevant to neutrophils in the Mrp8-Cre counterparts. Going forward, more studies will be done in the Mrp8-Cre strain as suggested by the reviewer.

      Reviewer #2 (Public review):

      The authors have utilised two main models to assess the function of S1PR1 in neutrophils in mice. The knockout of this receptor shows no conclusive effect on neutrophil numbers or functions; it was only the overexpression that resulted in significant alterations. Therefore, often the conclusions do not describe normal or disease physiology but could be useful in a bioengineering context.

      We agree with the reviewer that some of the key findings described in our manuscript, for example, neutrophil survival, spleen size, and ROS reduction, etc., were not observed in the S1pr1 KO strains. Our interpretation is that other receptors, for example S1PR4, could be involved in compensating for the loss of S1PR1. We will explain this better in the revisions.

      Strengths:

      From a bioengineering standpoint, this seems like an important study - showing enforced expression of S1PR1 in neutrophils has improved outcomes for influenza infection (Figures 6 and 7).

      We agree with the reviewer that overexpression of S1PR1 could be useful from “bioengineering standpoint”.

      Weaknesses:

      Although the strength is the influenza model, genetic modification of human neutrophils cannot be a strategy, and therefore, is there any way to increase this receptor for mouse, or more importantly, human neutrophils? This study only looks at mice with a non-physiological model of overexpression. It does not offer a real therapeutic option, which drastically hinders the importance of the study. I have other concerns with the data analysis and interpretation, which I detail on a figure-by-figure basis (and how it relates to conclusions) below:

      The reviewer's comments are acknowledged. However, at this early stage of discovery, we feel that it would be premature to address the issue of “a real therapeutic option”. This can be addressed in the future.

      Main specific issues:

      (1) Figure 2A+B: This is unconvincing; in the surface staining there seem to be real cells positive for the receptor (high staining in the histogram), but none of the transgenic protein is getting there? This undermines the idea that the effects of the transgene are related to S1P signalling. In the 'Total S1PR1' this is both underwhelming and misleading, as an isotype control (or better S1PR1 knockout) is missing, which would give a better representation of actual expression (flow cytometry autofluorescence famously increases in the red laser channels with fix/perm). The Imagestream chosen images are showing best-case scenarios - and aren't representative. What does the isotype/ KO look like here? All in all, the conclusion on receptor internalization is not well supported, especially when theoretically the TG overexpression should overload S1P availability. This also highlights the lack of another control - does overexpression of another random/non-functional protein have the same effect? To play devil's advocate, perhaps overloading of the ubiquitin-proteasome system is responsible?

      We will conduct S1PR1 antibody staining with knockout neutrophils as requested by the reviewer. The results will be shown in the revision. The comment of the reviewer on “ubiquitin-proteosome” system is not relevant in our opinion and does not impact the validity of our findings. The approach that we used is standard mouse genetics, and the results should be interpreted from that perspective and not from “devil's advocate”.

      (2) Figures 2E-H: In the text, the authors should fix the statement 'Additionally, surface CXCR2 was downregulated and CXCR4 upregulated in LysM-S1pr1 TG neutrophils across bone marrow, spleen, and blood (Fig. 2, E and F)' to better reflect that there is no significant difference in the bone marrow regarding CXCR4. Of note, the total MFI from this data would also be informative, another noticeable absence being the gating strategies for much of the data. Also, alter the statement: 'CD62L expression was largely preserved across compartments, with only a modest reduction in bone marrow neutrophils (Fig. 2G)'. A 50% reduction in CD62L is not modest.

      These statements will be changed in the revised manuscript, which we hope to submit soon.

      (3) Supplemental Figure 3. A common theme: the wrong statistics have been used here, which has led to a false conclusion. Megakaryocyte/erythrocyte progenitors (MEPs) were only elevated in 2/3 TG mice, and the numbers are so small that this is not significant by any measure of the word. This is certainly not statistically significant if the correct test of (log-normalized) two-way ANOVA is performed (with Sidak's post hoc test). Another acceptable test would be Kruskal-Wallis with Dunn's post-test just for MEPs.

      We will reanalyze these data with different statistical tests and discuss these in the revision.

      (4) Starting at Figure 3, the authors refer to 'S1PR1hi neutrophil accumulation'. Crucially, the authors must here and throughout be explicitly clear in which cells they are referring to, as this can be misleading - particularly as there are real S1PR1-high cells identified in Figure 2A surface staining. It is my understanding that the authors here mean the transgenic artificially high mice - a very large distinction.

      We will clarify and edit this terminology throughout the manuscript in the revision. S1PR1<sup>hi</sup> does not strictly imply receptor on the cell surface. It also includes internalized S1PR1 receptors.

      (5) Figure 3A: It is difficult to interpret the figure with the necessary details about the experiment. For instance, there is no mention that this is sterile inflammation or what caused it.

      This will be edited in the revision.

      (6) Figure 3B and C: It should be made clear whether these splenic neutrophils are related to the time course of peritoneal inflammation in 3A. Why are there so many apoptotic neutrophils in the spleen? The low numbers here suggest a processing issue rather than real death in vivo (which usually is absent).

      This will be addressed in the revision.

      (7) Figure 3D: This can also be misleading - the wrong statistics are again used. This should be a log-transformed two-way ANOVA. Regardless of this, the data is not strong enough to be conclusive, a minor effect at best that could also just be related to the type of cell tracker used.

      Same as #3 above.

      (8) Figure 5E: It is stated that 'LysM-S1pr1 TG mice exhibited a higher bacterial burden in the lungs than controls (Fig. 5E).' Again, misleading results, first the wrong statistical test was used (correct = log norm one-way ANOVA with Tukey's or Kruskal Wallis with Dunn's), secondly the only significance is between S1PR1(fsf) and the Mrp8-S1PR1, not with the LysM TG. 5F is also not strong, with only 2/7 values appearing outside the range of the control - P values can be misleading when poor statistics are used.

      Same as #3 above.

      (9) Figure 6G: Some discussion should be given for why Neutrophils are lower in BALF in the IAV model - even though higher in the lung in the non-IAC mice in Figure 1. In general, rather than focusing on the non-physiological differences, the discussion could better reflect the inconsistencies and more fully address the difference between the TG and KO and what this means going forward.

      Same as #5 above.

      Reviewer #3 (Public review):

      Summary:

      Using mice that overexpress S1PR1 in myeloid cells or specifically in neutrophils, the authors show that increased S1PR1 promotes neutrophil release from the bone marrow and accumulation in blood and peripheral tissues without causing baseline tissue injury. These cells acquire a CXCR4-high, CXCR2-low, CD101-low phenotype, survive longer, and display enhanced mitochondrial metabolism and mTOR signaling, together with reduced apoptotic, inflammatory, and ROS-related programs. Although phagocytosis is preserved, ROS production is markedly reduced. This is associated with impaired bacterial clearance in the lung but improved outcomes during influenza infection, including better survival, less weight loss, improved oxygenation, lower viral burden, and reduced lung inflammation. In contrast, myeloid S1PR1 deletion produces little detectable phenotype. The authors therefore propose that S1PR1 separates neutrophil persistence from inflammatory function, improving tolerance to viral lung injury at the expense of antibacterial defense.

      Strengths:

      This is a technically solid paper using novel mouse models to overexpress S1PR1 specifically in myeloid cells as well as neutrophils. The data are striking with respect to neutrophil expansion. The diverse roles of neutrophils and their population heterogeneity are an important scientific area that has led to many recent breakthroughs - PMC11785525; PMC12823425, thus this is a timely study.

      We thank the reviewer for positive comments.

      Weaknesses:

      The study mainly demonstrates what S1PR1 overexpression is sufficient to do, rather than establishing the physiological role of endogenous S1PR1. The conclusions should therefore be narrowed unless the authors provide stronger loss-of-function and physiological validation. As written, the abstract ("S1PR1 promotes mitochondrial fitness, enhances survival, and reduces inflammatory output") and the conclusion ("S1PR1 serves as a key regulatory axis") are sufficiency claims but should not be promoted as necessity claims. The honest sentence is: "Thus, a better conclusion would be that enforced S1PR1 expression is sufficient to reprogram neutrophils".

      We will edit the text to better reflect sufficiency versus necessity terms.

      The authors do not confirm efficient S1pr1 deletion in neutrophils. Furthermore, the knockout is examined only under steady-state conditions and limited in vitro stimulation, but not in the bacterial or influenza models where the transgenic phenotype is observed. Without these experiments, the study cannot establish whether endogenous S1PR1 is necessary for the reported functions.

      We will provide qRT-PCR data (that we have done already but did not include in the original version) confirming efficient deletion of the S1pr1 gene.

      The degree of S1PR1 overexpression is not quantified relative to normal physiological levels. The authors should determine whether naturally occurring S1PR1-high neutrophils display the same survival, metabolic, trafficking, and inflammatory features observed in the transgenic cells.

      We agree that the level of S1PR1 overexpression in our transgenic neutrophils should be quantitatively compared with endogenous S1PR1 expression. We will provide the quantitative result in the revision. Our analyses of independent human (GSE216009) and mouse (GSE243466 and GSE266518) single-cell RNA-seq datasets indicate that endogenous mRNA expression for this receptor is higher in specific neutrophil states associated with inflammatory, immature, and tissue-associated populations. Notably, these endogenous S1PR1-expressing neutrophils show transcriptional features involving altered oxidative/inflammatory programs, chemotaxis, and mitochondrial/metabolic regulation that partially overlap with the phenotype of S1PR1-transgenic neutrophils. We will provide additional quantitative and dataset analyses in the revised manuscript.

      Analysis of relevant human or mouse datasets, including sepsis, ARDS, viral infection, cancer, or aging, would also help establish whether this neutrophil state exists physiologically.

      As described above, we have analyzed independent human and mouse single-cell RNA-seq datasets from sepsis, cancer, and other inflammatory conditions to determine whether endogenous S1PR1-expressing neutrophil states occur naturally across biological and disease contexts. These analyses support the presence of distinct endogenous S1PR1-expressing neutrophil populations across multiple biological contexts. We will provide the expanded analyses and corresponding data in the revised manuscript.

      Surface S1PR1 expression appears similar between control and transgenic neutrophils, whereas total intracellular receptor is increased. This suggests that the phenotype may depend on receptor internalization or endosomal signaling. An internalization-deficient S1PR1 model, such as S1P1-S5A, would help distinguish sustained surface signaling from internalization-dependent signaling. The authors should also determine whether the phenotype requires ligand binding, Gi signaling, and mTOR activity.

      We agree that future experiments will use internalization-defective S1PR1 S5A knock-in neutrophils.

      The reduction in CXCR2 and decreased neutrophil accumulation in the airways could alone explain the protection from influenza-induced lung injury. The current experiments do not clearly distinguish neutrophil reprogramming from defective migration into the alveolar space.

      We agree. We did not examine the role of CXCR2 in the influenza experiments. Our interpretation is that reduced ROS from transgenic neutrophils reduced lung injury.

      Although this may be outside the scope of the current study, the authors should directly test whether CXCR2 inhibition reproduces the phenotype.

      This is a good point that the reviewer brought up. This can be addressed in a future study.

      The reported reduction in viral load should also be confirmed using plaque assay or TCID50, and the possible contribution of NET formation should be examined.

      This is again a good suggestion that can be addressed in a future study.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      This is an interesting paper, with the primary finding being that localizing MreB or PBP2 to the cell poles in E. coli is primarily demonstrated via an aggregate formed by expressing M. xanthus MreB.

      I have 2 main concerns:

      (1) First, the authors should clarify and adjust their interpretation of FDAA incorporation: As written, the authors interpret FDAA incorporation as being caused by incorporation during PG polymerization. However, in E. coli, FDAA incorporation does not result from the elongation of PG strands or their initial 4-3 crosslinking by DD transpeptidases following polymerization, but rather by the remodeling of L,D-transpeptidases.

      Thus, it is not accurate to refer to the FDAA incorporation as PG elongation, but rather the modification of crosslinks from 4-3 to 3-3 crosslinks at that location. To claim a link to PG polymerization, other experiments or substantial explanations are needed.

      (E. coli Cells Incorporate FDAAs by L,D-TPases in a Growth Independent Manner) - https://doi.org/10.1021/acschembio.

      We appreciate the reviewer for bringing up this important question. We will address this concern by further explaining our results in the revised manuscript.

      However, we respectfully disagree with the reviewer for two reasons:

      First, the paper by Kuru et al. (mentioned by the reviewer) revealed that (exact quote): “Our in vitro and in vivo data unequivocally demonstrate that these bacteria incorporate FDAAs using two extra cytoplasmic pathways: through activity of their D, D-transpeptidases, and, if present, by their L, D-transpeptidases…These mechanistic findings enabled development of a new, FDAA-based, in vitro labelling approach that reports on subcellular distribution of muropeptides, an especially important attribute to enable the study of bacteria with poorly defined growth modes” (Kuru et al., 2019).

      Thus, the paper by Kuru et al. identified two FDAA labeling patterns, a D, D-transpeptidase-dependent, concentrated labeling for PG growth and an L, D-transpeptidase-dependent, growth-independent labeling for PG modification along the entire cell envelope (Kuru et al., 2019). The polar FDAA foci we presented do not match the reported pattern of L, D-transpeptidase-dependent incorporation.

      Second, we provided the evidence in Fig. 3b that the polar FDAA labeling is due to the activity of PBP2, a D, D-transpeptidase, because mecillinam that inhibits PBP2 is sufficient to abolish polar FDAA incorporation.

      (2) Second, the evidence provided that this is polar elongation is not sufficient to prove elongation. In all images claiming polar growth, the FDAA focus appears as a single spot, which corresponds to the MreB aggregate visible in bright-field. That might indicate incorporation, but it does not demonstrate polar elongation. To prove this, the authors should do different-length pulses of FDAAs and demonstrate an increasing length of labeled PG along the cell. Cells with one focus should show increasing length; polar foci should elongate from both ends. I find the current 2-color labeling insufficient, as BADA labels the entire cell.

      We appreciate this comment and totally agree with the reviewer. The wide BADA labeling band could indeed come from L, D-transpeptidases. We will follow the reviewer’s recommendation to address this comment with additional staining experiments.

      Small points:

      (3) Lines 350- 302: "In this case, as the nonpolar region is no longer the growth zone, the established cylindrical PG structure is sufficient to maintain cell width, where MreB filaments become nonessential." Does polar elongation give robustness to rod shape? The authors should include an analysis of cell width and its variation within a cell and between cells.

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does give robustness to rod shape. We will follow the reviewer’s recommendation and provide the said analysis.

      (4) In the abstract: "This reprogrammed growth mode bypasses the requirement for MreB filaments, highlighting a plasticity of the Rod system that suggests polar elongation may have emerged through the evolutionary loss of MreB." This argument should not be made without evolutionary analysis or reference to such work indicating this is the case.

      We appreciate this comment and will remove this statement from our revised manuscript.

      (5) 365-367 "PG-depleted spheroplasts can spontaneously regenerate rod shape through curvature-dependent localization of MreB filaments [21, 57]". This should be amended. The Billing paper did indeed study spheroplasts, but the Hussain paper used teichoic acid-depleted cells that still had a cell wall.

      We thank the reviewer for pointing out this mistake. We will amend this statement in our revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Based on observations of localisation of MreBEc at the poles within an aggregate-like structure, upon heterologous expression of MreBMx, the authors set out to investigate how this non-canonical localisation of MreBs leads to a reprogramming of peptidoglycan synthesis to the poles. This is analogous to the polar growth observed in phyla which are not dependent on dispersed growth of PG, but only at the poles, and are MreB independent.

      The authors proceed to establish that PG synthesis is MreB-dependent, Rod enzyme-dependent, and requires the prior establishment of a pole.

      Strengths:

      (1) It is a very interesting idea to design experiments to demonstrate reprogramming of non-polar to polar growth based on the observation of localisation of a heterologously expressed MreB.

      (2) The experiments to demonstrate the factors that determine polar growth and the observation of the PG in each of these experimental situations are convincing.

      (3) I find the observation of an extra layer of PG in the heterologously expressed system very intriguing. It will be interesting to see if this layer merges with the other PG layer at some stage or branches from the non-polar growth near the poles.

      Weaknesses:

      (1) It is not clear what exactly the identity of the polar aggregates is and how much of this activity is an artefact of partially functional MreBs.

      We appreciate this comment and will address it by further explaining our results in the revised manuscript. While we can only say that the polar aggregates resemble inclusion bodies, we do believe that they cause polar PG growth because in the cells that express MreB<sub>Mx</sub>, polar PG growth does not occur at the poles that lack MreB aggregates.

      (2) I find it intriguing that the localisation and growth are predominantly at one pole only. It is unclear to me how this can be reconciled with growth and shape maintenance, and an increase in length and width. Is the increase in length and width a consequence of misshapen cells that are bulged in the absence of a normal PG layer?

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does not generate bulges and is thus sufficient for maintaining rod shape. We will follow the reviewer’s recommendation to clarify this.

      (3) The authors do not follow up on the observations in the first figure on the length and width changes and the extra peptidoglycan layer (which I feel are the most interesting aspects), and how this can be connected to the polar growth observed in the later sections of the manuscript.

      We appreciate this comment. We believe that the thickened PG patches are integral parts of the polar PG, rather than an extra layer, which is, however, technically challenging to prove. Thus, we will relay on fluorescence microscopy to visualize polar PG growth.

      (4) The claim that this could be a precursor of an MreB-independent polar growth mechanism appears to be a bit far-fetched, because the system is still dependent on having an established pole for PG synthesis to occur in the new place.

      We appreciate this comment and will remove such speculations from our revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Most rod-shaped bacteria grow by one of two mechanisms: growth from the pole or growth from the midcell. It is rare for a single species to utilize both modes of growth, although a few examples do exist. Here, the authors have artificially induced E. coli cells to grow from the poles, either by expressing mreB from Myxoccocus xanthus in E. coli, leading to the mislocalization of MreB to the poles in large aggregates, or by forcing the localization of major cell wall synthesis proteins to the cell pole. The fact that cells switched modes of growth suggests an evolutionary pathway from midcell to polar growing cells as well as suggests that there might be unknown conditions in nature when cells may switch growth modes.

      Strengths:

      (1) The authors use a strain that has replaced mreB with a functional fluorescent version at the native site. This eliminates any effects of having two copies of mreB. Because MreB is fluorescently tagged, they can monitor its localization when mreB from M. xanthus is expressed in E. coli. They notice that MreBec now forms bright polar foci and that there appear to be changes to the cell wall at the pole.

      (2) D-amino acids are specific to the cell wall, and fluorescent versions (FDAA) have been used to mark sites of new cell wall insertion. The authors use these FDAAs to determine how cell wall synthesis correlates to MreB and if that changes when MreBmx is expressed. Again, there is pretty clear evidence that cell wall synthesis follows MreB localization to the pole.

      (3) MreB itself does not synthesize the cell wall, but localizes the proteins, such as PBP2, that do. Using a published method to force proteins to the pole, the authors show that when they target PBP2 to the pole, they can phenocopy the polar growth seen when MreB is polar. Interestingly, these cells become resistant to A22, a drug that targets MreB, suggesting that localized growth at the pole does not require MreB and is sufficient to maintain rod shape.

      Weaknesses:

      (1) The authors do not show what the poles of control cells look like, making it difficult to determine if there is a change when MreBmx is expressed. However, the localization of both MreBec and MreBmx clearly forms bright foci at the pole.

      We appreciate this comment and totally agree with the reviewer. We will provide reference images in the revised manuscript.

      (2) While more quantification is needed, the authors show some evidence that RodZ, an MreB interaction partner, is needed for this polar growth, as cells lacking rodZ still form foci at pole-like regions when MreBmx is in the cell; however, these cells remain spherical and do not elongate from these foci.

      These results strongly support our conclusions. In the cells that express MreB<sub>Mx</sub>, MreB<sub>Mx</sub> causes MreB<sub>Ec</sub> to mislocalize, and mislocalized MreB<sub>Ec</sub> recruits Rod enzymes (RodA and PBP2) to poles through the connector protein RodZ. Here when we delete rodZ, while MreB<sub>Ec</sub> still forms aggregates, it is unable to recruit Rod enzymes to those aggregates.

      In contrast, when we directly relocalize PBP2 to cell poles through the PopZ tag, polar PG growth can bypass the requirement for MreB<sub>Ec</sub> filaments (Fig. 4).

      When MreB is deleted, and cells become spherical, the authors were unable to cause the polar growth mode. They suggest that this is due to the lack of a preexisting pole; however, experimental evidence to test this is missing.

      We appreciate this comment. We plan to localize PBP2 to cell poles in an mreB depletion strain to test if cells still remain rods when mreB is depleted.

      Conclusion:

      Overall, the authors do a good job of showing that E. coli can grow with a polar method rather than a midcell method of cell wall insertion. It is unclear why MreBec forms at poles when MreBmx is present and even if this MreB is functional. The foci look similar to inclusion bodies, which are normally aggregates of misfolded proteins that migrate to the poles. Past work has shown that when MreB is more polarly localized, branches form, which is not seen here. Importantly, the authors also show that there is feedback between the localization of MreB and PBP2 as both appear to regulate the localization of the other.

      Because CFP-labeled MreB<sub>Ec</sub> is still fluorescent in the aggregates, we believe that some MreB<sub>Ec</sub> molecules are still correctly folded there, at least on the surfaces of those aggregates.

      Reference

      Kuru, E., Radkov, A., Meng, X., Egan, A., Alvarez, L., Dowson, A., Booher, G., Breukink, E., Roper, D.I., Cava, F., Vollmer, W., Brun, Y., and VanNieuwenhze, M.S. (2019) Mechanisms of incorporation for D-amino acid probes that target peptidoglycan biosynthesis. ACS Chem Biol.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study leverages a large global dataset of tens of thousands of tuberculosis samples to place recurrent protein-coding mutations into their three-dimensional structural context, offering an expanded view of how antibiotic resistance emerges compared to traditional genetic analyses alone. The strength of evidence is convincing, supported by the scale and breadth of the dataset and the systematic structural analysis, although some of the assumptions made in the the modeling approach are only partially supported. Overall, the work will be of broad interest to researchers studying microbial evolution, antibiotic resistance, and structure-function relationships in pathogens.

      We thank the reviewers and editors for their careful critique of our work. We believe the work has been strengthened by addressing the comments and are delighted to submit a revised version. This version has a detailed discussion of prior literature on structure analysis of antibiotic resistance variants in Mycobacterium tuberculosis, more details on dataset origins and data processing to improve reproducibility, better explanations of the evolutionary assumptions underlying the scoring method we developed, and more discussion of proteins that show homoplasic and clustered mutational signals that are not known to confer antibiotic resistance. We include detailed responses to the comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Green et al. attempt to use large-scale protein structure analysis to find signals of selection and clustering related to antibiotic resistance. This was applied to the whole proteome of Mycobacterium tuberculosis, with a specific focus on the smaller set of known antibiotic-resistance-related proteins.

      Strengths:

      The use of geospatial analysis to detect signals of selection and clustering on the structural level is really intriguing. This could have a wider use beyond the AMR focussed work here and could be applied to a more general evolutionary analysis context. Much of the strength of this work lies in breaking ground into this structural evolution space, something rarely seen in such pathogen data. Additional further research can be done to build on this foundation, and the work presented here will be important for the field.

      The size of the dataset and use of protein structure prediction via AlphaFold, giving such a consistent signal within the dataset, is also of great interest and shows the power of these approaches to allow us to integrate protein structure more confidently into evolution and selection analyses.

      Weaknesses:

      There are several issues with the evolutionary analysis and assumptions made in the paper, which perhaps overstate the findings, or require refining to take into account other factors that may be at play.

      (1) The focus on antimicrobial resistance (AMR) throughout the paper contains the findings within that lens. This results in a few different weaknesses:

      (a) While the large size of the analysis is highlighted in the abstract and elsewhere, in reality, only a few proteins are studied in depth. These are proteins already associated with AMR by many other studies, somewhat retreading old ground and reducing the novelty.

      (b) Beyond the AMR-associated proteins, the proteome work is of great interest, but only casually interrogated and only in the context of AMR. There appears to be an assumption that all signals of positive selection detected are related to AMR, whereas something like cas10 is part of the CRISPR machinery, a set of proteins often under positive selection, and thus unlikely to be AMR-related.

      We agree that environmental pressures beyond AMR may impose positive selection. In response to the reviewer’s comment we have now included more results and a supplementary figure of the findings about Cas10. We have expanded our results about the proteins found with significant clustering that are not known AMR-associated proteins, and clarified in the discussion that we don’t believe positive selection is caused only by AMR.

      We note that data from homoplasic substitutions in proteins (Figure 1) does indeed support that AMR is the strongest driver of positive selection in Mtb. A challenge is that knowledge about protein function varies in depth by protein, and in Mtb the AMR proteins are among the most well-studied in the proteome. Thus, explanations from literature are most readily available for AMR proteins. Moreover, proteins may exhibit positive selection for more than one reason – for example, we find significant clustering in the proteins GlmM, GlmS, GlmU, and MurA, all of which are involved in amino sugar metabolism, a key component of the cell wall. These proteins could plausibly have a role in AMR via cell wall permeability mechanisms, and could plausibly have a role in adaptation to host environment via the same (or other) mechanisms.

      (2) The strength of the signal from the structural information and the novelty of the structural incorporation into prediction are perhaps overstated.

      (a) A drop of 13% in F1 for a gain of 2% in PPV is quite the trade-off. This is not as indicative of a strong predictor that could be used as the abstract claims. While the approach is novel and this is a good finding for a first attempt at such complex analysis, this is perhaps not as significant as the authors claim.

      (b) In relation to this, there is a lack of situating these findings within the wider research landscape. For instance, the use of structure for predicting resistance has been done, for example, in PncA (https://academic.oup.com/jacamr/article/6/2/dlae037/7630603, https://www.sciencedirect.com/science/article/pii/S1476927125003664, https://www.nature.com/articles/s41598-020-58635-x) and in RpoB (https://www.nature.com/articles/s41598-020-74648-y). These, and other such works, should be acknowledged as the novelty of this work is perhaps not as stark as the authors present it to be.

      We appreciate this comment and have made efforts to better situate our work in the wider context of structure-based prediction of antibiotic resistance. We have included description of and citation to these works and others in the introduction, results, and discussion. A differentiator between our work and previous is that we have trained a predictor across all proteins in the WHO catalogue of known resistance variants, not limited ourselves to a single protein at a time. We feel this makes the case that structure (and specifically proximity to known resistance-conferring variants) is a universally useful feature for resistance mutation prediction. We have also updated our abstract to reflect the exact performance of our method.

      Introduction: “Protein three-dimensional structure has shown utility as an input feature for identifying resistance-conferring variants in known resistance-conferring proteins such as RpoB,[26,27] PncA [28–30], and AtpE.[31]While past work has sought to reannotate parts of the M. tuberculosis proteome with computationally predicted protein structures using older structure prediction methods,[32] we can now infer a protein structure for nearly every protein in the proteome using AlphaFold,[33] leading to new works examining the 3D location of mutations in known and suspected resistance-conferring proteins.[34,35]”

      (3) The authors postulate that neutral AA substitutions would be randomly distributed in the protein structure and thus use random mutations as a negative control to simulate this neutral evolution. However, I am unsure if this is a true negative control for neutral evolution. The vast majority of residues would be under purifying selection, not neutral selection, especially in core proteins like rpoB and gyrA. Therefore, most of these residues would never be mutated in a real-world dataset. Therefore, you are not testing positive selection against neutral selection; you are testing positive against purifying, which will have a much stronger signal. This is likely to, in turn, overestimate the signal of positive selection. This would be better accounted for using a model of neutral evolution, although this is complex and perhaps outside the scope. Still, it needs to be made clear that these negative controls are not representative of neutral evolution.

      The goal of our negative control was to simulate the random accumulation of amino acid substitutions without the effects of selection, which we had originally referred to as “neutral evolution” but is better described as “randomly accumulating substitutions.” We agree that in the absence of antibiotics, essential proteins like RpoB and GyrA are probably under purifying selection, and thus will have depletion of mutations in their hydrophobic cores. We have revised our wording in the results section “A protein-level statistic to test for mutational clustering” to make it clear that our negative control is that of randomly accumulating substitutions. We have included possible extension to more realistic evolutionary scenarios in the discussion.

      As a side note, if we were to compare the observed mutation 3-D pattern to a purifying selection model, that may overestimate the signal compared with the randomly accumulating substitution model that we currently present in the manuscript. Purifying selection would tend to result in slower evolutionary rates than positive selection or randomly accumulation substitutions. So, a control based on purifying selection would have to have fewer mutations to account for the same evolutionary time, and this could lead to underestimating clustering.

      (4) In a similar vein, the use of 15 Å as a cut-off for stating co-localisation feels quite arbitrary. The average radius of a globular protein is about 20 Å, so this could be quite a

      large patch of a protein. I think it may be good to situate the cut-off for a 'single location' within a size estimator of the entire protein, as 15 Å could be a neighbourhood in a large protein, but be the whole protein for smaller ones.

      We interrogated the use of 15 Å as a cutoff and found that it is indeed not very stringent, and functions more as a filter to remove the most egregious examples of proteins lacking single-location clustering. We include a new supplementary figure showing the number of significant hits as the cutoff is varied from 2 Å to 40 Å, and summarize these results in the main text. We note that we in fact find a weak negative relationship between protein length and the distance between the top two residues with highest G-score (R<sup>2</sup> = 0.008, b = -3.1806, p-value = 0.048), the opposite of what would be expected under a scenario the distance between residues is simply driven by protein size and not a signal for clustering.

      Reviewer #2 (Public review):

      Summary:

      This is an important study that, for the first time, systematically places the homoplastic genetic variation observed in the coding regions in a large collection of >31,000 M. tuberculosis samples into the protein structural context. This should be much more informative when, e.g. predicting antimicrobial resistance. The authors imaginatively apply the Getis-Ord score, which originated in geographical spatial analysis but has also been used in human disease to demonstrate that missense mutations in M. tuberculosis known to be associated with antimicrobial resistance are clustered in space. That they are able to consider almost all of the proteome using a large dataset of 31,000 M. tuberculosis complex clinical samples, which makes the evidence convincing.

      Strengths:

      To my knowledge, this is the first study to place the homoplastic missense mutations from a large clinical dataset into their protein structural context and attempt to look for clustering in space, which could be indicative of a recent evolutionary pressure, such as the use of antibiotics. The field usually only views resistance through the genetic paradigm, so it is delightful to see a structural paradigm being brought to bear, as this should, in theory, be much more informative, as protein structure is much closer to function. In addition, the dataset used is large (>31,000 clinical M. tuberculosis samples), and the authors are able to consider almost all of the ORFs (3,687/3,996) in the M. tuberculosis reference, and hence the analysis is comprehensive.

      Weaknesses:

      It is not apparent at the time of this review if the study could be reproduced by other researchers as e.g. whilst the authors state that the raw sequencing files (FASTQ) underpinning the dataset of 31,428 M. tuberculosis isolates can be downloaded the table in the Supplement containing the sample and accession identifiers contains rows that do not contain NCBI accessions e.g. '01R0685' or 'IDR 1600023875' or '1479144813357T181715lib5022nextseqn0035151bp' instead of the expected form e.g. 'SAMEA1016138'. I have searched the NCBI SRA using these terms and got no results, so they cannot be used to download any FASTQ files. There is also no information in the preprint on how the reads were processed (which is a complex process) and the dataset of SNPs subsequently built. One can trace back through the references, but I cannot find anywhere where one can download the SNP dataset, which would permit researchers to reproduce at least the latter stages of the work -- one obvious option would be to make the SNP dataset available. Likewise, the authors have constructed a "M. tuberculosis structureome", which would be very useful for the community but does not appear to be publicly available. At the time of the review, not all the GitHub repositories were public, so these points may have been rectified when that was corrected.

      We have made a number of changes to improve reproducibility of the manuscript.

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      Second, we have added a new section to our methods about the origins of the Mtb genomic data used in our manuscript, and describing the process of variant calling and SNP dataset construction.

      Third, we have provided the data in a zenodo repository along with instructions for using the protein distance map files provided: 10.5281/zenodo.20766453

      Lastly, we have ensured that the github is publicly available.

      The authors correctly point out in the Introduction that supervised methods like GWAS or ML need datasets with matching genetic and phenotypic drug susceptibility data, which are much difficult/expensive to obtain, but don't then close the loop by comparing their results back to such supervised methods. They pick out RnJ as having previously been identified by a GWAS, but it would have provided a useful validation of their method to e.g. demonstrating that X% of the genes they identify were also identified by GWAS/ML studies, and therefore their method can achieve similar results but without having to collect pDST data.

      We agree that this is a compelling possible extension of the work but due to time constraints have chosen not to pursue it for this manuscript.

      Whilst the authors acknowledge that assuming all sites are equally likely to mutate in their random shuffling procedure is a shortcoming, a bigger weakness is, I suspect, that one should also only consider which amino acids could arise at each codon due to a SNP. Shuffling assumes any amino acid can arise at any codon which is only possible with multiple nucleotide changes, which is possible but highly unlikely.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which mutations occur. In this calculation we do not consider the identity of the amino acid change per se. We have now clarified in the methods that we are not explicitly simulating biochemical change of the wild-type amino acid to any given mutant, rather we are analyzing the wild type amino acid in its structural context. We have added the following text in the methods section: In Computing inter-residue distances, “The EVcouplings Python package was used to compute the distance between wild-type amino acid residues in all protein structures”; in Computing the Getis-Ord score for clustering of homoplastic mutations, “The two values input to the Getis-Ord statistic computation are a per-residue score x, here the per-amino acid homoplasy score, and a weight matrix W that contains the inverse of the inter-residue distances computed from the wildtype amino acids”; and in Preparing GeO score calibration data, “Note that we do not recompute inter-residue distances when simulating mutations in an amino acid, as the distances used as input to GeO score are the wild-type inter-residue distances.” We hope this addresses the reviewer’s concern.

      Finally, the authors implicitly assume that the mutations do not perturb the structure of the proteins, which is likely to be generally true for essential genes but less likely to be true for non-essential genes. This assumption underpins their entire approach and should be borne in mind when evaluating the results.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which missense mutations occur and are observable in a naturally evolving population. It is true that we have not undertaken an analysis of whether any given mutation does or does not perturb the protein structure in which it occurs. However, location alone is a useful piece of information to analyze, as the location of naturally occurring mutations gives a readout of what types of mutations are allowed to persist under natural selection. Among our findings is that mutations display significant clustering even in non-essential genes, which we address in our discussion, “for proteins where mutations that lead to loss of function are known to cause resistance, such as PncA and RsmG (GidB), it is not necessarily expected to find clustering of mutations. We suspect that the observed clustering is due to mutations in a certain region of the protein being more likely to cause loss of function.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be good if a more detailed description of the sequencing dataset origins were presented. The Supplementary Table 1, which is meant to hold these data, points to another paper, which in turn points to another paper which does not detail collection strategies for this data. This needs to be clear so that any bias in collection, which could over-inflate selection signals, can be assessed.

      [From public reviews]

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from the NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      (2) Line 351: The exact download date is missing here.

      We have added the exact download date

      (3) Line 363: Did you mean lowest e-value? Higher would be worse.

      Thank you for catching this, it is indeed the lowest e-value (confusion stemmed from looking at highest negative log e-value).

      Reviewer #2 (Recommendations for the authors):

      (1) The GitHub repo* was not public at the time of review (nor was it listed under user aggreen), so I could not check how reproducible the results are -- please make it public.

      Absolutely, this has been addressed.

      (2) The authors say "we envision structure being added as an additional feature in future work to predict resistance phenotypes from sequences" - this is not true, as some work has already been published predicting resistance in MBTC going back to 2019**

      We have addressed this with wording changes and citations to the mentioned work (see public responses)

      (3) Throughout the term 'non-synonymous' is used; that would include premature stop codons. Would 'missense' be more appropriate?

      You’re correct in pointing out that the term “non-synonymous” is too general for what we mean in this paper. Our analysis included missense mutations and in-frame indels, but did not include premature stop (nonsense) mutations or frameshift mutations. So, the term missense (alone) is narrower than what we wish to convey. We have made clarifications throughout.

      (4) The aminoglycosides are an important, albeit less used, class of antibiotics, and mutations arise in the ribosomal genes, e.g. rrs (which hence do not encode protein). They have therefore been excluded for obvious reasons, but it would help a reader from the tuberculosis field if this were acknowledged. Likewise, a reader might wonder why Rv0678 isn't in Figure 1 - I suspect it is because most of the samples were sequenced before the introduction of bedaquiline, but again, it would help if this were explained.

      See response to (5)

      (5) On a related note, it is not surprising that rpoC appears in Figure 1 due to its role in compensating for the fitness cost that arises when a rifamipicin-resistance mutation occurs in rpoB: obviously not central but a nice "oh yeah that makes sense" point for the reader if it were briefly mentioned.

      These are both great points about the relevance of our results to the Mtb community. We have added an additional paragraph interpreting the results of Figure 1 that mentions the reason for the appearance of RpoC and Cas10, and non-appearance of non-coding genes and genes relevant to resistance to newly introduced and repurposed drugs.

      (6) How is the "minimum coordinate difference" calculated? I assume all the structures are missing hydrogens as usual, so for two amino acids A and B, is it the smallest distance between any pair of heavy atoms from A and B? That would, I assume, introduce some bias for larger amino acids like Trp, or did you calculate from shared atoms like the backbone C_alpha atoms? That in turn will tend to make the distances a bit larger. It would be useful to know, as you explicitly mention a 1.5 nm threshold.

      We have explained this in a new section of the results, “The EVcouplings Python package was used to compute the distance between amino acid residues in all protein structures [49]. The package calculates the distance between all heavy (non-hydrogen) atoms in residue i and residue j, then returns the minimum of those distances.”

      (7) Given the reference used (H37Rv) is Lineage 4, one wonders about deeprooted/phylogenetic mutations, but then I suspect this sentence is doing a lot of that heavy-lifting: "We performed ancestral sequence reconstruction to determine the number of independent arisals of each mutation (homoplasy)". For the more general reader, it would be useful to touch on exactly what you mean and the importance of only considering homoplastic mutations.

      We have expanded our explanation in this section to better make the case for the use of homoplastic variants in our analysis, “Because analyzing the frequency of alleles in a population can be biased by oversampling of particular lineages, and by evolutionary recency, we chose to analyze the number of independent arrivals of each mutation (homoplasy) rather than their population-level frequency. This ensures that more recent evolutionary events are not underrepresented due to lack of time to spread in the population. To accomplish this, we used a previously compiled a dataset of genomes of 31,428 isolates from the Mycobacterium tuberculosis complex (MTBC), with ancestral sequence reconstruction to determine the number of independent arrivals of each mutation (Supplementary Data 1).”

      (8) Whilst this is true: "Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance without requiring resistance phenotype data", the dataset used to, e.g. build the second edition of the WHO catalogue of resistanceassociated variants has >50,000 samples and therefore is larger than the dataset you have analysed here. The last time I looked, there were >100k M. tuberculosis samples in NCBI, and therefore, to be valid, your approach should really use more samples than are available with WGS and pDST data. I appreciate, however, that this will not be possible for this manuscript, but it is an obvious criticism.

      We appreciate this critique and acknowledge that the number of isolates with both WGS and pDST has increased rapidly in recent years. The first edition of the WHO catalogue (2021) used 38,215 isolates, which increased to over 50k in the second edition. The dataset of homoplastic mutations on which we based this paper was originally published in 2021.

      To address your comment, we have softened our assertions in the introduction about the utility of evolution-based approaches, and emphasize instead the different nature of the underlying signal, “Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance by analyzing their mutational frequency and phylogenetic distribution”

      (9) Minor point, but the second sentence in the Introduction ignores that one can diagnose MDR-TB using phenotypic methods as well as genetic methods.

      We have added an additional citation and mention of laboratory phenotypic methods.

      (10) There are a few typos: "genic" "G-sore"

      Addressed.

    1. Author response:

      We thank all three reviewers for their detailed and constructive reviews, which will help us improve the manuscript. Concerning the simplicity of the model, we will elaborate on the limitations that arise from the modelling choices and better justify these choices. With that in mind, we believe this level of abstraction is well suited to the study’s objective. As a rate-coded model, it is designed to capture interactions between multiple circuit pathways, and lays the foundation to delve into neuronal dynamics in the future for further insight, once the network-level questions we address here - regarding the exploration-exploitation tradeoff and evasion of sub-optimal performance convergence - are well examined. For this purpose, we believe the model’s simplified network-level perspective is not a bug, but a feature. We will also explore in depth why and how the dual pathway model performs better in non-convex sensorimotor landscapes, building on prior analytical work that speaks to the question (Sankar, Leblois & Rougier, ICDL 2022). We will further substantiate and discuss our choice concerning the potential mechanisms underlying the proposed overnight synaptic volatility. We will include comparable approaches (including relevant machine learning analogues) in the introduction. Finally, we will clarify our interpretation of the sign of synaptic weights. We hope to incorporate all valuable feedback by the reviewers in the revised manuscript.

      References

      Remya Sankar, Arthur Leblois, Nicolas P. Rougier (2022). Dual pathway architecture underlying vocal learning in songbirds. In IEEE International Conference on Development and Learning, ICDL 2022, London, United Kingdom, September 12-15, 2022. pages 265-271, IEEE, 2022.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting behavioral paradigm and reveals interactive effects of social hierarchy and threat type on defensive behaviors. However, addressing the aforementioned points regarding methodological detail, rigor in behavioral classification, depth of result interpretation, and focus of the discussion is essential to strengthen the reliability and impact of the conclusions in a revised manuscript.

      Strengths:

      The paper is logically sound, featuring detailed classification and analysis of behaviors, with a focus on behavioral categories and transitions, thereby establishing a relatively robust research framework.

      Weaknesses:

      Several points require clarification or further revision.

      (1) Methods and Terminology Regarding Social Hierarchy:

      The study uses the tube test to determine subordinate status, but the methodological description is quite brief. Please provide a more detailed account of the experimental procedure and the criteria used for determination.

      We have included more details about how the tube test was performed in the revised manuscript. Social rank within each mouse pair was determined using a standard tube test paradigm. To minimize stress, mice were pair-housed for at least two weeks with a 15-cm tube placed in their home cage to allow voluntary exploration. Prior to rank assessment, mice were trained to traverse a 30-cm tube over two consecutive days (10 trials per day), with alternating entry from either end to prevent side bias. On the test day, each mouse pair underwent up to seven competitive trials in the same 30-cm tube. In each trial, the two mice were simultaneously released from opposite ends of the tube. A “win” was defined as one mouse successfully advancing through the tube while the opponent retreated completely out of the tube (all four paws outside) for at least 5 seconds. The first mouse to achieve four wins was designated as the dominant individual, whereas the opponent was classified as subordinate. Social rank stability was reassessed one day after threat exposure using the same criteria. Only pairs with consistent ranks were included in subsequent analyses.

      The dominance hierarchy is established based on pairs of mice. However, the use of terms like "group cohesion" - typically applied to larger groups - to describe dyadic interactions seems overstated. Please revise the terminology to more accurately reflect the pairwise experimental setup.

      Thanks for the comment. We have replaced the term “group cohesion” with “social engagement”.

      (2) Criteria and Validity of Behavioral Classification:

      The criteria for classifying mouse behaviors (e.g., passive defense, active defense) are not sufficiently clear. Please explicitly state the operational definitions and distinguishing features for each behavioral category.

      Passive defense was defined as an immobility-based defensive strategy characterized by suppression of locomotor activity, including freezing and tail rattling. Active defense was defined as movement- or posture-dependent defensive strategy, including approach, investigation, withdrawal, and stretch-attend. We have clarified these in the revised manuscript.

      How was the meaningfulness and distinctness of these behavioral categories ensured to avoid overlap? For instance, based on Figure 3E, is "active defense" synonymous with "investigative defense," involving movement to the near region followed by return to the far region? This requires clearer delineation.

      Defensive behaviors in the rat exposure paradigm were grouped into two categories: passive and active defense, each comprising distinct behaviors. All the manually annotated behaviors were mutually exclusive; that is, each video frame was assigned a single behavioral label to avoid overlap across behaviors. Active defense includes four behaviors: approach, investigation, withdrawal, and stretch-attend. We have clarified these points in the revised manuscript.

      The current analysis focuses on a few core behaviors, while other recorded behaviors appear less relevant. Please clarify the principles for selecting or categorizing all recorded behaviors.

      Thank you for pointing this out. In the current study, we focused primarily on defensive and social behaviors. We also included several neutral solitary behaviors related to anxiety and defensive state, such as sniffing, grooming, and rearing, which were consistently expressed across animals and closely linked to our main findings. We have clarified these in the revised manuscript.

      (3) Interpretation of Key Findings and Mechanistic Insights:

      Looming exposure increased the proportion of proactive bouts in the dominant zone but decreased it in the subordinate zone (Figure 4G), with a similar trend during rat exposure. Please provide a potential explanation for this consistent pattern. Does this consistency arise from shared neural mechanisms, or do different behavioral strategies converge to produce similar outputs under both threats?

      Thanks for bringing up this important question. The consistent increase in proactive bouts in dominant mice across both paradigms suggests a consistent rank-dependent reorganization of dyadic interaction under threats. We propose that this convergence reflect a shared neural mechanism that links defensive state with social-rank information, potentially involving top-down regulation from the mPFC to threat-specific midbrain and hypothalamic defensive circuits. We have expanded the discussion to incorporate this explanation.

      (4) Support for Claims and Study Limitations:

      The manuscript states that this work addresses a gap by showing defensive responses are jointly shaped by threat type and social rank, emphasizing survival-critical behaviors over fear or stress alone. However, it is possible that the behavioral differences stem from varying degrees of danger perception rather than purely strategic choices. This warrants a clear description and a deeper discussion to address this possibility.

      We thank the reviewer for this insightful comment. We agree that, in principle, behavioral differences could arise from variations in perceived danger rather than strategic choice. In humans, decisions can sometimes reflect value-based strategies that override perceived danger. In contrast, under naturalistic threat conditions, mice likely rely predominantly on danger perception to make behavioral decisions, and such responses are expected to be consistent with value-based strategies shaped by natural selection. In the revised manuscript, we have expanded the Discussion to address the role of threat perception and its relationship to decision-making in our behavioral paradigms.

      The Discussion section proposes numerous brain regions potentially involved in fear and social regulation. As this is a behavioral study, the extensive speculation on specific neural circuitry involvement, without supporting neuroscience data, appears insufficiently grounded and somewhat vague. It is recommended to focus the discussion more on the implications of the behavioral findings themselves or to explicitly frame these neural hypotheses as directions for future research.

      We have revised the Discussion to focus more directly on behavioral findings and added explicit neural hypotheses as potential future directions.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how dominance hierarchy shapes defensive strategies in mice under two naturalistic threats: a transient visual looming stimulus and a sustained live rat. By comparing single versus paired testing, they report that social presence attenuates fear and that dominant and subordinate mice exhibit different patterns of defensive and social behaviors depending on threat type. The work provides a rich behavioral dataset and a potentially useful framework for studying hierarchical modulation of innate fear.

      Strengths:

      (1) The study uses two ecologically meaningful threat paradigms, allowing comparison across transient and sustained threat contexts.

      (2) Behavioral quantification is detailed, with manual annotation of multiple behavior types and transition-matrix level analysis.

      (3) The comparison of dominant versus subordinate pairs is novel in the context of innate fear.

      (4) The manuscript is well-organized and clearly written.

      (5) Figures are visually informative and support major claims.

      Weaknesses:

      Lack of neural mechanism insights.

      The current study focused on behavior. In the revised manuscript, we have incorporated a discussion of potential neural mechanisms and highlight this as an important direction for future work.

      Reviewer #3 (Public review):

      Summary:

      This study examines how dominance hierarchy influences innate defensive behaviors in pair-housed male mice exposed to two types of naturalistic threats: a transient looming stimulus and a sustained live rat. The authors show that social presence reduces fear-related behaviors and promotes active defense, with dominant mice benefiting more prominently. They also demonstrate that threat exposure reinforces social roles and increases group cohesion. The work highlights the bidirectional interaction between social structure and defensive behavior.

      Strengths:

      This study makes a valuable contribution to behavioral neuroscience through its well-designed examination of socially modulated fear. A key strength is the use of two ethologically relevant threat paradigms - a transient looming stimulus and a sustained live predator, enabling a nuanced comparison of defensive behaviors. The experimental design is robust, systematically comparing animals tested alone versus with their cage mate to cleanly isolate social effects. The behavioral analysis is sophisticated, employing detailed transition maps that reveal how social context reshapes behavioral sequences, going beyond simple duration measurements. The finding that social modulation is rank-dependent adds significant depth, linking social hierarchy to adaptive defense strategies. Furthermore, the demonstration that threat exposure reciprocally enhances social cohesion provides a compelling systems-level perspective. Together, these elements establish a strong behavioral framework for future investigations into the neural circuits underlying socially modulated innate fear.

      Weaknesses:

      The study exhibits several limitations. The neural mechanism proposed is speculative, as the study provides no causal evidence.

      Establishing causal evidence for neural mechanisms is beyond the scope of the current behavioral study. We highlight this as an important direction for future work in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Clarify the definitions of all behavioral categories (escape, assessment, passive defense, active defense, etc.) in detail.

      We have clarified the definitions of all behavioral categories in detail in the revised manuscript.

      (2) Reduce speculative statements about SC-VMHdm-mPFC circuitry, and provide the data about this circuit, if possible.

      We have reduced the speculative statements about this circuitry in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      We commend the authors on a carefully executed and conceptually clear study that makes a valuable contribution to behavioral neuroscience. Below are some issues and suggestions for this manuscript:

      (1) Please add the following to the discussion: the reason why, compared to looming exposure, social behaviors were more often followed by defensive behavior during rat exposure.

      We have added this discussion to the revised manuscript. This difference reflects the distinct temporal characteristics of the two threats. Following the transient looming stimulus, the threat rapidly ceases once the stimulus ends, reducing the need for sustained defensive behavior. Consequently, social interactions primarily occur after threat termination and are less frequently interleaved with subsequent defensive behaviors. Consistent with this interpretation, looming exposure selectively increased the duration of social behavior in subordinate mice. In contrast, rat exposure represents a sustained multisensory threat that maintains a persistently elevated defensive state. Under these conditions, both the frequency and duration of social interactions increased in dominant and subordinate mice. Moreover, huddling emerged as the predominant social behavior, and the frequent transitions between freezing and huddling suggest that social interactions become integrated with ongoing defensive responses, potentially serving as a safety-seeking or cohesive defense during sustained threat.

      (2) Figures 1B and 1C showed that Grooming has significantly decreased. Please verify whether the statistical methods and results are correct.

      The grooming data in the original Figure 1C did not distinguish dominant and subordinate mice and did not show significant decrease. We have removed it in the revised manuscript. In original Figure 2I (current Figure 1P), the social modulation on grooming behavior was observed only in dominant mice (Two-way ANOVA with post hoc Tukey’s range test).

      (3) To investigate how social context modulates the expression and progression of defensive responses, the authors analyzed behaviors in two time windows: the early phase (0-5 seconds after stimulus onset) and the late phase (20-60 seconds after onset) in Figure2. What is the rationale for selecting 5 seconds as the cutoff for the early-phase behavioral analysis? Were the behaviors of mice between 5 and 20 seconds also analyzed?

      The reason to select 5 seconds as the cutoff for early-phase behavioral analysis is that most behavioral decisions are made within this time window. We also analyzed the behaviors of mice between 5 and 20 seconds and have integrated them into the revised manuscript.

      (4) Figure 2 demonstrated that the social context attenuates looming-evoked defensive behavior in a rank-dependent manner. Then, I would like to ask if the influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner? Discuss the similarities and differences between the social context's effect on looming-evoked defensive behavior and rat behavior.

      The influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner. Specifically, total stretch-attend (SA) time and SA frequency were increased only in dominant mice (revised Figures 2Q and 2R), whereas average SA duration was decreased only in subordinate mice (revised Figure 2S). Approach-investigation-withdraw (AIW) frequency was increased only in dominant mice while approach speed was increased only in subordinate mice (revised Figures 2U and 2V).

      Social context exerts a broad protective influence across threat types, but the form of this modulation differs depending on the nature of the threat. Similarity: social presence consistently alleviates threat-induced stress and reshapes defensive behavior in a rank-dependent manner, with dominant benefiting more strongly under both threats, suggesting higher social rank may associate with greater flexibility to integrate social modulation. Difference: the behavioral outcomes of social regulation are distinct across threat types. During looming, social presence primarily suppresses immediate defensive responses and alleviates post-looming anxiety, suggesting under transient and unpredictable threat, social context dampens acute defense and facilitates behavioral recovery. In contrast, during the sustained rat exposure, social presence promotes a shift in defensive strategy from passive to active defense, rather than mere suppression of defensive output. These results suggest that social context flexibly adjusts defensive behavior according to ecological demands by reducing excessive defense and anxiety in response to transient looming and facilitating active coping when threatened by a sustained live predator. These discussions have been included to the revised manuscript.

      (5) In Figure 4, looming exposure increased these two measures only in subordinate individuals, whereas rat exposure affected both ranks, indicating threat-specific modulation of social behaviors. The visual looming paradigm primarily simulates visual stimuli triggered by aerial predators, while the rat exposure stimulus involves not only visual cues but also other sensory inputs, such as olfactory signals for the experimental animals. This may lead to differences in the defensive behaviors of mice, particularly increasing the proportion of proactive behaviors in dominant individuals. If olfactory information transmission is blocked during rat exposure, can the behavioral phenotypes shown in Figure 4 still be observed?

      The rationale for incorporating both looming and rat exposure paradigms was to mimic the distinct, ethologically relevant predator encounters by rodents in natural environments. As the reviewer pointed out, the multimodality nature of the rat threat may contribute to differences in the defensive behaviors displayed by mice. This point is also acknowledged in our manuscript, where we state that rat exposure “imposes prolonged stress and elicits a broader repertoire of defensive behaviors”.

      We agree that systematically dissecting the contribution of specific sensory modalities (e.g., olfactory, visual, or auditory) to defensive behaviors during rat exposure is an interesting and important question. Based on prior literature, we speculate that olfactory cues likely play a major role. However, the main aim of the present study is to investigate how social regulation of defensive behaviors depends on dominance hierarchy and threat type, rather than isolating modality-specific sensory mechanisms. Within this framework, we preserved the multisensory features of the rat exposure to maintain its ecological validity and to emphasize its distinction from the visual-only looming threat. We thank the reviewer for raising this insightful question, but it is beyond the scope of current study. We will consider it as an important direction for future investigations.

      (6) Discuss the potential mechanisms for the similarities and differences in the impact of the looming threat and the rat threat on cohesive behavior?

      We have expanded the Discussion to address the potential neural mechanisms underlying both the similarities and differences in the effects of looming and rat threats on social cohesion. As discussed in response to issue (1), we propose that the behavioral differences between the two threats arise from their distinct nature. Here, we further discuss a potential circuit mechanism. Specifically, we propose that the mPFC serves as a common hub integrating social context and dominance-related information with threat processing, thereby contributing to the shared enhancement of social engagement under both threats. We further speculate that the distinct patterns of social behavior may arise from partially distinct defensive circuits. SC-centered visual threat circuits may facilitate rapid post-threat social engagement by recruiting vigilance-related networks, whereas VMHdm-centered predator-defense circuits may promote sustained social cohesion by engaging neural circuits that support coordinated coping during persistent defensive states. We emphasize that these are hypotheses, and that future studies combining circuit-level recordings and causal perturbations within the behavioral framework established here will be required to test them.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this work, the authors investigate the mechanisms of low-frequency synaptic depression at cerebellar parallel fiber to interneuron synapses using unitary recordings that allow direct quantification of synaptic vesicle release. They show that sparse stimulation can induce robust synaptic depression even in the absence of substantial vesicle consumption, and that this depressed state is rapidly reversed when stimulation frequency is increased. To account for these observations, the authors propose a model in which low-frequency depression reflects a redistribution of vesicles within the readily releasable pool, in particular, a reduction in docking site occupancy due to vesicle undocking.

      Strengths:

      I found the experimental work to be of high quality throughout. The use of simple synapse recordings to count individual vesicle release events is particularly powerful in this context and allows questions to be addressed that are difficult to approach with more conventional approaches. The demonstration that low-frequency depression can occur independently of prior vesicle release, together with the rapid recovery observed during high-frequency stimulation, places strong constraints on possible underlying mechanisms and represents a clear strength of the study.

      The modelling framework is clearly laid out and helps organize a broad set of observations across stimulation frequencies. Several of the experimental tests appear well-motivated by the model, including the recovery train experiments, the analysis of failures, and the use of doublet stimulation. Taken together, the data provide a coherent phenomenological description of low-frequency depression and its relationship to vesicle availability within the readily releasable pool.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      While the experimental results are strong, the manuscript would benefit from rebalancing the strength of the mechanistic conclusions drawn from the modelling in light of its limitations. The framework is clearly useful and provides a coherent interpretation of the data, but it is not uniquely constrained by the experimental observations, and alternative models or interpretations could plausibly account for the findings. The use of different model regimes concatenated across time, with substantially different parameter values, highlights the abstract nature of the approach. For these reasons, the model seems best presented as one plausible explanatory framework rather than a definitive biological mechanism. Clarifying the distinction between data-driven observations and model-based inferences would help readers assess which conclusions are strongly supported and which remain more speculative.

      The interpretation of the Ca2+-related experiments would benefit from more cautious wording. The absence of detectable changes in presynaptic Ca2+ signals does not exclude more localized or subtle Ca2+-dependent mechanisms, and conclusions regarding Ca2+ independence should therefore be framed accordingly. In addition, while low-frequency depression is still observed at reduced extracellular Ca2+, these experiments appear less diagnostic of the specific model-derived mechanism emphasized elsewhere in the manuscript - namely, a selective reduction in docking-site occupancy - and should be discussed with appropriate qualification in the text.

      Concerning Ca2+ signals, the Reviewer is right. While we found no change in Ca2+ signalling apart from a slow Ca2+ accumulation during long trains at 1 Hz, the possibility of an undetected change cannot be excluded. We have added a word of caution in this direction on p. 11. Concerning the 1.5 mM Ca2+ experiments, the Reviewer presumably alludes to the first recovery train (yellow) point in Supplementary Fig. 2C. This is also the last point (s11) of the slow train at 0.5 Hz because no delay at all was interposed between the slow train and the recovery train. We have now included one more experiment (with a present total number n = 6), and we have corrected Fig. S2C accordingly. In the new version the depression measured for s4-s10 vs s1 during the 0.5 Hz trains is 0.69 +/- 0.05 (p = 0.00058, paired one-tail t-test). The ratio of the s1 value of the recovery train compared to control s1 is 0.83 +/- 0.08 (p = 0.028, paired one-tail t-test).

      Major points:

      (1) Clarify and qualify mechanistic claims derived from the model.

      Throughout the manuscript, changes in model parameters are at times described as if they directly reflected underlying physiological mechanisms. As a result, the conceptual distinction between experimentally observed phenomena, model-derived variables, and biological interpretation is not always clear. Several conclusions in the Results and Discussion are phrased as mechanistic statements, although they rest on assumptions intrinsic to the modelling framework. The authors should systematically review the text and explicitly distinguish between (i) experimentally observed changes in synaptic responses and (ii) inferences about vesicle docking states or transitions within the model.

      In particular, statements implying that vesicle undocking is the mechanism underlying low-frequency depression should be rephrased to reflect that this is an interpretation within the proposed framework rather than a uniquely demonstrated biological process. For example, statements such as "Low-frequency depression is caused by synaptic vesicle undocking" should be replaced with formulations such as "Within the framework of our model, low-frequency depression is accounted for by a redistribution of synaptic vesicles away from docking sites" or "Our results are consistent with a model in which changes in vesicle docking-state occupancy contribute to low-frequency depression."

      A particularly problematic example is the statement that "these experiments further confirm that LFD only involves a decrease in δ, without accompanying changes in ρ or IP size." Here, an experimentally defined phenomenon (LFD) is directly equated with changes in model-derived variables. Such statements should be revised to make clear that δ, ρ, and IP size are inferred quantities within the model, and that the experimental data are interpreted through this framework rather than directly confirming changes in these parameters. Similarly, overgeneralizing statements such as "Undocking therefore represents the key mechanism controlling short-term depression across stimulation frequencies" should be softened to reflect that this conclusion emerges from the model rather than from direct experimental evidence.

      As suggested, we clarify the distinction in the revised version between experimental data and modelling, and we refrain from making definitive statements on underlying cellular mechanisms.

      (2) Address the biological interpretation of time-dependent model regimes.

      The model relies on distinct parameter regimes applied at different time points, with some transitions effectively suppressed in certain regimes. While this approach captures the data well, its biological interpretation remains unclear. The authors should either (i) expand the discussion to outline plausible biological processes that could give rise to such regime changes (for example, calcium-dependent modulation of transition rates or activity-dependent changes in vesicle state stability), or (ii) more explicitly frame this aspect of the model as a descriptive abstraction rather than a mechanistic proposal. This further underscores the need to clearly separate the descriptive role of the model from claims about underlying biological mechanisms.

      We thank the Reviewer for drawing our attention to this important point. Below 10 ms, rate constants are largely determined by the large-amplitude, fast-decaying Ca2+ signal occurring near voltage-dependent Ca2+ channels (‘Ca2+ nanodomain’). After 10 ms, the rate constants depend on the low-amplitude, slowly decaying Ca2+ signals averaged over the entire varicosity (‘volume-averaged Ca2+’). We explain this better in the revised version (Materials and Methods, p. 21).

      (3) Reframe conclusions drawn from calcium-related experiments.

      The calcium imaging data demonstrate no detectable changes in the measured presynaptic calcium signals under the tested conditions, but they do not rule out that calcium signals contribute in ways undetectable by the assay. Conclusions should therefore be revised to reflect this limitation, avoiding statements that exclude a role for calcium-dependent mechanisms. Wording such as "we did not detect evidence for..." would be more appropriate than conclusions implying the absence of an effect.

      Similarly, while low-frequency depression is still observed at reduced extracellular calcium (1.5 mM Ca<sup>2+</sup>), the specific mechanistic signature emphasized elsewhere in the manuscript - namely a selectively reduced first response during a high-frequency recovery train - is no longer apparent. These experiments should therefore be discussed as consistent with the proposed framework, but not as providing independent support for a selective reduction in docking-site occupancy. Explicitly acknowledging this limitation would improve clarity and avoid overinterpreting these data.

      This has been discussed above (‘weaknesses’).

      (4) Soften interpretations based on non-significant comparisons.

      In several places, comparisons that do not reach statistical significance are used to argue for equivalence between conditions (for example, comparisons involving failure versus non-failure trials or different LFD conditions). These conclusions should be revised to emphasize the limits of statistical power and framed as a lack of evidence for a difference rather than evidence of independence.

      We have amended this point in the revised version.

      Reviewer #2 (Public review):

      Summary:

      Silva and co-workers exploit their previously established methods of analyzing release events at single parallel fiber to molecular layer interneuron synapses. They observed synaptic depression at low transmission frequencies (< 5 Hz), which rapidly recovers during high-frequency transmission. Analysis of the time course of low-frequency depression revealed an initial rapid and a slow linearly increasing time course. Strikingly, the initial depression occurred even in the absence of preceding release, arguing against vesicle depletion as the underlying mechanism.

      Strengths:

      The main strength of the study is the careful demonstration of an interesting synaptic phenomenon challenging the classical vesicle-centered interpretation of synaptic depression.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      The finding of release-independent synaptic depression is important and would have widespread implications. Therefore, some more analyses to increase the confidence in these findings could be performed.

      My concern is whether rundown could explain the findings. If the rate of failures in s1 increases and at the same time the amplitude decreases during the experiments, an apparent depression in s2 could arise. The Supplementary Figure 5A addresses run-down, but the figure is not easy to understand, and, as far as I understood, it does not address the question of whether the release-independent depression could be caused by a rundown. To address this, the analysis of Figure 5 could be repeated by investigating the failure rate and amplitude separately or by analyzing the 1st and 2nd half of the recordings separately.

      The Reviewer makes a very important point that had escaped our attention. If the responses were declining over the course of an experiment, near the end of the recordings, a high proportion of failures would be associated with a weak response to the second AP. This could distort the relation between initial failures and amount of LFD, perhaps to the point of indicating LFD after failures when there were none. As suggested by the Reviewer, we tested this possibility by examining the stability of the synaptic responses during experiments. We found a mean s<sub>1</sub> value of 0.87 ± 0.13 for the first half of the experiments used in Fig. 5, and of 1.10 ± 0.17 for the second half (p > 0.05, n = 10). This analysis shows that there was no rundown during these experiments. We show in Author response image 1 a plot of s1 as a function of train number in these experiments, for two examples as well as for the average across frequencies and across experiments. These plots do not suggest any artefactual correlation between failures, mean s1, and rundown.

      Author response image 1.

      Plot of s1 as a function of train number for the experiments of Fig. 5

      In response to a request of Reviewer 2, Author response image 1 illustrates the evolution of s1 values as a function of train number for the experiments used to produce Figure 5. In each experiment, about 20 s1 values were obtained at two ISIs (either 10 ms and 500 ms, or 800 ms and 1600 ms). Author response image 1 shows two examples of s1 values as a function of train number (these values fluctuate widely between 0 and 3), and the average across cells and ISI values. There is no indication of a rundown of s1 values as a function of train number.

      Reviewer #3 (Public review):

      Summary:

      The manuscript builds on the observation that, at some synapses, low-frequency stimulation causes synaptic depression, which can be reversed by subsequent high-frequency stimulation. Such low-frequency depression (LFD) cannot be easily explained by the depletion of a single vesicle pool. Here, Silva and colleagues propose a model of activity-dependent vesicle trafficking to explain LFD at synapses between cerebellar granule cells and molecular layer interneurons.

      Strengths:

      Overall, LFD is interesting and worthy of examination, and the authors provide new experimental results that are of the high quality expected from this group.

      Weaknesses:

      The study proposes a novel model of vesicle trafficking that is not explained by known biological mechanisms, and the manuscript does not adequately compare or discuss alternative models.

      I have several concerns about how the authors interpret the data. First, the manuscript's primary conceptual advance is the idea that LFD involves vesicle undocking, rather than depletion. However, most experiments were performed under conditions that promote vesicle depletion (3 mM extracellular Ca2+). When experiments were repeated in physiological Ca2+, there appeared to be little or no LFD (stats are not provided). Second, the RS/DS/DU/undocking model, though not outside the realm of possibility, is not readily explained by known mechanisms and is only loosely supported by experimental findings. Third, when simulating LFD, the authors do not compare alternative models and use inappropriate language to imply that a model fit represents the truth (e.g., "the finding of identical experimental and simulated values confirms that the undocking mechanism accounts for LFD"). Finally, the model is presented in an overly complicated manner. The sheer amount of terms and nomenclature makes the manuscript confusing and difficult to read. Overall, the manuscript would benefit from added experiments and more statistics, a better justification and evaluation of the model, and more nuanced language.

      We respectfully disagree with these sweeping criticisms, as described in more detail below.

      Major concerns:

      (1) Most experiments were performed under conditions that exacerbate depletion

      In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal. As mentioned above, the authors only report significant LFD in elevated extracellular Ca2+. In a small number of experiments performed in more physiological Ca2+ (1.5 mM), there is no depression after a single stimulus, and it is not clear that there was statistically significant depression during a low-frequency train. Several studies cited in support of LFD share this problem:

      Abrahamsson et al., (2007) recorded from Schaffer collaterals in 4 mM Ca, 3-4X physiological Ca2+.

      Doussau et al., (2010) recorded from Aplysia synapses in 3X Ca compared to seawater.

      Rudolph et al., (2011) is cited as an example of LFD. However, this study performed experiments at high release probability cerebellar climbing fibers, and reported depression that increased monotonically with stimulation frequency, so it does not resemble the phenomenon studied in this paper. Lin et al., (2022) also largely describe monotonic depression at the calyx.

      The Reviewer suggests that LFD may only occur under non-physiological conditions, if the release probability has been increased by artificially elevating the extracellular Ca2+.

      The implication is that LFD is at best a curiosity with little or no significance for brain signalling. We disagree with this point of view for several reasons.

      Concerning the statement ‘In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal’: This is the purpose of the analysis shown in Fig. 5.

      The statement ‘the authors only report significant LFD in elevated extracellular Ca2+’ is inaccurate. Fig. S2C shows a clear LFD in 1.5 mM Ca2+, as acknowledged by Reviewer 1 (‘low-frequency depression is still observed at reduced extracellular Ca2+’). However, we failed to provide a p-value for the depression in the initial version of the paper (p was 0.004, n = 5 with this data set; paired t-test, one-tail). In the revised version, we document the 1.5 mM results more extensively, including the incorporation of the results of an additional experiment, and an explicit statistical analysis of the data (p = 0.00058, n = 6; paired t-test, one-tail).

      Concerning the statement ‘there is no depression after a single stimulus’: We find that the onset kinetics of LFD is slower in 1.5 Ca2+ than in 3 Ca2+ (respectively 1.8 ISI and 0.51 ISI, Fig. 2C and Fig. S2C). This explains that the PPR is not significantly <1 in 1.5 Ca2+ without implying any weakening of the extent of LFD at steady state.

      As explained in the manuscript (p. 5), in a previous work, we developed a method to ascribe changes in SV pools, within the RS/DS model, with specific modifications of s1, s2 and s5-s8 during test 100 Hz trains (Tran et al., 2022). This method was developed in 3 mM Ca2+ conditions, and for this reason, we performed most experiments for the present work in 3 mM Ca2+.

      Chiu and Carter (2024) demonstrated LFD in neocortical synapses; they performed their study in 1.2 mM Ca2+, not in elevated Ca2+.

      Rudolph et al., (2011) showed low-frequency depression not only in elevated external Ca2+, but also in 0.5 mM Ca2+. While Rudolph et al., (2011) did not make an explicit link between their observations and LFD, there is no reason to doubt that these observations are an example of LFD. They showed a biphasic depression when switching the stimulation frequency from 0.05 Hz to 2 Hz. In one of the founding papers of LFD, Doussau et al., (2010) describe a biphasic depression when switching the stimulation frequency from 0.025 Hz to 1 Hz; Fig. 1 of the two papers (Rudolph 2011 and Doussau 2010) are strikingly similar.

      Lin et al., (2022) would probably not agree with the statement that the depression at the calyx is ‘largely monotonic’, as they stress the finding of quasi-constant depression between 5 and 50 Hz.

      The authors note that their results differ from those of Atluri and Regehr, but do not mention that a possible reason for the difference is the increased release probability in their experiments.

      In fact, we clearly listed the difference in external Ca2+ as a likely source of the discrepancy by saying ‘This discrepancy presumably stems from differences in experimental conditions (room temperature, stimulation of multiple presynaptic PFs and 2 mM external Ca<sup>2+</sup> concentration in the previous work, vs. near-physiological temperature, single presynaptic stimulation and 3 mM external Ca<sup>2+</sup> here)’.

      The authors should provide statistics for the data obtained in 1.5 mM Ca, and discuss why LFD is increased in conditions that also elevate vesicle release probability.

      See our comments above: the revised version includes the requested statistics. On p. 6 of the manuscript, we do provide an explanation for the apparent lack of LFD at 1.5 Ca2+ and 2 Hz, namely a superimposition of LFD with facilitation. At 1.5 Ca2+ and 0.5 Hz, our LFD numbers are not weaker than at 3 mM Ca2+ and 0.5 Hz of 1 Hz.

      Altogether, it is correct that many LFD experiments have been carried out in high release probability synapses and/or under conditions of elevated Ca2+. However, the reasons underlying these choices are diverse (in our case, to build on the previous SV pool analysis developed in Tran et al., 2022 in 3 Ca2+ conditions) and do not imply a limitation to the phenomenon. LFD is present in physiological conditions for low-to-moderate release probability synapses (as shown in our work), and altogether, there is no reason to dismiss LFD as non-physiological.

      (2) Lack of biological mechanisms supporting the model

      The model is presented without compelling biological support. The evidence in support of vesicle undocking comes from experiments by the Watanabe lab, which showed fewer-than expected docked vesicles under EM when cultured synapses were stimulated immediately prior to high-pressure freezing. Kusick et al were careful to note that these vesicles may have been lost to fusion.

      The Watanabe lab showed an SV deficit at docking sites at times ranging from about 100 ms to several seconds (Kusick et al., 2020, their Fig. 5E). This corresponds to the ISI values where we see paired-pulse depression. In their Summary, Kusick et al., raise the possibility of SV fusion as an alternative to undocking at the 100 ms time point. But the same issue had previously been considered in Miki et al., 2018 with other techniques (their Fig. 2d), where it was shown that the SV deficit seen in paired-pulse experiments could not be explained by fusion. This leaves undocking as the most likely explanation, at least in our preparation. We have added a new paragraph on p. 14 to clarify this point.

      The putative undocking Kusick describes is immediate (< 5 ms after stimulation), and it was not shown to be Ca2+ sensitive. This manuscript describes "calcium-dependent undocking" that proceeds from 10 ms - 200 ms. Multiple studies from the Watanabe lab show that a single stimulus lowers the number of docked vesicles, and subsequently, there is a transient redocking of vesicles that can be blocked by EGTA or Syt7 knockout.

      This is not an accurate description of the Kusick results or of our results. In the Kusick paper, the SV deficit seen at <5 ms after stimulation is attributed to exocytosis, not to undocking. Clearly, it is Ca2+-dependent. Our manuscript describes potential calcium-dependent undocking not during the time 10 ms- 150 ms, during which our undocking rate is assumed to be calcium-independent, but starting at 150 ms and lasting a few hundred ms thereafter.

      I also question the rationale for the authors' model that 2 vesicles are coupled in series to a single release site. Previous papers from this lab cited EM studies from frog and neuromuscular that showed filamentous connections between vesicles (do these synapses show LFD?). Here, the authors primarily cite their previous models to support their arguments. I encourage them to continue searching for ultrastructural evidence for 2-vesicle-docking-units and to cite such studies.

      It is important to remember that our sequential two-step model was not based on EM data, but on a series of functional data including variance-mean analysis of summed SV release numbers; covariance analysis among subsequent SV release numbers; analysis of release latencies as a function of stimulus number during an AP train; analysis of SV release numbers under conditions of very high release probability. We note that the phenomenon of Ca2+-dependent docking that we proposed based on these observations has been consistent with flash-and-freeze or zap-and-freeze results from several laboratories. Concerning potential filamentous connections between SVs and the AZ plasma membrane at a distance of several 10s of nm, this has been seen not only in frog or mice neuromuscular junctions, but also at brain synapses (ex: Siksou et al., Journal of Neuroscience 2007; Cole et al., Journal of Neuroscience 2016; Fernandez-Busnadiego, Journal of Cell Biology 2010; 2013).

      (3) Comparison to other vesicle models

      The authors use overly assertive language to suggest that the model proves a mechanism. "Altogether, these results indicate that the slow phase of LFD ... reflects a δ decrease without significant changes in pr, in ρ or in IP size". Simulating data does not conclusively "indicate" the underlying mechanism, but the authors could state their data can be "explained by a model where..".

      Please see our response above to a similar point by Reviewer 1.

      However, LFD does not require activity-dependent undocking. Instead, the phenomenon has been explained by high-release probability, paired with an activity-dependent increase in either docking or release probability (Chiu and Carter, 2024; Doussau et al., 2017). Does the new model do a better job of replicating some facet of the data? If multiple models can explain the same data, how can we determine which model is correct? The "Alternative Presynaptic Depression Mechanisms" should be expanded to discuss these issues.

      We could not find statements in the Chiu and Carter paper or in the Doussau et al., paper explaining LFD ‘by high-release probability, paired with an activity-dependent increase in either docking or release probability’. As far as we can see, Chiu and Carter do not propose any specific mechanism for LFD, beyond saying that depression and facilitation must be separate. Doussau et al., (their Fig. 6) clearly frame their interpretation in a sequential two-step model. As in the preceding Miki et al., paper (which they cite extensively), they assume a rapid (a few ms), Ca-dependent transition between their ‘reluctant pool’ and their ‘fully-releasable pool’, respectively homologous to RS and DS. Thus, the Doussau et al., interpretation is close to that presented in our present work, even though significant differences exist. An important difference is that Doussau et al., did not use simple synapses, so that they did not have access to key synaptic parameters such as the number of docking sites or the release probability per docking site. Consequently, the model in Doussau et al., does not have the same level of detail as ours. The revised version explains better the differences and similarity between the models of Doussau et al., and that exposed in our work (new paragraph on p. 14).

      Recommendations for the authors:

      Reviewing Editor Comments:

      Three reviewers have seen the paper. They agree that the evidence for low-frequency depression (LFD) is solid, but they all suggest that the data with 1.5 mM extracellular calcium should be analyzed and discussed in more detail. The finding of release-independent depression is surprising, but further data analysis to investigate whether it could be caused by run-down is required. The reviewers think the study will be more complete if the authors add N's and stats for LFD in 1.5 mM Ca. In addition, the most important change would be adding clarity in the text about the limitations and simplifications of the model. Please edit the text to clearly state that the presented model provides one possible framework to account for the data, and that other models might also be plausible, and clearly describe the limitations and simplifications of the model.

      Thank you for your positive assessment of our work. As explained in more detail below, we have added the requested statistics for 1.5 mM Ca data. We have also thoroughly revised our text to separate more clearly facts and interpretation. Finally, we now explain in more detail the limitations of our model.

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) Statistical analysis and reporting.

      Please justify the use of one-tailed statistical tests and consider whether two-tailed tests would be more appropriate in some cases.

      In Materials and Methods (p. 17), we explain that one-tailed statistics were used when the sign of any possible deviation was predicted; otherwise two-tailed statistics were used.

      Exact p-values should be reported consistently rather than using threshold notation.

      We have corrected this as far as we could. Exact p-values that are still missing will be provided in the version of record.

      For analyses using rank-based statistics (e.g., calcium imaging and recovery experiments), please clarify which comparisons were paired and which were unpaired, and ensure that this is clearly indicated in the Methods and figure legends.

      We have now clearly indicated whether these statistics were paired or unpaired.

      (2) Figure clarity and presentation.

      In Figure 1, clarify whether the responses shown in panel C represent measured EPSCs or derived quantities, as this is not immediately clear from the labelling.

      Thank you for spotting this, the figure has been corrected.

      In Figure 2, please show the variability of control responses used for normalization.

      The variability of control responses used for normalization is now clearly explained in Materials and Methods (p. 18).

      In Figure 3, the use of similar color schemes for different model states and for data-model comparisons is confusing and should be revised.

      Thank you for spotting this; has been corrected.

      In Figures 4 and 5, consider reintroducing or clarifying schematic representations of the model states referred to in the text to aid reader comprehension.

      Figure 4 had already such representations. We have now modified these representations to make them simpler and clearer. We considered adding some model to Fig. 5 but could not come to any satisfactory option.

      (3) Controls and interpretation of pharmacological experiments.

      For experiments involving intracellular BAPTA, it would strengthen the interpretation to either include or explicitly discuss positive controls demonstrating that the manipulation was effective under the recording conditions. Alternatively, loosen the interpretation of those data.

      We have modified the text as suggested by the Reviewer.

      (4) Terminology and consistency.

      Ensure consistent use of terminology for vesicle states and model variables throughout the text and figures, and clarify whether different labels (e.g., RS/DS versus alternative notations) refer to equivalent states.

      We have revised our manuscript as suggested.

      (5) Model fit at short inter-stimulus intervals.

      In Figure 3B, the model appears to capture the overall trend of facilitation but shows a noticeable deviation from the data at the shortest inter-stimulus intervals. It would be helpful to clarify whether this reflects limitations of the current parameterization, known simplifications in the model (e.g., assumptions about fast calcium-dependent processes), or variability across recordings. Briefly commenting on this mismatch would help readers interpret the strengths and limits of the model in this regime.

      Because of the statistical nature of data (quantal fluctuations) and simulations (Monte Carlo trials), some differences between data and simulations are unavoidable. The low simulation point at 10 ms ISI in Fig. 3B may in addition result from an imperfect transition between the two time domains (time domains 1 and 2) of the simulation. We have added a sentence in Materials and Methods (p. 22) to explain this.

      (6) Model parameterization and fitting.

      Please clarify how model parameters were estimated, including the fitting procedure, cost function, and any constraints applied. It would be helpful to report confidence intervals or other measures of uncertainty for key fixed parameters, and to indicate how variable these parameters are across cells or recordings. In addition, a brief discussion of how sensitive the main conclusions are to variation in key parameters (such as release probability or transition rates) would help readers assess the robustness and identifiability of the model-derived interpretations.

      In addition, it is not entirely clear how model parameters and transitions differ across the concatenated temporal regimes used in the simulations, nor which parameters are held fixed versus altered or effectively suppressed. Summarizing, for each regime, the parameter values (or ranges) and active transitions-either in a table or schematic-would greatly improve transparency and help readers assess parameter identifiability and the robustness of the model-derived conclusions.

      For each group of experiments, results under control conditions (8-APs at 100 Hz) were averaged together, leading to the determination of one set of parameters (N, ⍴, δ, r, s, p<sup>r</sup>). Next the parameters for the second time domain were obtained by trial and error to optimize the fit for Fig. 1 (PPR) data. The parameters were constrained such the PPR recovered to 1 with a 10-second ISI. Unfortunately, the parameter domain was too complex to make an unsupervised fitting procedure practical. We have now added a paragraph on p. 22 to indicate which model parameters were heavily constrained by data and which ones were less constrained. Individual synapses were not simulated. Regarding the different time domains, the probabilities rb and sb are introduced for the second time domain and rf and sf are changed as shown in Fig. Supp. 3. The kinetic parameters are reset at the beginning of each simulated AP.

      (7) Clarify the scope of the model with respect to the slow component of low-frequency depression.

      The manuscript identifies a slow component of low-frequency depression that persists over long timescales. While this is an interesting observation, it is not fully clear whether this component is intended to be captured by the current model or represents an additional process outside its scope. Making this distinction explicit would help readers better grasp the scope and limitations of the model.

      As stated in our manuscript, modelling the slow component of LFD falls outside the scope of the present work. The set of rate constants in Fig. S3 does not produce any slow component, and we failed to find a combination of rate parameters that would produce a slow component. This is now clearly stated in the Materials and Methods section.

      Reviewer #2 (Recommendations for the authors):

      I have only the following suggestions to further improve the manuscript.

      (1) It is argued that LFD is alleviated when using doublets rather than singlets (Figure 8), but is the steady-state depression in s1 between singlets and doublets at the end of the experiment statistically different?

      It is unclear what should be considered steady state here. Average s<sub>1</sub> values are smaller for singlets than for doublets over the last 25 points (0.58 +/- 0.05 and 0.74 +/- 0.05, p = 0.02, unpaired t-test).

      (2) I do not understand Supplementary Figure 4. Does the value of zero in this plot indicate that the stimulation does not evoke any release anymore? How can this be differentiated from rundown? Can the LFD be reversed by a high-frequency stimulation?

      We have rewritten the description of this figure. We have also modified the figure labelling. Finally, we have changed the size and color of the symbol at zero delay to stress that the linear fit is constrained to include this point. Altogether we trust that these changes have clarified the presentation of these results.

      (3) The findings are related to Kusick et al., (2020). However, I think that in Kusick et al., (2020) as well as in Watanabe et al., (2013; doi:10.1038/nature12809), all tested time points ranging from a few milliseconds to several seconds after a stimulus show vesicle depletion except the time point of 14 ms in Kusick et al., (2020). I think these findings are, therefore, inconsistent with the much slower observed LFD.

      The increase in docked SV# at 14 ms is reported not only in Kusick et al., 2020 but also in Wu et al., 2023 as well as in Ogunmowo et al., 2025. Later, at 100 ms, the number of docked SVs drops again (Kusick et al., 2020, their Fig. 5 d-e, and Ogunmowo et al., 2025), to finally recover on a time scale of seconds (Kusick et al., 2020). These data fit with our observations if LFD is due to undocking, as we propose.

      (4) Page 13, 4th paragraph:

      Eshra et al., (2021) provide experimental evidence for "a Ca2+ -dependent rate of SV entry into the RRP".

      The Reviewer is right: Eshra et al., actually reported a weak Ca2+-dependence of the replenishment. We have therefore eliminated the Eshra reference in this sentence (note that this paper is mentioned elsewhere in the manuscript).

      Reviewer #3 (Recommendations for the authors):

      (1) Complicated terminology: I had a hard time understanding the model due to terms like δ, ρ, IP, DS, RS, RS gate, etc. Anything the authors can do to reduce the number of acronyms and simplify their description/depiction of the model will be helpful for readers.

      This resembles recommendation (4) of Reviewer 1. As stated in our response above, we have entirely revised the manuscript with these issues in mind. In addition, the abbreviation list at the onset of the manuscript should help the readers.

      (2) Figure 5 may represent the strongest evidence in support of an undocking model. This is the smallest figure, and I found it harder to understand. For example, I was initially confused because the blue "fail S1" traces seemed to suggest a very large PPR value. I suggest vastly expanding the size of this figure and the associated results section.

      Thank you for your interest in Fig. 5. We have explained the normalization procedure in more detail in the revised version. Also, we have modified the axis label of Fig. 5B to improve clarity. Finally, we have added a paragraph to explain the reason why the RS/DS model accounts for the results of Fig. 5.

      (3) Given the similar results but differing mechanistic conclusions of Doussau et al., (2017), that study merits more discussion in the manuscript.

      See our response to this point above (main comment (3) by the same Reviewer).

      (4) Several studies of LFD synapses have found a role for kinase activity (Silverman-Gavrila et al., 2005; Doussau et al,. 2010). This mechanism could be discussed or experimentally tested here.

      As mentioned in the Discussion section, the nature of the proteins involved in LFD, or in docking/undocking, remains uncertain. It is not surprising that broad-spectrum kinases and phosphatases affect LFD, as reported in the papers quoted by the Reviewer, but the full implications of these findings will only become clear

    1. Author response:

      The following is the authors’ response to the previous reviews

      Summary of revision for all reviewers:

      We are encouraged that the reviewers recognized the importance of the central question addressed by this study, the value of comparing acoustic- and cochlear-implant-evoked cortical responses within a common framework, and the potential relevance of these analyses for auditory neuroscience and neuroprosthetic design.

      At the same time, the reviewers made clear that the manuscript would be strengthened by:

      (1) Clearer calibration of claims regarding spatial organization 

      (2) Clarified explanation of the Figure 8 cross-modal analyses

      (3) More explicit discussion of the methodological and interpretive limitations of the present dataset, particularly the acute, anesthetized, and partially indirectly validated nature of the experiments

      In response, we substantially revised the manuscript to improve methodological clarity, narrow claims where appropriate, and align the title, abstract, results, discussion, figures, and legends with the precision of the data.

      Summary of major changes to revised manuscript:

      We revised the manuscript throughout to distinguish non-random spatial organization, coarse topographic/cochleotopic structure, and locally graded tonotopy/cochleotopy, and we softened claims accordingly, especially for the cochlear implant and TCA-derived analyses.

      We substantially clarified the Figure 8 cross-modal decoding framework, including:

      - The within-modality normal-hearing control

      - The normal-hearing-trained / cochlear-implant-tested analysis

      - The shuffled baseline used to estimate chance-level information transfer

      We expanded the methods and discussion to make the scope and limitations more explicit, including:

      - The acute and anesthetized nature of recordings

      - The use of monopolar stimulation

      - The lack of direct deafening validation in the main iEEG cohort

      - That ECAP forward-masking measurements were obtained in a separate acute cohort

      - The possibility that the observed acoustic–electrical mismatch is transient rather than fixed

      We also improved the consistency of our terminology, panel labeling, figure legends, and cohort descriptions.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank the reviewer for recognizing the importance of the question addressed by this study and for pushing us to sharpen the distinction between non-random spatial structure, coarse topography, and graded cochleotopy. These comments also helped us better calibrate our claims about TCA-derived maps and cross-modal generalization, and prompted clearer discussion of the limits of the present dataset.

      (1) The main weakness is that the evidence for spatial organization remains difficult to interpret. In Figure 2, the authors argue that both tone-evoked and cochlear implant-evoked responses are spatially organized, but the slope analyses are not significant for the cochlear implant condition. The revised vector-strength analysis supports the presence of non-random spatial structure, but this is not the same as demonstrating a clear graded cochleotopic organization. The manuscript would be strongest if it consistently distinguished between non-random spatial structure, coarse topography, and true graded tonotopy or cochleotopy.

      We agree. We revised the manuscript to distinguish these levels of interpretation more carefully and consistently. Specifically, we now reserve “non-random spatial organization” for results supported by the vector strength/shuffle analyses, and describe cochlear-implant-evoked maps as showing coarse and variable spatial structure rather than uniformly robust graded cochleotopy.

      We now also emphasize more explicitly that implanted animals varied considerably: some showed clear nonrandom organization, whereas others exhibited effectively random global maps. At the same time, the aggregate implant-evoked data remained non-random and showed decreasing spatial correlation with increasing electrode separation, which we interpret as evidence for coarse cochleotopic structure at the population level, rather than strong local graded cochleotopy in every animal.

      Accordingly, we revised the manuscript to distinguish:

      “Non-random spatial organization”

      “Coarse topographic/cochleotopic structure”

      “Locally graded tonotopy/cochleotopy”

      (2) A related issue is that some figure titles and interpretive statements still appear stronger than the data justify. For example, the TCA results in Figure 7 are described as revealing topographically organized latent spatial factors, but the statistical support appears strongest for normal-hearing high-gamma responses, with weaker or non-significant results in other conditions. These data remain interesting, but they would be better framed as evidence for weak or coarse spatial structure rather than robust topographic organization across all modalities.

      Now in our updated manuscript we revised the Figure 7 title, legend, and results to avoid implying robust topographic organization across all conditions. The manuscript now describes the TCA-derived spatial factors as showing coarse, non-random spatial structure, with the strongest support for local tonotopy in the normal hearing high-gamma condition and weaker or non-significant support for robust local topography in the other conditions.

      (3) The decoder analyses are improved, especially with the added tone-to-tone control. This control supports the conclusion that poor acoustic-to-CI transfer is not simply a failure of the TCA/LDA pipeline. However, the analysis remains model-dependent, and the absolute information transfer values are low. It would be helpful either to include an analogous analysis using raw ERP/high-gamma features or to explain more explicitly why the TCA-based approach is the appropriate primary test. The data support poor generalization between acoustic and implant-evoked cortical responses, but claims about perceptual qualities should remain speculative because perception is not directly measured in these experiments.

      Thanks, we now revised the Figure 8 results section to first explain that TCA is especially appropriate here because it constrains the decomposition into separable spatial, temporal, and trial factors. This makes it possible to learn spatial and temporal factors in the normal-hearing condition, hold those fixed, and re-optimize only trial factors in withheld normal-hearing data or cochlear-implant-evoked data, providing an interpretable test of within-modality recovery and cross-modal generalization.

      Thus, our emphasis on TCA is primarily conceptual, not merely computational. Raw-feature decoding could also be informative, but for the specific cross-modal question addressed here, TCA provides the more interpretable framework.

      We also now state more explicitly that the absolute information-transfer values are low, and we correspondingly limit our conclusion: normal-hearing-trained representations generalize poorly to acute cochlear-implant-evoked responses in this framework. We further and thoroughly revised the abstract, results, and discussion so that any perceptual implications remain explicitly speculative, since perception was not measured directly in these experiments.

      (4) Finally, although methodological reporting is much improved, some verification remains indirect. The authors provide useful implantation criteria and cite prior validation of their deafening approach, but the manuscript would be clearer if it explicitly distinguished between validation performed in the present animals and validation based on previous cohorts. This distinction is important because surgical variability, implantation efficacy, and deafening completeness can influence the interpretation of cochlear implant experiments.

      We revised the manuscript to make this distinction explicit. In the methods, we now state clearly that we did not obtain ABR measurements, hair-cell counts, or behavioral confirmation of deafening in the animals used for the present iEEG dataset, and that support for deafening efficacy in this cohort therefore relies on prior validation of the same procedure in separate cohorts.

      We also now clarify that the ECAP forward-masking measurements were obtained in a separate acute cohort (N=3) and were not recorded from the animals used in the main iEEG dataset.

      Our goal in these revisions was to make the provenance of each validation step fully explicit and to avoid implying that all validation measures were obtained in the same animals used for cortical recording.

      Reviewer #2 (Public review):

      We thank the reviewer for highlighting the value of the decoder-based analyses while also challenging us to frame the study’s novelty more precisely and to make Figure 8 substantially clearer. In response, we narrowed our novelty claims, improved the explanation of the cross-modal decoding framework, and clarified the logic of the shuffled baseline.

      (1) The observation that responses to cochlear implant stimulation (stimulation) is spatially organized is not new (e.g. Adenis et al. 2024)

      Thanks, good point. We revised the Introduction to make clear that prior studies have already demonstrated spatial organization of cochlear-implant-evoked cortical responses, including prior animal work cited in the manuscript. Our intended contribution is therefore not the demonstration of spatial organization per se, but rather the use of a shared analytical framework to test whether acoustic and electrical stimulation produce overlapping or transferable spatiotemporal cortical population representations, including single-trial decoding and cross-modal generalization analyses.

      (2) The claim that spatial and temporal dimensions contribute information about the sound is also not new there is a large literature on this topic.

      We revised the manuscript to avoid implying novelty for this general principle. Our intended point is narrower: in the present work, spatial and temporal dimensions were recorded simultaneously across a 60-channel cortical surface array and integrated into trial-by-trial population analyses that could be directly compared across normal-hearing and cochlear-implant conditions within a shared framework.

      (3) The analyses supporting the claim that there is a mismatch between cochlear implant and sound representation are still unclear, particularly in Fig. 8.

      We revised both the Figure 8 legend and the results section to explain the analysis step-by-step. In particular, we now distinguish clearly among:

      The normal-hearing to normal-hearing re-optimization control, which tests whether the TCA/LDA framework can recover stimulus-related structure when modality is unchanged;

      The normal-hearing-trained / cochlear-implant-tested analysis, which tests cross-modal generalization; the shuffled baseline, which estimates chance-level information transfer by destroying any structured relationship between predicted labels and actual implant channels while preserving matrix dimensions.

      We now also explain explicitly why the shuffled baseline can equal or slightly exceed the measured normal hearing-to-cochlear-implant transfer. Because observed cross-modal transfer was extremely small, shuffling does not restore meaningful structure; rather, it shows that the measured transfer lies at or below chance level. We believe that this substantially improves the clarity and interpretability of Figure 8.

      Reviewer #3 (Public review):

      We thank the reviewer for recognizing the strengths of the study design, the single-trial decoding analyses, and the potential clinical relevance of this approach, while also pressing us to better address the limitations of monopolar stimulation, possible place-frequency mismatch, and the acute nature of the recordings. These comments prompted important revisions to both interpretation and discussion.

      (1a) The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results as the authors ignored fundamental limitations of CI related stimulation. First, the authors stimulated in a Monopolar mode which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      We agree that monopolar stimulation is an important limitation, especially in the rodent cochlea, where current spread may be broader than in human subjects. We now acknowledge this more explicitly in the discussion.

      At the same time, our ECAP forward-masking measurements in a separate acute cohort provided supportive evidence for spatially and temporally tuned peripheral activation under the stimulus intensities used here. We therefore interpret the data as indicating that some peripheral selectivity was present, while also acknowledging that broader current spread under monopolar stimulation may have contributed to the coarse cortical organization observed in the implant condition.

      We also now state explicitly that the observed coarse cortical organization likely reflects a combination of peripheral current spread, downstream cortical pooling, and the spatial resolution limits of mesoscale surface iEEG, rather than any single factor alone.

      (1b) Comparing the averaged BF maps for iEEG (Fig-2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies might reveal a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Fig2F (and to some extend 2A) most of CI electrodes elicited responses around the 4kHz regions and averaged maps show a predominance of CI-3-4 across cortex (Fig-2C, H and Sup Fig. 3) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      We appreciate this point and agree that a precise one-to-one alignment between cortical best-frequency maps and intracochlear electrode position cannot be established from the present data. We now clarify this limitation in the manuscript, noting that the iEEG-based maps are relatively coarse and appear to overrepresent mid-frequency regions, limiting direct inference from implant electrode number to a precise acoustic-frequency equivalent.

      We also emphasize that our central conclusion does not depend on assigning each implant electrode a specific best frequency, but rather on comparing the spatial organization and cross-modal generalization of cortical population responses.

      (1c) Moreover, Supplemental figure 3 shows that only a couple of CI electrodes are predominately represented at the level of the cortex. Thus, it seems possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the placecoding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses.

      We agree that the predominance of only a subset of implant electrodes in the cortical maps is consistent with relatively coarse peripheral and/or cortical place coding. We revised the discussion to acknowledge more explicitly that possible current spread, broad peripheral excitation, and the limited spatial resolution of surface iEEG could all contribute to the coarse spatial organization observed in the implant condition.

      We also emphasize in the revised text that this possibility does not negate the evidence for non-random organization, but it does limit the strength of any claim about finely graded cochleotopy.

      (2) Second, although the authors acknowledge that post-lingual CI users always have an adaptation period, their conclusion is based on measurements that are relatively "early" in the CI-use timeline so to speak since iEEG were collected a) acutely right after mono-aural implantation and stimulation, b) under anesthesia, c) using unmodulated pulse train fixed at 900pps regardless of the electrode used and thus lacking any temporal information shifts in relationship to electrode cochleotopic placement. Basically, all CI electrodes had the same rate whereas you would expect basal CI electrodes to be amplitude modulated at higher frequencies than apical electrodes.

      We agree. We now emphasize more clearly that the present experiments probe only an early and simplified stage of cochlear implant use. The recordings were acute, performed under anesthesia, and used short constant-rate pulse trains chosen to isolate fundamental organizing features of primary cortical responses rather than to reproduce the full complexity of clinical stimulation.

      We therefore frame our conclusions specifically in terms of acute cortical encoding under these conditions, and now state explicitly that these data do not address how chronic use, wakefulness, adaptation, or more complex modulation-rich stimuli may alter more global cortical representations.

      (3) As much as the reviewer likes the overall approach with the use of PCA-LDA and TCA, and agrees that information transfer seems inexistant at time of measurement, authors should be more careful in their strong conclusion that two distinct encoding exist. The non-overlapping between sound and electric stimulation representations might exist only transiently and this should be acknowledged a bit more in the discussion. Without repetition of iEEG measurement at later period with chronic use of the CI, it is not possible to definitively claim that two distinct, non-overlapping coding co-exist at all times.

      We agree and appreciate this point. We revised the Discussion to make this much more explicit. Our data support poor overlap or poor cross-modal generalization at the time of acute implant activation, but they do not establish that this relationship is fixed over chronic implant use.

      We now state directly that the observed mismatch may be transient, and that future longitudinal studies will be required to determine whether cortical representations become more aligned with experience, whether downstream readout adapts to a novel code, or whether both processes contribute.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The authors have provided new analyses to support the claim that CI and NH representations do not match. However, the analyses performed are presented in an unclear manner. Most particularly, it is very hard to understand what has been done in Fig. 8D and Fig. 8G. The legend in D is incomplete. The legend for G does not explain why there are lines with different color and what are the different lines. The reasoning behind the shuffled CI dataset is not explained. Why does shuffling restore information transfer, this is very counter intuitive and raises again the question of the value of this analysis.

      Agreed, we have revised the Figure 8 legend and rewritten the Figure 8 results section to initially explain the analysis workflow more clearly.

      We now explain that the shuffled condition is used to estimate a chance-level baseline for information transfer in the normal-hearing-trained / cochlear-implant-tested confusion matrix. Specifically, we compute mutual information after randomizing the predicted labels in the confusion matrix, thereby destroying any structured relationship between predicted tone labels and actual implant channels while preserving matrix dimensions and marginal structure (results, Fig. 8 section).

      We also now state explicitly that while this shuffled condition can yield slightly higher mutual information than the non-shuffled NH→CI condition, the point is not that shuffling restores meaningful decoding. Rather, the point is that the observed NH→CI transfer is so low that it does not exceed this chance-level baseline (results, Fig. 8 section).

      (2) Note that the legends of fig. 8 are mislabel (goes up to H although the last panel is G, a legend for E is missing).

      Thank you for catching this error. We have corrected the Figure 8 legend so that all panels are accurately labeled and described.

      Reviewer #3 (Recommendations for the authors):

      (1) Fig. 2C and 2H are from 2 to 16kHz (as announced in the previous responses) but then Sup. Fig. 3 goes from 2 to 32kHz. Still on Sup. Fig. 3, color code for CI electrodes is inverted compared to the rest of the MS.

      Thanks, good catches. We have updated Figures 2C and 2H to represent the full tested frequency range, and we have corrected the inverted color code for cochlear-implant electrodes in Supplemental Figure 3.

      (2) Fig. 2A. says n=1, Fig. 2F should say the same. Fig. 2C and 2H should also have n=1 on bottom map. Fig. 2D and 2I should mention n=1. Same thing with Fig. 3A/3C, Fig. 3B/3D (n=7), Fig. 7A/7B/7C.

      We have revised the relevant figure legends to make clear that data are from a single animal unless otherwise noted, while minimizing visual crowding in the figure panels.

      (3) The reviewer once again thinks it would be better to not truncate the axis of Figs. 4C, 6C, 7D as it creates confusion with Figs. 5D and 8G, where suddenly, there are 8 electrodes. The legend justification isn't enough.

      We appreciate this concern and have clarified the issue in the revised manuscript. In some animals (N=3), some electrodes in the 8-channel array were non-functional prior to implantation, so those animals contributed six rather than eight usable implant channels. We now make this explicit in the manuscript and supplemental figures, including the number of channels used in the respective animals in the methods under “Cochlear implant programming” and in Supplemental Figure 3. 

      (4) Although the reviewer appreciated the justification of 15 PCA for their analysis, this justification should be provided as is in the methods, especially as the Sup. Fig. 4 does not even explain the significance of the linear regression of components 16 to 30. Some other reviewers would say that 10 components were more than enough without this justification.

      We agree and have provided further justification of the 15 PCA components to provide justification as is in the Methods “Principal component analysis” and in the results describing Figure 4. 

      (5) Sup. Fig. 1 and the eCAP measurements are a nice and needed addition to the MS but the authors should make it clearer that it came from 3 animals that are not part of the cohort presented in the rest of the paper / in Sup. Fig. 2. The method section certainly not make that distinction, giving the impression that all animals have been tested for channel interactions and temporal recovery.

      We agree and have revised the methods section to state clearly that the ECAP forward-masking measurements were obtained in a separate group of acutely implanted rats (N=3), rather than the main iEEG cohort in Supplemental Figure 1.

      (6) As said in the public review, the authors should acknowledge more the potential transiency of the nonoverlapping representation in the discussion, especially since most of the recordings are acute.

      We agree, and we have revised the discussion accordingly. As described in our public response above, we now state directly that the poor overlap between acoustic and electrical representations was observed under acute recording conditions and may not persist unchanged with chronic implant use or behavioral adaptation.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aimed at replicating two previous findings that showed (1) a link between prediction tendencies and neural speech tracking, and (2) that eye movements track speech. The main findings were replicated which supports the robustness of these results. The authors also investigated interactions between prediction tendencies and ocular speech tracking, but the data did not reveal clear relationships. The authors propose a framework that integrates the findings of the study and proposes how eye movements and prediction tendencies shape perception.

      Strengths:

      This is a well-written paper that addresses interesting research questions, bringing together two subfields that are usually studied in separation: auditory speech and eye movements. The authors aimed at replicating findings from two of their previous studies, which was overall successful and speaks for the robustness of the findings. The overall approach is convincing, methods and analyses appear to be thorough, and results are compelling.

      Weaknesses:

      Eye movement behavior could have presented in more detail and the authors could have attempted to understand whether there is a particular component in eye movement behavior (e.g., blinks, microsaccades) that drives the observed effects.

      Reviewer #2 (Public review):

      Summary

      Schubert et al. recorded MEG and eye tracking activity while participants were listening to stories in single-speaker or multi-speaker speech. In a separate task, MEG was recorded while the same participants were listening to four types of pure tones in either structured (75% predictable) or random (25%) sequences. The MEG data from this task was used to quantify individual 'prediction tendency': the amount by which the neural signal is modulated by whether or not a repeated tone was (un)predictable, given the context. In a replication of earlier work, this prediction tendency was found to correlate with 'neural speech tracking' during the main task. Neural speech tracking is quantified as the multivariate relationship between MEG activity and speech amplitude envelope. Prediction tendency did not correlate with 'ocular speech tracking' during the main task. Neural speech tracking was further modulated by local semantic violations in the speech material and by whether or not a distracting speaker was present. The authors suggest that part of the neural speech tracking is mediated by ocular speech tracking. Story comprehension was negatively related with ocular speech tracking.

      Strengths

      This is an ambitious study, and the authors' attempt to integrate the many reported findings related to prediction and attention in one framework is laudable. The data acquisition and analyses appear to be done with great attention to methodological detail. Furthermore, the experimental paradigm used is more naturalistic than was previously done in similar setups (i.e.: stories instead of sentences).

      Weaknesses

      While the analysis pipeline is outlined in much detail, some analysis choices appear ad-hoc and could have been more uniform and/or better motivated (other than: this is what was done before).

      Reviewer #3 (Public review):

      I thank the authors for their extensive revision of this paper, and I found some elements greatly improved.

      In particular, the authors do embrace a somewhat more speculative tone in the current version, which I think is fitting for this work, as the data seem (to me) to be not fully conclusive. The data set collected here is clearly valuable and unique (and I would encourage the authors to make it publicly available!), however, my overall impression is that the specific analyses reported here might not fully

      Despite the revised description of methods, results and figures, I still have trouble understanding many of the results and the authors conclusive interpretation of them. These are my main reservations:

      (1) Regarding "individual prediction tendency" - thank you for adding clarifying methodological details and showing the data in a new Figure (#2). Honestly, however, I still can't say that I fully understand the result. For example, why is there also a significant response in the random condition as well? And how do you interpret the interesting time-course (with a peak ~200ms prior to the stimulus, and a reduction overtime from there? Also (I may have missed this, but..) what neural data was used to train the classifier and derive the "prediction tendency" index? Was it just the broadband neural response? Is there a way to know which sensors contributed to this metric (e.g., are they predominantly auditory? Frontal?)? And is there a way to establish the statistical significance of this metric (e.g., how good the decoder actually was in predicting behavioral sensitivity?). I don't see any statistics in the results section describing the individual prediction tendency.

      (2) Regarding the TRF analysis - Thanks for clarifying the approach used to obtain 2-second long "segments" of speech tracking. This is an interesting approach, however I think quite new(?), and for me it raises a whole new set of questions, as well as additional controls and data that I would have liked to see, to be convinced that results are significant. I will elaborate:

      - Do I understand correctly that you segment the real and predicted neural response into 2-second long segments and then calculate the Pearsons' correlation between them to assess the goodness of the model? This is very unclear, since in the methods section you state only that "the same" analysis was performed as for the full data - but what exactly? Clearly, values will be very different when using such short segments. I feel that additional details are still required (and perhaps data shown) to fully understand the "semantic violation" analysis of TRFs.

      - I would like to reiterate my previous comment regarding the use of permutation tests to verify the validity of TRF-based measures derived. This would be especially important when using new approaches (such as the segmentation used here). The authors argue that this is not needed since this was not done in their previously published study. However, this sounds a bit like "two wrongs make a right" argument... why not just do it, and let us know that this 2-second segmentation approach allows estimating reliable speech tracking?

      - Following up on my previous comment that defining "clusters" as at least two neighboring channels (Figure 3) - the fact that this is a default in Fieldtrip is by no means sufficient justification! This seems quite liberal to me, especially given the many comparisons performed. Here too, permutations can help to determine the necessary data-driven threshold for corrections. This is of course critical for interpreting the result shown in Figures 3E&G that are critical "take home messages" of the paper - i.e., that the prediction-index from the first part of the experiment is related to speech tracking in the second part of the experiment. To my eyes, this does not look extremely convincing, but perhaps the authors can show more conclusive data to support this (e.g., scatter plots of the betas across participant?). - A similar point can be made for the effect of semantic violations (though here the scalp-level result is somewhat more clustered). The authors point out that the semantic effect is a "replication" of their result reported in Schubert et al. 2023, but if I am not mistaken the results there were somewhat different (as was the manipulation). It would be nice to explicitly discuss the similarity/difference between these effects.

      (3) Regarding the ocular-TRFs -

      - Maybe this is just me, but I believe that effects that are robust should be clearly visible in the data, without the need for fancy "black-box" statistical models. In the case of the ocular TRFs, it is hard for me to see how these time-courses are not just noise (and, again, a permutation test would have helped to convince me.). The inconsistent results for horizontal and vertical eye-movements vis a vis the experimental conditions (single vs. multi-speaker conditions) don't help either, despite the authors argument that these are "independent" - but why should this be the case, especially if there is nothing really to look at in this task? - I remain with this scepticism for the mediation-portion of the analysis as well... But perhaps replications from other groups or making the data public will help shed further light on this in the future.

      Minor

      - Thanks for adding information about the creation of semantic-violation stimuli. Since the violations and lexical-controls were taken from different audio recordings, it would have been nice to verify that differences between neural responses cannot be attributed to differences in articulations (e.g., by comparing their spectro-temporal properties)

      We are grateful to the Reviewing Editor and the three reviewers for the care and time they have devoted across the review rounds; their input has strengthened the paper. We are happy to confirm that we have carried out the two additional analyses the Assessment identifies as the path to a "solid" rating for methodological rigor. In brief:

      (1) Multiple-comparison correction for the prediction-tendency effect (Figure 3). We would like to be transparent about why our implementation differs from the specific procedure that was suggested. Testing the relationship between prediction tendency and neural speech tracking properly requires a mixed-effects model: the nested, repeated-measures structure of the data demands a per-subject random intercept (1|subject) to absorb the between-subject variability that the fixed effects do not capture. This is a standard and widely recommended approach for nested data, not an exotic one. Crucially, to our knowledge no implementation of a cluster-based permutation test exists for mixed-effects models — so the permutation/cluster procedure that was recommended cannot be applied to the very model the data structure requires. This constraint, rather than preference, is what led us to Bayesian estimation in the first place. We would add that the effect replicates Schubert et al. (2023), with a consistent left-frontal localisation across studies, which further speaks to its robustness.

      We fully share the concern underlying the recommendation, however — that a Bayesian model must equally guard against multiple comparisons across channels — and we have addressed it directly within the framework the data require. We tightened the HDI inclusion criterion to be analogous to a frequentist multiple-comparison correction (with boundaries comparable to a Bonferroni-corrected p-value) and re-estimated the model (encoding ~ prediction tendency * condition + (1|subject)) with 200,000 draws per chain (previously 2,000) to obtain stable posterior tails. The spatial pattern reproduces Figure 3E. At a Bonferroni-comparable per-channel boundary — itself widely regarded as overly conservative, just as the conventional 5% threshold is often criticised as arbitrary — no single channel survives in isolation; but considering the three-channel cluster jointly, the accumulated posterior places only ~0.7% of its mass at or below zero, i.e. a ~99.3% posterior probability of a positive effect (see Author response image 1). Rather than commit to a single arbitrary boundary, we report this accumulated posterior and invite the reviewers and readers to apply whatever criterion they consider appropriate.

      Author response image 1.

      (2) Circular-shift procedure for the mediation analysis (eye movements). Following Reviewer 1's suggestion, we replaced the previous shuffling control with a circularly shifted eye-movement predictor (shifted by half its total duration) as the control model. This isolates the genuine envelope contribution from any reduction in envelope weights that arises simply from adding a second predictor, and provides a stricter test of the specific temporal relationship between eye movements and the speech envelope. The mediation results are robust under this procedure (now Figure 5): vertical eye movements significantly mediate neural tracking of clear speech across all three principal components, and horizontal eye movements mediate tracking of target speech in the multi-speaker condition, with an early, sustained component peaking at ~0.18 s. Importantly, this stricter control did not materially change the interpretations we draw from these results: the pattern of mediation, and the conclusions about ocular speech tracking, remain consistent with those obtained under the previous approach.

      We have chosen to focus this revision specifically on these two points, which the Assessment identifies as what is needed to move the work from "incomplete" to "solid." We are sincerely grateful for the many further suggestions the reviewers raised, several of which we agree are thoughtful and point to worthwhile directions for future work. Since the previous round, however, the circumstances of the two authors who led this work have changed substantially: the corresponding author has since left academia altogether, and the shared first author's professional circumstances have likewise changed considerably. We are therefore not in a position to take on further analyses, or to prepare a point-by-point reply to the remaining comments, beyond the two above; we hope the editors and reviewers will understand that this necessarily defines the final scope of our revision. We remain sincerely grateful for the time and care they have invested in the manuscript.

    1. Author response:

      We are grateful to the reviewers and eLife editors for providing thoughtful and constructive commentary on our manuscript. The major concerns raised during review center on (A) the broader impact of methionine depletion on cellular metabolism beyond protein synthesis, including potential effects on cellular bioenergetics, and B) the extent to which assay conditions may induce cellular stress that influences downstream functional measurements. We will address both points during the revision period through experiments we have outlined below, most of which are already underway. 

      Related to (A), we will define the metabolic consequences of methionine deprivation on cellular metabolism with greater precision and are currently optimizing a third non-canonical alkyne amino acid, β-ethynylserine (β-ES), for labeling surface-exposed proteins in living cells. β-ES is a clickable threonine analogue that is efficiently incorporated into the proteome in the presence of physiological threonine and therefore does not require metabolic deprivation (1). We are in the process of optimizing β-ES for the MARBL workflow as a substitute for homopropargylglycine (HPG), which reduces the duration of methionine deprivation from six hours to two hours. Thus, β-ES eliminates the SAM-depletion concern for the baseline-translation half of the MARBL ratiometric measurement and constitutes a genuine methodological advance beyond the original submission. 

      Related to (B), we plan to generate additional data benchmarking MARBL against Seahorse and other established assays for measuring cellular bioenergetics, while also performing phenotypic and functional cellular characterization at key stages of the workflow. Experiments that are planned during the revision period are described in our point-by-point response below.

      Public Reviews:

      Reviewer #1 (Public review):

      The idea behind this paper is to have an alternate, reliable and quantitative approach to assess cell-cell metabolic heterogeneity. This study tries to achieve that using translationally-coupled energetic responses to metabolic stress. This is interesting because, in general, most quantitative measurements of metabolic outputs are 'bulk' and average for many cells. To overcome this, many recent studies use some read-outs of translation (presuming that translation is the single major energy sink in cells - however, this is objectively correct only in rapidly proliferating cells). That said, the authors take an interesting approach - to use two clickable methionine analogs, and assess baseline vs metabolically coupled translation within the same cell.

      The highlight is the methodology development where two distinct, clickable CMAs are used (to replace methionine in proteins). The labelling and approach are clever, and can be useful if carefully used. But there is going to be a challenge in using this, since the depletion of methionine itself (required for labelling), and a bias towards incorporation in some proteins (because these reagents are not highly permeable) will make it challenging to obtain precise metabolic state information - which otherwise can be obtained directly and far more precisely using a combination of other methods (ATP/flux measurement, respiratory capacity, translation rates etc). I therefore only make broader comments in my review below - to help structure this study better, clearly identify key limitations (and there are several that are clearly seen), and better clarify what MARBL may be useful for.

      (1) These reagents used for MARBL are not highly cell permeable/transported into cells, and largely work on the surface proteome. Which means the measurements related to changes in translation are indirect - quantified based on changes on the surface proteome (seen with labelling), and not the entire proteome.

      We appreciate the reviewer’s careful consideration of the MARBL labeling strategy and the opportunity to clarify an important feature of the method. We interpret this comment to mean that the detection reagents used in MARBL are not cell-permeable, rather than that the methionine analogues AHA and HPG are inefficiently transported into cells. HPG and AHA are transported through the ubiquitously expressed sodium-dependent neutral amino acid transporter SLC1A5 (2), and are incorporated into the proteome by intracellular translational machinery. The bias identified by the reviewer arises from the click detection step, not from metabolic labeling. We agree with the reviewer that MARBL measures the appearance of a subset of newly synthesized proteins where the amino acid analogue is accessible at the cell surface rather than the nascent production of the entire proteome. However, this is an intentional feature of MARBL and represents a key technical advance. It enables MARBL to generate a translation-dependent signal in live cells without requiring destructive interventions like fixation and permeabilization to access the entire proteome, as is required for other approaches such as BONCAT, THRONCAT, CENCAT, and SCENITH (1,3–5). Importantly, we empirically demonstrate that the surfaceaccessible fraction of the nascent proteome provides sufficient signal for quantitative measurements of translation that respond predictably to inhibition of protein synthesis as well as to metabolic perturbations that alter cellular energetics (Main Figure 1D-G, 2C). MARBL is also readily compatible with fixation and permeabilization, allowing the same labeling strategy to quantify analogue incorporation when a measure of total protein synthesis is desired and there is no need for the recovery of live cells. In the revised manuscript, we will include data directly comparing MARBL surface labeling with total nascent protein synthesis measured following fixation and permeabilization to show that these signals track with one another.

      (2) A major limitation - which will confound any interpretation- is the need to use methionine-free media; this can be a problem beyond protein synthesis/met incorporation since this will almost instantly deplete SAM pools in cells. There is little data provided on the impact of using this approach on SAM pools (time kinetics, how quickly SAM pools are affected, how much of the impact on metabolism comes from purely that, etc).

      This is important to establish because (i) of the continuous, very high flux of SAM -> SAH (eg. in the folate pathway, other methylations), and a constant need for SAM synthesis from methionine. This information will set the limits of capabilities of this method, as well as help delineate how much you can interpret results related to metabolic states between compared cells/states, etc. The labelling process is ~2 hours, while effects on SAM can be seen within minutes of methionine starvation in media, in metabolically active cells.

      The reviewer is correct in pointing out that cellular SAM pools are dynamic and adapt rapidly to methionine availability within hours of its deprivation or add-back (6,7). In the current version of the manuscript, we observe an ~100-fold decrease in intracellular SAM with 2 hours of methionine-free incubation in Jurkat cells, corresponding to the same time window over which we label with AHA in MARBL (currently presented as a min-max scaled heat map in Supplementary Figure 1A-B). We agree this needs to be presented more clearly as a potential limitation and quantified more explicitly. To address this, we plan to (i) reformat the existing SAM/ SAH/ methionine kinetics as absolute peak areas (Author response image 1A-C), (ii) add a novel dual-isotope simultaneous measurement of ATP turnover and SAM turnover across a panel of human cell lines to define the extent to which changes in SAM metabolism during the MARBL labeling workflow are likely to influence cellular energy demand (using <sup>13</sup>C<sub>5</sub>-methionine and H<sub>2</sub><sup>18</sup>O), and (iii) add β-ES as a threonine-based alternative to HPG for measuring baseline translation that substantially reduces the duration of methionine deprivation required for MARBL. β-ES is a clickable, alkyne-modified threonine analogue that is incorporated into newly synthesized proteins by mammalian cells in complete medium (1,4). Compared to methionine, which is a major part of the methionine cycle and 1-carbon metabolism that are required for cell proliferation and survival, threonine plays a less extensive role in mammalian intermediary metabolism. We plan to optimize and validate β-ESàAHA as well as AHAàβ-ES dual labeling and determine whether this modified MARBL workflow preserves the dynamic range and metabolic responsiveness of the HPGàAHA approach as in the original submitted manuscript. These experiments are currently underway, and we have already found that β-ES produces robust surface signal above background within 2-4 hours of labeling using our established MARBL protocol in the presence of normal threonine levels (Author response image 2A-D). We are currently optimizing the surface labeling protocol to further reduce the duration of baseline β-ES incubation in the revised manuscript. We are unable to eliminate methionine depletion entirely from the MARBL workflow. This is because endogenous methionyl-tRNA synthetases possess drastically higher affinities for canonical methionine (8–10), which prevents the incorporation of methionine analogues into proteins. However, these planned experiments will better define the impact of methionine limitation while also providing an alternative MARBL implementation that restricts methionine withdrawal to the shorter AHA-labeling window.

      Author response image 1.

      Dynamics of intracellular methionine and its derived metabolites. (A-C) LC-MS/MS raw peak areas in Jurkat cells of methionine (A), SAM (B), and SAH (C) after 2, 4, 6, or 8 hours of incubation in Met-free RPMI media supplemented with Met, AHA, or HPG. Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, SAM = Sadenosylmethionine, SAH = S-adenosyl-L-homocysteine. Statistics: Graphs display mean ± SD (A-C).

      Author response image 2.

      Extension of live surface labeling protocol using clickable threonine analogue to monitor surface translation. (A) Chemical structures for threonine and its clickable analogue β-ES. (B-C) Representative flow cytometric histograms (B) and corresponding gMFIs of extracellular Azd-647 signal after 1, 2, or 4 hours of β-ES incorporation. (D) gMFIs of Alk-647 (corresponding to AHA) or Azd-647 (corresponding to HPG or β-ES) after 1, 2, or 4 hours of incorporation. Abbreviations: β-ES = β-Ethynylserine, AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = Geometric mean fluorescence intensity. Statistics: Graphs display mean ± SD (C-D).

      (3) What is the effect on overall adenylate charge/ATP due to shifting to methionine-free media + label addition? How does it vary between the cells tested (suspension vs adherent)? This should be established before data related to Fig. 2. Does it correlate with extent of AHA incorporation?

      The primary conclusion that this method suitably reflects overall changes in energetics comes from the titration of 2DG/glycolytic inhibition.

      We thank the reviewer for raising these important points, which are related to point #2 above, as both ask how methionine-free labeling conditions alter cellular metabolism. We agree that a more comprehensive set of experiments exploring methionine deprivation will better delimit interpretations that can be drawn from a MARBL assay related to cellular energetics. As described in our response to point #2, we will measure baseline adenylate energy charge across a panel of adherent and suspension cell lines under methionine-replete, methionine-free, and methionine-free +AHA/HPG labeling conditions and determine its relationship to AHA incorporation. We will also perform analogous experiments during β-ES supplementation to determine how this alternative labeling condition affects cellular energetic state and β-ES incorporation. 

      We also wish to clarify that the adenylate energy charge measurements in Main Figure 2D and the measurements of newly synthesized ATP by H<sub>2</sub><sup>18</sup>O turnover in Main Figure 2E were performed in methionine-free media supplemented with either AHA or methionine. We apologize that this was not made clear in the figure key and legend. In addition, we inadvertently covered the methionine-replete control in the original figure with the key. This has been corrected in Author response image 3A-B and will be updated in the revised manuscript. In Jurkat cells, both adenylate energy charge and ATP synthesis are similar in the presence and absence of methionine.

      We note, however, that cells with intact energy-generating systems maintain adenylate energy charge within a relatively narrow range (11). Given this buffering capacity, we do not expect methionine deprivation to produce a large change in adenylate energy charge in most cell types. 

      Author response image 3.

      Validation of surface translation as a readout of cellular energetics in methionine-free and AHA-supplemented media. (A) Correlation between changes in adenylate energy charge versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.9044, Pearson’s R<sup>2</sup> = 0.8179, p = 0.0052). (B) Correlation between changes in newly synthesized ATP versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.8854, Pearson’s R<sup>2</sup> = 0.7840, p = 0.008). Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = geometric mean fluorescence intensity, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Pearson’s correlation coefficient was used for correlation and significance (A-B). The Met + DMSO condition was excluded from the Pearson correlation analysis. Graphs display mean ± SD (A-B).

      (4) Relatedly, if this label incorporation experiment is carried out (for ~2 hrs), and subsequently there is a washout/replacement with fresh, methionine-supplemented medium, (how quickly) do the cells recover and restore their energetic allocations?

      If we observe substantial changes in adenylate energy charge associated with methionine restriction, as assessed in the experiments planned in Points #2 and #3 above, we will perform methionine add-back experiments to define the kinetics of recovery. In addition, our plan to provide an alternative workflow that replaces HPG with β-ES will reduce the total duration of methionine depletion and further address this concern.

      (5) One possible advantage of a system like this can be to address questions in single cells/study cell metabolic heterogeneity. However, these are best done if the attaching moiety has a (selective) fluorescence increase and/or other read-out that can be quantitatively obtained at a single cell level. Largely, using AHA or HPG effectively only leads to bulk estimates (which can be sub-sorted towards single-cell estimates indirectly). This means that this method cannot really be used to study cell-cell metabolic heterogeneity effectively - compared to far simpler approaches, for example using a mitochondrial potentiometric dye with high fluorescence, or reporters for glycolytic activity, etc. This would also be related to Figure 5 - at best, this approach may complement existing approaches towards identifying heterogeneous sub-populations of cells. 

      However, I do agree that MARBL is flexible, stable, and can be internally normalised and used through flowbased platforms. It can supplement existing approaches to perturb bioenergetics, and also supplement existing approaches to understand metabolic state in live cells, particularly in suspension cells.

      We thank the reviewer for raising this thoughtful point, which we will address with textual changes in the discussion. We note that MARBL provides a quantitative single-cell readout when flow cytometry is used as the analytical endpoint, because the dual-color labeling scheme provides an internal control for baseline translation that accounts for variation in protein synthesis rates independent from cellular energetics. This permits a single metabolic resilience index to be calculated per cell and interpreted either at single-cell resolution or after grouping cells into sub-populations as a bulk estimate. Both approaches have utility depending on the biological question and intended downstream application. We agree with the reviewer that MARBL can be paired with other reporters for metabolism to obtain a more comprehensive understanding of bioenergetic heterogeneity and appreciate the reviewer’s recognition of its value as a complementary approach for studying metabolic state in live cells.

      Reviewer #2 (Public review):

      Summary: 

      Delacruz et al. describe a new method, called “MARBL” (Methionine Analogues for Ratiometric Bioenergetics in Live cells) to measure metabolic activity in single cells. The concept is similar to the SCENITH (anti-puromycin flow cytometry) assay to measure energy metabolism by measuring protein translation activity, yet offers, in theory, two advantages: 1) it keeps cells alive for downstream biological assays and 2) it is a ratiometric measurement, measuring both baseline translation and translation in the presence of metabolic inhibitors to correct for inherent cell-to-cell translation differences.

      Specifically, this method takes advantage of two click-chemistry-active methionine analogs, and then clicks fluorophores onto newly-synthesized surface proteins that have incorporated these analogs to measure translational activity. One methionine analog is given to cells for 2-4 hours to measure baseline translational activity, then metabolism is blocked using 2-deoxyglucose and oligomycin and the second methionine analog given to measure “metabolically-linked” translation activity. The authors establish this technique and show that mouse T cells polarized as pathogenic Th17 cells are more translationally active (“resilient”) compared to nonpathogenic Th17 cells, and when sorted, the resilient cells produce more interferon-gamma. This latter finding requires live cells after the metabolic measurement assay, showcasing findings that are inaccessible to the SCENITH assay.

      Strengths:

      The approach used is conceptually clever. It is appealing to measure metabolic/translational activity and to then be able to carry out further assays on sorted cell populations with different degrees of metabolic activity. This would indeed represent a useful advance.

      We thank the reviewer for recognizing the conceptual strengths of MARBL and its potential to enable downstream analysis of live cell populations with distinct metabolic states.

      Weaknesses:

      In principle, one key benefit of this technique is that cells can be used for biological assays after the metabolic measurement. Indeed, this would represent a valuable tool in the field.

      However, in this technique, cells are subjected to methionine deprivation, addition of non-natural methionine analogs, click chemistry, and high doses of toxic metabolic inhibitors 2-deoxyglucose and oligomycin. Indeed, the authors show in Figure S5F that 1/3 more of the post-MARBL cells die relative to cells not subject to this technique (60% viability in unclicked control, 40% in MARBL-measured cells). This data suggests that cells after this technique may be stressed and not reflective of the biological function of unmanipulated cells. More controls on viability and cell function (e.g. cytokine production) at more time points after the MARBL assay would have been valuable to address this issue.

      The reviewer makes important points regarding the conditions needed for the workflow of a MARBL assay that may impact cellular fitness. Perturbational methods, particularly techniques that interrogate bioenergetics, inherently require media-based or pharmacologically induced stress to evaluate cellular responses. Still, we agree that comprehensively profiling the fitness of cells following a MARBL assay and sorting is important since our technique aims to link cellular bioenergetics to functional outcomes. First, we would like to highlight that the MARBL-processed pathogenic and non-pathogenic TH17 cells depicted in Main Figure 5 were rested overnight after Fluorescence-Activated Cell Sorting (FACS) in complete RPMI medium prior to the restimulation assay. This is in line with standard practice to allow cells to recover from the shear stress of sorting prior to subsequent experiments (12,13). In addition, both pathogenic and non-pathogenic TH17 cells maintained the expression of lineage-defining transcription factors throughout the MARBL workflow, as shown by analyzing rested cells stained with antibodies for T-bet (expressed by pathogenic TH17) as well as RORγt (expressed in both cell types) by flow cytometry (Author response image 4A-C). To address this in the revised manuscript, we plan to repeat our non-pathogenic and pathogenic TH17 dual-MARBL and sorting experiment (Main Figure 5A-B) and rest sorted cells in RPMI with IL-2 for longer periods of time (24 or 48 hours) before restimulation for viability and cytokine analysis. These controls will provide more information on cellular fitness throughout the MARBL workflow, and we appreciate the reviewer’s suggestion.

      Author response image 4.

      Expression of lineage-defining transcription factors is maintained post-MARBL processing and fluorescence-activated cell sorting (FACS). (A) Experimental schematic. Ex vivo differentiated pTH17 and npTH17 cells were stained with CD45.2 antibodies conjugated to different color fluorophores, processed via MARBL, mixed at a 1:1 ratio, sorted, and then re-cultured in IL-2-supplemented RPMI media. After resting overnight, the expression of RORγt and T-bet were evaluated by intracellular staining and flow cytometry, distinguishing pTH17 from npTH17 cells based on prior CD45.2 staining. (B-C) Frequency of RORγt (B) and T-bet (C) positivity in murine Th17 cells post-MARBL processing, sorting, and overnight rest in IL-2-supplemented media. Abbreviations: npTH17 = non-pathogenic TH17, pTH17 = pathogenic TH17, HPG = Homopropargylglycine, AHA = Azidohomoalanine, Met = Methionine, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Graphs display mean ± SD (B-C).

      Another weakness of the paper is limited benchmarking against established metabolic assays in the field. The main assays used currently in the field are SCENITH and Seahorse. The authors do not compare their findings to SCENITH. They do compare their results to Seahorse, but the data shown don't address the key question: how does energy production measured by Seahorse, say in unmanipulated vs 2dg+oligomycin-treated cells, compare to the MARBL measurement? (Instead, they show a calculated "glucose dependence" metric in cells subjected to low vs high inhibitor dose, not showing the underlying data or cells that didn't receive an inhibitor).

      We appreciate the opportunity to perform additional benchmarking to define how the MARBL signal compares to existing methods for measuring cellular energetics. For our initial validation, we chose to benchmark the single-color MARBL signal against direct LC-MS/MS quantification of adenylate energy charge as well as newly synthesized ATP (Main Figure 2A-E). Although Seahorse and SCENITH are widely used standards in the field, these assays still provide indirect measures of cellular energetics, whereas LC-MS/MS quantifies the high-energy nucleotide pools that dictate cellular energy status. Both the adenylate energy charge and newly synthesized ATP decreased in response to increasing concentrations of 2-deoxy-D-glucose (2DG) and oligomycin A (Omy) treatments that impair ATP regeneration, and this energetic response displayed a linear relationship with the AHA click signal (Main Figure 2D-E). Nevertheless, we agree that additional benchmarking suggested by the reviewer will strengthen the methodological foundation of MARBL and help users understand how MARBL measurements relate to those obtained using more established metabolic assays. For this purpose, it is important to account for differences in what each assay measures. Seahorse resolves oxidative and glycolytic activity through simultaneous measurements of OCR and ECAR, whereas MARBL (as well as SCENITH and CENCAT) integrate the energetic contributions of these pathways into a single translation-dependent readout. This is the reason why we compared MARBL with Seahorse using the glucose-dependence calculation employed by SCENITH in the current version of the manuscript (Main Figure 2FH) (5). As suggested by the reviewer, we will assess how OCR and ECAR measured by Seahorse vary relative to the MARBL signal across different oligomycin and 2-deoxyglucose treatment conditions, providing a more comprehensive view of the bioenergetic responses captured by MARBL. 

      References

      (1) Ignacio BJ, Dijkstra J, Mora N, Slot EFJ, van Weijsten MJ, Storkebaum E, et al. THRONCAT: metabolic labeling of newly synthesized proteins using a bioorthogonal threonine analog. Nat Commun. 2023 Jun 8;14(1):3367. doi:10.1038/s41467-023-39063-7 PubMed PMID: 37291115; PubMed Central PMCID: PMC10250548.

      (2) Pelgrom LR, Davis GM, O’Shaughnessy S, Wezenberg EJM, Van Kasteren SI, Finlay DK, et al. QUAS-R: An SLC1A5-mediated glutamine uptake assay with single-cell resolution reveals metabolic heterogeneity with immune populations. Cell Reports. 2023 Aug 29;42(8):112828. doi:10.1016/j.celrep.2023.112828

      (3) Dieterich DC, Link AJ, Graumann J, Tirrell DA, Schuman EM. Selective identification of newly synthesized proteins in mammalian cells using bioorthogonal noncanonical amino acid tagging (BONCAT). Proceedings of the National Academy of Sciences. 2006 Jun 20;103(25):9482–7. doi:10.1073/pnas.0601637103 PubMed PMID: 16769897.

      (4) Vrieling F, van der Zande HJP, Naus B, Smeehuijzen L, van Heck JIP, Ignacio BJ, et al. CENCAT enables immunometabolic profiling by measuring protein synthesis via bioorthogonal noncanonical amino acid tagging. Cell Rep Methods. 2024 Oct 21;4(10):100883. doi:10.1016/j.crmeth.2024.100883 PubMed PMID: 39437716; PubMed Central PMCID: PMC11573747.

      (5) Argüello RJ, Combes AJ, Char R, Gigan JP, Baaziz AI, Bousiquot E, et al. SCENITH: A flow cytometry based method to functionally profile energy metabolism with single cell resolution. Cell Metab. 2020 Dec 1;32(6):1063-1075.e7. doi:10.1016/j.cmet.2020.11.007 PubMed PMID: 33264598; PubMed Central PMCID: PMC8407169.

      (6) Mentch SJ, Mehrmohamadi M, Huang L, Liu X, Gupta D, Mattocks D, et al. Histone Methylation Dynamics and Gene Regulation Occur through the Sensing of One-Carbon Metabolism. Cell Metab. 2015 Nov 3;22(5):861–73. doi:10.1016/j.cmet.2015.08.024 PubMed PMID: 26411344; PubMed Central PMCID: PMC4635069.

      (7) Chen Z, Chen W, Reheman Z, Jiang H, Wu J, Li X. Genetically encoded RNA-based sensors with Pepper fluorogenic aptamer. Nucleic Acids Res. 2023 Sep 8;51(16):8322–36. doi:10.1093/nar/gkad620 PubMed PMID: 37486780; PubMed Central PMCID: PMC10484673.

      (8) Kiick KL, Saxon E, Tirrell DA, Bertozzi CR. Incorporation of azides into recombinant proteins for chemoselective modification by the Staudinger ligation. Proc Natl Acad Sci U S A. 2002 Jan 8;99(1):19–24. doi:10.1073/pnas.012583299 PubMed PMID: 11752401; PubMed Central PMCID: PMC117506.

      (9) Beatty KE, Liu JC, Xie F, Dieterich DC, Schuman EM, Wang Q, et al. Fluorescence visualization of newly synthesized proteins in mammalian cells. Angew Chem Int Ed Engl. 2006 Nov 13;45(44):7364–7. doi:10.1002/anie.200602114 PubMed PMID: 17036290.

      (10) Kiick KL, Weberskirch R, Tirrell DA. Identification of an expanded set of translationally active methionine analogues in Escherichia coli. FEBS Lett. 2001 Jul 27;502(1–2):25–30. doi:10.1016/s0014-5793(01)02657-6 PubMed PMID: 11478942.

      (11) De la Fuente IM, Cortés JM, Valero E, Desroches M, Rodrigues S, Malaina I, et al. On the dynamics of the adenylate energy system: homeorhesis vs homeostasis. PLoS One. 2014;9(10):e108676. doi:10.1371/journal.pone.0108676 PubMed PMID: 25303477; PubMed Central PMCID: PMC4193753.

      (12) Pollizzi KN, Patel CH, Sun IH, Oh MH, Waickman AT, Wen J, et al. mTORC1 and mTORC2 selectively regulate CD8<sup>+</sup> T cell differentiation. J Clin Invest. 2015 May 1;125(5):2090–108. doi:10.1172/JCI77746 PubMed PMID: 0.

      (13) Roth TL, Puig-Saus C, Yu R, Shifrut E, Carnevale J, Li PJ, et al. Reprogramming human T cell function and specificity with non-viral genome targeting. Nature. 2018 Jul;559(7714):405–9. doi:10.1038/s41586-018-03265

  3. Aug 2026
    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors are trying to characterize the sources of H. pylori in an island system. They find that, like the humans, the bacteria are admixed, but there is no correlation within the island between human ancestry and bacterial ancestry.

      We thank the reviewer for their consideration of our work.

      Strengths:

      The study has taken particular care to characterize the humans from which isolates were obtained. Thus, it is a particularly convincing demonstration of a bacterial "melting pot".

      Weaknesses:

      The GWAS is highly confounded by population structure. In particular, there is a large group of strains that lack the cag pathogenicity island, and also differ in frequencies of other genes. So it's not clear that these differences are other than in cag status.

      We ran the mentioned GWAS using the Cabo Verdean European lineage as controls group and the European gastric cancer lineage as cases to investigate the genomic differentiation between the Cabo Verdean European lineages and gastric cancer-associated lineages. Our objective was not to replicate previously reported associations or to identify novel H. pylori loci associated with gastric cancer. Therefore, we will revise the manuscript to make our objective clearer and remove statements that suggest direct associations with gastric cancer. In addition, because we employed a linear mixed model that accounts for population structure through the estimation of a genome-wide kinship matrix, we will include the genomic inflation factor. These measures will provide readers with a clearer assessment of the extent of population structure influencing the association analysis.

      Figure 1 seems very inconclusive.

      We agree that Figure 1 shows overlap between H. pylori-seropositive and seronegative distributions; in addition, its visual interpretation might be also influenced by a number of outlying observations. Nevertheless, the figure illustrates differences in pepsinogen I and II concentrations and pepsinogen I/II ratios between H. pylori-seropositive and seronegative individuals. To improve clarity,  we will include statistics in the plot, and we will revise the text to provide a more complete description of the key findings shown in Figure 1. If the reviewers and editor consider that these descriptive data are not central to the main conclusions of the study, we would be happy to move Figure 1 to the Supplementary Material, where it can provide supporting context without detracting from the presentation of the principal results.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the population structure and ancestry of Helicobacter pylori in Cabo Verde, where the human population has mixed West African and European ancestry. The authors combine a population survey, serum markers, bacterial genome analysis, and paired human-bacterial ancestry data. They report high H. pylori seropositivity, several distinct bacterial groups, and limited correlation between human ancestry and bacterial ancestry. They also identify one European-derived bacterial group that appears to have undergone a recent expansion and carries fewer well-known virulence-related genes.

      The study is interesting, and the dataset is valuable, especially because population-based H. pylori genomic data from Cabo Verde and West Africa are limited. The results provide useful information on bacterial diversity, historical migration, and host-bacterial ancestry. However, some of the main conclusions are stronger than the evidence currently supports, particularly the claims of host adaptation, increased transmission, and reduced virulence.

      We thank the reviewer for their comments and careful consideration of our work.

      Strengths:

      (1) A major strength is the study population. Participants were recruited from the general population and not only from patients with gastrointestinal disease. This gives a broader view of H. pylori diversity in Cabo Verde than studies based only on hospital patients.

      (2) The number of participants tested for H. pylori antibodies is substantial, and the authors also obtained a relatively large number of bacterial genomes. The combination of human and bacterial genomic information is another important strength. This allows the authors to directly examine whether human ancestry is related to the ancestry of the colonising bacteria.

      (3) The population genetic analyses are extensive. The authors use several different approaches, and these generally support the existence of African-derived and European-derived bacterial groups in Cabo Verde. The identification of two low-diversity European-derived groups is also interesting and suggests a relatively recent expansion.

      (4) The addition of new strains from Ghana and Portugal improves the reference dataset. The results may help future studies of H. pylori population structure in Africa, Europe, Cabo Verde, and populations affected by historical Atlantic migration.

      (5) The finding that human ancestry and bacterial ancestry are only weakly related in this population is potentially important. It suggests that the long-term relationship between human and bacterial ancestry may be less stable in recently admixed populations.

      Weaknesses:

      The main weakness is that several biological conclusions are based on indirect evidence. The genomic results support recent expansion of one bacterial group, but they do not directly show that this expansion was caused by adaptation to local hosts or by increased transmission. Founder effects, population history, geographic clustering, household transmission, or random expansion may also explain the pattern. The wording should therefore be more cautious.

      We thank the reviewer for underlining the limitations of inferring the causes of lineage expansion from indirect genomic patterns alone. We will review the manuscript to more carefully reflect the uncertainty surrounding the drivers of this lineage expansion. We will also expand the Discussion to outline analytical approaches that may help distinguish between demographic and selective scenarios in a highly recombining species (where high levels of recombination may have obscured local signatures of selection). Where feasible, we will explore additional analyses of the genomic data to assess the extent to which the observed patterns are consistent with these alternative hypotheses.

      The conclusion of reduced virulence is also not fully supported. The expanded lineage often lacks the cag pathogenicity island and carries less virulent forms of vacA, which suggests lower virulence potential. However, this does not prove that the strains cause less gastric damage or lower disease risk. There are no endoscopic or histological data, and serum pepsinogen values are only indirect markers.

      We will review our manuscript to adopt a more cautious description of these strains. However, we emphasize that, as show in  in Figure 4 – source data3, we could not find strains belonging to this expanded lineage carrying an active cagPAI,  or a virulent vacA allele. Our GWAs analyses show high divergence between these strains and gastric cancer strains both at a genome-wide level, and at previously identified virulence genes (as sabA and BabA, in addition to the ones already mentioned above). Finally, serum pepsinogen serum analysis, although indirect, have been shown to compare with  to the “gold standard” method, histopathological biopsy microscopy (ex: Telaranta-Keerie et al 2010, 10.3109/00365521.2010.487918; Kitamura et al 2015, 10.1111/jgh.12987; Miftahussurur et al, 2020, 10.1371/journal.pone.0230064; please see more about this in our following comment). Taken together these results provide minimal support for the hypothesis that the expansion of this lineage is driven by increased virulence. We will review the manuscript in order to reflect this.

      The description of the study population as having limited gastric inflammation is too strong. Serum pepsinogen measurements are useful for estimating gastric atrophy, but they do not directly measure the degree of histological gastritis. In addition, participants were recruited independently of symptoms, but this does not mean that they were all asymptomatic.

      The epidemiological estimate is based on antibody testing. This measures seropositivity and cannot clearly distinguish current from previous infection. Therefore, terms such as active infection or colonisation should be used carefully.

      As mentioned above, serum pepsinogen serum analysis has been widely shown to reliably detect both chronic and atrophic gastritis (e.g. Miftahussurur et al, 2020, 10.1371/journal.pone.0230064). We adopted conservative cut-offs for serum pepsinogen values after a careful review of the literature in our analysis. However, we acknowledge that these cut-offs may vary between populations. Therefore, we agree with the reviewer that a more precise assessment would have included a sensitivity and specificity analysis of  serum pepsinogen measurements against histopathological biopsy microscopy in a subset of the individuals. Taking this is into consideration, we will review the manuscript to include this discussion and to adopt a more cautious description of these results. Although we acknowledge that seropositive does not distinguish active from past infection in the Discussion, we will bring this discussion into the Results section as well.

      The proposed new West-Central African bacterial group is based on a small number of reference strains from Ghana and Nigeria. The result is interesting, but broader sampling from African countries is needed before this group can be considered firmly established.

      We thank the reviewer for pointing the need to assess higher African diversity, which is an issue that we also mention in the Discussion. Although the new West-Central African group comprised fewer isolates than the other comparison populations, the clustering pattern is unlikely to be explained solely by sample size because the analysis was based on a normalised chromosome-painting coancestry matrix, which reduces the influence of unequal donor numbers. In our review, we will provide additional analyses to confirm the observed clustering pattern: 1) we will provide fineSTRUCTURE MCMC tree assignments; 2) we will repeat the Chromopainter/ fineSTRUCTURE analyses under random downsampling of the larger groups; 3) as suggested in further reviewer recommendations, we will repeat the Chromopainter/ fineSTRUCTURE analyses after removing the new Ghanaian sequenced strains.

      The interpretation related to the trans-Atlantic slave trade is plausible, but the data mainly show patterns consistent with known historical migration. They do not directly demonstrate when or how the bacterial lineages moved.

      We thank the reviewer for highlighting this point. The chromosome-painting analysis identifies shared ancestry and gene flow between populations but does not directly estimate the timing of those events. However, several lines of evidence suggest that the majority of the strains have been present in Cabo Verde for an extended period rather than representing recent introductions. First, the isolates were obtained from individuals who, as well as both of their parents, were born in Cabo Verde. Second, Cabo Verdean African strains show an excess of European ancestry relative to their putative parental African populations. And vice-versa: Cabo Verdean European strains exhibit an excess of African ancestry compared with their putative parental European populations. To further investigate this question, as also suggested in reviewer recommendations, we will explore the feasibility of identifying clonal or near-clonal relationships between Cabo Verdean and assess whether dating approaches such as BactDating can provide additional insights into the timescale over which these lineages have diversified and admixed within Cabo Verde.

      The gastric cancer comparison may also be affected by bacterial population structure. Differences between the Cabo Verdean lineage and gastric cancer strains may reflect ancestry or lineage differences rather than disease association alone.

      We thank the reviewer for highlighting this point. As outlined in our response to Reviewer 1, the primary objective of this analysis was to investigate genomic differentiation between the Cabo Verdean and gastric cancer-associated lineages. Therefore, the strong contribution of ancestry and lineage effects is central to the interpretation of this comparison. We will make this clearer in our final manuscript.

      Overall, the study achieves its main aim of describing H. pylori diversity and ancestry in Cabo Verde. The evidence is strong for the population structure and ancestry findings, but less strong for the proposed mechanisms of adaptation, transmission, and reduced disease-causing potential. The work will be useful to the field, but the main conclusions should be stated more carefully.

      We thank  the reviewers and the editor for their time and effort in assessing the manuscript. We will submit a revised version that carefully addresses these points and incorporates the suggested changes.

    1. Author response:

      We are grateful to the editor and reviewers for providing their time and expertise in the assessment of this article. We are glad that the overall evaluation is supportive.

      Reviewer #1 (Public review):

      This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

      Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

      Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

      We thank the reviewer for their positive assessment of our work. We agree all these questions are important. Our revision will provide repeatability estimates for all our indices, and use the repeatability of SSRT and T0 to estimate the proportion of SSRT variance that remains unexplained. We will mention that it is theoretically possible that shared neural processing speed between T0 and inhibitory processes contributes to the reported effect but, since the neural pathways are anatomically largely distinct, the available cognitive neuroscience literature suggests such contribution is unlikely to be strong enough to explain the observed relationship. The impact of correcting group differences in SSRT for T0 will entirely depend on where differences in SSRT between these groups come from. If it comes exclusively from peripheral delays, then the effect may indeed disappear. If it is central, then stronger differences may be revealed after correction. We appreciate these are key questions we would like answers to already, but identifying and reanalysing suitable datasets to answer it will be the purpose of future papers.

      Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

      By good quality data, we mean enough trials at the optimal RT and SOA combinations, and an adequate preprocessing pipeline that excludes non-standard trials (poor fixation, pre-emptive responses, large undershoot …). Participants with increased intra-individual variance or low compliance will need more trials overall to get enough trials around divergence times for these to be accurately estimated. Supplementary figure 5 illustrates how low trial numbers lead to an overestimation of T0, which can be mitigated by pooling across participants. We expect noisy data to have a similar effect, although its impact may differ across clinical conditions based on where the additional variability comes from. We shall have more clarity on the reliability of T0 in clinical populations once relevant datasets have been reanalysed, and use this to define constraints for future data collection. Again, this will need to wait for future papers.

      In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

      Reviewer #2 (Public review):

      Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

      Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

      As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

      The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

      We agree that our manuscript should acknowledge that Logan and Cowan (1984) explicitly stated that internal and ballistic delays contribute to SSRT and thank the reviewer for highlighting the need to clarify the relationship between our work and the original Logan and Cowan framework. We will clarify that the novelty of our claim is not that peripheral delays contribute to SSRT, which is a logical necessity, but that differences in peripheral delays (across conditions or people) contribute to differences in SSRT, sometimes to a large extent. Such differences in SSRT are very widely assumed to reflect inhibitory control in the large corpus of work that followed this initial literature. This corpus has essentially ignored the message about stimulus processing and ballistic delays and their implications for individual differences or changes across conditions. Therefore, we maintain that SSRT differences may have been widely misinterpreted. We reference the articles where these implications were clearly spelled out: these empirical and modelling studies were based on a few monkeys or human participants, and therefore unable to provide the large-scale demonstration we provide here.

      The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

      We thank the reviewer for pointing out that their original approach was mechanistically agnostic, which we will explicitly clarify in revision. We will remove the words “top-down” from our 4th sentence and reword the quoted sentence into "Changes in peripheral delays must be ruled out before linking changes in SSRT to inhibition or cognition". It remains the case that many hundreds of studies have since interpreted “ability to inhibit” and “latency of control” as specific to inhibition and control, and therefore have assumed that changes in SSRT directly reflect an inhibitory cognitive process. We agree that we do not currently offer a definition of what we mean by cognitive processes, except that they do not include incompressible sensory and motor delays. Based on the reviewer’s clarification, this common shortcut in the literature appears inconsistent with the spirit of the initial work, which we seek to rectify.

      Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

      We acknowledge our inhibition functions are narrow, but we do not believe this undermines the main conclusions, for several reasons. On the question of sensitivity to the stop signal, average spans for inhibition functions after participant exclusions were 37% for manual and 29% for saccades. This limited span is mainly attributable to our fixed SOA design, which was a necessary feature of a direct comparison between manual and saccadic behaviours. Figure S3 covers only 80 ms spread of SOA. The slopes are commensurate with most other studies, which cover much wider differences in SOA. If one selects the central 100 ms (where the slopes are steepest) from the figures in most previous papers, one will find stopping accuracy changes of around 30%. Therefore, sensitivity is similar.

      On the question of some functions not crossing 50%, the correlations between SSRT and T0 remain the same if we only keep those participants who crossed the 50% point (R(26)=0.64 for manual, R(11)=0.4 for saccades, same statistical significance levels). We will additionally rerun our SSRT and T0 correlation with stopping accuracy as a covariate. As stopping accuracy affects SSRT but not T0, it is unlikely to drive our results.

      The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

      We agree that the issue raised by the reviewer would be important if no difference were present between the distributions. In our dataset, however, all participants showed a measurable distributional difference. We believe the difference in perspective that ‘many participants…produce RTs on ignore trials essentially indistinguishable from RTs on no-stop trials’ may be attributed to differences in the way we analyse data (RT distributions versus mean RT).

      Nearly all participants in our final sample had mean RTignore – RTgo > 10 ms (significant at the individual level), except for 2 in the manual condition (and none for saccades). Following Bisset & Logan (2014), a lack of clear mean RTignore - RTgo difference in these 2 participants might have been interpreted as reflecting a different strategy. However, all our participants showed clear dips between go and ignore RT distributions when locked on signal onset, and these two manual participants were no exception. Therefore, accounting for response probability and RT at each SOA in our distributional analysis revealed clear ignore versus go differences, masked when relying on mean RT. The lack of mean RT difference for these two participants in the manual modality had no impact on our main hypothesis testing because T0 was extracted by comparing signal-absent and signal-present trials (pooling ignore and stop), while TS was extracted by comparing ignore and stop (not ignore and go).

      In terms of interpretation and whether participants employ a pause-then-go strategy, we understand performance in the selective stopping task as reflecting a combination of automatic activation and interference, and endogenous activation and inhibition. Our interpretation is that the pause component primarily reflects automatic interference triggered by stimulus onset, although strategic factors may also contribute in some circumstances. Individual differences in mean RTignore - RTgo could reflect both automatic interference and endogenous pausing (both of which would increase the difference between ignore and go trials), as well as subsequent failing to go on ignore trials (omissions, which would decrease differences by removing longer latency responses from ignore distributions just as in stop signal distributions). Although some of this can be described as strategic, some won’t be, and we therefore refrain from inferring strategy based on mean RTignore – RTgo.

      Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

      The second mode indicates that participants occasionally ignore the stop signal, which is why it peaks at the same latency as the rebound for the ignore distribution. These are not unusual, in particular in selective stopping designs, but are not as easily seen on cumulative functions, which are the standard way of plotting the results in this field.

      Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

      We will save figures showing the individual distributions in the OSF folder. Note that these figures can be produced by running the code we shared, so each step is already fully transparent, but we will create tidy versions that also highlight manually corrected indices.

      In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions, so that the paper can be revised to address the concerns.

      Reviewer #3 (Public review):

      Summary:

      Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

      Strengths:

      The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

      Weaknesses:

      The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

      We will add separate regression lines and R-values for manual and saccadic modalities on Fig.2A and an inset showing the variance only driven by individual differences across all archival data (i.e. z-scored per condition and datasets, R(215)=0.43, p<0.001). We will also state more explicitly that the overall correlation may not be the relevant one for researchers specifically interested in individual differences.

      While a large portion of SSRT literature is about individual differences, there are also many studies about differences between conditions, including comparing different response modalities. Therefore, it is a general question whether differences of any kind in SSRT reflect differences in inhibitory control or differences in sensory-motor delays. At a conceptual level, most users of the SSRT are intending to measure control, and have a conceptual model in which control is separate from the modality of response or the exact characteristics of stimulus delivery. Thus, it is important to point out that their measure of control is very much dependent on these things, and in fact to a much larger degree than the more subtle individual differences, group differences or conditions of interest.

      We agree that repeatability is critical for interpreting null results and will provide split-half repeatability for all our indices, including T0 and SSRT, and use these to correct their correlations. We thank the reviewer for this suggestion.

      A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T<sub>0</sub> and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T<sub>0</sub> across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T<sub>0</sub> commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

      We fully agree that the presence of a correlation is indeed entirely expected and obvious in our own account, as conveyed early on in the manuscript (“From Fig. 1D, it seems clear that SSRT and T<sub>0</sub> are inevitably connected”). We agree the main question is what is left for SSRT to explain. We will update our analysis of the Campbell et al. (2017) alcohol study as suggested. Future work can then focus on other “independent variables”.

      The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T<sub>0</sub>, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

      We fully agree with all this and will update the wording surrounding the lack of correlation between ΔT and the other measures in light of its repeatability

      In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T<sub>0</sub> and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

      All divergence times (T0, T0,stop, T0,ignore and TS) were confirmed by two of the authors. Each index for each modality is plotted on a separate figure (showing all individuals). It is technically easy to compare, say, T0 and TS, for one individual, but we refrained from doing this (and indeed ended up with many T0 > TS). We agree that blinding and inter-rater reliability are important safeguards and that our current analysis fell short of this. We will explore whether this can be done retrospectively, and report on this exercise alongside guidance on criteria used for manual corrections.

      This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

      We will make figures available in the OSF folder with individual distributions that supported the extraction of each index, flagging those that got manually corrected. This will make apparent that the vast majority of dips were very clear, for both T0 and TS. Unclear dips led to missing indices, and were therefore excluded from our hypothesis testing.

      One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

      Indeed, blockwise feedback would have been hard to implement reliably for saccades. As manual and saccadic blocks were interleaved, our hope was that participants could use the feedback received for manual to adjust their strategy for both modalities. We will check how often the feedback was triggered for manual blocks and, if more than negligible, we will note the reviewer’s suggestion as a possible driver for modality differences.

      A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T<sub>0</sub> and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T<sub>0</sub> varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T<sub>0</sub>, and manual and saccadic T<sub>0</sub> differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T<sub>0</sub> is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T<sub>0</sub> would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

      We agree with all this. We will use the repeatability of SSRT and T0 to disattenuate the slopes and consider alternatives to subtraction to correct for visuo-motor deadtime. Before we can elaborate on the modality effect on the speed of inhibition, we need to simulate the effect of motor variability on T0, TS and SSRT. If motor noise affects TS or SSRT more or less than it affects T0, this will affect our conclusions. We will explore this issue in the revision and report any analyses that bear on the robustness of the modality effect.

      These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this paper, Chen et al. identified a role for the circadian photoreceptor CRYPTOCHROME (CRY) in promoting wakefulness under short photoperiods. This research is potentially important as hypersomnolence is often seen in patients suffering from SAD during winter times. The mechanisms underlying these sleep effects are poorly known.

      Strengths:

      The authors clearly demonstrated that mutations in cry lead to elevated sleep under 4:20 Light-Dark (LD) cycles. Furthermore, using RNAi, they identified GABAergic neurons as a primary site of CRY action to promote wakefulness under short photoperiods. They then provide genetic and pharmacological evidence demonstrating that CRY acts on GABAergic transmission to modulate sleep under such conditions.

      Weaknesses:

      The authors then went on to identify the neuronal location of this CRY action on sleep. This is where this reviewer is much more circumspect about the data provided. The authors hypothesize that the l-LNvs which are known to be arousal promoting may be involved in the phenotypes they are observing. To investigate this, they undertook several imaging and genetic experiments.

      While the authors have made improvements in this resubmitted manuscript, there are still multiple concerns about the paper. I think the authors provide enough evidence suggesting that CRY plays a role in sleep under short photoperiod. The data also supports that CRY acts in GABAergic neurons. However, there are still major issues with the quality of the confocal images presented throughout the paper. In many cases it appears that the images are oversaturated with poor resolution, making it hard to understand what is going on. In addition, none of the drivers used in this study are specific to the neurons the authors aim to manipulate. Therefore, the identity of the GABAergic neurons involved in this CRY dependent sleep mechanism remains unclear. Similarly, whether l-LNvs are the target of this GABA mediated sleep regulation under short photoperiod is not fully demonstrated. The data presented suggests that but does not prove it.

      Major concerns:

      (1) While the authors provided sleep parameters like consolidation or waking activity for some experiments. These measurements are still not shown for several experiments (for example Figures 2E, 3, 4, 5, and 6). These data are essential, these metrics must be reported for all sleep experiments.

      These metrics have now been added to Fig.2 S4 and 5, Fig.3 S1-3, Fig.4 S2 and 3, Fig.5 S2 and 4, as well as Fig.6 S1.

      (2) Line 144 "We fed flies with agonists of GABA-A (THIP) and GABA-B receptor (SKF-97541) (Ki and Lim, 2019; Matsuda et al., 1996; Mezler et al., 2001). Both drugs enhance sleep in WT," The proper citation is needed here, Dissel et al., 2015 PMID:25913403. Both THIP and SKF-97541 were used in that paper.

      Thank you for pointing this out. We have modified our manuscript accordingly.

      (3) Figure 2C and 2F: it appears that the control data is the same in both panels. That is not acceptable.

      Thank you for pointing this out. We are now using data from control flies that were monitored in the same experiments as the experimental groups.

      (4) Figure 4A: With the quality of the images, it is impossible to assess whether GABA levels are increased at the l-LNvs soma.

      We apologize for the poor quality. Unfortunately, the GABA immunostaining does not work very well in our hands and thus the background is high. We have now commented on this issue in the fourth paragraph of discussion and have toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      (5) Fig 4 S1A shows colabeling of l-LNvs and Gad1-Gal4 expressing neurons. They are almost 100% overlapping signals. This would indicate that the l-LNvs are GABAergic themselves, or that there is a problem with this experiment.

      Fig 4 S1A demonstrates the expression pattern of SYT-GFP driven by Gad1GAL4, which should label the synaptic terminals of GABAergic neurons. Therefore, the labeling observed at l-LNvs suggest that GABAergic neurons project to l-LNvs. This is further validated by the GRASP and trans-Tango experiments.

      (6) Fig 4 S1B: Again, I can see colabelling of the GFP and PDF staining, suggesting that Gad1-Gal4 expresses in l-LNvs.

      Fig 4 S1B demonstrates anatomical sites where GABAergic neurons project to and form synaptic connections with PDF neurons. Therefore, GFP signals at the l-LNvs suggest that these cells receive synaptic inputs from GABAergic neurons, echoing the results shown in Fig 4 S1A.

      (7) Line 184: "Consistently, knocking down Rdl in the l-LNvs rescues the long sleep phenotype of cry mutants (Figure 4-figure supplement 1D)." This statement is incorrect as the driver used for this experiment, 78G01-GAL4 is not specific to the l-LNvs, so it is possible that the phenotypes observed are not coming from these neurons.

      Thank you for pointing this out. We have modified our manuscript to note this.

      (8) Figure 4G-K: None of these manipulations are specific to the l-LNvs. The authors describe 10H10-GAL4 and 78G01-GAL4 as l-LNvs specific tools, but this is not the case. Why not use the SS00681 Split-GAL4 line described in Liang et al., 2017 PMID: 28552314? It is possible that some of the effects reported in this manuscript are not caused by manipulating the l-LNvs.

      Thank you for pointing this out. We have now modified our manuscript to avoid misleading remarks. We have used SS00681 Split-GAL4 to express TrpA1 but did not observe any substantial effect on sleep duration under short photoperiod. Therefore, we did not use this line for further experiments.

      (9) Similarly for the manipulation of s-LNvs, the authors cannot rule out effect that are coming from other cells as R6-GAL4 is not specific to s-LNvs.

      We have now modified our manuscript to avoid misleading remarks.

      (10) The staining presented in Fig 5 S1 is not very convincing. Difficult to see whether Gad1-GAL4 only expresses in the s-LNvs.

      We have now quantified the GFP signal in the l-LNvs and s-LNVs in Fig.5 S1B and D. As can be seen, the s-LNvs show prominent signal above the background while the l-LNvs do not.

      Reviewer #3 (Public review):

      Summary:

      In humans, short photoperiods are associated with hypersomnolence. The mechanisms underlying these effects is however, unknown. Chen et al. use the fly Drosophila to determine the mechanisms regulating sleep under short photoperiods. They find that mutations in the circadian photoreceptor cryptochrome (cry) increase sleep specifically under short photoperiods (e.g. 4h light: 20 h dark). They go on to show that cry is required in GABAergic neurons and that the effects of the cry mutation on sleep are mediated by alterations in GABA signalling. Further, they suggest that the relevant subset of GABAergic neurons are the well-studied small ventral lateral neurons that they suggest inhibit the arousal promoting large ventral neurons via GABA signaling

      Strengths:

      Genetic analysis to show that cryptochrome (but not other core clock genes) mediates the increase in sleep in short photoperiods, and circuit analysis to localise cry function to GABAergic neurons.

      Weaknesses:

      The authors' have substantially revised their manuscript, and the manuscript is better for the revisions. However, the conclusion that the sLNvs are GABAergic is unfortunately still not well supported by the data. A key sticking point remains the anti GABA immunostaining, and specific driver lines for sLNvs and lLNvs.

      The authors should tone down their conclusions to reflect the fact that their data, as presented, does not support the model that cry acts in sLNvs to modulate GABA signalling onto lLNvs and thus modulate sleep.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #4 (Public review):

      Summary:

      Short photoperiod is an important experimental manipulation in neurobiology, endocrinology, and metabolism studies. However, the molecular mechanisms by which short photoperiod gives rise to behavioral phenotypes that are seen in seasonal affective disorders remain unknown. Using the classic circadian model organism Drosophila, this study examines short photoperiod-induced hypersomnolence and identifies the circadian photoreceptor cryptochrome as a regulator of GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions. The discovery has broad implications for understanding how short photoperiod modulates neural inhibition in circadian circuits in regulating sleep.

      Strengths:

      The Drosophila model provided a powerful platform to dissect the molecular mechanisms underlying short photoperiod-induced hypersomnolence. A battery of behavioral, imaging, circuit-manipulation approaches was employed to test the novel hypothesis that the circadian photoreceptor cryptochrome modulates GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions.

      Weaknesses:

      The current model proposed by the authors suggests that the small ventral lateral neurons of the Drosophila clock circuit are GABAergic; however, this remains unclear. At present, the field lacks sufficient data and validated reagents to definitively establish the GABAergic identity of these neuropeptidergic neurons.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      Recommendations for the authors:

      The manuscript has improved after revisions. However, the evidence in support of the claim that the sLNVs secrete GABA onto the lLNvs remains unconvincing. The evidence that loss of cry in GABAergic neurons modulates sleep is solid. However, the authors' claim that the sLNVs are the relevant GABAergic neurons is not sufficiently backed up by the evidence presented. We suggest that the authors tone down their conclusions to reflect this.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #3 (Recommendations for the authors):

      Minor points:

      (1) The authors suggest that the effects of cry on sleep and mediated by the Rdl receptor, and use the GABA agonist THIP as support of this argument. However THIP acts on the Lcch3 and Grd receptors, not Rdl

      Thank you for pointing this out. We have modified relevant discussion accordingly.

      (2) In several instances (e.g. line 66, line 148), the authors use 'consistently' in the sense of 'consistent with previous data'. It would be better if they rephrase this.

      This has been fixed.

      Reviewer #4 (Recommendations for the authors):

      It is my pleasure to serve as a reviewer for this revised manuscript. The authors have carefully revised the manuscript in response to the critiques raised by all previous reviewers and have used all the available reagents to conduct additional experiments to assess the GABAergic properties of the small ventral lateral neurons (sLNv). Although it remains unclear in the field whether sLNvs co-transmit GABA, this study raises this possibility within an interesting biological relevant context. I recommend this manuscript for final publication.

      Thank you for your comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      (1) The study emphasizes H3K4me2, which often serves as a precursor to H3K4me3, a well-studied modification during early development. Analyzing the new H3K4me2 dataset alongside published H3K4me3 data is crucial for comprehensively understanding epigenetic reprogramming post-fertilization and the interplay between histone modifications. However, the current analysis is preliminary and lacks depth.

      We fully agree with this valuable suggestion. Our research group has previously systematically profiled H3K4me3 dynamics in human and mouse early embryos, and the relevant results have been published in Science (2019). The core objective of the current study is to explore the erasure, re-establishment and biological functions of H3K4me2 during mammalian parental-to-zygote transition. To enrich our analysis, we have now integrated our H3K4me2 data with publicly available H3K4me3 datasets for joint analysis. The results clearly demonstrate that H3K4me2 is not merely a precursor of H3K4me3. These two histone marks present distinct genome-wide distribution patterns and perform independent regulatory roles in embryonic epigenetic reprogramming. We have added the joint analysis results and relevant discussions in the revised manuscript to elaborate the crosstalk between H3K4me2 and H3K4me3.

      Manuscript Revisions

      All supplementary analyses and discussions are located in the Results and Discussion sections, Results section (Page 6, Lines 156–165) (Page 7, Lines 197–201) (Page 10, Lines 274–278) and Discussion section (Page 13- 14, Lines 374–382) (Page 15, Lines 424–429.

      (2) Tranylcypromine (TCP) is known as an irreversible inhibitor of monoamine oxidase and LSD1. While the authors suggest TCP inhibits the expression of LSD2, this assertion is questionable. Given TCP's potential non-specific effects in cells, conclusions related to the experiments using TCP should be made with caution.

      We highly appreciate this important reminder about the off-target effects of TCP. We have supplemented two classic literatures (J Am Chem Soc, 2010; Mol Cell, 2010) which have proven that TCP acts as an irreversible inhibitor targeting both LSD1 (KDM1A) and LSD2 (KDM1B). According to our transcriptome and protein detection data, the endogenous expression level of LSD1 is extremely low in mouse early embryos. Therefore, LSD2 is the primary functional target of TCP in our embryonic experimental system. All conclusions derived from TCP treatment experiments are described prudently in the full text to avoid over-interpretation.

      Manuscript Revisions

      Relevant supplements are made in the Results section (Page 7, Lines 197–201).

      (3) Some batches of H3K4me2 antibody are known to cross-react with H3K4me3. Has the H3K4me2 antibody used in CUT&RUN been tested for such cross-reactivity? Heatmaps in the figures indeed show similar distribution for H3K4me2 and H3K4me3, further raising concerns about antibody specificity.

      Thank you for raising this critical question regarding antibody specificity, which is essential for the reliability of CUT&RUN experiments. The H3K4me2 antibody used in this study was purchased from Millipore (Cat. No. 07030). Based on the manufacturer’s product specification and our internal verification, this antibody has very low cross-reactivity with H3K4me3.The similar distribution shown in heatmaps is not caused by antibody contamination. Instead, it reflects the inherent spatial correlation between H3K4me2 and H3K4me3 on chromatin in early embryos.

      (4) Certain statements lack supporting references or figures (examples on page 9 can be found on line 245, line 254, and line 258).

      We apologize for the inadequate citation in the original manuscript. We have comprehensively checked the full text and added standard peer-reviewed references to all statements without literature support. Specifically, we have supplemented corresponding references for the content on Page 9, Line 259 and Line 266 as suggested. We also completed a full-text inspection to fix similar problems in other positions.

      Manuscript Revisions

      References are supplemented on Page 9, Line 259 and Line 266.

      (5) Extensive language editing is recommended to clarify ambiguous sentences. Additionally, caution should be taken to avoid overstatement - most analyses in this study only suggest correlation rather than causality.

      We fully accept this suggestion. We have thoroughly revised all ambiguous, redundant and grammatically problematic sentences throughout the manuscript to improve readability and academic rigour. Furthermore, we have carefully modified all overstated expressions. For all experimental results and bioinformatics analyses, we only use words such as correlate with, suggest, indicate to describe correlative relationships. All inappropriate causal inferences have been completely removed to ensure objective presentation of our data.

      Manuscript Revisions

      Full manuscript is polished and revised.

      Reviewer #2 (Public Review):

      (1) The authors claim that the Cut & Run worked for MII oocytes, zygotes, and the 2-cell embryos. However, it is unclear if H3K4me2 is erased during the stage or if the Cut & Run did not work for these samples. To support the hypothesis of the erasure of H3K4me2, the authors conducted immunofluorescence staining, and H3k4me2 was undetected in the MII oocyte, PN5, and 2-cell stage. However, the published papers showed strong staining of H3K4me2 at the zygote stage and 2-cell stage ((Ancelin et al., 2016; Shao et al., 2014)). The authors need to cite these papers and discuss the contradictory findings.

      The authors used 165 MII oocytes and 190 GV oocytes for the Cut & Run. The amount of DNA in MII oocytes is halved because of the emission of the first polar body. Would it be a reason that H3K4me2 has fewer H3K4me2 peaks in MII oocytes?

      Thank you for putting forward these thoughtful questions. Firstly, we have cited two published literatures (Ancelin et al., 2016; Shao et al., 2014) in the revised manuscript and discussed the inconsistent immunofluorescence results. The main reason for the discrepancy lies in different confocal microscope parameters including laser power, gain and exposure time adopted by different laboratories. In our study, we used unified imaging parameters to continuously observe samples from GV oocytes to blastocysts, so weak H3K4me2 signals at zygote and two-cell stages could not be detected. When we adjust parameters specifically for these stages, weak fluorescence signals can be observed. We have elaborated this point in the Discussion section.

      Secondly, we clarify that the reduction of H3K4me2 peaks in MII oocytes is not caused by decreased DNA content. Although MII oocytes extrude the first polar body during maturation, we collected the polar body together with oocytes in all CUT&RUN experiments, so the total DNA content of MII samples is not reduced. Combined with previous studies on human oocytes, we confirm that the loss of H3K4me2 peaks from GV to MII stage is a real physiological epigenetic change accompanying oocyte meiotic maturation and chromatin remodeling.

      (2) The authors claim that Kdm1a is rarely expressed during mouse embryonic development (Figure 4A). However, the published paper showed that KDM1a is present in the zygote and 2-cell stage using immunostaining and western blotting ((Ancelin et al., 2016)). Additionally, this paper showed that depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage, and therefore, KDM1a is functionally important in early development. The authors should have cited the paper and described the role of KDM1a in early embryos.

      We apologize for the ambiguous expression in the original manuscript. What we described is a relative expression level: in mouse early embryos, the expression of KDM1A is lower than KDM1B, rather than the absolute absence of KDM1A.

      (3) The authors used the published RNA data set and interpreted that KDM1B (LSD2) was highly expressed at the MII stage (Figure S3A). However, the heat map shows that KDM1B expression is high in growing oocytes but not at 8w_oocytes and MII oocytes. The authors need to interpret the data accurately.

      We sincerely apologize for the data misinterpretation caused by improper data normalization in the original heatmap. We have completely re-normalized the RNA-seq data and redrawn Supplementary Figure S3A.

      The updated heatmap clearly shows that KDM1B is highly expressed in growing oocytes, while its expression decreases in 8-week oocytes and MII oocytes. Combined with Figure 4A, we have rewritten the description of KDM1B expression trends across different oocyte stages, and all textual descriptions are now consistent with the corrected data.

      Manuscript Revisions

      Supplementary Figure S3A is remade; data interpretation is revised in the Results section. Supplementary Figure S3A (remade); Results section (Page 42).

      (4) All embryos in the TCP group were arrested at the four-cell stage. Embryos generated from KDM1b KO females can survive until E10.5 (Ciccone et al., 2009); therefore, TCP-treated embryos show a more severe phenotype than oocyte-derived KDM1b deleted embryos. Depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage ((Ancelin et al., 2016)). The authors need to examine whether TCP treatment affects KDM1a expression. Western blotting would be recommended to quantify the expression of KDM1A and KDM1B in the TCP-treated embryos.

      We dig the transcriptome data to confirm the specificity of TCP to KDM1b. In addition, the intervention of TCP on the whole fertilized egg in this study increased the H3K4me2 content, and the embryo development retarding effect was more significant than that obtained by crossing with normal paternal lines after knocking down KDM1B from the mother.

      (5) H3K4me2 is increased dramatically in the TCP-treated embryos in Figure 4 (the intensity is 1,000 times more than the control). However, the Cut & Run H3K4me2 shows that the H3K4me2 signal is increased in 251 genes and decreased in 194 genes in the TCP-treated embryos. The authors need to explain why the gain of H3K4me2 is less evident in the Cut & Run data set than in the immunofluorescence result.

      Thank you for this valuable question. The inconsistent data performance between immunofluorescence (IF) and CUT&RUN is determined by the essential differences between the two technical principles.

      Immunofluorescence is a global semi-quantitative method, which reflects the total content of H3K4me2 in the whole nucleus. The 1000-fold increase refers to the overall fluorescence intensity of the nucleus. In contrast, CUT&RUN combined with high-throughput sequencing is a locus-specific quantitative method, which detects H3K4me2 enrichment changes at individual gene loci. Different analytical models and threshold settings also lead to differences in final data presentation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper asks how the NK cell receptor KIR2DL4 binds HLA-G and undergoes endocytosis. The authors propose that an allosteric disulfide-bond switch controls whether the receptor is in a ligand-binding or non-binding state, and they support this model using mutagenesis, imaging, mass spectrometry, and structural prediction.

      Strengths:

      A major strength is the use of diverse, complementary approaches to validate the central claim. The authors combined unbiased random mutagenesis to identify key residues, confocal microscopy to track cellular localization, and mass spectrometry to quantify the redox states of specific disulfide bonds. These methods consistently support a single model: an allosteric disulfide switch. The transition between a Cys10-Cys28 bond and a Cys28-Cys74 bond serves as a functional switch that controls whether the receptor resides at the plasma membrane to bind ligand or remains inactive in endosomes.

      Weaknesses:

      (1) The core model is interesting, but some of the strongest mechanistic claims still rely heavily on structure prediction rather than direct structural evidence, especially the proposed HLA-G contact surface in Figure 6 (now in Figure 7).

      The crystal structure of KIR2DL4 has a D0 domain in the C10-C28 disulfide configuration [1]. The AlphaFold prediction is different, having a C28-C74 disulfide bond in the D0 domain. It is understood that any prediction could be wrong. Nevertheless, the AlphaFold structure did point to the possibility that KIR2DL4 exists in two different disulfide-bonded forms. We went on to demonstrate experimentally that these two forms coexist in human cells. This conclusion is independent of the structure predicted by AlphaFold.

      The second difference predicted by AlphaFold is an allosteric change in a loop distant from the disulfide bond, suggesting the possibility that it could control binding of HLA-G. Again, this prediction could be wrong. Nevertheless, considering that the KIR2DL4 used to obtain a crystal structure was in a C10-C28 bond configuration and did not bind HLA-G [1], we wondered if HLA-G would bind to KIR2DL4 in a C28-C74 configuration.

      New experiments included in the revision have shown that a purified Cys10Leu KIR2DL4 mutant binds HLA-G (new Figure 6). Solving the structure of a KIR2DL4–HLA-G complex would be ideal, but this has not been possible thus far. The difficulty in crystallizing KIR2DL4 may be due, in part, to its propensity to form oligomers [1], as shown in the new Figure S6.

      The addition of both direct binding of HLA-G to KIR2DL4 and functional data showing that KIR2DL4 induces an ISG response when in the C28-C74 but not in the C10-C28 configuration strengthens the conclusion that disulfide switching controls ligand binding and downstream signaling relevant to NK cell interactions with HLA-G in early pregnancy.

      (2) The paper supports an effect of the disulfide state on trafficking and uptake, but the case for direct KIR2DL4-HLA-G binding still feels somewhat indirect. The manuscript itself notes that direct binding had not been previously shown, and the current explanation partly depends on inference about which disulfide state is present.

      Direct binding and affinity measurements of HLA-G bound to the Cys10Leu KIR2DL4 mutant (in a C28-C74 disulfide form) are in the new Figure 6. This crucial result is also consistent with functional data. New experiments (new Figure 5D) have shown that the ability of HLA-G to stimulate a transcriptional interferon-stimulated gene (ISG) response occurred with the C28-C74 form, but not the C10-C28 form of KIR2DL4.

      Surface plasmon resonance data showed for the first time direct binding between the C28-C74 form of KIR2DL4 and soluble HLA-G, with a K<sub>D</sub> of 1.6 mM (new Figure 6). Binding to WT KIR2DL4, which is in both configurations, C10-C28 and C28-C74, was also detected but with a lower affinity (K<sub>D</sub> = 19.4 mM). The purified WT KIR2DL4 formed oligomers (new Figure S6). In addition, binding of HLA-G to KIR2DL4 depended on the sequence of the peptide presented by HLA-G, as only one out of three peptides tested was compatible with KIR2DL4 binding.

      This new data was obtained in the laboratory of Jamie Rossjohn at Monash University, Victoria, Australia. He, along with Jan Peterson and Priyanka Chaurasia are new co-author on our revised manuscript.

      (3) Most of the main experiments are done in transfected 293T cells, so it is still not fully clear how strongly this mechanism carries over to the more relevant NK-cell setting discussed in the paper.

      Primary resting NK cells are not amenable to transfection. Despite this technical hurdle, we have included two key findings with primary NK cells in the revision.

      (1) As in the 293T transfected cell system, we have shown that inhibition of PDI caused reduced uptake of HLA-G in primary resting NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      (2) We have shown that cell-surface C28-C74 KIR2DL4 on primary NK cells, as detected by mAb 2388, decreased upon inhibition of PDI, again consistent with the role of PDI in maintaining a pool of C28-C74-bonded KIR2DL4 at the cell surface (New Figure 5E, F). As shown in the original Figure 3, PDI could reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      (4) The cellular evidence for the PDI story is not specific, since it depends a lot on inhibitor and blocking experiments that could affect the broader extracellular redox environment.

      Using inhibitors that target PDIA1 selectively, namely Rutin (PDI-specific up to 30 microM), and a PDI-specific monoclonal antibody, we found that HLA-G uptake by primary NK cells was inhibited (Figure 4C, D). We admit that pCMPS and thiol blockade by DTNB (Figure 4A, B) affect the extracellular redox environment. Data that were obtained without PDI inhibitors include the reduction of the C10-C28 bond by PDI in WT KIR2DL4 in vitro (Figure 3F), direct binding of HLA-G to KIR2DL4 in a C28-C74 disulfide conformation (Figure 6), and a functional transcriptional response to HLA-G by C28-C74 KIR2DL4 and not with the C10-C28 KIR2DL4.

      Reviewer #2 (Public review):

      Summary:

      Rajagopalan et al show how extracellular domain features regulate KIR2DL4 internalization. The trafficking phenotypes of cysteine mutants are logically organized, and well-summarized in a Table. The disulfide mapping and differential alkylation strategy are appropriate and provide strong support for alternative disulfide configurations in D0. The higher accessibility or more selective reduction of Cys10-Cys28 as compared to Cys28-Cys74 by PDI is a key mechanistic anchor.

      Strengths:

      The identification of a conformational switch in KIR2DL4 is conceptually novel. Experimental elegance, detailed and well-written.

      Weaknesses:

      Most of the mechanistic work was shown in HEK293. The authors should exhibit relevance using primary NK cells (using primary NK)

      As primary NK cells are not amenable to transfection, it is difficult to dissect the role of each disulfide form of the receptor KIR2DL4.

      Instead, we have now included PDI inhibition experiments using primary NK cells and shown that PDI inhibition reduces HLA-G uptake by primary NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      Furthermore, inhibition of PDI caused a decrease of C28-C74 KIR2DL4 at the cell surface of primary NK cells (New Figure 5E, F). This data is consistent with a requirement for a switch from C10-C28 to C28-C74 catalyzed by PDI, which maintains a pool of C28-C74 KIR2DL4 at the cell surface for HLA-G binding and internalization. As shown in the original Figure 3, PDI can reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      Recommendations for the authors:

      Reviewing Editor Comments:

      To improve the strength of the evidence and the overall impact of the paper, please address the following major points:

      (1) Validation in Primary Cells:

      The central biological framing of the paper involves decidual NK cell responses to soluble HLA-G. We strongly recommend performing a critical experiment using primary NK cells to test whether PDI inhibition or thiol blockade alters KIR2DL4 surface retention and HLA-G uptake in a manner consistent with your observations in 293T cells.

      We have added new experiments with primary, resting NK cells, as described in our response to the major point 3 of reviewer #1, and to the weakness raised by reviewer #2.

      Briefly, we have included experiments in the revised manuscript on the effect of PDI inhibition on HLA-G uptake in primary NK cells (New Figure 4E, F, G) and on transient accumulation of KIR2DL4 at the cell surface (in a C28-C74 bonded form) of primary NK cells (New Figure 5E, F). The data showed that HLA-G endocytosis by primary NK cells and the presence of KIR2DL4 at the plasma membrane of primary NK cells were reduced after inhibition of PDI.

      (2) Clarification of the "Switching" Mechanism:

      The current data points toward the coexistence of the Cys10-Cys28 and Cys28-Cys74 states. Please clarify or provide evidence regarding whether a dynamic conversion occurs (e.g., prior to binding, upon ligand engagement, or during trafficking) versus a model of stable coexistence of two distinct receptor pools.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively, on fetal trophoblasts that encounter maternal NK cells in the decidua. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua and encounter decidual NK cells selectively express HLA-C, HLA-E, and HLA-G. Strong inhibition signals by LILRB1 and NKG2A-CD94 could prevent activation through KIR2DL4. However, KIR2DL4 signaling, which occurs in endosomes [4] where signaling is sustained [5], can bypass these inhibitory signals at the plasma membrane.

      (3) Specificity of the PDI Model:

      Please elaborate on the relevance of extracellular PDI. Specifically, how does PDI perturbation affect the relative abundance of the two disulfide forms in a cellular context?

      We show in Figure 5E that two mAb for KIR2DL4 recognize different forms of the receptor. While mAb #33 recognizes only the C10-C28 form of the receptor, which is not at the cell surface, mAb 2238 recognizes both forms of the receptor. This allowed us to examine the effect of PDI on surface expression of the C28-C74 form of KIR2DL4 as detected by mAb 2238. We show that PDI inhibition reduces surface staining of C28-C74 (new Figure 5F), consistent with a model whereby a switch from C10-C28 to C28-C74 is catalyzed by PDI.

      A quantitative assessment of the relative abundance of the two forms of KIR2DL4 upon inhibition by PDI in a cellular context would have to be carried out by mass spec analysis of the two forms before and after treatment. That would be a very challenging experiment to perform with intact cells rather than purified proteins.

      (4) Agonist Antibody Mechanism:

      The manuscript mentions mAb #33 as a KIR2DL4 agonist. It would be highly informative for the reader if you could elaborate on whether this antibody activates the receptor by stabilizing a specific disulfide state or by driving internalization independently of HLA-G.

      We have shown that the agonist mAb #33 recognizes only the C10-C28 form (Figure 5E). We do not yet understand how it activates KIR2DL4. We do know that mAb #33 is not driving internalization considering that the receptor internalizes constitutively and is predominantly located in endosomes in the absence of HLA-G. Instead, it is the C10-C28 form of the receptor that carries mAb #33 into endosomes. Understanding how mAb #33 may function as a receptor agonist will require crystallization of the antibody bound to the receptor and is beyond the scope of this study. Structural studies of KIR2DL4 have been very difficult, due in part to its isoforms and tendency to form oligomers. It is not possible to answer your interesting question at this time.

      Minor Revisions:

      (1) Imaging Quantification:

      Ensure all figure legends include the number of independent experiments (n), specific statistical tests used, and precise alignment with the Methods section.

      This information is now included in the Methods section.

      (2) Textual Flow:

      To enhance engagement, please integrate the logic of Table 1 more explicitly into the main text of the Results section.

      This has been done.

      (3) Structural Discussion:

      Acknowledge the limitations of using structure prediction for the binding interface and discuss how these models align with existing literature on KIR-ligand interactions.

      We have described the use of AlphaFold solely as a tool to make predictions. Predictions can be wrong. Even so, they can generate new and useful hypotheses, as they did here. Existing, traditional KIR-ligand interactions are not informative in the context of the D0 domain in KIR2DL4 for the following reasons:

      The KIR2DL1/2/3 receptors with 2 Ig domains (hence 2D) have a D1 and a D2 domain. A comparison with KIR2DL4, which has a D0 and a D2 domain, may not be informative.

      The KIR3D receptors have the three domains, D0, D1 and D2. A structure of KIR3DL1 bound to HLA-B has been solved [6] by our collaborator for the revision, Dr. Jamie Rossjohn. As shown and mentioned in our manuscript (Fig. S7C and Legend), “predicted” contacts of the KIR2DL4 D2 domain with HLA-G involve residues conserved in the heavy chains of HLA-B and HLA-G and residues conserved in the KIR3DL1 and KIR2DL4 D2 domains. It is therefore likely that the KIR2DL4 D2 domain contacts HLA-G in a similar way.

      As for the KIR2DL4 D0 domain, it is very different. Due to the similarity between D2 domains of KIR3DL1 and KIR2DL4, and to the lack of a D1 domain in KIR2DL4, the KIR2DL4 D0 domain is in a completely different space than the D0 domain of KIR3DL1. “Predictions” by AlphaFold show that there could be interactions between the KIR2DL4 D0 domain and HLA-G (Figures 7 and S7). These predictions could be wrong. Nevertheless, the disulfide switch in the KIR2DL4 D0 domain correlates with a predicted change elsewhere on D0 at a position compatible with proximity to HLA-G. Furthermore, the KIR2DL4 isoform with a Cys28-Cys74 bond is “predicted” to be more aligned with a potential binding site than the Cys10-Cys28 isoform. Having no structural guide as a reference on how KIR2DL4 D0 domain may interact with HLA-G, such predictions may generate testable hypotheses.

      As we clearly state in the manuscript: “Structures of KIR2DL4–HLA-G complexes obtained experimentally are required to determine how HLA-G distinguishes the D0 domain in the alternative disulfide-bonded configurations.” (Results), and “Rules that dictate HLA-G binding to KIR2DL4 await further studies and structures of KIR2DL4–HLA-G complexes.” (Discussion).

      In the revised manuscript, we have now included SPR binding data for KIR2DL4 with HLA-G. We also show a higher affinity of HLA-G for the C28-C74 form of KIR2DL4. This has strengthened the study as it validates our model whereby switching to the functional form of the receptor allows binding of HLA-G. In this regard, we also include data showing that only the C28-C74 form of KIR2DL4 can respond to HLA-G to induce transcription of an ISG response. This provides a functional correlate to the role of the different disulfide forms of the receptor.

      Reviewer #2 (Recommendations for the authors):

      Major points to address:

      (1) Exhibit relevance using primary NK cells (using primary NK). The central biological framing is decidual NK responses to soluble HLA-G during early pregnancy, yet most mechanistic work is in 293T transfectants. The authors can perform one of the critical experiments using primary NK cells with soluble HLA-G stimulation. They should test whether PDI inhibition/thiol blockade similarly alters KIR2DL4 surface retention and HLA-G uptake in primary NK cells

      These experiments have been performed in primary NK cells and are described in the new Figure 4E, F, G and Figure 5F.

      (2) The authors should detail more about the relevance of extracellular PDI and the effect of PDI perturbation on the abundance of the two disulfide forms in cells. They should also provide evidence or discuss whether switching occurs prior to ligand binding, upon ligand engagement, or during trafficking.

      Such experiments would be very challenging. The predicted structural change is minor and may not be detectable by changes in proximity of labeled reporters. Ligand is not required for switching. We do know that PDI can convert C10-C28 into C28-C74, presumably by accessibility to the KIR2DL4 Cys28 when bonded in a C10-C28 configuration (Figure 2). How ligands (mAb #33 or HLA-G) impact KIR2DL4 structure is unknown. Data are compatible with the possibility of a stabilization of C10-C28 by mAb #33 and of C28-C74 by HLA-G.

      (3) The authors should elaborate on whether mAb #33 activates by stabilizing or by driving internalization independent of HLA-G. This is very interesting to the reader, given mAb #33 as a KIR2DL4 agonist.

      The question is undeniably interesting. mAb #33 is not required for internalization but is required for signaling. The C10-C28 KIR2DL4 configuration to which it binds internalizes constitutively and resides mainly in endosomes. How mAb #33 internalization by KIR2DL4 (not the reverse) results in signaling is not known. Nor is it known for the alternative form, C28-C74, which binds HLA-G, internalizes it, and signals for a transcriptional response very similar to that of C10-C28 bound to mAb #33 [2]. The C28-C74 KIR2DL4 configuration is retained, probably transiently, at the cell surface, to be available for HLA-G binding and internalization.

      Minor points to address:

      (1) The authors should ensure that all imaging quantifications include n, the number of experiments, and statistical treatment. Some are described in the Methods section. Please align figure legends with the method in detail.

      Details of the imaging experiments are provided in the Methods section.

      (2) Please summarize the Table 1 logic in the main text for enhanced reader engagement.

      This has been done.

      (3) The authors identify both Cys10-Cys28 and Cys28-Cys74 states in human cells. The data points towards coexistence rather than towards dynamic conversion. Please provide clarity on switching versus stable coexistence of two forms.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua express HLA-E and HLA-G and encounter decidual NK cells that express LILRB1 and NKG2A-CD94. Strong inhibition signals induced by these two receptors could prevent activation through KIR2DL4. KIR2DL4 signaling in endosomes protects it from these inhibitory signals and benefits from the sustained signaling property of endosomal signaling platforms [5].

      (1) S. Moradi et al., The structure of the atypical killer cell immunoglobulin-like receptor, KIR2DL4. J Biol Chem 290, 10460-10471 (2015).

      (2) S. Rajagopalan et al., The fetal trophoblast cell marker HLA-G activates a type I interferon response in primary NK cells through the receptor KIR2DL4. Sci Signal 19, eadv2400 (2026).

      (3) E. O. Long, H. S. Kim, D. Liu, M. E. Peterson, S. Rajagopalan, Controlling natural killer cell responses: integration of signals for activation and inhibition. Annu Rev Immunol 31, 227-258 (2013).

      (4) S. Rajagopalan et al., Activation of NK cells by an endocytosed receptor for soluble HLA-G. PLoS Biol 4, e9 (2006).

      (5) M. Miaczynska, L. Pelkmans, M. Zerial, Not just a sink: endosomes in control of signal transduction. Curr Opin Cell Biol 16, 400-406 (2004).

      (6) J. P. Vivian et al., Killer cell immunoglobulin-like receptor 3DL1-mediated recognition of human leukocyte antigen B. Nature 479, 401-405 (2011).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      We thank the reviewer for pointing out that, as E/DXXD has been considered a conserved caspase-3 cleavage site, TEVDA has also been validated as a caspase-6 cleavage site, which we have missed. We will further confirm this using site-mutated expression vectors in S2 cells.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      We thank the reviewer for the suggestion. We will attempt to label Xkr-mCherry under apoptosis with other vesicles and change the ER marker to Cnx99A (Calnexin ortholog in Drosophila).

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      We thank the reviewer for their careful reading, and we will correct the mistakes in the revised manuscript.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      We agree with the reviewer’s opinion. These mutants from the BDSC are not completely deleted but partially deleted; therefore, the RT-qPCR assay may not be very accurate. We will detect the mRNA levels of tm9sf4, dorp9, and sac1 using RT primers from different cDNA regions to make the results more convincing.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      We thank you for raising this intriguing hypothesis. Indeed, PM localization of Lact:C2 was reduced in xkr<sup>ko</sup> cells, and the distribution could not be rescued after apoptosis. Unlike xkr<sup>ko</sup>, tm9sf4, and sac1 RNAi-treated cells displayed weak PS disorder, which may be due to the efficiency of knockdown. However, the statistical results indicated that the ratio of PM/Cyto was reduced in tm9sf4 and sac1 RNAi-treated cells.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      We thank for the reviewer’s suggestion. We will add the statistical results of Fig. 3I in the revised version.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      We thank you for raising this intriguing hypothesis. We will delete the relevant statement to eliminate unnecessary misunderstandings.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      We thank for the reviewer’s suggestion. We will move this to the main text and provide quantification in the revised version.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

      We thank for the reviewer’s suggestion. We will weaken the statement in the Results section of Figure 6 and perform osbpl9 knockdown to observe efferocytosis in Raw264.7 cells, as OSBPL9 interacts with Xkr8 strongly.

      Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

      Thank you for your comprehensive and generous assessment of our work and for the time and expertise you have devoted to reviewing our manuscript. We will revise the manuscript accordingly and provide a point-by-point response in the revised version.

      Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      We sincerely thank you for your careful reading and accurate summary of our manuscript. We appreciate the time, effort, and expertise you have dedicated to evaluating our work, and we will try our best to improve our manuscript according to your suggestions.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      We thank for the reviewer’s suggestion. We used AcD-treated S2 cells and untreated S2 cells to perform mass spectrometry. To clarify this, we will add a detailed method description in the method section.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      We thank for the reviewer’s suggestion. We used anti-HA magnetic beads to perform immunoprecipitation, and single HA-TM9SF4 was used as a negative control.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      We thank for the reviewer’s suggestion, we will add the negative control for BiFC results in the revised version.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      We sincerely thank you for the careful consideration of our statistical results, which were obtained from 3 repeats. We will choose another flow cytometry histogram of tm9sf4 and sac1 to make the data and images more consistent.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      We sincerely thank the reviewer for their careful and thoughtful observations. Given the results we have observed, we will add this to the discussion section in the revised version.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      We thank for the reviewer’s suggestion. We will overexpress Xkr with or without AcD treatment to observe whether the localization or amount of Lact:C2 changes and to re-evaluate the role of Xkr in PS exposure.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      We thank the reviewer for the suggestion, and we believe that the TIRF technique can help us demonstrate the ER-PM signal. Since our college has no TIRF microscope, we will try our best to seek cooperation from other colleges to achieve this experiment.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

      We sincerely thank the reviewer for this careful and thoughtful observation, and we agree that the interpretation in Line 285-288 is overstated. To explain this, we plan to detect the Lact:C2 and MAPPER signals in S2 and xkr<sup>ko</sup> cells with or without AcD to confirm how Xkr regulates PS via ER-PM under apoptotic conditions.

    1. Author response:

      We are pleased that the reviewers found the study conceptually novel and the analytical framework rigorous. In response we have substantially revised the manuscript to clarify methodological details, temper several interpretations, expand discussion of alternative explanations, and include additional analyses using the existing dataset. We have deliberately revised the manuscript so that our conclusions are limited to those directly supported by the data, namely that physiologically identified RVM pain-modulatory neurons exhibit structured dynamics spanning multiple temporal scales. We do not interpret the slow fluctuations as evidence for a specific intrinsic oscillator or for a causal role in physiological state regulation. We have also expanded the rationale for the lightly anaesthetized preparation, emphasizing that it provides both the recording stability required for prolonged single-unit recordings from sparse neurons in the deep RVM and a controlled physiological setting in which the baseline temporal organization of the circuit can be characterized while minimizing ongoing sensory, motor, and behavioral influences.

      Regarding the rationale for the lightly anaesthetized preparation, these experiments take advantage of the well-validated lightly anaesthetized Sprague-Dawley rat in which much of the foundational data concerning physiology and function of RVM neurons was obtained. This “middle-out” strategy [1] has allowed direct connections between the activity and pharmacology of identified RVM neurons and altered nociceptive behavior. This protocol demonstrably spares the essential links between brainstem pain-modulating neurons and nociceptive transmission pathways. Although the focus here was on ongoing activity, precluding the repeated nociceptive testing needed to link neuronal activity to nociceptive threshold, previous work has demonstrated that ongoing activity of OFF and ON-cells is correlated with nociceptive sensitivity [2] and that alterations in OFF- and ON cell firing in response to pharmacological manipulation and in models of persistent pain states, stress, and sickness have behavioral relevance [3–6,6–24]. Further, conclusions from work in lightly anaesthetized rats have repeatedly been found to be congruent with behavioral observations by other groups in awake rats and mice [19,25–36] and with functional imaging evidence in humans [37–40]. The lightly anaesthetized model has thus established a circuit-level explanatory framework for behavioral findings obtained in several species in multiple laboratories.

      A further consideration for the present study is that the lightly anaesthetized preparation allows us to examine the underlying temporal organization of the RVM under controlled conditions, without the additional factors that would necessarily come into play in an awake animal. Dynamics would inevitably be influenced by ongoing sensory input, behavioral priorities, arousal and other internal state changes. These factors would make it difficult to distinguish the intrinsic dynamics of the descending pain-modulatory system from the effects of the animal’s constantly changing experience.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors hypothesized that “RVM neurons operate across multiple temporal scales, integrating fast responses associated with reflex-linked control with slower fluctuations reflecting ongoing network or state-dependent modulation”. The hypothesis was tested with the established ON/OFF-cell model and probabilistic modeling. The study is conceptually interesting and methodologically sophisticated. The findings build toward the conclusion that pain-control circuits operate across multiple timescales.

      Strengths:

      The use of Bayesian regression and Gaussian process modeling to quantify and characterize recovery dynamics and ongoing oscillatory activity.

      The authors show that slow rhythmic activity appears preferentially in ON- and OFF-cells but not in NEUTRAL-cells, suggesting that the oscillations are related to pain-modulatory circuitry rather than being a generic feature of all recorded neurons.

      The observation that some oscillatory activity is coherent with autonomic measures aligns with broader views of the RVM as a hub integrating nociceptive and homeostatic regulation.

      Some pitfalls are appreciated and discussed by the authors, including the functional significance of slow fluctuations, the influence of anesthetics on global brain-state dynamics, the molecular profiles of the studied ON- and OFF-cells, and the heart rate as a covarying signal of RVM neuronal activity.

      Weaknesses:

      A general weakness is that the work is mostly descriptive and relies on anesthetized preparations. Whether the observed rhythms occur in awake animals and are linked to fluctuations in pain behavior needs to be confirmed in future studies.

      The study measures limited autonomic variables. The causal relationship between “slow fluctuations” and “ongoing physiological state” is unclear and overstated, since the data presented appear correlational.

      ON- and OFF-cells in the RVM are identified by their responses correlated with reflexive activity. The significance of the observed oscillations in spontaneous pain conditions is unclear.

      It is uncertain whether the observed rhythms truly reflect intrinsic RVM organization rather than anesthesia-dependent phenomena; the authors appreciated this pitfall, though.

      The Gaussian process analysis suggests predictability and quasi-periodicity, but predictability alone does not necessarily imply a true biological oscillator.

      Conclusion:

      The results support the authors’ hypothesis. The findings provide a compelling conceptual message about the multiscale organization and dynamics of descending pain-control circuits and encourage further studies on the topic.

      We thank Reviewer R1 for their thoughtful and balanced assessment of our work. We are grateful for the reviewer’s positive evaluation of the conceptual framework, the analytical methodology, and the conclusion that the results support the hypothesis that RVM pain-modulatory neurons operate across multiple temporal scales. We agree with the reviewer’s central assessment that the present study is primarily descriptive and that several important questions regarding the origin and functional significance of the slow dynamics remain unresolved. We also appreciate the reviewer’s emphasis on clearly distinguishing observations directly supported by the data from their mechanistic interpretation. Although the manuscript already acknowledged that the present findings do not establish the mechanistic origin of the slow dynamics, we agree that this distinction could be made more explicit. We have therefore revised the manuscript to clarify that the approximately 5-minute fluctuations represent structured, quasi-periodic activity whose underlying origin cannot be determined from the present experiments. Throughout the manuscript we now explicitly acknowledge that these dynamics may arise from interactions between RVM circuitry and broader physiological or network processes, including anaesthesia-related state modulation, autonomic regulation, or other slow network influences. We also emphasise that the relationship between RVM activity and heart rate is correlational and does not establish a causal interaction. Finally, we now discuss more explicitly the complementary roles of controlled lightly anaesthetized and awake preparations. The present preparation was chosen to characterize baseline RVM dynamics under controlled sensory and behavioral conditions, whereas future awake studies will be important for determining how these dynamics are expressed and modulated during ongoing behavior, sensory experience, and chronic pain.

      (1) In the abstract, the ”timescales” are vaguely stated as ”rapid activation”, ”fast recovery dynamics,” and ”slow dynamics”. Quantifying these expressions with approximate ranges (milliseconds, seconds, tens of seconds, minutes, etc.) whenever possible would benefit readers.

      We thank the reviewer for this helpful suggestion. We have revised the Abstract to provide approximate timescales for the different phases of neuronal activity, distinguishing the rapid stimulus-evoked response (sub-second), recovery dynamics (seconds to hundreds of seconds), and ongoing quasi-periodic fluctuations (approximately 5 minutes). We believe these revisions improve the clarity of the Abstract and better convey the central findings of the study.

      (2) The abstract states that ”Effective pain therapies increasingly target neural circuits...” The connection to therapy is not clarified in the manuscript. A brief statement about how temporal dynamics might influence neuromodulation, analgesic interventions, or chronic pain could strengthen translational impact.

      We appreciate this suggestion. We have revised both the Abstract and Discussion to better explain the potential translational relevance of our findings. Rather than making a broad statement regarding pain therapies, we now briefly discuss how understanding the temporal organisation of descending pain-modulatory circuits may ultimately inform the design and timing of neuromodulatory interventions. We also emphasise that these implications remain speculative and require future investigation.

      (3) To address inter-animal and inter-neuron variability. Are the effects consistent across animals? Are all neurons oscillatory? Are the reported timescales driven by a subset of cells?

      We thank the reviewer for raising this important point. We have expanded the Results and Discussion to clarify the degree of variability observed across neurons and animals. In particular, we now emphasise that the slow quasi-periodic dynamics are not uniformly expressed across all neurons, but rather represent a structured population-level phenomenon with variability in predictability and modulation strength between cells, particularly within the OFF-cell population. We also clarify the consistency of the observed timescales across animals and discuss this variability as an important feature of the underlying circuitry rather than evidence for a single homogeneous oscillatory process.

      (4) Discuss what circuit mechanisms generate the oscillations. Are they driven by inputs from the PAG or intrinsic to RVM?

      We agree that the mechanisms underlying the slow temporal dynamics are an important question. We have expanded the Discussion to consider several possible sources of these dynamics, including intrinsic RVM circuitry, descending inputs from higher-order structures, and broader physiological or brain-state fluctuations. We emphasise that the present experiments cannot distinguish between these possibilities and have revised the manuscript to make this limitation more explicit while highlighting it as an important direction for future work.

      Reviewer #2 (Public review):

      Using electrophysiological recordings in a well-characterized animal model of acute pain, and analytical and modeling methods, the authors show that descending pain-modulatory neurons in the rostral ventromedial medulla (RVM) operate across various timescales. They have both rapid multi-phase responses to noxious stimuli that unfold over tens of seconds, with distinct fast and slow recovery dynamics. Additionally, they generate slow quasi-periodic oscillations with approximately 5-minute periods during ongoing activity. These oscillations are statistically predictable and cell-type specific, demonstrating that descending pain control is organized through structured temporal dynamics that encompass immediate stimulus-evoked responses and slower fluctuations associated with physiological state.

      A novel discovery is a 5-minute quasi-periodic oscillation in ongoing ON- and OFF-cell activity. This oscillation, along with its coherence with heart rate, forms the basis for the claim that descending pain circuits exhibit intrinsic multi-timescale organization. However, it’s crucial to demonstrate that this periodicity is independent of external experimental cycles such as methohexital infusion pharmacokinetics, servo-controlled temperature regulation, or slow autonomic feedback loops, all of which operate on similar timescales. For instance, the 300-second period closely matches typical drug infusion cycling and thermoregulatory feedback intervals. Therefore, heart-rate coherence peaks at multiples of this period could equally reflect a shared external driver rather than intrinsic RVM organization. Although the absence of this cyclic structure in Neutral cells argues against this possibility, the authors might want to explicitly discuss this potential confound.

      The findings are important and novel in that they characterize an intriguing structure in the activity of ON and OFF neurons in the RVM. However, in the absence of a causal manipulation causality can only be inferred. That there is no phase-dependence of withdrawal latency argues against a causal role. The author are encouraged to qualify their conclusions (and their title) accordingly. Because anesthesia can affect global dynamics, this might affect the oscillations reported. Without awake validation, it remains uncertain whether these rhythms reflect an intrinsic property or an anesthesia-induced regime. Again, the absence of oscillations in Neutral cells argues against this possibility, but it is still possible that ON/OFF cells are embedded in different circuits that are affected differently by anesthesia.

      We thank Reviewer R2 for their careful and constructive assessment of our work and for recognising the novelty of identifying structured multi-timescale dynamics in physiologically characterised RVM neurons. We particularly appreciate the reviewer’s thoughtful consideration of alternative explanations for the observed low-frequency temporal structure.

      The reviewer raises an important question regarding the extent to which the approximately 5-minute quasi-periodic dynamics reflect processes generated within descending pain-modulatory circuitry versus broader physiological or experimental influences. As discussed in the original manuscript, the present experiments cannot determine the precise mechanistic origin of these dynamics, and we have revised the Discussion to make this distinction more explicit. We now consider possible contributions from autonomic regulation, thermoregulatory processes, anaesthesia-related state modulation, and other slow physiological influences. We also clarify that the NEUTRAL-cell population argues against a uniform global effect acting similarly across all RVM neurons, but cannot exclude systemic influences that preferentially engage ON- and OFF-cell circuitry.

      We have additionally expanded the rationale for the lightly anaesthetized preparation. This preparation was not used solely for technical convenience. Stable single-unit recordings from physiologically identified ON- and OFF-cells are technically challenging because the RVM is a deep brainstem structure and these functional cell classes are relatively sparse; suppression of spontaneous movement therefore permits substantially greater recording stability over the prolonged epochs required here. Importantly, the preparation also provides a controlled physiological setting in which the underlying temporal organization of RVM activity can be examined while reducing the continuously changing sensory, motor, arousal, and behavioral influences that would necessarily contribute to RVM activity in an awake animal. Awake preparations are essential for determining how RVM neurons respond during ongoing behavior and natural sensory experience, but that is a complementary question to the one addressed here: whether physiologically identified RVM neurons exhibit structured temporal dynamics under controlled conditions.

      We have therefore revised the manuscript to present the lightly anaesthetized preparation as both a methodological choice and an important boundary condition on interpretation. We continue to acknowledge that anaesthesia may influence slow network dynamics, and that future awake recordings will be required to determine how the temporal structure identified here is expressed in the behaving animal. We have also revised the title and several sections of the manuscript to ensure that our conclusions consistently reflect the correlational nature of the data and do not imply mechanistic or causal interpretations beyond those directly supported by the experiments.

      (1) Analyses of many of the ON-cells had longer training windows (> 1 sec) compared to those for the NEUTRAL cells. Could this have reduced the ability to fit and validate periodicity for the latter cell type?

      We thank the reviewer for raising this important point. The difference between ON/OFFand NEUTRAL-cell analyses reflects the available recording durations rather than differences in the Gaussian process fitting procedure. All cell classes were fitted using the same GP model and training strategy; however, some NEUTRAL-cell recordings were shorter (960 s versus 1500 s for ON- and OFF-cells), resulting in correspondingly shorter training segments. Whilst the minimum frequency recoverable from a 960 s training segment is 0.00104 Hz, meaning that the 0.0033 Hz frequency observed in ON- and OFF-cells would be recoverable if present in a 960 s recording. We agree that shorter recordings could, in principle, reduce the ability to estimate slow periodic structure. We have therefore clarified this point in the Methods and Discussion. Importantly, the absence of predictable low-frequency dynamics in NEUTRAL-cells is supported not only by GP prediction performance but also by the independent power spectral analysis, the low latent GP variance, and the near-flat phase-normalised reconstructions, suggesting that the difference between cell classes is not solely attributable to recording duration. To address this concern, we will revise the manuscript to repeat the GP analysis after truncating the ON- and OFF-cell recordings to match the duration of the NEUTRAL-cell recordings. We will also include a 960-second-long simulated NEUTRAL-cell recording with periodic structure, to demonstrate that this would be located by our method if present.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript by Ashworth and colleagues, the authors investigate the temporal dynamics of the rostral ventromedial medulla (RVM), a key output node in a major descending pain-modulation circuit. Using data from extrasellar single-unit recordings of RVM ON, OFF, and NEUTRAL cells in lightly anesthetized rats, the authors’ computational modeling yielded two major findings: (1) heat-evoked ON burst and OFF pause, followed by exponential recovery components in10s of seconds; and (2) ON and OFF cells exhibit periodic fluctuations in 5-minute cycles that are statistically predictable.

      Strengths:

      The manuscript’s concept is innovative, offering the first quantitative analysis of multitimescale dynamics in physiologically characterized RVM pain-modulating neurons. This advances a field that has mostly depended on qualitative or single-timescale descriptions. The authors use contemporary Gaussian process and probabilistic models to capture statistically predictable slow dynamics. The study is further strengthened by identifying ON-, OFF-, and NEUTRAL-type cells using well-established criteria grounded in decades of RVM research. The combination of rapid reflex-related responses and slower ongoing rhythms supports a dual-timescale framework, providing a more integrated understanding of how these neurons may regulate reflex activity and state-dependent processes.

      Weaknesses:

      Several limitations are noted. Incomplete characterization of light anesthesia during recording sessions, such as methohexital stability and clear criteria for identifying “lightly anesthetized” states. While the NEUTRAL cell control is helpful, it does not fully address concerns about circuit specificity or systemic confounds. The findings are male-dominant, which may limit their generalizability. The synchrony between ON and OFF cells was suggested but not directly tested. The heart rate coherence with ON, OFF, and NEUTRAL cell activity results is intriguing but does not fully clarify how these neurons influence heart rate, particularly within the “lightly anesthetized” model.

      We thank Reviewer R3 for their thoughtful and constructive assessment of our work. We appreciate the reviewer’s emphasis on providing additional methodological detail and placing the findings within the context and limitations of the experimental preparation. In response, we have substantially expanded the Methods to provide a more complete description of the lightly anaesthetized preparation, the methohexital infusion protocol, physiological monitoring, and the rationale for the ongoing recording paradigm.

      We have also clarified why this preparation was appropriate for the question addressed here. In addition to enabling stable long-duration single-unit recordings from sparse, physiologically identified neurons in the deep RVM, the lightly anesthetized preparation provides a controlled physiological setting in which baseline temporal dynamics can be characterized while minimizing ongoing sensory, motor, and behavioral influences. We nevertheless acknowledge that anaesthesia may alter slow brain-state dynamics, and we now make this limitation more explicit throughout the manuscript. We have also revised the Discussion to more clearly acknowledge the predominantly male sample, the interpretation of the NEUTRAL-cell population as a comparison group rather than a definitive control for systemic effects, the limitations of inferring synchrony from pseudo-population data, and the correlational nature of the heart-rate coherence analysis.

      (1) How does the lightly anesthetized preparation affect evoked and oscillation activity modeling? Given that cell activities can be highly influenced by the state of sedation and the pharmacology of methohexital, detailing how light anesthesia was achieved and determined can help interpret the limitations of the current model. For example, did the methohexital rate adjustments occur during the ongoing activity period used for GP modeling? What specific criteria defined “lightly anesthetized” beyond stable paw withdrawal latency, such as stable respiratory rate, EMG (reflex vigor?), and core temperature? Given RVM activity coupled to autonomic/thermoregulatory circuits, data on these variables should be reported, or their absence should be acknowledged.

      We thank the reviewer for this important comment. We have substantially expanded the Methods and Discussion to describe both the rationale for the lightly anaesthetized preparation and the criteria used to maintain it.

      The preparation offers both technical and conceptual advantages for the present question. Technically, the RVM is a deep brainstem structure and physiologically identified ON- and OFF-cells are relatively sparse. Prolonged extracellular recordings therefore depend on maintaining stable electrode–neuron contact, which is readily disrupted by spontaneous movement. Light methohexital anaesthesia suppresses spontaneous movement while preserving nocifensive withdrawal responses and the canonical physiological response patterns used to identify ON-, OFF-, and NEUTRAL-cells. Conceptually, the aim of the present study was to characterize the baseline temporal organization of identified RVM neurons rather than to determine which sensory, cognitive, or behavioral events drive their activity in an awake animal. An awake preparation would necessarily introduce continuously changing sensory input, motor activity, arousal, behavioral priorities, and other internal-state variables, all of which are known to influence RVM activity. These are important influences in their own right, but for the present question they would make it more difficult to distinguish underlying temporal structure from activity driven by ongoing experience. We therefore view controlled lightly anaesthetized and awake preparations as complementary: the former is useful for identifying foundational circuit dynamics under controlled conditions, whereas the latter will be essential for determining how those dynamics are modified and expressed during natural behavior.

      This preparation has also been extensively used to establish the canonical relationship between ON-/OFF-cell activity and nocifensive responses, pharmacological modulation of the RVM, and top-down control from structures including the hypothalamus and amygdala, with many of these functional relationships subsequently confirmed in awake behavioral experiments. We have added this context to the revised manuscript.

      With respect to physiological monitoring, core temperature was continuously monitored and maintained at 36–37 °C, heart rate was monitored by EKG, and EMG was recorded to monitor withdrawal responses. Light anesthesia was defined functionally by preservation of a stable nocifensive withdrawal response in the absence of spontaneous movement. Respiratory variables were not recorded, and we now acknowledge this explicitly as a limitation. We have also clarified in Methods that the Methohexital rates were not adjusted during the recording windows used for the gaussian process analysis.

      We therefore agree that the findings must be interpreted within the context of the lightly anaesthetized preparation, but we do not view awake recordings as a direct substitute for the present experiment. Rather, awake studies provide the important next step of determining how the structured dynamics identified under controlled conditions are modulated by sensory experience, behavioral state, and ongoing cognition.

      (2) It is unclear how the absence of slow oscillations in NEUTRAL cells can be used as an internal control for anesthesia and systemic drift. It is unlikely that NEUTRAL cells are identified in every single-cell recording session for them to be used as a consistent internal control. Also, as the authors suggested that the shared modulatory inputs to ON/OFF cells explain the coordinating mechanism for ON/OFF rhythmicity, the lack of rhythmicity or coherence in majority of the NEUTRAL cells may indicate that they do not receive the same modulatory inputs as ON/OFF cells. Would this make NEUTRAL cells insensitive to systemic changes throughout the recording sessions? Do rhythmic vs. non-rhythmic cells differ in location within the RVM?

      We appreciate the reviewer’s important distinction. We agree that NEUTRAL-cells should not be considered a definitive internal control for anaesthesia or systemic physiological drift. NEUTRAL-, ON-, and OFF-cells were not necessarily recorded simultaneously within the same session, and the functional classes may differ in the systemic or modulatory inputs they receive. Our intended inference is therefore narrower: the absence of comparable low-frequency temporal structure in most NEUTRAL-cells argues against a uniform global process that imposes the same temporal pattern on all RVM neurons. It does not exclude anaesthesia-related, autonomic, thermoregulatory, or other systemic processes that preferentially influence ON- and OFF-cell circuitry. We have revised the manuscript throughout to make this distinction explicit and now refer to NEUTRAL-cells as an informative comparison population rather than as a definitive control for systemic influences.

      Indeed, as the reviewer suggests, differential sensitivity to common modulatory inputs could itself contribute to the distinction between ON/OFF- and NEUTRAL-cell dynamics. This interpretation is also compatible with the observation that a subset of NEUTRAL-cells shows low-frequency coherence with heart rate despite lacking the structured approximately 5-minute temporal dynamics observed in the ON/OFF populations.

      We additionally examined the reconstructed recording locations and found no obvious anatomical segregation between neurons showing stronger versus weaker low-frequency structure within the sampled RVM region. We now state this in the revised manuscript. We appreciate the reviewer’s point that there may be locational differences between RVM rhythmic and non-rhythmic cells, which should be addressed in future work; however, determining this would require a substantially larger sample size, for example with multichannel probe recording, for a valid analysis.

      (3) It is important to acknowledge that findings are effectively male-only (77 M and 6 F). Although a recent publication demonstrated that RVM ON and OFF cell activities do not differ substantially on an individual level between male and female rats, sex differences in RVM population dynamics remain unexplored. The current finding may not be generalizable to females.

      We thank the reviewer for highlighting this important limitation. We now explicitly acknowledge in the Discussion that the present dataset is predominantly male (77 males, 6 females) and therefore does not permit meaningful assessment of sex differences in population dynamics. Although previous studies suggest that individual ON- and OFF-cell responses are broadly comparable between sexes, the generalisability of the present findings to female animals remains unknown and should be addressed in future work.

      (4) It was suggested that strong synchrony exists within each functional population (e.g., ON and OFF cells). However, phase-relationship or coherence analyses were lacking. Since there were recoding sessions with > 2 cells/animals, were there enough recordings that contain simultaneous ON/OFF pairs to allow for these analyses?

      We thank the reviewer for this helpful suggestion. We agree that direct analyses of synchrony between simultaneously recorded neurons would provide valuable additional information. However, the number of simultaneous recordings containing identifiable ON/OFF-cell pairs was insufficient to support a robust phase or coherence analysis. We have therefore revised the Discussion to avoid implying that synchrony has been directly demonstrated and instead describe the results as evidence for consistent low-frequency temporal structure across recordings. We also identify direct analysis of synchrony in larger simultaneously recorded neuronal populations as an important direction for future work.

      (5) It was intriguing that the ON-cell population’s ongoing activity shows a predictive structure, while the OFF-cell population does not (Figure 5). However, this interesting asymmetry in ongoing activity between two cell classes was not adequately explained in the discussion. For example, since shared modulatory inputs were proposed as the coordinating mechanism for ON/OFF rhythmicity, how may this difference in ON and OFF rhythm predictivity occur?

      We appreciate the reviewer drawing attention to this interesting observation. We have expanded the Discussion to consider possible explanations for the greater predictability observed in ON-cells relative to OFF-cells. Although both populations exhibited similar dominant timescales, ON-cells displayed larger latent GP variance and more consistent predictive performance, whereas OFF-cells exhibited greater heterogeneity across recordings. We now discuss several possible explanations for this asymmetry, including differences in intrinsic cellular properties, network coupling, or modulation amplitude, while emphasising that the present data do not allow these possibilities to be distinguished.

      (6) The relationship between RVM activity oscillations and cardiac rhythms appears to be covariate but may not support the ”physiologically meaningful” claim with the current analysis. Additional discussion could help clarify the findings of a) how the RVM oscillation period of 300s relates to the heart rate peak/oscillation period of 600s (Figure 6d) and b) how NEUTRAL cells show heart rate coherence but lack rhythmicity.

      We thank the reviewer for this thoughtful comment. We have revised the Discussion to more carefully interpret the heart-rate coherence analysis. In particular, we now emphasise that the observed coherence demonstrates shared low-frequency temporal structure but does not establish a causal relationship between RVM activity and cardiac dynamics. We also discuss the relationship between the approximately 300-s RVM timescale and the broader low-frequency components observed in the heart-rate spectrum, noting the limited frequency resolution available at these timescales. Finally, we expand our discussion of the NEUTRAL-cell results, clarifying that significant coherence in NEUTRAL-cells despite the absence of comparable structured low-frequency firing dynamics is consistent with shared physiological influences acting on multiple cell classes without implying that the slow temporal structure originates within NEUTRAL-cells.

      References

      (1) Noble, D. The Music of Life: Biology beyond the Genome (Oxford University Press, 2006).

      (2) Heinricher, M. M., Barbaro, N. M. & Fields, H. L. Putative Nociceptive Modulating Neurons in the Rostral Ventromedial Medulla of the Rat: Firing of On- and Off-Cells Is Related to Nociceptive Responsiveness. Somatosensory & motor research 6, 427–39 (1989).

      (3) Barbaro, N. M., Heinricher, M. M. & Fields, H. L. Putative Nociceptive Modulatory Neurons in the Rostral Ventromedial Medulla of the Rat Display Highly Correlated Firing Patterns. Somatosensory & Motor Research 6, 413–425 (1989).

      (4) Heinricher, M. M., Haws, C. M. & Fields, H. L. Evidence for GABA-mediated control of putative nociceptive modulating neurons in the rostral ventromedial medulla: Iontophoresis of bicuculline eliminates the off-cell pause. Somatosensory & Motor Research 8, 215–225 (1991).

      (5) Heinricher, M. M. & Kaplan, H. J. GABA-mediated inhibition in rostral ventromedial medulla: Role in nociceptive modulation in the lightly anesthetized rat. Pain 47, 105–113 (1991).

      (6) Heinricher, M. M. & Tortorici, V. Interference with GABA transmission in the rostral ventromedial medulla: Disinhibition of off-cells as a central mechanism in nociceptive modulation. Neuroscience 63, 533–546 (1994).

      (7) Heinricher, M. M., McGaraughty, S. & Grandy, D. K. Circuitry Underlying AntiOpioid Actions of Orphanin FQ in the Rostral Ventromedial Medulla. Journal of Neurophysiology 78, 3351– 3358 (1997).

      (8) Heinricher, M. M., McGaraughty, S. & Farr, D. A. The role of excitatory amino acid transmission within the rostral ventromedial medulla in the antinociceptive actions of systemically administered morphine. Pain 81, 57–65 (1999).

      (9) Heinricher, M. M., McGaraughty, S. & Tortorici, V. Circuitry Underlying Antiopioid Actions of Cholecystokinin Within the Rostral Ventromedial Medulla. Journal of Neurophysiology 85, 280–286 (2001).

      (10) Heinricher, M. M., Schouten, J. C. & Jobst, E. E. Activation of brainstem N-methyl-daspartate receptors is required for the analgesic actions of morphine given systemically. Pain 92, 129–138 (2001).

      (11) McGaraughty, S. & Heinricher, M. M. Microinjection of morphine into various amygdaloid nuclei differentially affects nociceptive responsiveness and RVM neuronal activity. Pain 96, 153–162 (2002).

      (12) Heinricher, M. M. & Neubert, M. J. Neural Basis for the Hyperalgesic Action of Cholecystokinin in the Rostral Ventromedial Medulla. Journal of Neurophysiology 92, 1982–1989 (2004).

      (13) Heinricher, M. M., Martenson, M. E. & Neubert, M. J. Prostaglandin E2 in the midbrain periaqueductal gray produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Pain 110, 419–426 (2004).

      (14) Heinricher, M. M., Neubert, M. J., Martenson, M. E. & Gonc¸alves, L. Prostaglandin E2 in the medial preoptic area produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Neuroscience 128, 389–398 (2004).

      (15) Kincaid, W., Neubert, M. J., Xu, M., Kim, C. J. & Heinricher, M. M. Role for Medullary Pain Facilitating Neurons in Secondary Thermal Hyperalgesia. Journal of Neurophysiology 95, 33–41 (2006).

      (16) Ortiz, J., Heinricher, M. & Selden, N. Noradrenergic agonist administration into the central nucleus of the amygdala increases the tail-flick latency in lightly anesthetized rats. Neuroscience 148, 737–743 (2007).

      (17) Xu, M., Kim, C. J., Neubert, M. J. & Heinricher, M. M. NMDA receptor-mediated activation of medullary pro-nociceptive neurons is required for secondary thermal hyperalgesia. PAIN 127, 253 (2007).

      (18) Ortiz, J. P., Close, L. N., Heinricher, M. M. & Selden, N. R. α2-Noradrenergic antagonist administration into the central nucleus of the amygdala blocks stress-induced hypoalgesia in awake behaving rats. Neuroscience 157, 223–228 (2008).

      (19) Edelmayer, R. M. et al. Medullary pain facilitating neurons mediate allodynia in headache-related pain. Annals of Neurology 65, 184–193 (2009).

      (20) Martenson, M. E., Cetas, J. S. & Heinricher, M. M. A possible neural basis for stress-induced hyperalgesia. Pain 142, 236–244 (2009).

      (21) Heinricher, M. M., Maire, J. J., Lee, D., Nalwalk, J. W. & Hough, L. B. Physiological Basis for Inhibition of Morphine and Improgan Antinociception by CC12, a P450 Epoxygenase Inhibitor. Journal of Neurophysiology 104, 3222–3230 (2010).

      (22) Heinricher, M. M., Martenson, M. E., Nalwalk, J. W. & Hough, L. B. Neural basis for improgan antinociception. Neuroscience 169, 1414–1420 (2010).

      (23) McGaraughty, S., Farr, D. A. & Heinricher, M. M. Lesions of the periaqueductal gray disrupt input to the rostral ventromedial medulla following microinjections of morphine into the medial or basolateral nuclei of the amygdala. Brain Research 1009, 223–227 (2004).

      (24) Rogness, V. M. et al. Descending Facilitation of Nociceptive Transmission From the Rostral Ventromedial Medulla Contributes to Hyperalgesia in Mice with Sickle Cell Disease. Neuroscience 526, 1–12 (2023).

      (25) Smith, D. J. et al. Dose-Dependent Pain-Facilitatory and -Inhibitory Actions of Neurotensin Are Revealed by SR 48692, a Nonpeptide Neurotensin Antagonist: Influence on the Antinociceptive Effect of Morphine1,2. The Journal of Pharmacology and Experimental Therapeutics 282, 899–908 (1997).

      (26) Hurley, R. W. & Hammond, D. L. The Analgesic Effects of Supraspinal µ and δ Opioid Receptor Agonists Are Potentiated during Persistent Inflammation. Journal of Neuroscience 20, 1249–1259 (2000).

      (27) Kovelowski, C. J. et al. Supraspinal cholecystokinin may drive tonic descending facilitation mechanisms to maintain neuropathic pain in the rat. Pain 87, 265–273 (2000).

      (28) Porreca, F. et al. Inhibition of Neuropathic Pain by Selective Ablation of Brainstem Medullary Cells Expressing the µ-Opioid Receptor. Journal of Neuroscience 21, 5281–5288 (2001).

      (29) Zhang, Y. et al. Identifying local and descending inputs for primary sensory neurons. Journal of Clinical Investigation 125, 3782–3795 (2015).

      (30) Franc¸ois, A. et al. A Brainstem-Spinal Cord Inhibitory Circuit for Mechanical Pain Modulation by GABA and Enkephalins. Neuron 93, 822–839.e6 (2017). URL https://www.ncbi.nlm. nih.gov/pmc/articles/PMC7354674/.

      (31) Kim, J.-H. et al. Yin-and-yang bifurcation of opioidergic circuits for descending analgesia at the midbrain of the mouse. Proceedings of the National Academy of Sciences 115, 11078–11083 (2018).

      (32) Nguyen, E. et al. Medullary kappa-opioid receptor neurons inhibit pain and itch through a descending circuit. Brain 145, 2586–2601 (2022).

      (33) Jiao, Y. et al. Molecular identification of bulbospinal ON neurons by GPER, which drives pain and morphine tolerance. The Journal of Clinical Investigation 133, e154588 (2023).

      (34) Nguyen, E., Grajales-Reyes, J. G., Gereau, R. W. & Ross, S. E. Cell type-specific dissection of sensory pathways involved in descending modulation. Trends in Neurosciences 46, 539–550 (2023).

      (35) Fatt, M. P. et al. Morphine-responsive neurons that regulate mechanical antinociception. Science 385, eado6593 (2024).

      (36) Wang, Q. et al. Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain. Nature 1–10 (2026).

      (37) Brooks, J. C., Davies, W.-E. & Pickering, A. E. Resolving the Brainstem Contributions to Attentional Analgesia. The Journal of Neuroscience 37, 2279–2291 (2017).

      (38) Mills, E. P. et al. Brainstem Pain-Control Circuitry Connectivity in Chronic Neuropathic Pain. Journal of Neuroscience 38, 465–473 (2018).

      (39) Mills, E. P., Keay, K. A. & Henderson, L. A. Brainstem Pain-Modulation Circuitry and Its Plasticity in Neuropathic Pain: Insights From Human Brain Imaging Investigations. Frontiers in Pain Research (Lausanne, Switzerland) 2, 705345 (2021).

      (40) Oliva, V., Hartley-Davies, R., Moran, R., Pickering, A. E. & Brooks, J. C. Simultaneous brain, brainstem, and spinal cord pharmacological-fMRI reveals involvement of an endogenous opioid network in attentional analgesia. eLife 11, e71877 (2022).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Liao et al. present SCOPE (Spatial reConstruction via Oligonucleotide Proximity Encoding), a method for reconstructing spatial organization from diffusion-defined DNA barcode interactions without the use of optical imaging. In SCOPE, hydrogel beads bearing unique DNA barcodes contain both "sender" and "receiver" oligonucleotides. Upon enzymatic release, sender oligos diffuse locally and hybridize to receiver oligos on neighboring beads, forming chimeric molecules that encode spatial proximity. Sequencing these products yields an interaction matrix, which is then used to reconstruct a spatial coordinate map.

      The authors demonstrate reconstruction of synthetic two-dimensional shapes, a large multicolor Snellen eye chart, and the interior surface of three-dimensional molds. The work expands the conceptual and experimental landscape of optics-free spatial sequencing.

      Thank you for this accurate summary of the work.

      Strengths:

      SCOPE employs bidirectional sender and receiver oligonucleotides on every bead, rather than using asymmetric transmitter-receiver architectures found in other diffusion-based methods. The symmetric design may improve detection sensitivity and reconstruction strategies, and represents a meaningful variation on optics-free spatial encoding.

      A notable strength of this study is the physical scale achieved. The authors reconstruct a Snellen chart spanning approximately 704 mm² and demonstrate molded 3D structures on the order of 75-100 mm³. Although some larger-scale warping is evident, and is discussed as potentially due to non-uniform diffusion, the relative local positioning across these large areas appears impressively accurate.

      The authors extend reconstruction beyond two-dimensional arrays to three-dimensional molded surfaces. This demonstrates that the assay and the computational methods for interpreting proximity graphs can support nonplanar spatial relationships, expanding the scope of optics-free spatial inference.

      Thank you for highlighting these strengths of SCOPE.

      Weaknesses:

      Although the method is discussed in the context of spatial genomics and potential tissue applications, it is currently demonstrated only on engineered two-dimensional bead arrays and three-dimensional shapes fabricated in molds. It remains unclear how SCOPE would perform in heterogeneous biological environments, where diffusion may exhibit additional non-uniformities. A biological proof-of-concept, even limited in scope, would help define the method's strengths and limitations more clearly.

      We concur with the reviewer that a biological proof-of-concept is a key next step, and that diffusion will be more heterogeneous in this more complex environment. To this end, we are actively working to further develop SCOPE for use in tissue sections, with the goal of capturing transcriptomes, accessible chromatin, and genomes. As part of this work, we also hope to systematically explore a range of tissue permeabilization and tissue clearing approaches to mitigate the impact of heterogeneity on performance.

      The reconstruction of three-dimensional structures lacks strong sampling from volume interiors. This is speculated to be due to several possible factors; however, this limitation constrains the method to reconstruction of volume surfaces rather than comprehensive three-dimensional profiling.

      Thank you for highlighting this important limitation. The 3D reconstructions are indeed constrained by undersampling of volume interiors. We anticipate that this might be addressed via relatively minor adjustments to the protocol, e.g. using light- or base-labile linkers to trigger oligo release, with the expectation that this will improve reaction consistency throughout the volume. However, even if we are unable to resolve this issue, we note that surface-resolved reconstructions may be useful for some goals, e.g. embedding a bead-packed gel within a tissue lumen, such as the gut. This could enable surface beads to capture RNA transcripts from adjacent cells, while bead–bead associations serve to define the surface topology.

      The reconstruction workflow involves multiple preprocessing steps and embedding choices. While these appear to work well for synthetic shapes with known geometry, it is less clear how parameter choices would be made in contexts where ground truth is unknown. Clarifying how reconstruction robustness is assessed without prior knowledge of spatial structure would help readers understand how the method could be practically deployed, particularly in more heterogeneous tissue contexts.

      Thank you for the opportunity to clarify. The computational pipeline used for 2D SCOPE reconstruction is designed to operate on a standardized input format and can be applied to arbitrary datasets without prior knowledge of spatial structure. For example, as shown in Figure 3, both the circle and “swoosh” geometries were reconstructed using the same algorithm and identical initial parameters. While certain hyperparameters are pre-specified (e.g. the number of k-nearest neighbours used to compute the pairwise distance matrix for UMAP), these are fixed across datasets. Other parameters, such as UMAP’s “min_dist,” are selected via an automated heuristic grid search that proceeds without user intervention. The agreement with ground truth in these controlled settings, together with the reproducibility of stochastic reconstructions (see Figure 3E-F), supports the robustness of the approach.

      Importantly, there was one exception. Reconstruction of the Snellen eye chart dataset required a manual step, involving an initial 3D UMAP embedding followed by a 2D projection to “flatten” the result. We suspect this reflects radial non-uniformities in sender/receiver oligo diffusion at larger spatial scales. Addressing such confounders algorithmically by explicitly modelling diffusion heterogeneity represents an important area for future work, with the goal of entirely eliminating the need for manual intervention.

      Finally, we note that these benchmark shapes represent somewhat contrived examples, and the geometries encountered in practice may often be much less complex. For example, in conventional spatial genomics, the geometry consists of a bead monolayer forming a flat, regular surface on a rectangular slide of known dimensions. Regardless of the tissue architecture overlaid on this surface, the reconstruction problem is defined by the bead monolayer itself, inferred through sender-receiver interactions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It would be helpful to further clarify the limitations in interior sampling of three-dimensional structures by providing a more explicit comparison with Qian and Weinstein's volumetric DNA microscopy (UMI-UEI) approach. In particular, do the authors anticipate that the current limitation in SCOPE's interior sampling can be mitigated through experimental optimization, or might this represent an inherent challenge associated with hydrogel-based bead scaffolds relative to substrate-free approaches? A more detailed discussion of this point would help readers understand whether the observed volumetric constraint is technical and potentially solvable, or structural to the platform design.

      We thank the reviewer for this suggestion and agree that a more explicit comparison to volumetric DNA microscopy helps clarify the origin of the current limitation in interior sampling. Based on our experiments to date, we view this constraint as primarily technical and, in principle, addressable.

      A key distinction between SCOPE and volumetric DNA microscopy(Qian and Weinstein, 2025), using the UMI– UEI framework, lies in the recovery of recorded molecules from the interior of the sample. In volumetric DNA microscopy, the hydrogel-embedded specimen can be fully digested and treated with Proteinase K, enabling efficient liberation and recovery of molecules throughout the volume for downstream sequencing(Qian et al., 2026). In contrast, in the current implementation of SCOPE, we do not dissolve the polyacrylamide hydrogel, and recovery therefore relies largely on diffusion of chimeric molecules out of the scaffold. This likely biases against molecules generated in the interior and leads to reduced sampling of internal regions. This interpretation is supported by a control experiment in which barcoded beads were allowed to settle in solution at the bottom of a tube in the absence of a polymerized hydrogel scaffold. In this setting, 3D UMAP reconstruction yielded a solid, non-hollow structure consistent with the expected conical geometry of the tube bottom, indicating that SCOPE is capable of recovering volumetric structure when recovery is not diffusion-limited. Taken together, these observations suggest that the apparent “hollowing” in current 3D reconstructions reflects a limitation in molecule recovery from hydrogel scaffolds, rather than an inherent constraint of the SCOPE framework itself.

      We are currently exploring potential solutions, including the use of reducible crosslinkers to enable hydrogel dissolution and/or mechanical shearing of the gel. If these experiments are successful, we would plan to include the results in revisions to the manuscript, together with an appropriately edited version of the paragraph above. If they are unsuccessful and major experimental effort is going to be required to address this issue, we would likely move forward with textual changes only, incorporating the points made in the paragraph above into the discussion.

      References

      Qian N, Li J, Yasser R, Yu M, Weinstein JA. 2026. Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions. Nat Protoc. doi:10.1038/s41596-025-01329-3

      Qian N, Weinstein JA. 2025. Spatial transcriptomic imaging of an intact organism using volumetric DNA microscopy. Nat Biotechnol 1–11.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study by Akhtar et al. aims to investigate the link between systemic metabolism and respiratory demands, and how sleep and the circadian clock regulate metabolic states and respiratory dynamics. The authors leverage genetic mutants that are defective in sleep and circadian behavior in combination with indirect respirometry and steady-state LC-MS-based metabolomics to address this question in the Drosophila model.

      First, the authors performed respirometry (on groups of 25 flies) to measure oxygen consumption (VO2) and carbon dioxide production (VCO2) to calculate the respiratory quotient (RQ) across the 24-hour day (12h:12h light-dark cycle) and assess metabolic fuel utilization. They observed that among all the genotypes tested, wild type (WT) flies and per0 flies in LD and WT flies in DD exhibit RQ >1. They concluded the >1 RQ is consistent with active lipogenesis. In contrast, the short-sleep mutants fumin (fmn) and sleepless (sss) showed significantly different RQ; the fmn exhibits a slight reduction in RQ values, suggesting increased reliance on carbohydrate metabolism, while sss exhibits even lower RQ (0.94), consistent with a shift toward lipid and protein catabolism.

      The authors then proceeded to bin these measurements in 12-hour partitions, ZT0-12 and ZT12-24, to assess diurnal differences in average values of VO2, VCO2, and RQ. They observed significant day-night differences in metabolic rates in WT-LD flies, with higher rates during the day. The diurnal differences remain in the short-sleep mutants, but the overall metabolic rates are higher. WT-DD flies exhibit the lowest respiratory activity, although the day-night differences remain in free-running conditions. Finally, per01 mutants exhibit no significant change in day-night respiratory rates, suggesting that a functional circadian clock is necessary for diurnal differences in metabolic rates.

      They then performed finer-resolution 24-hour rhythmic analysis (RAIN and JTK) to determine if VO2, VCO2, and RQ exhibit 24-hour rhythmic and if there are genotypespecific differences. Based on their criteria, VCO2 is rhythmic in all conditions tested, while VO2 is rhythmic in all conditions except in fmn-LD. Finally, RQ is rhythmic in all 3 mutants but not in WT-LD and WT-DD. Peak phases for the rhythms were deduced using JTK lag values.

      The authors proceeded to leverage a previously published steady-state metabolite dataset to investigate the potential association of RQ with metabolite profiles. Spearman correlation was performed to identify metabolites that exhibit coupling to respiratory output. Positive and negative lag analysis were subsequently performed to further characterize these associations based on the timing of the metabolite peak changes relative to RQ fluctuations. The authors suggest that a positive lag indicates that metabolite changes occur after shifts in RQ, and a negative lag signifies that metabolite changes precede RQ changes. To visualize metabolic pathways that exhibit these temporal relationships, a clustered heatmap and enrichment analysis were performed. Through these analyses, they concluded that both sleep and circadian systems are essential for aligning metabolic substrate selection with energy demands, and different metabolic pathways are mis regulated in the different mutants with sleep and circadian defects.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strength:

      The research questions this study explores are significant, given that metabolism and respiratory demand are central to animal biology. The experimental methods used, including the well-characterized fly genetic mutants, the newly developed method for indirect calorimetry measurements, and LC-MS-based metabolomics, are all appropriate. This study provides insights into the impact of sleep and circadian rhythm disruption on metabolism and respiratory demand and serves as a foundation for future mechanistic investigations.

      We thank the reviewer for the positive comments.

      Weaknesses:

      There are some conceptual flaws that the authors need to address regarding circadian biology, and some of the conclusions can be better supported by additional analysis to provide a stronger foundation for future functional investigation.

      At times, the methods, especially the statistical analysis, are not well articulated; they need to be better explained.

      Thank you for this suggestion and have revised and we expanded the Methods and figure legends to improve transparency and reproducibility.

      Specifically, we have:

      (i) Strengthened the rhythmicity description by specifying that rhythmicity was assessed in Nitecap using the RAIN algorithm with FDR-adjusted p-values (significant p ≤ 0.05; trending 0.05–0.1), and that period and peak phase were estimated with JTK_CYCLE (JTK lag = peak phase), with per-genotype period, phase, and p-values reported in Table 1;

      (ii) Clarified sample size and replication by adding these lines to the methods section and indicating sample sizes in figure legends. Figure legends now report n (chambers) and SEM.

      “Each genotype measurement represents an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). The chamber was treated as the experimental unit for all analyses.”

      In addition, we have expanded the description of metabolomics-respirometry correlation analyses to include dataset structure, time-matching across ZT, normalization steps, the use of Spearman correlations, and interpretation of lagged associations.

      (iii) We have added these lines to the methods section:

      “To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite time points (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Reviewer #2 (Public review):

      This is an innovative and technically strong study that integrates dual-gas respirometry with LC-MS metabolomics to examine how sleep and circadian disruption shape metabolism in Drosophila. The combination of continuous O<sub>2</sub>/CO<sub>2</sub> measurements with high-temporal-resolution metabolite profiling is novel and provides fresh insight into how wild-type flies maintain anticipatory fuel alignment, while mutants shift to reactive or misaligned metabolism. The use of lag-shift correlation analysis is particularly clever, as it highlights temporal coordination rather than static associations. Together, the findings advance our understanding of how circadian clocks and sleep contribute to metabolic efficiency and redox balance.

      We thank the reviewer for the positive comments.

      However, there are several areas where the manuscript could be strengthened.

      The authors should acknowledge that their findings may be gene specific. Because sleep deprivation was not performed, it remains uncertain whether the observed metabolic shifts generalize to sleep loss broadly or are restricted to the fmn and sss mutants. This concern also connects to the finding of metabolic misalignment under constant darkness despite an intact clock.

      We agree that our findings should be framed as genotype- and condition-specific. The phenotypes we report arise from chronic, genetically encoded sleep loss (fmn, sss) and clock loss (per01); because acute sleep deprivation was not performed, we do not claim these effects generalize to sleep loss broadly. This also bears on the reviewer's point about constant darkness: the metabolic misalignment we observe in WT-DD occurs despite an intact clock, and we therefore interpret it as a consequence of removing external light-dark cues under our conditions. We have scoped the claims accordingly in the Abstract, Results, and Discussion (subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness, DD”).

      The text now reads as follows:

      “We restrict our conclusions to the genotypes and conditions tested (fmn, sss, and per<sup>01</sup>), and we do not generalize these effects to acute sleep deprivation because sleep deprivation was not performed in this study. Accordingly, the ‘metabolic misalignment’ observed in constant darkness (DD) likely results from a decrease in synchrony due to the removal of external light:dark cues.”

      The conclusion that external entrainment is essential for maintaining energy homeostasis in flies may not translate to mammals. It would help to reference supporting data for the finding and discuss differences across species. Ideally, complementary circadian (lightdark cycle disruption) or sleep deprivation (for several hours) experiments, or citation of comparable studies, would strengthen the generality of the findings.

      Thank you. We have tempered the interpretation and expanded both the discussion and its citations. We now (i) avoid stating that external entrainment is universally “essential” for energy homeostasis, (ii) explicitly discuss fly-mammal differences (sleep architecture, thermoregulation, feeding control, and entrainment mechanisms), and (iii) anchor the translational comparison to the mammalian circadian-misalignment and sleep-loss literature already integrated in our Discussion (refs [3, 43-46]), noting that establishing cross-species generality will require additional paradigms (constant light, acute sleep deprivation).

      The text now reads as follows:

      “These phenotypes parallel mammalian systems, where sleep loss and circadian misalignment are linked to elevated basal metabolic rate, a shift toward carbohydrate oxidation and lipid/protein catabolism, and blunted, phase-shifted respiratory oscillations [3, 43-46]; physiological differences between flies and mammals nonetheless caution against direct mechanistic extrapolation. In constant darkness, our DD data show that endogenous free-running regulation persists but that removing external light-dark cues degrades temporal coordination between respiration and metabolism; however, we acknowledge that the coupling may be different in mammals.”

      Figures 1-4 are straightforward and clear, but when the manuscript transitions to the metabolite-respiration correlations, there is little description of the metabolomics methods or datasets, which should be clarified.

      Thank you for noting this. We agree that the transition to the metabolite–respiration correlation analyses required clearer description of the metabolomics datasets and processing. We have revised the Methods and the corresponding Results text to briefly summarize the metabolomics dataset parameters and workflow, including how metabolomics and respirometry measurements were time-matched across ZT, the normalization procedures applied prior to analysis, the use of Spearman rank correlations, and how we interpret lagged relationships between metabolite abundance and respiratory outputs.

      The text now reads as follows:

      Methods:

      “Metabolomics-respirometry integration and lag analysis

      To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite timepoints (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Results:

      “Temporal Profiling of Respiratory Quotient in Wild-Type Flies Under Light-Dark Conditions

      RQ values corresponding to each 2-hour Zeitgeber Time (ZT) point were extracted from the 5-minute-binned dataset. Building on this alignment, we explored temporal relationships by systematically shifting the RQ time series by −120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes relative to the metabolite dataset. The continuous respirometry time series showed an oscillatory day-night pattern in RQ; metabolomics was then used to relate time-matched and lagged metabolite dynamics to RQ patterns (Figure 4).”

      The Discussion is at times repetitive and could be tightened, with the main message (i.e., wild-type flies align metabolism in advance, while mutants do not) kept front and center.

      Thank you for this helpful suggestion. We have revised the Discussion to reduce repetition and improve focus by keeping the central takeaway explicit throughout, and by consolidating overlapping paragraphs into a more streamlined narrative.

      We added this revision at the start of the Discussion, in the opening subsection “Temporal Misalignment Alters Fuel Utilization and Respiratory Rhythms.”

      The Discussion now reads as follows:

      “Across the manuscript, the central takeaway is that wild-type flies under LD exhibit anticipatory alignment of fuel selection with time of day, whereas short-sleep mutants (fmn, sss) and clock-disrupted flies (per01) show reactive or misaligned metabolism under our conditions. We therefore focus the Discussion on loss of temporal coordination between respiratory output and pathway-level metabolism, rather than reiterating rate changes alone.”

      Terms such as "anticipatory" and "reactive" should be defined early and used consistently throughout.

      Thank you for this suggestion. We agree and have revised the manuscript to define these terms early (at first use) and apply them consistently throughout. We added this definition in two places:

      (i) In the Results, at the start of the metabolomics-respirometry integration section where we first introduce the lag analysis, and

      (ii) In the Methods, within the paragraph describing the lag analysis workflow, using identical wording.

      The text now reads as follows:

      “We define ‘anticipatory’ as metabolite changes that precede the associated respiratory shift (negative lag) and ‘reactive’ as changes that follow or coincide with the respiratory shift (positive lag), and we use these terms consistently throughout.”

      Overall, this is a strong and novel contribution. With clarification of scope, refinement of presentation, and a more focused Discussion, the paper will make a significant impact.

      We again thank the reviewer for the positive comments.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate how sleep loss and circadian disruption affect whole-organism metabolism in Drosophila melanogaster. They used chamber-based flow-through respirometry to measure oxygen consumption and carbon dioxide production in wild-type flies and in mutants with impaired sleep or circadian function. These measurements were then integrated with a previously published metabolomics dataset to explore how respiratory dynamics align with metabolic pathways. The central claim is that wild-type flies display anticipatory coordination of metabolic processes with circadian time, while mutants exhibit reactive shifts in substrate use, redox imbalance, and signs of mitochondrial stress.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strengths:

      The study has several strengths. Continuous high-resolution respirometry in flies is challenging, and its application across multiple genotypes provides good comparative insight. The conceptual framework distinguishing anticipatory from reactive metabolic regulation is interesting. The translational framing helps place the work in a broader context of sleep, circadian biology, and metabolic health.

      We thank the reviewer for the positive comments.

      Weaknesses:

      At the same time, the evidence supporting the conclusions is somewhat limited. The metabolomics data were not newly generated but repurposed from prior work, reducing novelty.

      Thank you for raising this point. We now make the provenance of the metabolomics dataset explicit in the manuscript. Importantly, the current study uses this dataset in a new analytical context by integrating it with continuous VCO<sub>2</sub>/VO<sub>2</sub> respirometry through timematched and lag-aware analyses. This approach allows us to evaluate dynamic relationships between respiratory output and metabolite profiles across circadian time, which was not addressed in the original metabolomics study. We have clarified this point in the Introduction and Methods.

      The text now reads as follows:

      In the Introduction:

      “To provide a more comprehensive perspective on metabolic regulation, we complemented newly generated respiratory measurements with steady-state metabolomic profiling using liquid chromatography-mass spectrometry (LC-MS) data previously published from our group [27]. This integrative framework enabled timematched and lag-aware analysis of respiratory output and metabolite profiles across Zeitgeber time in the LD cycle….”

      In the Methods:

      “The metabolomics dataset analyzed in this study was previously published and is publicly available, as described in detail in [27, 31]. In the present study, these data were integrated with respirometry measurements to assess temporal relationships between metabolite abundance and respiratory output.”

      The biological replication in the respirometry assays is low, with only a small number of chambers per genotype.

      Thank you for highlighting this concern. We suggest that this is a lack of clarity in our initial description of the design and that the replication structure should be stated more explicitly. We have revised the Methods and all relevant figure legends to clearly report biological replication using the chamber as the experimental unit, including n (number of chambers) per genotype and the associated error structure. We also clarify sampling depth by stating that each genotype measurement reflects an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). This information is now reported consistently to make the unit of analysis transparent.

      We added this clarification in the Methods under “Respirometry Setup” (where chamber loading and experimental design are described) and ensured that each relevant figure legend explicitly reports n (chambers) and SEM.

      The text now reads as follows:

      “Each genotype measurement represents an average of ~300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), with the chamber as the experimental unit; n (chambers) and SEM are reported in each figure legend.”

      Importantly, respiratory parameters in flies are strongly influenced by locomotor activity, yet no direct measurements of activity were included, making it difficult to separate intrinsic metabolic changes from behavioral differences in mutants.

      A detailed timing comparison between behavior (feeding and locomotion) compared to respirometry is given in our master response to Reviewer 1, Major comment 4.

      In addition, repeated claims of "mitochondrial stress" are not directly substantiated by assays of mitochondrial function.

      Thank you. We agree that “mitochondrial stress” requires direct functional evidence. We therefore directly assayed mitochondrial respiration (baseline gut-tissue OCR in fmn and per01 versus iso31 controls), added a new Methods subsection and Figure 9, and reframed our wording from “mitochondrial stress/impairment” to altered (elevated) baseline mitochondrial respiration. The experimental details are now in the Methods and the result in the Results (both quoted below). sss was not assayed, so we removed the functional mitochondrial-stress claim for sss; we retain Vaccaro et al. (2020) as prior support for fmn.

      Methods — new subsection “Gut Tissue Respirometry” now reads: “Oxygen consumption rate (OCR) was measured in dissected gut tissue from iso31, fmn, and per01 flies using the Resipher System (Lucid Scientific, GA, USA). Baseline OCR (fmol/mm<sup>2</sup>/s) was averaged over a 24-hour window following a 12-hour acclimation; values from 3 independent runs were pooled, median-normalized to iso31 within each run, and log2(x+10)-transformed. A single iso31 outlier (third per01 run) was excluded; no other values were removed. Each mutant was compared with iso31 using the MannWhitney test (GraphPad Prism 10), with significance at p < 0.05 (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 after excluding one iso31 outlier).”

      The text has been added to Results:

      “Because pathway-level metabolomics implicated mitochondrial pathways in the sleep and circadian mutants, we directly assayed mitochondrial respiration by measuring baseline oxygen consumption rate (OCR) in dissected gut tissue from fmn and per01 relative to iso31 controls. Both fmn and per01 guts showed significantly elevated baseline OCR (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 across 3 runs; Mann-Whitney test, p<0.05; Figure 9A,B). Because only baseline OCR was measured, we interpret this as altered (elevated) baseline mitochondrial respiration rather than reduced capacity or a specific coupling defect (Figure 9).”

      Discussion- fmn:

      “These interpretations are further supported by gut-tissue respirometry showing elevated baseline mitochondrial respiration in fmn relative to iso31 controls. Together with prior evidence of ROS accumulation (oxidative stress) in fmn (Vaccaro et al., Cell, 2020), these functional data indicate that chronic sleep loss in fmn is associated with altered mitochondrial respiration.”

      Discussion- per01:

      “Gut-tissue respirometry in per01 likewise showed elevated baseline mitochondrial respiration, providing functional evidence consistent with the metabolomic signatures of disrupted redox balance and mitochondrial metabolism.”

      The study also excluded female flies entirely, despite well-documented sex differences in metabolism, which narrows the generality of the findings.

      Thank you for raising this point. We agree that sex is an important biological variable in metabolic regulation. While our Methods state that male flies were collected, we have now made this explicit and unambiguous by stating that only males were used for respirometry and metabolomics integration, and we have added this as a limitation in the Discussion, noting that sex-specific physiology could influence the magnitude and/or timing of the effects we report. We also highlight inclusion of females as an important future direction.

      The text now reads as follows (Methods):

      “Only male flies were used for all respirometry experiments and for integration with the metabolomics dataset. This was done to reduce variability introduced by female reproductive status (e.g., mating/egg production) and associated metabolic differences, enabling a clearer comparison across genotypes and lighting conditions.”

      The text now reads as follows (Discussion):

      “Because only males were analyzed, our conclusions may not generalize to females, which can show sex-specific metabolic physiology. Inclusion of female flies and direct sex comparisons across LD and DD conditions will be an important future direction. More specifically, females carry a higher reproductive and biosynthetic load (egg production) that typically raises metabolic rate and can shift RQ toward lipogenesis and alter the amplitude and phase of diurnal respiratory rhythms; females might therefore show larger or differently-timed effects than the males studied here.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors appear to be using WT-DD as a condition to disrupt circadian rhythm (line 216). Although rhythmicity is often dampened in DD compared to in LD, circadian rhythm is defined as rhythm in a constant condition (e.g., DD) after entrainment. WT-DD is not a condition that the authors should use if they want to disrupt circadian rhythms in flies. WTLL would be a better condition to use since flies become arrhythmic in LL, not DD.

      Thank you for this clarification. We agree that DD is not a circadian-disrupting condition; circadian rhythmicity is defined by rhythms that persist under constant conditions after entrainment. Our intent was not to treat WT-DD as arrhythmic, but to use DD to assess free-running circadian regulation in the absence of external light–dark cues. We have revised the manuscript to clearly distinguish diurnal rhythms under LD from free-running circadian rhythms under DD, and to avoid implying that DD abolishes rhythmicity.

      The Results text now reads as follows:

      “We used constant darkness (DD) to assess free-running circadian regulation in the absence of external light–dark cues.”

      In the Discussion, in the subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness,” we revised the interpretation of WT-DD to clarify that the observed metabolic effects reflect removal of external light–dark cues under our experimental conditions, while avoiding overgeneralization.

      The Discussion text now reads as follows:

      Our DD data show that endogenous free-running circadian regulation persists but that removing external light–dark cues degrades the temporal coordination between respiration and metabolism, indicating that external entrainment normally strengthens this coordination under our conditions; however, we acknowledge that the coupling may be different in mammals.

      (2) The authors need to be more precise in the use of the term "circadian" throughout the manuscript. When describing 24h rhythmicity in the LD condition in flies, they can use the diurnal rhythm. Circadian rhythm is an endogenous rhythm without external time cues (e.g., DD rhythm).

      Thank you for this point. We agree and have revised the manuscript to use terminology consistently: rhythms measured under LD are now referred to as diurnal (LD) rhythms/patterns, and the term circadian is reserved for endogenous free-running rhythms in DD. We updated the Methods, Results, Discussion and figure legends throughout to correct instances where LD rhythmicity was previously labeled as “circadian.”

      We added this clarification in the Methods (“Drosophila Strains, Entrainment and Collection”) and in the Results/figure legends where LD time courses are described.

      The text now reads as follows (Methods):

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).”

      We revised the Results text where WT-DD and LD time courses are described, replacing imprecise references to ‘circadian disruption’ or ‘circadian cycle’ with ‘free-running conditions in DD,’ ‘24-hour cycle,’ or ‘diurnal pattern under LD,’ as appropriate. We also revised the Discussion to avoid describing LD patterns as circadian and to avoid implying that DD disrupts circadian rhythmicity.

      (3) Lines 253-256: The authors' interpretation of this data is not accurate. The authors observed significant day-night differences in VO2 and VCO2 in WT-DD. This suggests there is circadian control over metabolism rhythms. The authors noted there is "limited circadian control".

      Thank you for pointing this out. We agree that the significant day-night differences in VO<sub>2</sub> and VCO<sub>2</sub> in WT-DD support persistent endogenous (circadian) control of respiratory rhythms under constant darkness. We have revised the Results text (lines 253-256) to remove the statement implying “limited circadian control” and instead describe the WTDD effect as maintained rhythmicity with altered amplitude and/or phase relative to LD, rather than loss of rhythmic regulation.

      We added this revision in the Results section under “Diurnal Variation in CO<sub>2</sub> Production, O<sub>2</sub> Consumption, and Respiratory Quotient Across Genotypes” (WT-DD description; lines 253-256).

      The text now reads as follows:

      WT-DD flies, maintained in constant darkness, exhibited the lowest overall respiratory activity. Despite the absence of environmental light cues, VCO<sub>2</sub> and VO<sub>2</sub> retained significant day-night differences (Figure S4), consistent with persistent free-running circadian control in constant darkness (with altered amplitude and/or phase relative to LD).

      (4) Besides sleep, metabolic rates are known to be affected by food consumption. Measuring food consumption of the sleep and circadian mutants might provide insights into whether the metabolic rates are more affected by changes in sleep profile or food consumption. This might also be important given fumin displays impaired dopamine transport function and defective dopamine reuptake, and dopamine is known to affect eating behavior. This was not considered and/or discussed.

      It is technically challenging to measure feeding during respirometry, so we acknowledge it as a limitation. To test whether behavioral timing could explain our results, we compared our respiratory rhythms with the feeding and activity rhythms reported in Malik et al. (2026); the comparison and its interpretation are now in the Discussion (quoted below). This is our master response to the activity/feeding concern and is crossreferenced from Reviewer 3’s public review and Recommendation 1; feeding and activity were not measured for sss.

      We added this clarification in the Discussion (Limitations/confounds) where we address potential behavioral contributors (activity/feeding) to respirometry outcomes.

      The text now reads as follows:

      “Locomotor activity and feeding could not be measured during the respirometry recordings, so genotype differences in respiratory parameters should be interpreted with caution. To assess whether behavioral timing could account for these differences, we compared our respiratory rhythms with the feeding and activity rhythms reported for these genotypes in Malik et al. (2026): wild-type feeding peaked at ZT ~3.25 and fmn feeding was phase-delayed to ZT ~4.5, while fmn also showed elevated locomotor activity, particularly during the dark period. In our data the fmn VCO2 rhythm peaks at ZT ~3.25 (RQ at ZT ~4.25), so the respiratory peak slightly precedes the feeding peak; the phase of the fmn respiratory rhythm is therefore not driven by feeding, although the elevated activity of fmn may contribute to its higher overall metabolic rate. Feeding and activity were not measured for sss.”

      (5) It is not clear whether the authors are simply analyzing the SAME dataset in Figure 1, Figure S4, and Figure 2-3, but just with different resolutions. They need to better articulate this point.

      We thank the reviewer for pointing this out. While the source data for Figures 1, 2-3, and Figure S4 use the same underlying respirometry dataset, they present different analyses to address specific questions. Figure 1 shows the full time-course traces (fine time bins), Figures 2-3 extract rhythmicity metrics (e.g., period/phase) from those same time-series, and Figure S4 collapses the same data into simple day (ZT0-12) vs night (ZT12-24) averages.

      We added this clarification in the Figure legends for Figure 1, Figures 2-3, and Figure S4, and also noted it in the Methods where the respirometry analysis outputs (time-series binning, rhythmicity analysis, and day/night averaging) are described.

      (6) Figure 3 and page 14: What are their criteria for differentiating "conserved" vs "distinct" phase alignment? It is not clear whether this conclusion is supported by any statistical analysis.

      We now specify that the “conserved” vs “distinct” phase descriptions refer to early- vs late-peaking rhythms, and we state the criterion explicitly: a phase was called “conserved” when it fell within ±3 h of the WT-LD peak. The revised text reads: “VO<sub>2</sub> peaked … suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”; “RQ peaked … indicating distinct phase alignment, with the sleep-mutant phases ~8 h apart (Figure 3, Table 1).”

      The text now reads as follows:

      “VO<sub>2</sub> peaked …suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”

      “RQ peaked … indicating distinct phase alignment compared to respiratory output, with phases of the sleep mutants 8 h apart (Figure 3, Table 1).”

      (7) Figure 4: The authors need to provide more details as to how the metabolite dataset was utilized to generate this figure and how they made the conclusion that their analysis "revealed a distinct circadian rhythmicity in RQ, characterized by oscillatory patterns indicative of coordinated substrate utilization across the day-night cycle".

      We have clarified how the respirometry and metabolomics data are used for Figure 4 and corrected the overstated rhythmicity claim:

      (1) The RQ patterning in Figure 4 is derived from the continuous respirometry time series, not from the metabolomics dataset, which is used only for the time-matched and lagged correlation analyses that relate metabolite dynamics to RQ. (2) We removed the statement that the analysis “revealed a distinct circadian rhythmicity in RQ”: RQ was not statistically rhythmic in WT-LD or WT-DD (Table 1), and Figure 4 instead shows the day-night RQ pattern that serves as the reference for the lag-based metabolite correlations. We revised the Methods, the Results paragraph introducing Figure 4, and the Figure 4 legend accordingly.

      The text now reads as follows:

      “Respiratory quotient (RQ) was recorded continuously and averaged into 5-minute bins. To integrate with metabolomics collected every 2 hours, we extracted the corresponding 2-hour ZT RQ values and performed a lag analysis (−120 to +120 min) to relate metabolite dynamics to RQ patterns; rhythmicity of RQ itself was assessed from the respirometry time series.”

      (8) It is unclear why the examples of hydroxyhexadecenoylcarnitine and quinolinate were chosen to be presented in Figure 5a. The authors should clarify their choice of these two examples. Also, the authors should generate a supplemental table with the "several metabolites demonstrating strong correlations (line 294).

      Thank you for this suggestion. Hydroxyhexadecenoylcarnitine and quinolinate were selected as representative examples, and we agree that the metabolites supporting the strong RQ-associated correlations should be provided more explicitly. We have now added a Supplementary Table listing the metabolites demonstrating strong correlations with RQ across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions, using the same selection criterion applied in the heatmap analyses (|ρ| ≥ 0.7, p < 0.05).

      We also revised the Results text near the statement describing strong metabolite-RQ correlations to direct readers to this new table.

      The text now reads as follows:

      Hydroxyhexadecenoylcarnitine and quinolinate are among the strongest positively- and negatively-lagged RQ-correlated metabolites in WT-LD (ρ = +0.78 at +120 min and ρ = −0.77 at −120 min; Supplementary Table 1), illustrating the two opposite lag directions of the workflow. The full set of metabolites showing strong RQ-associated correlations across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions is provided in Supplementary Table 1.

      (9) Although Spearman correlation analysis suggests some correlation between RQ and the two metabolites shown in Figure 5, the correlation shown in Figure 5b does not appear to be compelling. Results shown in Figure 5b do not provide confidence that conclusions based on clustered heatmap analysis shown in Figures 6 to 8 are meaningful. In addition to Spearman correlation, the authors might consider performing additional statistical methods to provide further support.

      Thank you for this comment. We suggest that Fig. 5b was not explained clearly and have clarified. Figure 5b is a lag analysis, not a separate correlation result: it shows how the Spearman correlation changes when the RQ time series is shifted forward or backward in time relative to the metabolite timepoints. The goal is to illustrate lead-lag timing (which shift gives the strongest association), rather than to present a single “strong” correlation as standalone proof.

      To address the concern about confidence in the heatmap-based results (Figs. 6-8), we have strengthened the reporting by providing effect sizes (Spearman ρ) and lag for the metabolite-respirometry associations (now included as Supplementary Table 1). This allows readers to evaluate the statistical support underlying the clustering, beyond the visual patterns in the heatmaps.

      We additionally report multiple-testing–corrected significance for the metabolite–RQ correlations (Benjamini–Hochberg FDR) alongside nominal p in Supplementary Table 1, using the same correction already applied to the pathway enrichment in Supplementary Table 2.

      Minor comments:

      (1) Line 82: The authors should clarify what they mean by "circadian collection". Except for WT-DD, my interpretation is that they collected their samples in LD, so that would not be "circadian collection".

      Thank you for catching this. We agree that “circadian collection” was imprecise. We have revised line 82 to clarify that samples collected under LD were collected across diurnal (LD) time (ZT), and we now reserve “circadian” specifically for collections under constant conditions (DD). We also updated the wording throughout the manuscript to maintain this distinction consistently. We added this clarification in the Methods section “Drosophila Strains, Entrainment and Collection” (line 82).

      The text now reads as follows:

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).

      (2) The authors cited Frayn 1983 to indicate how the RQ value can be used to reflect metabolic fuel utilization. Is this interpretation accepted for all animals, including flies?

      RQ is widely used in indirect calorimetry as an index of relative substrate utilization, including in small model organisms, but we agree that it should be interpreted with appropriate caveats. We have revised the manuscript to clarify that we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.

      We made this change in two places: in the Introduction, where RQ is introduced and Frayn is cited, and in the Methods, under “Carbon Dioxide and Oxygen Analysis and Calculations,” where RQ is defined.

      The text now reads as follows:

      Introduction: “These measurements allow for the estimation of energy expenditure and respiratory quotient (RQ), which can provide an index of relative substrate utilization, with appropriate caveats[13, 14]. In this study, we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.”

      Methods: “RQ was used as an index of relative shifts in substrate utilization over time and between genotypes, interpreted conservatively rather than as a precise measure of absolute carbohydrate versus lipid oxidation as this has not been directly characterized in flies.

      (3) Figure S4: Since the authors are comparing day-night differences, they should plot them in the same graph to make it easier to compare.

      We have revised Figure S4 to plot day and night within the same graph/panel for each metric (VCO<sub>2</sub>, VO<sub>2</sub>, RQ), using side-by-side day vs night groupings per genotype to facilitate direct visual comparison, and we updated the legend accordingly.

      (4) Line 252: When comparing diurnal differences, the authors should not use the word "rhythm". Pattern or profile might be a better word to use.

      We appreciate the suggestion. Where we use “rhythm” we refer specifically to 24-hour oscillations established statistically by JTK_CYCLE and RAIN (Figures 2-3, Table 1); for the coarser day-versus-night comparisons we agree “pattern” or “profile” is preferable and have adopted it there. We have gone through the manuscript to apply this distinction consistently.

      (5) Figure 6a: larger font labels are necessary for the metabolites.

      Thank you. We have revised Figure 6a to increase the metabolite label font size for improved readability.

      Reviewer #3 (Recommendations for the authors):

      (1) Activity controls: To strengthen the paper, include or reference direct measures of locomotor activity (e.g., DAM system). This would allow the separation of metabolic changes from behavioral differences and would enable better analysis of circadian patterns.

      This activity/feeding confound is addressed in full in our response to Reviewer 1, Major comment 4; we cross-reference it here to avoid repetition.

      (2) Mitochondrial function: "mitochondrial stress" should be supported by additional assays such as mitochondrial enzyme activities, high-resolution respirometry, or reactive oxygen species measurements.

      See our full response to the mitochondrial point in Reviewer #3's public review above, including new Figure 9 and the Gut Tissue Respirometry Methods.

      (3) Sex differences: Provide a clear justification for excluding female flies. If feasible, incorporate female data or explicitly discuss how sex differences could alter metabolic outcomes.

      This is addressed in the public review for reviewer 3.

      (4) Statistical presentation: Increase n. Clarify in figure legends whether error bars represent SD or SEM and ensure consistency across all figures (Figure legends).

      We have revised all figure legends to clearly state that error bars represent SEM and ensured this is applied consistently across all figures. We also explicitly report the n for each genotype (with chambers as the experimental unit) in the relevant legends.

      The text now reads as follows:

      “Error bars represent SEM, and n denotes the number of chambers (experimental units) per genotype.”

      (5) Sample size and replication: Indicate more clearly that the chamber, not the individual fly, is the experimental unit. Discuss limitations of replication and statistical power in the text.

      Thank you for this comment. We have clarified throughout the Methods and figure legends that the chamber (not the individual fly) is the experimental unit. We now state that each genotype measurement reflects an average of 300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), and we report n (number of chambers) and the error structure (SEM) for each genotype.

      Added to Discussion:

      “We also acknowledge the limits of this replication: with n = 12 chambers per genotype, statistical power to detect small-magnitude differences and subtle phase shifts is limited, and negative calls (e.g., arrhythmicity) should be interpreted with this caveat.”

      (6) Writing and clarity: (a) Streamline the Discussion to focus on mechanistic themes (anticipatory vs reactive alignment, substrate shifts, redox imbalance).

      (b) The opening phrase ("Precise temporal regulation of metabolism by sleep and circadian rhythms is essential for dynamic energy homeostasis") is dense, vague, and difficult to interpret. Consider rephrasing to something more concrete, for example: "Sleep and circadian rhythms tightly control when and how the body uses energy, but we do not yet know exactly how this timing connects to oxygen use and breathing needs."

      Thank you for these helpful suggestions. We have revised the Discussion to reduce repetition and improve focus by organizing it around the key mechanistic themes raised by our data, including anticipatory vs reactive metabolic alignment, substrate-use shifts, and redox/mitochondrial imbalance. We revised the opening sentence of the Abstract to make the biological question more concrete and accessible, following the reviewer’s suggestion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Some of the authors proposed in a PNAS paper in 2016 the occurrence of the EntnerDoudoroff (ED) pathway in cyanobacteria and plants, on the basis of several lines of biochemical and genetic evidence. However, more recent results indicated that one of the two specific enzymes of the ED pathway (EDD) is missing in Synechocystis PCC 6803. The authors carried out additional experiments, which demonstrated that EDD is missing, and one of the enzymes (ED aldolase) is a promiscuous enzyme which seems to be involved in proline metabolism and is not actually participating in the ED pathway as initially believed. The results described in this paper are strong evidence that this new interpretation is appropriate, and therefore, it corrects the previous proposal, providing an honest description of the reasons why the authors had reached the wrong conclusion about the existence of the ED pathway in cyanobacteria and plants.

      We thank Reviewer 1 for the summary and comments. Based on our finding that EDA is a promiscuous aldolase that, in addition to the cleavage of KDPG to GAP and pyruvate (a reaction of the ED pathway) catalyzes other reactions in vitro, we proposed potential in vivo functions of EDA, including its involvement in proline metabolism. However, these assumptions require further experimental testing. We do not yet have definitive findings regarding the function of the promiscuous aldolase EDA in Synechocystis in vivo, but respective studies are currently underway.

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations in light of more recent findings and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides, they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause of the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; this could have been easily prevented by confirming the correct generation of the mutant by DNA sequencing.

      We thank the reviewer for this kind comment. We agree that biochemical characterization of enzymes as well as DNA sequencing to check deletion mutants, are important and valuable tools. As outlined in the manuscript and additionally in more detail in a recently submitted article, which is available at bioRxiv (Theune et al. 2026, doi: https://doi.org/10.64898/2026.04.08.717167) and is currently under review at PLOS One, we suggest that genome sequencing of deletion mutants in combination with complemented strains as controls are required to minimize the risk of misinterpretation based on secondary mutations (1). During the early stages of our research on the ED pathway, and later as well when we were already trying to resolve the conflicting results that had accumulated concerning the ED pathway, genome sequencing for Synechocystis mutants was not affordable as a routine procedure (2-4). Therefore, we could not have easily prevented this misconception based on this technique at that time. However, we strongly encourage genome sequencing of deletion mutants (in combination with complemented strains) as routine procedures these days (1).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      As stated above and in the manuscript, we agree that the in vivo role of EDA requires further experimental work which is in progress. However, our results demonstrate that EDA splits KDPG into GAP and pyruvate in vitro, but we assume that this reaction does not play a role in vivo due to the absence of its substrate.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but it certainly would confer additional robustness to the enzymatic results.

      Synechocystis is easily amenable to genetic manipulation, and we agree that expression and purification of all enzymes from this host would have been ideal. However, the first characterization of Synechocystis EDA was performed with proteins that were purified from Synechocystis and showed activity on KDPG at comparable rates as proteins that were purified from E. coli in this study (2). Moreover, most biochemical characterizations of EDAs from archaea, bacteria and plants were performed after recombinant expression in E. coli and yielded highly active enzyme as in the case of Synechocystis is this study (5-7). Therefore, we currently have no reason to worry that the expression in E. coli might affect the enzymatic activity of EDA. The main reason for utilizing E. coli as an expression strain in this study was to gain higher yields of protein for in-depth analyses.

      Bibliography:

      I think the list of papers used in this manuscript is complete and up to date. However, I do miss recent papers which addressed one aspect that was proposed in the original 2016 PNAS paper: the authors wrote, "We therefore suggest that Prochlorococcus might oxidize glucose via the ED pathway under mixotrophic conditions, as shown for Synechocystis." Recent studies checked this hypothesis and have shown that the ED pathway seems to be also missing in Prochlorococcus and marine Synechococcus, and I think this manuscript is a good place to cite them, since these results are consistent with the findings of this paper.

      We will include a references from Moreno-Cabezuelo et a. 2023 (DOI: 10.1128/spectrum.03275-22) in which the proteomes of three marine Prochlorococcus and three marine Synechococcus strains were investigated upon exposure to glucose (8). Protein levels of EDA were either downregulated or not affected while proteins involved in OPP pathway and CBB cycle were upregulated. The authors of this study conclude that this indicates that the latter processes rather than the ED pathway are involved in photomixotrophy in these strains. However, flux analyses are still missing.

      Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner-Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data from different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data are presented to reconcile contradicting data of earlier studies and to serve as a basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetition, and nevertheless misses some specifications.

      We thank Reviewer 2 for the kind summary and comments and will improve the materials and methods part accordingly in a revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Some additional aspects that could be improved:

      (1) L182-184: Are there any known in vitro attempts to determine whether some DHADs can accept 6PG as substrate? If not, did the authors check this possibility in the lab?

      To the best of our knowledge, as mentioned in the manuscript, some DHADs are tested for their activity towards gluconate and some other substrates but not 6PG. It was suggested that since gluconate is smaller it might fit in the catalytic site of DHADs normally occupied by DHIV (7). We discussed this point in the manuscript (Line 177-185).

      In this study we tested the DHAD Slr0452 from Synechocystis and DHAD from Synechococcus with 6PG as substrate but the enzyme did not catalyze 6PG dehydration. The result for Slr0452 is shown in Figure S3.

      (2) L234-241: This paragraph shows how important it is to avoid relying only on sequence alignments to assign functions to proteins (and enzymes in particular). The authors stress this idea elsewhere in the paper, but I think it should receive even more attention in a manuscript like this. Physiological characterization of the protein function is paramount, especially in the case of enzymes. Also, the GDH1 overexpression mutant was done in E. coli; this adds an additional layer of uncertainty, since the protein processing in E. coli might not be entirely identical to that carried out in cyanobacteria, as I mentioned above. This, in turn, could be one of the reasons for not finding the expected enzymatic activity. This comment is also relevant for the results shown in Table 1 (page 12).

      As outlined following lines L234-241 we tested crude cell extracts from Synechocystis WT, a Synechocystis strain overexpressing putative GDH1 (Sll1709) and a Dzwf deletion mutant which can be assumed to upregulate a GDH/GK bypass if present in Synechocystis for GDH activity but did not find any. This strongly indicates that sll1709 does not code for an active GDH in Synechocystis. We thereafter also overexpressed GDH1 in E. coli and again could not detect any GDH activity. In case of GDH1, enzyme activity was therefore tested both in Synechocystis and E. coli and yielded similar results.

      (3) L309-310: "we currently have no explanation for the gluconate that was detected in previous IC-ESI-MSMS measurements in Synechocystis". This is a serious issue, given how important this evidence was for the conclusions of the 2016 PNAS paper and the hypothesis of the ED pathway in cyanobacteria. I would suggest that the authors provide some possible explanations, even if it is based on studies from other teams.

      It would be very speculative and therefore in our eyes not helpful to search for explanations in this case as indeed the measurements were done in another lab. One explanation might be that 6P gluconate got dephosphorylated and yielded gluconate as an artifact. However, as we have no experimental validation for this idea and furthermore cannot test it. We therefore prefer to not further comment on this aspect.

      (4) L347: The presence of an insertion in the sequence of zwf in the ∆gnd mutant is a welcome explanation for the observed results: abolition of 6PG production in this mutant. This is another important message to stress in this manuscript: construction of mutants should always be double checked by DNA sequencing of the relevant genomic regions, to ensure that these kinds of side problems do not appear, leading to confusing results.

      We agree with the reviewer and would get even one step further and rather suggest combining whole-genome sequencing with complemented mutants as it is difficult to know all relevant genomic regions that should be sequenced. We discuss this issue in even more detail in another manuscript that is currently available as online preprint in bioRxiv and is under review (9). This work was now added as a citation in line 566 and the end of the following statement (line 564-566): Routine complementation of deletion mutants and sequencing of selected genes or the entire genome are effective means of identifying secondary mutations that can lead to misleading phenotypes (9).

      (5) L418-419 and Table 2: The results observed for the authors (i.e., that EDA could also catalyze reactions with OAA and KHG, albeit with substantially lower catalytic efficiency than with KDPG), is a matter for concern: if their hypothesis is correct (meaning that this EDA is fundamentally involved in the PEP-pyruvate-OAA node and/or proline metabolism, and not with the ED pathway), then why should it keep in the evolution of these organisms such a strong preference for KPDG, when it is not being used physiologically for the ED pathway)? Furthermore, the Km for KDPG is lower than for OAA or KHG, leading to a very big difference in the Kcat/Km values.

      To solve these questions further respective studies are underway. As EDD is absent from Synechocystis no KDPG should be available in the cells so that catalytic activity on KDPG should be irrelevant in vivo. The in vivo role of Eda requires further clarification.

      Hereafter, I will mention some aspects, following the instructions of eLife, which are related to suggestions for improved experiments/data/analyses, improvements of writing and presentation, and minor corrections to text/figures.

      (1) L81: I would modify the text to "Accordingly, this raises further questions...".

      Thanks for this hint. We modified the text accordingly.

      (2) L189: Add "pages" after "following".

      Thanks for pointing this out. We replaced “following” by “below”.

      (3) Page 7: The whole beginning of the Results section is actually more discussion than description of results, and the first mention of figures appears in L207 of page 8. Given the content of the paper, I think the authors might reconsider using "Results and Discussion" rather than different, specific sections for Results and Discussion. This is one of the papers where I think the combined use of both makes sense and will allow an easier understanding of the message.

      We thank the reviewer for this valuable suggestion and changed the heading to Results and Discussion.

      (4) L181: Add reference regarding the llvD/EDD superfamily.

      We added the following references in lines 172-176 and 185-189:

      (1) Melse, O., Sutiono, S., Haslbeck, M., Schenk, G., Antes, I., & Sieber, V. (2022). Structure Guided Modulation of the Catalytic Properties of [2Fe− 2S]-Dependent Dehydratases. ChemBioChem, 23(10), e202200088.

      (2) Ren, Y., Vettenranta, E., Penttinen, L., Jänis, J., Rouvinen, J., & Hakulinen, N. (2025). The engineered dimer of L-arabinonate dehydratase from Rhizobium leguminosarum bv. trifolii: The role of intersubunit interactions in IlvD/EDD family. Biochemical and Biophysical Research Communications, 757, 151610.

      (3) Ahmed, H., Ettema, T. J., Tjaden, B., Geerling, A. C., Van Der Oost, J., & Siebers, B. (2005). The semi-phosphorylative Entner–Doudoroff pathway in hyperthermophilic archaea: a reevaluation. Biochemical Journal, 390(2), 529-540. ff

      (4) Bräsen, C., Esser, D., Rauch, B., & Siebers, B. (2014). Carbohydrate metabolism in Archaea: current insights into unusual enzymes and pathways and their regulation. Microbiology and Molecular Biology Reviews, 78(1), 89-175.

      (5) Figure 2: The data shown in column plots in Fig 2A, B and C, and 4B, could be presented in tables, which would save space while providing the same information: basically, very little/no activity in some cases vs high levels of activity in others.

      We would like to keep the column plots as we still think that they visualize our data well.

      (6) L255 "Unfortunately, we were not able to overexpress putative GDH2". It would be interesting to give more details about the possible reasons for this fact.

      We tested different growth conditions for recombinant GDH2 expression. The expression culture was incubated at 37 °C for 3 hours as well as overnight at 18 °C for overnight after induction. Both experiments did not resolve the expression problem.

      (7) L409-410: I think this sentence should include a brief part explaining the kind of essay used to test this activity.

      We added now that the LDH-coupled continuous assay was used (see line 403).

      (8) Figure 6E: Please give the specific activity in U/mg, as in Fig 6F, instead of percents.

      100% is given in U/mg units in the figure legend as “control without effector (100 %; specific activity of 4.3 U/mg)”. For easy comparison of effectors, the relative activity (%) is often used. We would therefore prefer to keep the current data presentation.

      (9) L576: Provide the origin of the utilized PCC 6803 strain, given there is a certain level of variability in this strain (glucose tolerance, etc). Also, even if there are some methods which are very widely used, I think the Materials and Methods section should either properly describe them or else cite the source. For instance, BG11 medium is mentioned, but no further information is given.

      We included the information that the glucose-tolerant Synechocystis strain was utilized and added the receipt of and a citation for BG11 medium (10).

      (10) L582 Generation of mutants: This section mentions the Gibson Assembly cloning method, but I miss further information to allow the reader to reproduce the methodology with as many details as possible, or at least cite papers which do so.

      We added a reference in which Gibson Assembly is described (11). Together with the primers listed in Table S3 the generation of mutants is now reproducible.

      (11) L609 Please give information in g, not rpm, for centrifugation. Also, mention the model and brand of the centrifuge and rotors used. Also, immunoblotting is very loosely described. This is also valid for other sections, for instance, L618, L636.

      We now added the following information: Cells were harvested by centrifugation at an RCF (relative centrifugal force) of 3,992 x g in a Beckman Coulter with a JLA-8.1000 Rotor for 20 minutes at 4°C. We now added a reference (12) in which immunoblotting is described in more detail.

      (12) L613: Describe the "small scale purification".

      We now added the information that the small-scale purification was performed using a 50-ml aliquot of the large culture which was treated as described below for the remaining sample.

      (13) L619-620: Describe the composition of the lysis buffer.

      The composition of the lysis buffer is already described as follows: lysis buffer (50 mM NaPO<sub>4</sub> pH=7.0; 250 mM NaCl; 1 tablet complete protease inhibitor EDTA-free

      (Roche) per 50 mL)

      (14) L691: Specify which amounts of auxiliary enzymes in coupled enzymatic assays were used.

      Thanks for pointing this out. We have now integrated the information that 1U of each of the auxiliary enzymes was utilized in the coupled enzymatic assays.

      (15) L716: The authors mention several times using a "double beam spectrophotometer". Please provide the model and brand.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany).

      (16) L717 and 718: define "∆absorption".

      In line 715, the information is given that absorption was measured at 340 nm, "∆absorption" is accordingly the ∆absorption at 340 nm. This information was added.

      (17) L723: "Synechocystis" should be in italics.

      Synechocystis is now written italics.

      (18) L749-759: This section should be described in more detail: preparation of protein extracts, SDS, immunoblotting, etc, or cite references of the same team where these methods were properly described.

      In this section the listed methods are already described in detail.

      (19) 798-799: "frozen cell pellets". Please provide numbers to specify the amount of material used.

      Thanks for pointing this out. We now added the information that frozen cell pellets with a wet weight of 2.4 g wet weight were resuspended.

      (20) L871-872: "It was ensured that auxiliary enzymes were not rate-limiting. One unit (1 U) of enzyme activity is defined as 1 µmol substrate consumed or product formed per minute" is repeated several times in the manuscript (L 907-909, L936-938). I would advise using it the first time, and on other occasions, refer to the same conditions as described above.

      We have accordingly circumvented the repetition of 1 U definition from the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Interpunctuation, especially comma placement, should be improved.

      We improved interpunctuation, especially comma placement to the best of our knowledge.

      (2) Line 63: delete the first "which".

      “Which” was deleted.

      (3) Lines 81/82: revise sentence.

      We revised the sentence to: Accordingly, this raises further questions about the presence of the ED pathway in cyanobacteria and plants.

      (4) Line 164: "presumed" instead of "presumes".

      The word was changed accordingly.

      (5) Line 228: by others? especially in reference 1?

      We deleted by others as the reference is given.

      (6) Line 240: "or" instead of "no".

      “no” was replaced by “or”

      (7) Figure 2: The axes are not well visible, and the explanation for the positive control in panel B is missing in the legend.

      Axes from figures 2, 3 and 4 were enlarged. For Figure 2B the following information was added in the legend: As a positive control, 0.05 U glucose dehydrogenase from Pseudomonas sp. was added to Δzwf cultures and to purified putative GDH1 (Sll1709). Axes from figures 2, 3 and 4 were enlarged.

      (8) The investigation on the general absence/presence of the GDH/GK bypass in cyanobacteria may not be necessary for this study.

      We included this data in this manuscript as the mistaken assumption that the GDH/GK bypass exist in Synechocystis lead among other observations to the misinterpretation of an existing ED pathway in Synechocystis. We would therefore prefer to keep these data in the manuscript.

      (9) Line 316: delete "on".

      “on” was deleted.

      (10) Line 352: delete "or".

      “or” was deleted.

      (11) Lines 352/353: ZWF expression level appears to be reduced accordingly. This should be stated.

      We agree that Zwf expression might be lower, however, we are not entirely sure if this is truly valid and would rather test this assumption further as described in the following lines.

      (12) Figure 4: The axes are not well visible.

      Axes from figures 2, 3 and 4 were enlarged.

      (13) Line 363: values are not only normalized to protein content, but also to activity found for the WT.

      We now added: The values are normalized to Zwf enzyme activity found in the WT based on protein content.

      (14) Line 408: no separate subsection required.

      The title for a new subsection was deleted.

      (15) Lines 437-439: These are results descriptions, which should not be placed in the legend, but in the main text, as is partially done.

      We deleted these result descriptions in the legend.

      (16) Lines 441/442: formatting: one or no bracket pair.

      The brackets were corrected.

      (17) Lines 443/444: refer to Table 2 instead of giving the values in the legend to avoid duplication.

      We deleted the values and now refer to Table 2.

      (18) Line 503: delete "identified".

      We deleted the second “identified” in the sentence and changed the wording to: Apart from four identified cyanobacteria that possess potential EDDs. In addition, we also added the names of the four cyanobacteria that were found including the sequence IDs of the putative EDDs.

      (19) Lines 529/530: revise sentence and format.

      We added one sentence and revised the following sentence: In contrast to Synechocystis EDA, EDA from Synechococcus prefers OAA over KDPG. The catalytic efficiency of Synechococcus EDA on oxaloacetate is rather low (OAA 0.437 s<sup>-1</sup> mM<sup>-1</sup>), however, its activity can be enhanced by NADP<sup>+</sup>(13).

      (20) Line 542: revise sentence.

      We revised the sentence to: It remains to be investigated whether this reaction could play a role in vivo, with KDPG potentially acting as a regulatory metabolite at low concentrations.

      (21) Line 577: The glass tubes used for cultivation should be specified.

      We now added the following information: Custom-made glass tubes were placed in a photobioreactor (manufactured by Willi Hilke, Uslar, Germany).

      (22) Lines 584-585: unclear, was the resistance cassette not placed in the gene to be deleted?

      Yes, the resistance cassette was placed in the gene to be deleted and was fused for homologous recombination to two DNA fragments approximately 200 bp directly upstream and downstream of the gene. This information was now added.

      (23) Line 607: cultivation equipment to be specified.

      The following information was now added: For the purification of GDH1 from Synechocystis, a 6 L photoautotrophic culture of the P3-His-GDH1 overexpression strain was grown in a 10 L glass flask at 28°C, illuminated with constant light (50 µmol m<sup>-2</sup> s<sup>-1</sup>) and gassed with filter sterilized ambient air to an OD<sub>750</sub> of about 1.

      (24) Line 670: GTS should be specified.

      Thank you for this hint. This was a typo. GST was meant not GTS. This was now corrected and GST was specified as Glutathione S-Transferase.

      (25) Line 679: delete "gluconate kinase (GK) and".

      The second gluconate kinase (GK) was deleted and sentence was revised to:

      For gluconate kinase (GK) activity measurements in Synechocystis crude cell extracts the GK reaction was enzymatically coupled to 6-phosphogluconate dehydrogenase (GND) reaction, the latter providing NADP<sup>+</sup> reduction, which was monitored photometrically at 340 nm.

      (26) Line 686: again GK activity? Difference unclear. Was the previously described procedure for GND activity determination?

      GK activity measurements in Synechocystis crude cell extracts and GK activity measurements with recombinant enzyme that was expressed in E. coli were done in two different labs with different protocols. Therefore, the first description refers to measurements with Synechocystis while the second measurement refers to measurements with E.coli. This is now specified more clearly.

      (27) Type/supplier of spectrophotometers and centrifuges used should be given.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany). As this study was performed in two different labs over a period of 8 years including one lab moving to a new location, it is now impossible to specify all centrifuges that were utilized. Even though we agree that it would be good to provide this information, we now would have difficulties to be specific.

      (28) Consider the referencing of published methods to streamline the materials and methods section.

      We now streamlined the materials and methods section by deleting repetitions as outlined below. However, as protein expression, protein purification and enzymatic tests were performed in different labs, in some cases several protocols are given.

      (29) Line 757: give specifics of anti-rabbit antibody and define PBS-T and PBS-T Cytiva.

      Specifics were added to the text.

      (30) Lines 761ff: It is not given for all genes used how they were derived. All synthesized?

      We now added detailed information for all genes.

      (31) Line 762: codon-optimized for? E. coli?

      The information was added that genes that were expressed in E. coli were codon-optimized for E. coli.

      (32) Lines 782-786: Rationals for experimental strategies do not belong to materials and methods sections, but to results sections.

      The part was deleted in the materials and methods section and transferred to the results section.

      (33) Lines 818/819: repetitive.

      We removed the repetition and refer to the purification method as stated above in the materials and methods section.

      (34) Lines 847-851: True for all EDA-type assays? Kinetic parameters are shown in Table 2 rather than Table 1.

      Yes, true for all EDA-type assays. We changed the Table number to 2.

      (35) Line 863: delete "in".

      “in” was deleted

      (36) Lines 888-893: sounds repetitive.

      The lines were modified accordingly.

      (37) Lines 908/909: repetitive.

      The repetitive comment on the definition of 1U was deleted.

      (38) Lines 928-938: repetition

      The repetition was deleted.

      References

      (1) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in <em> Synechocystis </em> PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (2) X. Chen et al., The Entner–Doudoroff pathway is an overlooked glycolytic route in cyanobacteria and plants. Proceedings of the National Academy of Sciences 113, 5441-5446 (2016).

      (3) D. Schulze et al., GC/MS-based 13C metabolic flux analysis resolves the parallel and cyclic photomixotrophic metabolism of Synechocystis sp. PCC 6803 and selected deletion mutants including the Entner-Doudoroff and phosphoketolase pathways. Microbial Cell Factories 21, 69 (2022).

      (4) A. Makowka et al., Glycolytic Shunts Replenish the Calvin–Benson–Bassham Cycle as Anaplerotic Reactions in Cyanobacteria. Molecular Plant 13, 471-482 (2020).

      (5) V. Zaitsev et al., Insights into the Substrate Specificity of Archaeal Entner–Doudoroff Aldolases: The Structures of Picrophilus torridus 2-Keto-3-deoxygluconate Aldolase and Sulfolobus solfataricus 2-Keto-3-deoxy-6-phosphogluconate Aldolase in Complex with 2-Keto-3-deoxy-6-phosphogluconate. Biochemistry 57, 3797-3806 (2018).

      (6) J. S. Griffiths et al., Cloning, isolation and characterization of the Thermotoga maritima KDPG aldolase. Bioorg Med Chem 10, 545-550 (2002).

      (7) S. E. Evans et al., Plastid ancestors lacked a complete Entner-Doudoroff pathway, limiting plants to glycolysis and the pentose phosphate pathway. Nature Communications 15, 1102 (2024).

      (8) J. Moreno-Cabezuelo, G. Gómez-Baena, J. Díez, J. M. García-Fernández, Integrated Proteomic and Metabolomic Analyses Show Differential Effects of Glucose Availability in Marine Synechococcus and Prochlorococcus. Microbiol Spectr 11, e0327522 (2023).

      (9) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in Synechocystis sp. PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (10) R. Y. Stanier, R. Kunisawa, M. Mandel, G. Cohen-Bazire, Purification and properties of unicellular blue-green algae (order Chroococcales). Bacteriol Rev 35, 171-205 (1971).

      (11) D. G. Gibson et al., Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature Methods 6, 343-345 (2009).

      (12) M. Boehm et al., Comprehensive study on ferredoxin isoforms in the cyanobacterium Synechocystis sp. PCC 6803. bioRxiv 10.64898/2026.04.08.717189, 2026.2004.2008.717189 (2026).

      (13) N. Xie, C. Sharma, K. Rusche, X. Wang, Phosphoketolase and KDPG aldolase metabolisms modulate photosynthetic carbon yield in cyanobacteria. The Plant cell 10.1093/plcell/koae291 (2024).

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

      Strengths:

      The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

      Weaknesses:

      The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

      (1) Improve the framing of the paper:

      The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

      Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

      We agree that the reviewer makes a valid point that while we were focusing on explicit hypothesis testing in our statements that the other papers mentioned are making predictions based off ecological theory. We will remove line 87 and make sure to limit these comments in the manuscript. We will reframe the introduction and discussion to better reflect this and change the verbiage throughout to focus on behavioural ecology, the impacts of anthropogenic noise and the conservation implications of the study as well as the novel arboreal aspect of the work. We will also update the title to better fit this framing of the study.

      (2) Methods:

      The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

      We do touch on the limitation of not identifying individuals in the discussion (lines 252-258), but we will draw this into a specific section which will address this and the other limitations mentioned here.

      Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

      For one group, there were only 13 videos without a playback where a marmoset was present. So, we used these 13 videos for this group, and selected 13 videos at random from the other groups to match sample size. For these groups, we assigned each video without a playback but with a pygmy marmoset a sequential number, and then used a random number generator to select 13 videos for analysis. We refer to these videos as controls as they are negative controls, and the reviewer is correct – they do measure basal levels of behaviour. In contrast, the playback of control sounds is a procedural control (Bui et al., 2022). We will change how we refer to these controls in the manuscript to reflect the types of control they are. We use the comparison between negative controls and videos with playbacks in experiment 1 to assess the impact of the playback procedure itself, though the reviewer is correct that our conclusions would be better supported if we explicitly compared the intervention playbacks with the procedural control. We will add post-hoc tests to make this comparison explicit.

      Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

      We edited the audios so that they were at similar volume levels. We will update the methods with this information and touch on this in the updated limitations section discussed above.

      Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys?

      Although we do not have information about exposure rates to predators, we do mention levels of exposure of human noise in our methods section on lines 318-320 and we also touch on this in our discussion lines 206-212.

      We will update the methods section to be clearer that all groups are exposed to high levels of anthropogenic noise due to their proximity to the ecotourism lodge and community. We will also expand on this in the discussion as a limitation as having a broader array of groups with varying levels of exposure to humans would allow us to see the broader behavioural reactions to these playback stimuli.

      For the predators we mention in the methods section line 382 “All four species have been found in the study area (Barker and Papworth, 2024)” but we will further expand on this to provide information on the density of raptors in the area based on our previous study.

      Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

      The number of videos was the total number of times the camera trap triggered. These included videos triggered by another animal or foliage movement where there was no marmoset present, and includes both videos with and without a playback.

      We will make this clearer in our description of the results and will report the number of unsuccessful playbacks.

      (3) Analysis:

      A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

      For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

      We used various tests from the performance package (e.g. check_overdispersion) to test the fit of different distributional models (e.g. Poisson, negative binomial) to the same data. Although some models were not overdispersed and did not show evidence of zero-inflation, they did have singular fits, and/or a conditional R<sup>2</sup> of 1.0 suggesting overfitting. We therefore did not choose these models. We will clarify this and provide further details of our approach in the manuscript.

      The models for the behaviours not reported did not fit well using any distributional model. These behaviours had very low occurrence (0 seconds in most videos), and so there was very sparse data for generating estimates, and very low variation within / between groups and conditions. Therefore, these models generated the errors ‘Model nearly unidentifiable’ and warnings about singular boundaries. These suggest we would not be able to reliably generate model estimates, so we do not report the results. We will change the manuscript to make this reasoning more explicit.

      Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

      We will update the results reporting to be standardised as the reviewer has suggested.

      Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

      We will include video number in a reanalysis of the data.

      Reviewer #2 (Public review):

      Summary:

      The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

      Strengths:

      This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

      Weaknesses:

      Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

      We thank the reviewer for their comments and we will update the discussion to add more detail on the ecological relevance of the behaviours that we observed.

      References

      Bui, S., Madaro, A., Nilsson, J., Fjelldal, P.G., Iversen, M.H., Brinchmann, M.F., Venås, B., Schrøder, M.B. and Stien, L.H. 2022. Warm water treatment increased mortality risk in salmon. Veterinary and Animal Science, 17, 100265. https://doi.org/10.1016/j.vas.2022.100265

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Li et al. used genetically engineered murine intestinal organoids to investigate how the temporal order of oncogenic mutations influences cell state and tumourigenicity of colorectal epithelial cells. By sequentially introducing Apc and Trp53 loss-of-function mutations in alternate orders within a Kras^G12D background, the authors generated isogenic organoid lines for both in vitro and in vivo characterisation. Bulk RNA-seq reveals expected transcriptional changes with relatively modest differences between the two triple-mutant configurations (KAT vs KTA). The key finding emerges from transplantation assays: while KAT and KTA organoids show equivalent tumourigenic potential in immunodeficient mice, only KAT organoids form tumours in immunocompetent hosts (5/10 vs 0/10), suggesting that mutation order shapes susceptibility to immune-mediated clearance. The experiments are well-executed, and the conclusions are generally supported by the data.

      Strengths:

      The experimental system is well-designed for the question. By combining a Kras^G12D transgenic background with sequential CRISPR-mediated knockout of Apc and Trp53 in alternate orders, the authors generated truly isogenic organoid lines that differ only in mutational sequence. This is technically non-trivial and provides a clean platform for dissecting order effects, a question otherwise difficult to address experimentally.

      The authors performed comprehensive baseline characterisation of these organoids, including morphological and histological assessment, quantification of organoid-forming efficiency and proliferation, and bulk RNA-seq profiling. While these analyses revealed no major differences between KAT and KTA organoids, and the observed enhancement of epithelial stemness upon Apc loss and proliferative advantage conferred by Trp53 loss are largely expected, the systematic nature of this characterisation establishes a useful methodological template for future organoid-based studies.

      The authors further investigated the functional impact of mutational order using subcutaneous transplantation assays. By comparing tumour formation in immunodeficient versus immunocompetent hosts, the authors uncover a genuinely unexpected finding: KAT and KTA organoids behave equivalently in the absence of adaptive immunity, but diverge dramatically when immune pressure is applied (KAT: 5/10; KTA: 0/10). This observation is arguably the most compelling aspect of the study and opens an interesting line of inquiry.

      We greatly appreciate your comments on this study.

      Weaknesses:

      The authors acknowledge that initiating with Kras^G12D does not reflect the typical human sporadic CRC trajectory, where APC loss is usually the first event. While this design choice was pragmatic, it means the observed order effects are contextualised within an artificial starting point. It remains unclear whether the Apc/Trp53 order would matter in a Kras-wild-type background, or whether the Kras-driven cellular state is a prerequisite for these phenotypes to emerge.

      We agree with the reviewer that initiating tumorigenesis with Kras<sup>G12D</sup> does not fully recapitulate the most common trajectory of sporadic human CRC, where APC loss typically occurs first. We had noted this point in the original Discussion and further clarified it more explicitly in the Introduction part of the revised manuscript as shown in Line 97–103.

      Our experimental design was intended to establish a controlled and genetically tractable system to interrogate the principle of mutation order effects. In this context, Kras<sup>G12D</sup> activation provides a stable oncogenic baseline that facilitates sequential genome engineering and comparison of isogenic lines.

      Although APC loss is frequently the initiation event, a recent study has suggested that Kras<sup>G12D</sup> priming can reshape the selective landscape for subsequent driver events, including Apc alterations (PMID: 41339549). Consistent with this notion, our data indicate that Kras<sup>G12D</sup> activation induces a permissive oncogenic cellular state that may influence the phenotypic consequences of later mutations. We therefore speculate that the Kras<sup>G12D</sup>-primed context may contribute to the observed order-dependent effects.

      We agree that testing Apc Trp53 order in a Kras-wild-type background would be an important future direction, and we have pointed this out explicitly in the revised Discussion as shown in Line 549–554.

      Subcutaneous implantation provides a tractable readout of tumourigenicity, but the cutaneous immune microenvironment differs substantially from that of the intestinal mucosa. Given that the central claim concerns immune-mediated selection, orthotopic transplantation would more directly test whether the observed order effects hold in a physiologically relevant context.

      In the present study, we employed subcutaneous transplantation as a widely used platform to assess tumorigenic potential under controlled immune conditions. This approach offers high reproducibility, straightforward tumour monitoring, and has been broadly applied in organoid-based cancer studies in both immunodeficient (PMID: 23273993, 23776211, 32209571, 33055221) and immunocompetent (PMID: 32209571, 33055221, 41672595) settings.

      Importantly, our primary goal was to determine whether mutation order influences susceptibility to immune-mediated clearance, rather than to model the full complexity of the intestinal niche. The clear divergence between KAT and KTA specifically in immunocompetent hosts supports the existence of intrinsic mutation order-dependent immune vulnerability.

      Nevertheless, we fully agree with the reviewer that orthotopic transplantation would provide a more physiologically relevant immune microenvironment and represents also an important direction for future investigation. We have explicitly discussed this limitation and highlight orthotopic validation as an important future direction in the revised Discussion as shown in Line 563–571.

      The ssGSEA comparison involves only 14 ATK tumours, and the key comparisons (Figure 6E) yield borderline significance (p=0.052). More fundamentally, since mutation order cannot be inferred from the clinical samples, the authors are correlating organoid-derived IFN signatures with tumour immunophenotypes without direct evidence that these patients' tumours followed a KAT-like trajectory. The reasoning becomes circular: KAT organoids define the signature used to identify KAT-like clinical tumours.

      We thank the reviewer for raising this important point. We would like to clarify that our intention was not to infer the actual mutation order in clinical samples, which indeed cannot be reliably reconstructed from bulk tumour RNA-seq data.

      Instead, our goal was to determine whether the transcriptional programs distinguishing KAT and KTA organoids could be observed in human CRC cohorts. In this context, the organoid-derived IFN-related signature was used as a molecular reference to assess potential clinical correlation, rather than to classify tumours by evolutionary trajectory.

      We agree that the statistical significance in Figure 6E is modest (p = 0.052), and we have revised the text (Line 478–480) to present this analysis more cautiously as a suggestive trend rather than definitive evidence. We also clarified this limitation explicitly in the revised manuscript (Line 537–542) to avoid overinterpretation.

      Furthermore, the most striking finding of the study, that KTA organoids fail to form tumours in immunocompetent hosts while KAT organoids can, lacks a mechanistic follow-up. The transcriptomic differences between KAT and KTA are modest when cultured as monocultures, yet their in vivo fates diverge dramatically. The authors do not address why these subtle intrinsic differences translate into such divergent immune susceptibility, nor do they characterise the immune response adequately (beyond limited CD4/CD8 IHC at tumour peripheries).

      We thank the reviewer for this important point. We agree that the mechanistic basis underlying the differential immune susceptibility between KAT and KTA remains incompletely resolved.

      A practical limitation of the current study is that KTA grafts failed to establish tumours in immunocompetent hosts, which precluded downstream histological and immune profiling of established lesions. As a result, our in vivo immune characterization of KTA grafts is nearly impossible.

      Nevertheless, our transcriptomic analyses indicate that KAT and KTA organoids differ in interferon-response and immune-related programs prior to transplantation, and those differentially expressed genes were consistently preserved in tumour cells derived from immunodeficient hosts. These results suggest the presence of intrinsic tumour-cell-autonomous differences may influence immune recognition or clearance.

      We have expanded the Discussion to outline several non-mutually exclusive mechanisms that could account for this phenotype, including altered interferon responsiveness, differential antigen presentation capacity, and changes in tumour cell-intrinsic immune escape programs (Line 527–533). These hypotheses are consistent with the transcriptional differences observed prior to transplantation and provide a framework for future mechanistic investigation. We agree that deeper immune profiling (e.g., immune infiltrate composition, antigen presentation status, and functional immune assays) will be important to fully elucidate the mechanism and represents a key direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study addresses an important and timely question in colorectal cancer biology by systematically examining the effects of the common driver mutations APC, KRAS G12D, and TP53 in murine colorectal organoids, with particular emphasis on how the order of APC and TP53 acquisition influences tumor phenotype. These mutations are well known to be frequent, truncal, and often co-occurring in colorectal cancer. While it is increasingly appreciated that mutational order can shape tumor behavior, studies directly comparing the phenotypic consequences of alternative APC-TP53 mutation orders remain rare. This work, therefore, addresses a relevant and timely question.

      Strengths:

      A major strength of the study is its focus on previously unexplored biology, combined with the generation of multiple isogenic murine organoid models with controlled mutational sequences. The authors employ careful and robust quality control of the CRISPR-mediated alterations, and the inclusion of both in vitro and in vivo experiments strengthens the relevance of the work.

      We greatly appreciate your comments on this study.

      Weaknesses:

      There are, however, several limitations that should be considered when interpreting the findings. First, KRAS G12D activation is used as the initiating alteration, whereas APC loss is generally believed to be the initiating event in most human colorectal cancers.

      We sincerely thank the reviewer for their insightful comments regarding the initiation of tumorigenesis with a Kras mutation rather than the more canonical Apc loss, which was also raised by the reviewer #1. We fully agree that the Apc-first represents the most prevalent sequence in human colorectal cancer (CRC), We have more clearly explained the rationale for our experimental design in the revised Introduction part as outlined in our response to reviewer #1.

      Second, the analysis is restricted to comparing only two mutation orders (KAT versus KTA), which limits the breadth of conclusions that can be drawn about mutation ordering more generally.

      We thank the reviewer for this critical concern, which we agree is essential for strengthening the robustness and generality of our findings. However, as a proof-of-concept study of Apc and Trp53 loss, two major oncogenic events in CRC, serves as a biologically meaningful starting point for dissecting order-dependent effects. Although it is of great significance to compare all six possible mutation orders of these three driver genes, generating and thoroughly characterizing all genotypes (with identical replicates) represents a substantial undertaking beyond the scope of this initial study.

      Finally, key RNA-sequencing and in vivo experiments rely on a single isogenic line, which substantially constrains interpretability.

      The aim of the study was to systematically investigate how mutation accumulation and order influence colorectal cancer initiation. While the data suggest that the relative timing of APC and TP53 loss may be particularly important for tumor initiation, the absence of biological replication makes it difficult to draw robust conclusions. Engraftment efficiency and tumor behavior can be influenced by many factors for a single clone, including additional passenger mutations acquired during culturing, as well as epigenetic differences that are independent of the engineered mutations.

      We thank the reviewer for this concern. We apologize that we have not made a clear presentation of our data source. Indeed, for all major in vitro and in vivo assays of double and triple mutants, we analyzed at least two independently derived clones per genotype. These independent clones harbour distinct mutations in target genes and were treated as biological replicates throughout the study.

      To improve clarity and transparency, we have revised the relevant figure legends and further provided a Table S5 to explicitly indicate the clonal origin of each data point throughout the study.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) It would be useful to have a timeline of the study, much like the one provided for the water maze protocol. On that timeline, please include the sample sizes examined, ages at exposure, and other pertinent procedures, indicating which animals remained alive for testing, etc.

      We agree with the reviewer that a timeline would be helpful for clarifying the PAE exposure paradigm and the subsequent sample collection and processing. Sample sizes vary across the different analyses; therefore, we have indicated the sample size for each experiment in the corresponding figure. To further improve clarity, we will add Author response image 1, which will include a schematic of the overall experimental timeline, including the PAE regimen, collection time points, and sample processing, as shown below.

      Author response image 1.

      (2) What were the attrition rates for each study group?

      We thank the reviewer for raising this important point. There was no attrition of animals within the experimental groups; the number of animals included at the beginning and end of the study remained the same. However, we did observe a reduction in litter size following prenatal alcohol exposure (PAE). In our 3xTg-AD colony, litters typically consisted of approximately 8 pups under control conditions, whereas PAE litters occasionally contained 4–6 pups. Thus, the reduction in animal numbers reflects decreased litter size associated with PAE rather than attrition during the study. We would also like to clarify that the primary scope of this study was not to provide a terminal/end-point analysis of disease progression, but rather to investigate the emergence of Alzheimer’s disease (AD)-related phenotypes during early adulthood following PAE. Accordingly, our longitudinal experimental design focused on identifying the earliest molecular, synaptic, behavioral, and neuropathological alterations that emerge during this period. This approach allowed us to examine whether PAE accelerates or precipitates the onset of AD-related symptomatology in the 3xTg-AD model, rather than following the animals through advanced disease stages. We will clarify this rationale in the revised manuscript.

      (3) What are the human age equivalents of the maternal mice?

      We appreciate the reviewer’s question regarding the age of the maternal mice. The dams used in our study were young adult females (2 to 3 months of age) at the time of breeding. Because chronological age does not translate linearly between mice and humans, particularly during development and reproductive maturation, we have avoided assigning a precise human-age equivalent. Based on established comparative developmental frameworks and calculations, these animals represent a 20 years old young-adult in the reproductive stage, rather than an advanced maternal-age condition [1]. We will clarify this point in the revised manuscript.

      (4) It would be helpful to see a graph of the BECs of each animal relative to the doses given. That would clarify how the alcohol exposure amount and timing are the same and where they are different for all exposed mice, given that the alcohol levels were somewhat different by group, as noted in the Methods. Were these BEC differences at all related to group differences in outcome measures or memory performance?

      We appreciate the reviewer’s suggestion to provide a more detailed representation of the BEC data. We agree that displaying the BECs for individual animals would provide additional clarity regarding the consistency of alcohol exposure across groups. We have therefore included the individual BEC values in the new Figure 1. We observed some variability in BECs between the 3xTg-AD and B6129 groups, as noted in the Methods. Importantly, however, the average alcohol consumption was comparable between the two genotypes, indicating that the difference in BECs was not due to differences in the amount of alcohol consumed. All dams in the PAE groups reached BECs above 0.08 g/dL, the commonly used legal blood alcohol concentration limit in the United States, supporting the use of our paradigm as a binge-like alcohol exposure model. We further examined whether the variability in BECs was associated with the differences observed in outcome measures, including memory performance. We did not find evidence that the differences in BECs accounted for the group differences in behavioral or molecular outcomes. Thus, although some intergroup variability in BECs was present, the overall alcohol exposure was comparable, and the observed phenotypic differences were not attributable to differences in alcohol consumption.

      Reviewer #2 (Public review):

      (1) Some figures lack prenatal alcohol treatment in the 3xTg-AD mice.

      We appreciate the reviewer’s observation and agree that the rationale for the different experimental groups across the figures should be clarified. The primary focus of this study is to characterize the effects of prenatal alcohol exposure (PAE) in wild-type B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference to determine whether the effects observed following PAE in wild-type animals overlap with or resemble features of AD pathology. Accordingly, the initial figures focus on the effects of PAE in B6129 mice and include the non-exposed 3xTg-AD group as a reference for the AD phenotype. The last two figures specifically address the effects of PAE in the 3xTgAD model, with the 3xTg-AD mice becoming the experimental subject of interest rather than serving solely as a disease reference. For this reason, the PAE-3xTg-AD group is not included in the earlier figures, whereas it is included in the final two figures where the effect of PAE on the AD model is directly evaluated. We will clarify this experimental rationale in the revised manuscript and figure legends.

      (2) Some overstatements should be tempered. For instance, one cannot conclude that the changes in CTFs are driving the changes in learning and memory (as suggested in the last line of the abstract) without a direct intervention testing this. For instance, though PAE caused a more robust learning deficit at 6 mo in WT, the impact on CTFs was less than it was at 3 mo. PAE did not significantly change CTFs or learning/memory in 3xTg-AD mice at 4 months, suggesting the genotype effect takes over at this point. The text should be adjusted to reflect this.

      We appreciate the reviewer’s careful consideration of this point. We agree that the relationship between APP CTF accumulation and learning and memory deficits should not be interpreted as causal in the absence of a direct intervention experiment. We were careful in choosing the wording throughout the manuscript to describe these findings as associated changes rather than evidence of a causal interaction. Our data demonstrate the presence of APP CTF accumulation and learning and memory deficits following PAE, but they do not establish that CTF accumulation directly drives the behavioral phenotype. We therefore will temper the language in the Abstract and throughout the manuscript to avoid overstatement. We also acknowledge that the relationship between these phenotypes is not necessarily linear across age: although PAE produced a more pronounced learning deficit at 6 months in B6129 mice, the magnitude of APP CTF accumulation was greater at the earlier time point. Importantly, we consider the possibility that the greater APP CTF accumulation observed at earlier ages may represent an early molecular insult whose functional consequences become evident later in life. In this context, the temporal dissociation between the molecular and behavioral phenotypes could be consistent with a “two-hit” model, in which an early-life insult induced by PAE creates or primes a pathological vulnerability that subsequently manifests as cognitive dysfunction with ageing [2]. We recognize, however, that this interpretation remains a hypothesis and would require longitudinal mechanistic studies to establish. This temporal relationship may also contribute to the broader concept of early-life origins of AD/ADRD, suggesting that prenatal environmental exposures may initiate molecular alterations during neurodevelopment that remain detectable or predispose the brain to later-life dysfunction. Similarly, the absence of significant changes in APP CTFs or learning and memory in 4-month-old 3xTg-AD mice suggests that the effects of the AD genotype may become dominant at this stage. Consistent with this interpretation, we state in the Discussion that future studies are needed to identify and experimentally test the direct molecular pathways affected by PAE that ultimately contribute to learning and memory impairment. We will revise the text accordingly to make this distinction clear while highlighting the potential significance of an early molecular insult preceding the later emergence of behavioral phenotypes.

      Reviewer #3 (Public review):

      (1) It is unclear as to whether there are sex differences, particularly in the adult cohort.

      We appreciate the reviewer’s comment regarding potential sex differences. We agree that considering sex as a biological variable is important, particularly for the adult cohorts. To address this point, we will include identifying marks for male and female animals separately in our plots where sample size permits. This will allow the reader to better evaluate potential sex-dependent effects of PAE and to determine whether the observed phenotypes are consistent across sexes. We will also clarify this approach in the revised manuscript.

      (2) More clarity is needed on sample size per cohort and whether mice that were used for anatomy and biochemical analyses were previously used for behavior. Including a table and noting any overlap would be useful.

      We appreciate the reviewer’s suggestion and agree that greater clarity regarding the sample sizes and use of animals across analyses is important. The sample size for each cohort and experimental group is indicated in the corresponding figures, and we will make this information more explicit in each figure legend to facilitate interpretation. Animals that underwent behavioral testing were subsequently used for biochemical analyses, allowing us to examine molecular changes in the same animals in which behavioral phenotypes were characterized. In contrast, for the neonatal cohort, we performed both anatomical and biochemical analyses, to assess the distribution and extent of APP CTF accumulation across the brain during this early developmental period. This approach was selected to provide a broader assessment of the spatial distribution of APP CTF accumulation at birth. We will clarify these experimental details in the revised Methods and figure legends.

      (3) In many instances, two-way ANOVAs with treatment (PAE vs vehicle) and genotype as factors will be useful to report (e.g., Figure 1).

      We appreciate the reviewer’s suggestion regarding the use of two-way ANOVA with treatment and genotype as factors. However, we respectfully disagree that this approach is appropriate for all of the analyses presented in this manuscript. Our experimental design and the specific biological questions addressed in each experiment were not uniform across cohorts. In particular, the primary objective of the study was to characterize the effects of PAE in B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference, while the effects of PAE in the 3xTg-AD model were specifically examined in the later experiments. Therefore, combining genotype and treatment as factors across all datasets would not always reflect the experimental questions or the structure of the cohorts. In addition, some experiments did not include all four groups, making a two-way ANOVA inappropriate for those analyses. Nevertheless, we agree that a two-way ANOVA may be informative for experiments in which both genotype and treatment are fully represented and the experimental design supports this analysis. We will therefore consider and apply two-way ANOVA, where appropriate, to those datasets, including evaluation of the main effects of genotype and treatment and their interaction. We will clarify the statistical approach and its rationale in the revised Methods and figure legends.

      (4) In Figure 5 and line 253, it is stated that older mice have more severe deficits, but there are no direct statistical comparisons with younger AD mice.

      We appreciate the reviewer’s observation. We agree that, in the absence of a direct statistical comparison between age groups, the statement that older mice have “more severe deficits” may be too strong. Our intention was to describe the apparent progression of the phenotype across age rather than to imply that we had statistically demonstrated an age-dependent increase in severity. We have therefore revised the text in Figure 5 and at line 253 to use more cautious language, describing the greater magnitude of the observed deficits in older mice without implying a direct statistical comparison between age groups. We agree that a formal conclusion regarding age-dependent progression would require a statistical analysis, which we will include in the revised manuscript.

      (5) Lines 270-271 refer to mice as "presymptomatic", but these mice do have behavioral symptoms. Do the authors mean no neuropathology yet? Any data showing lack of robust neuropathology would be useful.

      We appreciate the reviewer’s careful observation. We agree that the term “presymptomatic” was not sufficiently precise, particularly because the mice already exhibit measurable behavioral alterations at this age. Our intention was not to suggest that these animals were free of phenotypic abnormalities, but rather that they were at an early stage of disease progression, before the emergence of robust neuropathological features. We have therefore revised the terminology to avoid referring to these mice as “presymptomatic.” Instead, we describe them as being in an early stage of disease development, characterized by emerging behavioral and molecular alterations but without the extensive cardinal neuropathology typically associated with later stages of the 3xTg-AD phenotype. We agree that the distinction between behavioral symptoms and neuropathological progression is important. In this study, our focus was on the emergence of early AD-related phenotypes during young adulthood, rather than on establishing the absence of neuropathology. We have therefore avoided making a definitive claim regarding the lack of neuropathology and have revised the text to more accurately reflect the scope of our data.

      Additional References

      (1) Dutta, S. & Sengupta, P. Men and mice: Relating their ages. Life Sci. 152, 244–248 (2016).

      (2) Gunn, J. S. et al. Exploring the ‘Multiple-Hit Hypothesis’ of Neurodegenerative Disease: Bacterial Infection Comes Up to Bat’. Frontiers in Cellular and Infection Microbiology | www.frontiersin.org 1, 138 (2019).

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank you for the time you took to review our work and for your feedback! We have performed the following additional analyses:

      (1) We analyzed the optomotor mismatch response as a function of time spent in the optogenetic closed-loop session.

      (2) We show the correlation between optomotor mismatch response and visuomotor mismatch response for functionally identified PE neurons.

      The main changes to the manuscript are:

      (3) Clarification of our terminology (moved and refined definition of the teaching signals).

      (4) A new figure panel summarizing the functional influence patterns we identified.

      (5) Clarification of our argumentation for the JEPA-inspired proposal.

      All comments are addressed individually in the following.

      Public Reviews:

      Reviewer #1 (Public review):

      Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closedloop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two-photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

      The authors conclude that both versions of predictive processing architectures they analyze are likely invalid, and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing, engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction error neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

      Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

      In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

      The authors find that the functional impact of L5 neurons on L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

      They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Figure 7).

      All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new, exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

      I have several questions about the interpretations of some of the claims and suggestions for potential additional experiments and analyses.

      We thank the reviewer for their comments. We address them below.

      (1) Some of the pieces of the puzzle remain to be identified and demonstrated: the existence of internal representation neurons in L2/3 and ascertaining that the L5/6 neurons analyzed function indeed as internal representation neurons. The authors find that stimulation of L2/3 positive prediction error neurons enhances activity of L5 neurons...If L5 neurons hold a latent representation that serves as a teaching signal for L2/3 neurons (as the authors posit), wouldn't one expect that the input they receive from the positive prediction neurons be suppressive, such that the error is further minimized?

      Not necessarily - this depends on the model one has for how cortex works. In the hierarchical predictive processing model, PE+ neurons are expected to suppress the local internal representation neurons in L5. Our data, however, are not consistent with this model. This is one of the key arguments we build the idea on that JEPA is a better model for cortex than hierarchical predictive processing. In JEPA we would not expect the L2/3 prediction error to update L5 directly, but instead update the prediction of L5 activity (the source of this signal remains to be identified; see our speculations on the origin of prediction signals in the comments to question 3), and drive plasticity in the local L5 encoder network. That something acts like a teaching signal for the L2/3 comparator does not, by itself, determine how the outcome of the comparison influences the source of the teaching signal.

      (2) Do the authors envision any specific differences between the representations of the two encoder networks posited to exist in L2/3 and L5 in the JEPA-like implementation? Are they synchronous/offset in their temporal representations, or any other features?

      Given theoretical work (Mohammadi et al., 2025), one would expect to find differences in learning rates between the predictor and the encoder networks. Assuming the predictor network is also implemented in L2/3, we would expect to see faster learning rates in L2/3 compared to L5. Implementations inspired by related theoretical works also predict differences in learning rate between the two encoder networks (Grill et al., 2020), again with higher learning rates expected in L2/3 compared to L5. Beyond that, however, we are not aware of any experimentally observable differences one might expect to find. We are hoping the computational community will remedy this soon.

      (3) Where is the prediction coming from onto L2/3 neurons? Is it emerging locally in L2/3 from the putative internal representation neurons, or is it long-range - as work from the authors previously proposed? Or a mix of both?

      We expect the predictions to come from long-range inputs. In classical (hierarchical) JEPA one would expect these to be the lateral communication within L2/3. In a non-hierarchical implementation that is capable of operating on arbitrary graphs (as would be necessary for it to work in cortex), we suspect that the L2/3 network will use both long-range L2/3 and long-range L5 input. In JEPA terminology: the encoder A network uses information from non-local sources of both networks to predict local activity of encoder B network (assuming that information is statistically useful in predicting that activity).

      (4) What is the role of the indiscriminate L4 input that appears to enhance activity of both positive and negative prediction error neurons in L2/3?

      The short answer is, we don’t know. We would have indeed expected to find some asymmetry of influence. There are a few options: A) We might be failing to activate a specific interneuron that mediates feedforward inhibition on this pathway by the artificial stimulation of Scnn1a neurons. B) We know that Scnn1a neurons are only a subset of L4 neurons – there are other populations of L4 neurons that might exhibit the opposing influence. C) The Scnn1a population is more than 1 layer of a JEPA away from the comparator and not part of the predictor that forms the actual representation compared against L5 (e.g. L4 could provide an input to L2/3 internal representation neurons, and thus only indirectly influence L2/3 PE neurons). D) The JEPA analogy is wrong.

      (5) Does Figure 7D change in a meaningful manner if the authors plot the correlation between optomotor mismatch response and visuomotor mismatch response specifically for the negative prediction error neurons in L2/3 (Adamts-2) rather than for all L2/3 cells sampled?

      We might be misunderstanding. If the reviewer means genetically identified negative prediction error neurons (Adamts2), we do not have the data to address this question, as we did not perform any recordings of molecularly defined Adamts2 population in L2/3. If the reviewer means functionally identified negative prediction neurons, this would be the two rightmost data points in Figure 7D (the x-axis in this panel is the visuomotor mismatch response strength we use to functionally identify PE- neurons). These two data points on the right-hand side of the plot correspond exactly to what we classify as negative prediction error neurons throughout the rest of the manuscript (15% of the most responsive neurons to visuomotor mismatch).

      Does the reviewer mean, is there also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that we identify as PE-? If so, the answer is yes (Author response image 1). Interestingly, while there is a positive correlation, there is also nonuniformity in response patterns. We think that this is expected, given that our visuomotor coupling paradigm captures only a very small subspace of stimuli that prediction error neurons are tuned to. Bulk stimulation of L5, in contrast, might work better to separate all PE neurons. In other words, we speculate that L5 stimulation is a better predictor of the functional role of an L2/3 neuron.

      Author response image 1.

      Optomotor mismatch response as a function of visuomotor mismatch response for neurons that are functionally classified as PE. Red line shows a linear fit estimated with a bootstrap approach.

      (6) Do the optomotor mismatch responses in L2/3 neurons depend on how long the closed-loop coupling of optogenetic stimulation of Tlx3 L5 neurons and locomotion speed has been in place for?

      No, not that we can measure. We performed an analysis in which we split the optogenetic closed-loop session into two equal parts, early and late. We then quantified the average optomotor mismatch response independently for early and late parts of the session (Author response image 2A). Based on this quantification, we find no evidence of a change as a function of time in the closed-loop session. We also performed a sliding window analysis on a shorter timescale. While it did look like the responses may be smaller in the first few minutes, we do not have sufficient data to address this. None of the differences in response size were significant (Author response image 2B).

      Author response image 2.

      Optomotor mismatch response as a function of experience with artificial closed-loop coupling. (A) Mean L2/3 population response to optomotor mismatch in the first half (Early MM) and in the second half (Late MM) of the optogenetic closed-loop session. (B) Mean L2/3 population response to optomotor mismatch as a function of time in the optogenetic closed-loop session. Error bars indicate SEM. Differences between the mean values were estimated by hierarchical bootstrap and are not significant.

      Reviewer #2 (Public review):

      This manuscript reveals the functional connectivity of two different classes of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

      We thank the reviewer for their comments. We address them below.

      General comments:

      (1) The methods of statistical testing are insufficiently described. I did not understand the description in lines 1105-1119. The authors should provide sufficient details so the reader can reproduce their analyses. For example, it may be helpful to provide specific details of the testing procedure for one of the comparisons (e.g. the first comparison in Table S1).

      We assume the reviewer is not familiar with hierarchical bootstrapping in general. If so, the explanation below would summarize the procedure. This is the procedure with particular emphasis on its application to neuroscience data is described in the paper we reference in that part of the methods (Saravanan et al., 2020). Given that the analysis has become relatively standard (and is described in the reference provided), we think it might be an overkill to add the full explanation below to the manuscript. In addition to the general procedure, there are only 2 pieces of information relevant to fully reconstructing the analysis:

      (1) What are the “levels” (mice, recording sites, neurons, trials)?

      (2) What is the number of bootstrap samples used.

      Thus, we think all information is already provided in the methods. We now also explicitly added the levels when describing the nested structure of the data in the manuscript to increase clarity.

      Hierarchical bootstrap analysis:

      Our data are naturally nested: multiple neurons are recorded within a single mouse, and multiple mice are tested within an experimental group. Standard bootstrapping (sampling with replacement from the entire pool of neurons) fails because it assumes all observations are independent. In reality, neurons from the same mouse are more similar to each other than to neurons from a different mouse. The hierarchical bootstrap (or multi-level bootstrap) preserves this nested structure, ensuring your confidence intervals are not artificially narrow due to pseudoreplication.

      To illustrate the problem, assume you have 100 neurons from Mouse A and 10 neurons from Mouse B, a simple bootstrap will be heavily biased toward Mouse A. Furthermore, the simple bootstrap ignores the fact that the true variance in your population comes from two sources:

      (1) Between-mouse variance (differences in surgery, genetics, or behavior).

      (2) Within-mouse variance (differences in tuning or activity between individual cells).

      Hierarchical bootstrap addresses this problem, and is implemented as follows: To estimate the mean response while accounting for different sample sizes per mouse, a two-level resampling scheme is used

      (1) Resample the higher level (Mice)

      First, you account for the variability between animals.

      Suppose you have N mice in total.

      Randomly draw N mice with replacement from your original pool.

      Note: Because this is with replacement, a single mouse’s data might be included multiple times in one bootstrap iteration, while another mouse might be left out entirely.

      (2) Resample the lower level (Neurons)

      For each mouse selected in Step 1, you must now account for the variability within that specific animal.

      Look at the number of neurons actually recorded from that mouse (let’s call it k<sub>i</sub>).

      Randomly draw k<sub>i</sub> neurons with replacement from that mouse’s specific pool of recorded cells.

      This step is crucial: you always resample the same number of neurons that were originally recorded for that specific mouse. This maintains the "weight" or "influence" that animal had in the original dataset.

      (3) Calculate the resampled statistic

      Calculate the mean of all neurons collected in this “bootstrap sample”.

      (4) Iterate

      Repeat Steps 1–3 many times (typically B = 1,000 or 10,000 iterations).

      The distribution of these B bootstrap means represents your sampling distribution and is used to calculate confidence intervals and p-values. 

      (2) The authors should clarify how the problem of multiple comparisons was addressed for comparisons performed in multiple moments of time, where significance is indicated by a black bar (e.g. in Figure 2F).

      There is no family-wise error correction in cases of comparing response time courses implemented in our analysis. If the reviewer has a good suggestion for how to implement family wise error correction, we would be happy to implement it. We are not aware of anything that is better than what we currently do (the time bin-wise comparison). To briefly explain the problem: In most neuroscience papers, response curves are compared by choosing a time window (e.g. 0.5 to 1 s following a trigger) and calculating mean values of the curves in these windows. This hides the problem of multiple comparisons that arises from the fact that the experimenter is free to choose a response window used for analysis. Note, this also creates a strong incentive to “optimize” choice of an analysis window – a part of the analysis that a reader is typically completely blind to. We could of course also choose an analysis window that sounds reasonable and yields significant differences for our analyses. However, to provide a more unbiased image of the data, we have come to do bin-wise comparisons with a fixed p-value (typically 0.05). This is used in all of our papers at the moment. Given that samples from neighboring timepoints are correlated via a combination of actual responses and a subset of noise sources, the samples are not independent. We now also implemented a correction for spurious positive values by requiring at least 2 neighboring bins to have a p-value below 0.05 to be shown. If the samples were independent, this would mean a false positive rate of 0.0025. Given that they are not, this is a lower bound only. Additionally, any family-wise error correction would be a function of the number of time bins we show in the plot. This would mean that our choice of the time window shown in a plot (-1s to +4s, or +5s, etc.) would change the p-value we consider significant. Thus, there is no explicit family-wise error correction, and we have come to the conclusion that the bin-wise comparison with a fixed p-value is the most unbiased representation of the data we can provide. The alternative would be to additionally plot z-scores or the p-values as a function of time, but in our experience these types of plots are even harder to read for readers not used to it.

      (3) It would be helpful to add a figure in the Discussion summarising the functional connectivity suggested by all experiments.

      We now added a panel that summarizes the functional connectivity observed in our experiments to Figure S10. 

      (4) Throughout the manuscript, the authors use the term "teaching signals", but I am unclear what they mean by it: after reading the definition in lines 45-46, I thought that they corresponded to values (as they are compared to sensory signals). Later (428-430), the text suggests that they correspond to error neurons. But then lines 605-607 say it is not an error signal. The authors should define teaching signals very precisely or remove this term.

      The formal definition of the teaching signal was in footnote 1 of the manuscript. We assume the reviewer may have missed this. We now moved this definition into the main text to increase clarity.

      We use the term teaching signal to mean exactly this definition throughout the manuscript. We suspect, a second source of confusion may come from ambiguity in regards to anatomical vs. functional definitions. We have attempted to emphasize that the definition of teaching signal is a functional one, not an anatomical one (as is the case for ‘prediction’ – predictions are functionally defined, not anatomically – hence it makes sense to ask questions of the form “what are the potential sources of predictions” etc.). A teaching signal is a signal that is compared against a prediction (the ‘ground truth’ the prediction is compared and trained against). In different circuit implementations of predictive processing, different inputs function as predictions and teaching signals. In the hierarchical implementation, activity of the internal representation neuron at the lower level serves as a teaching signal for a prediction signal that is formed by the internal representation neuron from the higher level. In the non-hierarchical implementation, external inputs from the thalamus or other cortical areas serve as teaching signals for the respective prediction signals that are formed by internal representation neurons. This terminology is most intuitive when thinking from the perspective of a prediction error neuron, since a prediction error neuron computes the difference between two inputs signals – one of which functions as a teaching input, and the other one as a respective prediction. Hence, lines 428-430 specify that layer 5 input onto prediction error neurons of layer 2/3 serves as a teaching input.

      Reviewer #2 (Public review):

      Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

      The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

      This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

      Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

      We thank the reviewer for their comments. We address them below.

      While the work is overall convincing and significantly advances our understanding of the circuit-level implementation of predictive processing, there are a few weaknesses that should be addressed or discussed:

      (1) The authors define putative positive prediction error neurons as the 15% of neurons most responsive to grating onset and putative negative prediction error neurons as the 15% most responsive to visuomotor mismatch. While this selection would be expected to overlap with negative and positive prediction error neurons, the criterion is not sufficiently stringent (independent of the exact percentage chosen). In particular, classification of a neuron as a prediction error neuron should ideally be accompanied by evidence that it does not exhibit a significant increase in activity when the prediction matches the sensory input or teaching signal.

      We understand the reviewer’s intuition. We can indeed use other stimuli to identify prediction error neurons, like the relative suppression of responses in closed-loop running onset vs open-loop running onsets. This was the reason behind including Figure S1, to show that our selection results in expected pattern of running onset responses. We don’t typically use the running onset responses as they have an additional confound we have not fully understood. This is that running onset always tends to result in an increase of calcium activity in all neurons. This could have a variety of reasons: A) Contamination of hemodynamic occlusion signals (blood vessels tend to constrict at running onset, making it appear like an increase in calcium activity) – see Yogesh et al., 2025. B) The virtual coupling in our VR is not good enough to provide a true “closed loop” experience. Humans typically notice lags of larger than 30ms – in our VR it is approximately 100 ms. C). Running onset in head-fixed animals is not accompanied by a vestibular input. D) Predictive processing is wrong. We tend to think it is a combination of the three.

      We can also use combinations of the two criteria to select neurons - if the reviewer has a specific selection criteria in mind (top XXX% MM responsive AND top XXX% closed-loop suppressed, etc.) we are happy to repeat the analysis for that specific set of criteria, but the fundamental problem that we are using a functional response to select these neurons does not go away. We know that our functional selection criteria mean we select a population of neurons that is enriched for prediction error neurons. If the enrichment is too weak, we would expect to find no effects in terms of functional influence. It is hard to explain, however, how a weak enrichment could result in a strong effect on functional influence. Our arguments in more lengthy form, for why the visuomotor mismatch is a good stimulus to identify negative prediction error neurons can be found here: Attinger et al., 2017; Jordan and Keller, 2020; Leinweber et al., 2017; O’Toole et al., 2023; Vasilevskaya et al., 2022; Zmarz and Keller, 2016.

      (2) The authors "speculate that the prediction error responses in layer 2/3 may not be computed with respect to sensory input, but with respect to layer 5 activity as a teaching signal." However, it is unclear how this perspective differs from earlier statements in the manuscript. In the Introduction, the authors note that "these signals, typically referred to as sensory signals, we will refer to as teaching signals," and later describe the hierarchical variant as one "in which internal representation neurons act as a source of the teaching signal." Given this framing, it is difficult to identify what is conceptually novel in the updated view. Is the key distinction that layer 2/3 neurons are now proposed to generate predictions in an internal representation space rather than in sensory input space, as briefly suggested in the Discussion? Or are the authors introducing a distinction between an external (sensory) and an internal (cortical) teaching signal? If so, this distinction should be made explicit. Clarifying this point would considerably strengthen the manuscript.

      There might be a misunderstanding regarding our usage of the term teaching signal. In hierarchical predictive processing the teaching signal is typically referred to as a sensory signal, as e.g. in: “prediction error neurons compare predictions to sensory input”. In non-hierarchical predictive processing, or far away from the sensory input (think prefrontal cortex), or for cross-modal interactions “sensory” input is misleading. Also, in non-predictive-processing type models (like JEPA), sensory input has a different functional role. Thus, we operationally define teaching signal as the signal that is compared against the prediction by prediction error neurons.

      The two primary options we are comparing are:

      (1) Is the teaching signal to the L2/3 comparator a bottom-up input to V1 (as one would expect in predictive processing). 

      (2) Is the teaching signal to the L2/3 comparator L5 input (as one would expect in JEPA).

      Our data argue in favor of option 2. We have rephrased parts of the manuscript to try to make this clearer. 

      (3) The authors propose that "L2/3 neurons predict L5 activity, hence making predictions in the internal representation space rather than the input space," and further suggest that, since both deep and superficial cortical layers receive thalamic input, the cortex may function like a JEPA. This idea appears closely related to the model introduced by Nejad et al. (2025), which effectively implements a JEPA-like architecture: L5 activity serves as a target against which L2/3 predictions are compared in a selfsupervised manner, with both L5 and L2/3 (via L4) receiving thalamic input. It would be helpful for the authors to clarify how their framework differs from that model, and to specify the key conceptual or mechanistic distinctions between the present proposal and the approach described by Nejad et al.

      The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

      (1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.

      (2) Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 (‘the learning-driving error signal originates in L5’). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.

      (3) The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework – silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.

      (4) The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.

      (5) Lastly, there is a time delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.

      We expect that the most useful future models should move beyond JEPA, with the emphasis on nonhierarchical models capable of operating on arbitrary graphs. We think cortex functions according to principles of a JEPA (predictions in latent space), and that the role of cell types and the exact computational organization remain to be constrained.

      REFERENCES

      Attinger, A., Wang, B., Keller, G.B., 2017. Visuomotor Coupling Shapes the Functional Development of Mouse Visual Cortex. Cell 169, 1291-1302.e14. https://doi.org/10.1016/j.cell.2017.05.023

      Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020. Bootstrap your own latent: A new approach to self-supervised Learning. https://doi.org/10.48550/arXiv.2006.07733

      Jordan, R., Keller, G.B., 2020. Opposing Influence of Top-down and Bottom-up Input on Excitatory Layer 2/3 Neurons in Mouse Primary Visual Cortex. Neuron 108, 1194-1206.e5. https://doi.org/10.1016/j.neuron.2020.09.024

      Leinweber, M., Ward, D.R., Sobczak, J.M., Attinger, A., Keller, G.B., 2017. A Sensorimotor Circuit in Mouse Cortex for Visual Flow Predictions. Neuron 95, 1420-1432.e5. https://doi.org/10.1016/j.neuron.2017.08.036

      Mohammadi, A.G., Halvagal, M.S., Zenke, F., 2025. Understanding cortical computation through the lens of joint-embedding predictive architectures. https://doi.org/10.1101/2025.11.25.690220

      O’Toole, S.M., Oyibo, H.K., Keller, G.B., 2023. Molecularly targetable cell types in mouse visual cortex have distinguishable prediction error responses. Neuron 111, 2918-2928.e8. https://doi.org/10.1016/j.neuron.2023.08.015

      Saravanan, V., Berman, G.J., Sober, S.J., 2020. Application of the hierarchical bootstrap to multi-level data in neuroscience. Neurons Behav. Data Anal. Theory 3, https://nbdt.scholasticahq.com/article/13927-application-of-the-hierarchical-bootstrap-tomulti-level-data-in-neuroscience.

      Vasilevskaya, A., Widmer, F.C., Keller, G.B., Jordan, R., 2022. Locomotion-induced gain of visual responses cannot explain visuomotor mismatch responses in layer 2/3 of primary visual cortex. https://doi.org/10.1101/2022.02.11.479795

      Yogesh, B., Heindorf, M., Jordan, R., Keller, G.B., 2025. Quantification of the effect of hemodynamic occlusion in two-photon imaging of mouse cortex. eLife 14, RP104914. https://doi.org/10.7554/eLife.104914

      Zmarz, P., Keller, G.B., 2016. Mismatch Receptive Fields in Mouse Visual Cortex. Neuron 92, 766–772. https://doi.org/10.1016/j.neuron.2016.09.057

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the Reviewing Editor for their careful evaluation of our manuscript and for their constructive and insightful comments. Their suggestions have helped us to improve the clarity, rigor, and presentation of our work. In response to these comments, we have substantially revised the manuscript and performed several additional analyses and experiments, as summarized below.

      Major additions and modifications made during revision

      In response to the reviewers' comments, we have substantially revised the manuscript and performed several additional analyses and experiments:

      New analyses

      - Quantification of MyoF recovery following auxin washout using MyoF-mAID-HA immunofluorescence (Figure 8B, Figure S10D).

      - Quantification of maternal MIC2 fluorescence intensity following 24 h MyoF depletion and subsequent redistribution after auxin washout (150 micronemes per condition; Figure S10A,B).

      - Pearson correlation analysis of ANKER1-Halo and HDEL-GFP localization (Pearson's R = 0.92 ± 0.04; n = 30 parasites).

      - Additional probability-based analysis supporting regulated microneme inheritance.

      - Expanded analysis of microneme redistribution across larger replication stages (Figure S5D).

      New figures

      - Figure S6: Dual-labelling analysis of additional Group 2 organelles (ER, apicoplast, glideosome).

      - Figure S10A, B: MIC2 fluorescence intensity analysis following RB retention and redistribution.

      - Figure S10D: Correlation between MyoF recovery and phenotype rescue.

      - Figure S11: Schematic overview of quantification and analysis workflow.

      Additional experimental efforts

      - Generation of a MIC2-Halo / IMC1-mKATE / Cb-Emerald parasite line to improve visualization of RB-associated trafficking.

      - Multiple attempts to perform higher-temporal-resolution live-cell imaging. However, prolonged acquisition resulted in severe phototoxicity, replication arrest, and parasite death, preventing reliable long-term recordings.

      Textual and methodological revisions

      - Expanded Materials and Methods section with detailed descriptions of fluorescence quantification, colocalization analyses, and statistical procedures.

      - Re-evaluation of statistical analyses using two-tailed tests throughout.

      - Revision of manuscript text to clarify the evidence supporting RB-associated trafficking and to better acknowledge current limitations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags, this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting, and metabolic compartments) is divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to a better understanding of the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites.

      Weaknesses:

      In addressing the question of residual body participation in sorting of organelles, it would be useful to clearly define this structure and when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. This might seem like a moot point, but it would give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation, and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? In subsequent cell replications (4, 8, 16... stages), is there a separation between the residual body that links them all and the posterior of each new 'mother cell', and if so, when is this distinction lost? This is important because without a definition, we might be confusing different processes.

      We thank the reviewer for this excellent and thoughtful question. The residual body (RB) emerges at the end of the first replication cycle, where it is delineated by the basal complex and persists as an IVN-associated compartment connecting all daughter parasites through both plasma membrane and cytoplasm. Previous EM and live-cell studies, including ours, have shown that the RB is not a passive remnant but a dynamic structure dependent on F-actin and unconventional myosins, supporting recycling, inter-parasite connectivity, and synchronous growth (Delbac et al., 2001; Muñiz-Hernández et al., 2011; Frénal et al., 2017; Periz et al., 2017).

      In the present study, the MyoF reversibility experiment provides strong support for a model in which RB functions as an active recycling hub. Upon MyoF depletion, maternal microneme and rhoptry proteins accumulate within the RB. Following restoration of MyoF expression, this material is redistributed to daughter organelles. We interpret this reversible phenotype as evidence that the RB represents a distinct and regulated trafficking intermediate rather than simply a by-product of late daughter cell formation.

      We agree with the reviewer that mother cells retain a functional invasion apparatus until very late during daughter formation, and that the delayed release of micronemes and rhoptries likely contributes to their late trafficking toward the cell posterior. However, our data indicate that once released, these organelles transit through a defined RB compartment that actively participates in their recycling rather than merely reflecting positional constraints. This has been previously well illustrated for micronemes, which are trafficked along F-actin filaments within the residual body (Periz et al., 2019).

      At later rounds of replication (4, 8, 16 parasites), previous studies have demonstrated the presence of multiple residual body centres within the same vacuole. However, the precise temporal and structural distinction between the RB linking parasites within the vacuole and the posterior of newly formed mother cells remains insufficiently resolved and is beyond the scope of the present study. Importantly, available ultrastructural and live-cell imaging supports the persistence of shared RB compartments connecting parasites within a vacuole, arguing against a simple conflation of posterior membranes and residual body material.

      While the primary aim of the current work was to investigate the RB's role in organelle recycling, we fully agree that a more precise definition of when and how the RB is formed, remodelled, and ultimately resolved during successive replication cycles will be essential to distinguish recycling from spatial constraints. We have revised the Discussion to better acknowledge this limitation and to avoid overinterpreting the role of the RB in organelle inheritance.

      Are rhoptries/micronemes that originate in one 'mother' able to be sorted to the 'daughters' from a distinct mother in this syncytium? If so, this would make it a sorting centre, but otherwise we could be just capturing the activities at the posterior of any given cell during replication. The authors' further thoughts on this would be very interesting.

      We agree with the reviewer that our current data do not definitively demonstrate whether rhoptries or micronemes originating from one “mother” parasite can be redistributed to daughters derived from another mother within the same syncytial vacuole. Nevertheless, our MyoF chase experiments are consistent with a model in which the RB/IVN functions as an active recycling and sorting hub rather than simply representing posterior trafficking events associated with individual parasites.

      Upon MyoF depletion, maternal micronemes accumulated within the RB. Following restoration of MyoF expression, these accumulated micronemes were subsequently redistributed to daughter parasites. This reversible redistribution is more consistent with an active recycling process than with passive accumulation alone.

      To further support this interpretation, we expanded the analysis presented in Figure S5 by including additional vacuoles and larger replication stages (new panel D). These analyses show that maternal micronemes are redistributed broadly and relatively evenly among daughter parasites. We additionally performed a probability-based analysis demonstrating that the recurrent and homogeneous redistribution patterns observed are highly unlikely to arise from stochastic capture events occurring independently at the posterior end of each parasite during replication. Together, these analyses support the interpretation that microneme redistribution is a regulated process.

      Direct demonstration of recycling between all parasites within a vacuole would require a system allowing simultaneous differential labeling of (i) daughter parasites derived from a specific mother cell and (ii) the maternal organelles originating from that same mother during a subsequent replication cycle. To our knowledge, such an approach is not currently technically feasible. Nevertheless, our live-cell imaging experiments provide additional support for communal redistribution, as microneme material accumulated within the RB was subsequently observed redistributing, albeit unevenly, across multiple tachyzoites within the same vacuole.

      The Group 2 structures are described as those that are divided between daughters and receive newly synthesised proteins that add to the maternal protein of these compartments. While this is a logical conclusion for several that are mentioned, where the maternal protein signal is seen to be depleted with replication (including for the apicoplast, ER, glideosome, and Golgi). Data for the addition of new proteins to these existing structures is actually only presented in direct support of this for the Golgi.

      We thank the reviewer for this important clarification. We initially selected the Golgi as a representative example because its morphology and restricted localization provide the clearest visualization of the dual-labeling dynamics. However, the same experimental approach was applied to all Group 2 organelles analyzed in this study. To address the reviewer's concern more directly, we have now included a new supplementary figure (Figure S6) showing that the same pattern is also observed for the apicoplast, ER, and glideosome.

      We would also like to clarify that the maternal protein signal is not lost during replication but instead becomes progressively diluted as these organelles expand, are partitioned into daughter parasites, and incorporate newly synthesized proteins. The Golgi was originally highlighted because these dynamics are most readily visualized in this compartment; however, the same principle applies to all Group 2 organelles analyzed in this study, as now illustrated in Figure S6.

      Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of Toxoplasmosis. Parasite invasion into host cells, intracellular replication, and then egress, which results in the destruction of the infected cell, is central to pathogenicity. This manuscript focuses on understanding how maternal resources (in this case, cellular organelles) are shared between daughter parasites during cell division. Many organelles are single copy, meaning that division and inheritance by the daughters is crucial for successful replication. The major strength of this study was the use of a Halobased pulse chase assay to characterize patterns of organelle inheritance. The results show that both microneme and rhoptries (secretory vesicles) previously thought to be synthesized de novo are inherited by daughter parasites. Thus, this paper adds new insight to our understanding of cell division in this important parasite.

      Strengths:

      This study demonstrated that pulse labeling of proteins can be used to monitor protein synthesis, turnover, and movement. This approach will be of great interest to the field. Using this method, the authors demonstrate three main modes of organelle inheritance.

      (1) Organelles, where there are multiple copies (such as secretory vesicles, micronemes, and rhoptries), are divided between the daughter parasites, with additional contribution of newly formed vesicles. New and old material remain as separate entities in the cell.

      (2) Single-copy organelles, which are expanded to include newly synthesized material prior to division, such as the Golgi and apicoplast.

      (3) Cytoskeletal structures that are synthesized anew during each round of division. These studies provide more refined insight into patterns or organelle inheritance and demonstrate that secretory organelles are not made de novo during each round of division as previously thought. The paper has a logical flow, and overall, the data is presented in a clear and organized fashion.

      Weaknesses:

      (1) Descriptions of methodology and statistical analysis were incomplete.

      We agree with the reviewer that the description of the methodology and statistical analyses required further clarification. To address this, we have added a new supplementary figure (Figure S11) illustrating the experimental workflow, quantification strategy, and analysis pipeline. We have also expanded the Materials and Methods section to provide detailed descriptions of the experimental design, fluorescence quantification procedures, statistical analyses, and the number of biological replicates. These revisions provide a clearer and more comprehensive description of the methodology and data analysis.

      (2) There are inconsistencies between the data in Figures 1 and 5. In Figure 1, a small amount of maternal IMC is visible in stage 2 parasites. Although this is a ~90% reduction, these parasites should be quantified as parasites with material IMC. However, the graph in Figure 5C indicates that no material parasites have GAPM1a, given that graph 5C is a binary measure (present vs. absent), one would expect a non-zero percent of parasites to have maternal material.

      We agree with Reviewer 2 that, based on the raw fluorescence signal, one might expect a non-zero percentage of parasites to retain maternal IMC material after the first replication. The apparent discrepancy between Figures 1 and 5 reflects our thresholding strategy rather than inconsistent data.

      Figure 5C presents a binary analysis (presence versus absence) using a threshold calibrated from stage 1 parasites and applied uniformly across all markers. Under this criterion, the residual GAPM1a signal after the first replication falls below the detection threshold, resulting in 0% positive vacuoles. Although normalization to stage 2 parasites would detect this weak residual signal, such a protein-specific threshold would compromise direct comparison across the dataset.

      To clarify this point, we have updated the Figure 5C legend to explain the analytical approach and the asterisk associated with GAPM1a. The residual maternal IMC signal visible in Figure 1 represents a rare example selected to illustrate the remaining ~10% signal and is consistent with the absence of detectable maternal IMC1 after replication in Figures 2C and 5E.

      (3) The conclusion from Figure 6 was not justified based on the data. I agree with the author's conclusion that the accumulation of micronemes and rhoptries in the residual body was timedependent. In Figure 6A, the signal observed in the residual body at times 6:30, 13, and 14 hours is not observed in subsequent time points. However, the fate of these micronemes and rhoptries is unclear. It cannot be concluded that these vesicles are recycled back to the mother. They could also have been degraded. In fact, the graphs of microneme inheritance in Figure 2B show a decrease in maternal signal from 100% to 80% between stages 1 and 2, indicating that some microneme degradation is taking place.

      We agree with the reviewer that Figure 6 alone does not definitively establish the fate of micronemes and rhoptries accumulating within the residual body (RB), and that both recycling and degradation remain possible interpretations. Our conclusion that maternal micronemes are predominantly recycled is therefore based on the integration of Figure 6 with our MyoF depletion and recovery experiments, additional quantitative analyses, and previous work demonstrating F-actin-dependent microneme trafficking through the RB (Periz et al., 2019).

      Consistent with this model, MyoF depletion results in the accumulation of maternal micronemes within the RB, whereas restoration of MyoF expression following auxin washout leads to their redistribution across multiple tachyzoites within the same vacuole (Figure 8). Furthermore, maternal microneme signal remains detectable even after prolonged MyoF depletion (up to 48 h) and multiple rounds of replication (Figures 7 and 8), arguing against extensive degradation.

      To further address this possibility, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes retained within the RB after 24 h of MyoF depletion and following redistribution after auxin washout (150 micronemes per condition). No significant difference in fluorescence intensity was observed compared with control maternal micronemes (Figure S10A,B), indicating that maternal microneme signal is preserved during RB retention and redistribution.

      We therefore interpret the decrease in maternal microneme signal observed between stages 1 and 2 in Figure 2B primarily as a consequence of redistribution and dilution rather than degradation, consistent with the stable fluorescence intensity of individual micronemes (Figure 3). Regarding rhoptries, we note that the majority (~90%) are incorporated into daughter parasites before budding is complete, limiting their accumulation within the RB and suggesting that RB-associated trafficking primarily reflects redistribution rather than bulk degradation.

      (4) To convincingly demonstrate that the redistribution of micronemes and rhoptries was due to recovery of MyoF protein levels after auxin washout, a Western blot should be performed to show MyoF protein levels over time. In addition, the decrease in mMIC2 protein levels in the residual body in Figure 8F should be measured and normalized for photobleaching. Both apical and basal signals appear to be reduced over the time course of imaging.

      We agree with the reviewer that demonstrating MyoF recovery following auxin washout is important. Rather than performing a Western blot, we monitored MyoF recovery by immunofluorescence using the HA tag in the MyoF-mAID-HA strain, allowing direct correlation between MyoF reappearance and microneme redistribution at the single-vacuole level. These data are now included in Figure 8B, with the corresponding MyoF presence–phenotype association analysis presented in Figure S10D.

      Regarding photobleaching, we agree that fluorescence loss during time-lapse imaging is an important consideration. However, in this experiment, changes in fluorescence intensity reflect not only photobleaching but also biological redistribution of micronemes and movement of parasites in and out of the imaging plane. In the absence of a stable internal reference fluorophore, applying a standard photobleaching correction could therefore introduce additional inaccuracies. For this reason, we did not quantify fluorescence intensity during the redistribution phase.

      Instead, to assess whether maternal microneme signal is lost during RB retention and redistribution, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes following 24 h of MyoF depletion and subsequent auxin washout (Figure S10A,B). No significant difference was observed compared with control maternal micronemes, supporting the conclusion that redistribution occurs without substantial loss of the maternal microneme pool.

      Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughterderived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns, which they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strengths:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are compelling. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weaknesses:

      (1) The evidence provided to support the claim that microneme and rhoptry inheritance specifically traffics through the residual body does not sufficiently substantiate the claim. The temporal resolution of the imaging is inadequate to precisely trace the path of microneme and rhoptry inheritance. From the data shown in the manuscript, it can be concluded that at least some of the micronemes and rhoptries might be recycled through the residual body, but it is unclear whether many or most of these organelles do so.

      We thank the reviewer for this important comment and refer also to our response to Reviewer 1 above.

      Previous work has demonstrated F-actin-dependent trafficking of micronemes within the residual body (RB) (Periz et al., 2019). Consistent with these findings, our data support a model in which RB-mediated trafficking contributes to maternal microneme recycling. We acknowledge, however, that the temporal resolution of our imaging does not allow continuous tracking of every individual organelle throughout the entire replication process.

      In contrast, our observations indicate that the majority of maternal rhoptry material is incorporated into daughter cells before replication is complete and therefore does not necessarily transit through the RB under normal conditions (Figure 6). Nevertheless, rhoptry inheritance remains dependent on the actin–MyoF trafficking machinery, as MyoF depletion results in the accumulation of rhoptry material within the RB (Figures 6 and 7).

      Taken together, our data support a model in which the RB serves as an important recycling hub for maternal micronemes and can, under conditions of impaired trafficking, also transiently accommodate rhoptry material. However, our imaging resolution does not allow us to conclude that all microneme or rhoptry inheritance obligatorily transits through the RB, and we have revised the manuscript to reflect this limitation more explicitly.

      (2) The absence of specific markers for the residual body brings into question whether microneme inheritance occurs through a discrete residual body or simply via the basal end of the maternal parasite. The authors need a robust way to visualize and define the residual body to claim that micronemes and rhoptries are specifically transported through this structure.

      We agree with the reviewer that the absence of a dedicated residual body (RB) marker remains a limitation and that such a tool would improve the precision of our analyses. To date, no specific RB marker has been identified (see also our response to Reviewer 1). The most reliable proxy currently available is the F-actin chromobody, which labels the dense F-actin network associated with the RB. Using this approach, previous work demonstrated F-actin-dependent trafficking of micronemes within the RB (Periz et al., 2019).

      Building on these findings, our data support a model in which RB-associated trafficking contributes to maternal microneme recycling, whereas rhoptries are more frequently incorporated directly into daughter cells without obvious RB transit. In addition, functional perturbation of the actin–MyoF transport machinery, through MyoF depletion and subsequent recovery, supports the interpretation that the RB represents a discrete actin-associated compartment involved in organelle redistribution. Nevertheless, we acknowledge that our current imaging resolution does not allow us to determine the extent to which all microneme or rhoptry inheritance occurs through the RB, and we have revised the manuscript accordingly.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Comments for revision where either the clarity or accuracy could be improved:

      (1) The methods could do with some further detail with respect to the fluorescence intensity measurement. For example, where the Z-series was taken, and for measurements, were maximum projects taken or single Z-planes? Were all measurements made on unprocessed images, or was any deconvolution, etc, undertaken?

      We updated the Materials and Methods for better clarity and now read as: “Parasites were labeled as described above and allowed to replicate for 24 h on HFF-coated Ibidi live-cell dishes. Approximately 15 fields of view were imaged using Z-stacks spanning 3 μm centred on the vacuoles. For each replication stage (1, 2, 4, and 8 parasites per vacuole), individual tachyzoites were sampled across multiple vacuoles.

      Maximum-intensity projections were generated from non-deconvolved images. Fluorescence intensity (FI) was quantified as the maximum grey value measured within regions of interest (ROIs) drawn on individual tachyzoites from each vacuole stage. ROIs excluded overlapping parasites, neighbouring vacuoles, and regions with atypical signal intensity. For each biological replicate, up to 25 tachyzoites per replication stage were analyzed, and the mean FI value was calculated for each stage. The highest mean FI observed among the stages within a replicate was defined as 100%, and FI values for the other stages were expressed relative to this maximum. Relative FI values were then averaged across three independent biological replicates. Data are presented as mean ± SD.”

      (2) Line 93: Is more JF549 added at each of stages 2, 4, and 8? I assume so to get the progressive increase, but it would help to clarify this here.

      As illustrated in Figure 1a, the second dye is added only once, after 24 h of replication, and is then washed off before imaging. It is not maintained throughout the replication steps. The observed increase in signal is related to the presence of newly synthesized (de novo) proteins generated during replication. The Halo tags of these new proteins are initially free of ligand, as no ligand is present during replication. During the second labelling step, all free Halo tags can bind the dye. The signal intensity at each replication step therefore reflects the amount of de novo material present, which is higher at step 4 than at step 2, because the proteins generated de novo during stage 2 are also present in stage 4. The signal will be determined by both the number of newly produced molecules and their concentration at the same localization.

      (3) Line 103: Check that the Carruthers and Sibley, 1997, ref for Tic20 in the apicoplast is correct. I don't think this could be correct given the date.

      The reference it will be corrected to G van Dooren et al. 2008.

      (4) Figure legends: It would be useful to state what form of microscopy was used in each figure.

      Following the reviewer's advice the legends has been updated.

      (5) Figure 2, S2: How are single organelles tracked, such as rhoptries? I'd assume that a cell will gain new de novo organelles as well, and that this would reduce the signal per cell. Stage 2 has a rhoptry signal in both daughters, so I'd expect the signal to be roughly half for the whole cell, unless the authors can resolve individual rhoptries (which would surprise me with this microscopy). If individual rhoptries were resolved, how was this done, and what was the confidence in this (were there controls?)

      We thank the reviewer for this important point. We do not resolve individual rhoptries with the imaging conditions used in this study. Instead, fluorescence intensity was measured as the maximum grey value within a representative region of interest (ROI) encompassing the apical rhoptry signal while excluding overlapping parasites and regions with atypical fluorescence intensity. The same ROI selection strategy was applied consistently across all replication stages, allowing direct comparison with stage 1 parasites.

      The analyses presented in Figures 3 and S2 show an increasing proportion of tachyzoites lacking detectable maternal rhoptry signal as replication progresses, while the fluorescence intensity of the remaining maternal signal remains relatively stable. Together, these observations are consistent with the redistribution of intact maternal rhoptries rather than a progressive loss of rhoptry fluorescence. 

      (6) Line 150: it is unclear what is meant by 'regulated partitioning'. The more equal inheritance of micronemes versus rhoptries might not indicate a 'regulated partitioning' but just a more uniform distribution, given the larger number of micronemes versus rhoptries.

      We agree with the reviewer that this statement required clarification, and we have revised the text accordingly. Our intention was to emphasize that microneme inheritance appears to be a regulated process, rather than to directly compare it with rhoptry inheritance. The observed differences between these organelles are likely influenced, at least in part, by their different abundances.

      Across successive rounds of replication, daughter parasites consistently inherit comparable amounts of maternal micronemes, even at later replication stages (Figure S4). Given that a single mother parasite contains approximately 30–40 micronemes, whereas successive rounds of endodyogeny can generate up to 32 daughter parasites, a purely stochastic segregation would be unlikely to produce the relatively uniform distribution observed (~1–2 maternal micronemes per tachyzoite). To support this interpretation, we performed an additional probability-based analysis, which indicates that the observed redistribution patterns are unlikely to arise by chance alone. We therefore interpret these findings as supporting the existence of mechanisms that promote balanced microneme inheritance during parasite replication.

      (7) Line 184: How is the maternal signal measured without detecting the internal daughter signal? Is this an average signal for the full parasite, or just for a cross-section of the IMC? And if the latter, how are the different profiles of mother and daughter accounted for? Also, the abbreviation in the brackets doesn't make sense here.

      We thank the reviewer for this important point. We have revised the Materials and Methods section to provide a clearer description of the fluorescence intensity (FI) measurements and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, FI measurements were performed on maximum-intensity projections generated from Z-stack images without deconvolution. A representative region of interest (ROI) was selected, and the maximum grey value was used for quantification. This approach minimizes variability arising from differences in ROI size and provides a robust metric for comparison across replication stages.

      Daughter cell fluorescence was measured using the same approach while excluding overlapping signals from neighboring daughter cells and the maternal IMC. Maternal and daughter signals were distinguished based on their spatial localization and fluorescence labeling. Finally, the abbreviation in brackets has been corrected for clarity.

      (8) Line 191: Subheading a bit unclear. Distinct from other organelles, or are miconeme and rhoptry pathways distinct from each other?

      We agree with the reviewer and have updated the subheading to “Whole-organelle inheritance of micronemes and rhoptries occurs via distinct recycling pathways”

      (9) Line 193: The site of disassembly of the IMC (suggested RB here) might not be the same as the site of degradation. I suggest using 'disassembly' instead here.

      We agree that “disassembly” is an appropriate term to describe the breakdown of the IMC at the residual body (RB). However, we also believe that the RB represents the primary site of IMC degradation, for two reasons. First, if IMC material were not degraded at this site, we would expect to detect Halo-positive signal elsewhere following IMC collapse, which we do not observe. Second, transport of IMC material to an alternative degradation site would be required, but no IMC-positive vesicles are observed, arguing against significant redistribution. Together, these observations support the conclusion that the RB is both the site of disassembly and degradation of maternal IMC.

      (10) Line 215: The conclusion for a difference in timing of microneme and rhoptry segregation is not clearly supported by the data presented. Also, if there are more micronemes than rhoptries, then the frequency of observing a microneme being trafficked through the RB would need to be higher than for rhoptries if the mechanisms were the same. So, a difference in frequency here cannot be used to argue for a different mechanism.

      We agree with the reviewer and the text have been edited to soften our conclusion. 

      (11) Line 234: 'segregation' might be a better term than 'recycling' here because it is actually the sorting into daughter cells that is the important process.

      The text have been edited

      (12) Line 235: I don't think this can be what the authors intend to say. If the maternally-inherited rhoptries are not trafficked through the RB (every time), then how do they get into the daughters? Perhaps this is a case where a clear definition of the RB is required.

      Our observations indicate that maternally inherited rhoptries are frequently incorporated into daughter cells before collapse of the mother cell and establishment of the residual body (RB). Although F-actin is enriched within the RB, an actin network is also present throughout the parasite cytoplasm, where MyoF is likewise localized. We therefore propose that, unlike micronemes, rhoptries do not necessarily transit through the RB during every replication cycle but can be incorporated directly into developing daughter cells while still relying on the same actin–MyoF-dependent trafficking machinery.

      (13) The MyoF Rescue, the experimental plan is not fully described in order to be clear. If the endomembrane architecture was disrupted by MyoF depletion, and this secondary effect caused the segregation phenotype, restoration of MyoF might also simply restore the endomembrane system. So a direct role for MyoF doesn't seem to have been tested in this case.

      We appreciate the reviewer's concern that the segregation phenotype could, in principle, arise indirectly from disruption of endomembrane architecture following MyoF depletion. However, although Golgi morphology is altered in MyoF-depleted parasites, its core functions appear largely preserved. This is supported by the normal biogenesis of de novo micronemes, their correct targeting to the apical pole, their efficient secretion, and the previously reported preservation of parasite invasion. In addition, Golgi markers are not detected in the residual body, where maternally inherited micronemes accumulate, arguing against Golgi-mediated trafficking as the primary cause of the segregation phenotype.

      Taken together, these observations support the interpretation that the segregation defects are more likely to reflect a direct role of MyoF in organelle trafficking and inheritance than a secondary consequence of generalized disruption of endomembrane organization.

      (14) Line 279: Why call it a checkpoint? What is the evidence for its presence here being sensed before a further process is activated, which is what a checkpoint does?

      We agree the reviewer that checkpoint is a misleading term and have been updated to trafficking hub. 

      (15) Line 287 confuses replication of the daughters from cytokinesis, which only happens when each cell loses cytoplasmic connectivity with the other.

      We will clarify this point. In Toxoplasma gondii, cytokinesis represents the final step of daughter cell formation, during which the two fully assembled daughter parasites separate from the mother cell following collapse of the maternal cytoplasm. Historically, the residual body was proposed to arise simply as leftover material from this process. However, multiple studies have now shown that residual body formation is an active and regulated process, dependent on specific cytoskeletal and trafficking factors. Importantly, although cytokinesis marks the physical separation of daughter cells from the mother, parasites within a vacuole remain connected via the residual body and continue to share cytoplasmic and plasma membrane components until egress. 

      (16) Line 298: I don't think there is direct evidence of degradation in the RB. There might be disassembly, but degradation implies proteolysis, which hasn't been tested for.

      We agree with the reviewer that our data do not provide direct biochemical evidence of proteolysis within the residual body (RB) and primarily demonstrate disassembly of the maternal IMC at this site. However, several observations are consistent with local degradation. Following IMC collapse, we do not detect Halo-positive signal elsewhere in the parasite, nor do we observe IMC-positive vesicles or other structures that would suggest transport to a distinct degradation compartment.

      In addition, previous work identified the E3 ubiquitin ligase CSAR1 as a mediator of protein turnover within the RB, supporting the idea that this compartment is associated with degradation-related processes (O'Shaughnessy et al., 2023). While we cannot formally demonstrate proteolysis, these observations support a model in which IMC disassembly is closely coupled to local degradation within the RB.

      (17) The paragraph structure gets a bit confusing at times. See single sentence paragraph, Line 224. Does this sentence justify its own paragraph?

      The text has been edited.

      (18) Make sure Toxoplasma gondii is in italics throughout.

      The text has been edited

      (19) Line 279 cites Figure 10. But there is none.

      The text has been edited

      (20) I advocate introducing a few new acronyms, like DCs. I find that this ultimately reduces the ease with which readers read the work if they don't learn them all quickly.

      We agree that excessive use of acronyms can negatively impact readability. In the present manuscript, all abbreviations used in the text are introduced at their first occurrence in the Introduction, including DCs (line 32), IMC (line 41), PV (line 35), ER (lines 38–39), RB (line 50), and IVN (line 49). We have carefully limited the use of abbreviations to commonly used terms in the field and to those that recur frequently throughout the manuscript, with the aim of balancing clarity and readability. Nevertheless, we are happy to reduce or remove specific abbreviations if the reviewer feels this would further improve clarity.

      Reviewer #2 (Recommendations for the authors):

      (1) Descriptions of methodology and statistical analysis were incomplete as follows:

      (1a) It was unclear how the fluorescence intensity measurements (used to evaluate inheritance vs. new synthesis) were carried out. The y-axis on the graph is labeled average fluorescence intensity (% of max intensity). However, it does not state what was averaged (average fluorescence per vacuole?) and what was max intensity (max pixel intensity in each image or time point with the highest average intensity, relative to the other time points?)

      We agree with the reviewer that the original description of the fluorescence intensity (FI) measurements lacked clarity. We have therefore revised the Materials and Methods section and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, vacuoles were imaged as Z-stacks, and maximum-intensity projections were used for analysis. FI was quantified as the maximum grey value measured within representative regions of interest (ROIs) drawn on individual tachyzoites, rather than as an integrated fluorescence signal across the vacuole. This approach minimizes variability arising from differences in ROI size and allows direct comparison between replication stages.

      For each biological replicate, up to 25 tachyzoites per replication stage were analyzed. The mean FI for each stage was normalized to the highest mean value within that replicate, and data from three independent biological replicates were subsequently averaged.

      (1b) Given the uncertainties with how these measurements were performed, it is difficult to interpret the data. For example, one would expect that the fluorescence intensity of newly synthesized IMC1 in 8-parasite vacuoles would be 4 times higher than that of a 2-parasite vacuole; however, based on the graph in Figure 1B, the measured increase was only 30%.

      We agree that the original description of the fluorescence intensity (FI) measurements required further clarification and have revised the Materials and Methods accordingly. As the reviewer correctly notes, a fourfold increase in FI between 2- and 8-parasite vacuoles would be expected if total IMC fluorescence across the entire vacuole had been measured. However, this was not the parameter quantified.

      Instead, FI was measured as the maximum grey value within representative regions of the daughter IMC, providing a per-cell rather than a whole-vacuole measurement. Using this approach, FI increases between the 2- and 4-parasite stages and then reaches a plateau.

      This behavior is consistent with the biology of IMC biogenesis. Although the total amount of IMC per vacuole increases with parasite number, the amount of IMC protein incorporated into each daughter parasite remains relatively constant. Consequently, once daughter IMCs are fully assembled from de novo-synthesised material, additional rounds of replication increase the total IMC content per vacuole but not the fluorescence intensity measured for individual parasites.

      (1c) T. gondii replicates in an asynchronous manner, so that at the 24-hour time point, a single dish can contain vacuoles containing 2, 4, and 8 parasites. This should be stated explicitly so readers unfamiliar with T. gondii's growth patterns can understand how the experiment was performed.

      We agree with the reviewer and have updated the text line 93. “As Toxoplasma gondii replicates in an asynchronous manner, after 24 of replication, vacuoles containing 1, 2, 4, and 8 parasites can be observed in a single dish.”

      (1d) Colocalization package in Fiji used for ANKER1-Halo with HDEL-GFP and MIC2/RON2 with CbEmeraldFP should be specified.

      We thank the reviewer for this suggestion. Following this recommendation, we performed Pearson correlation analysis for the ANKER1–HDEL-GFP experiment using the Coloc 2 plugins of FiJi. ANKER1Halo and HDEL-GFP showed a strong spatial correlation (Pearson's R = 0.92 ± 0.04, n=30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER.

      We note, however, that this analysis should be interpreted as evidence for co-distribution within the same organelle rather than direct molecular colocalization, as ANKER1 is a transmembrane protein whereas HDEL-GFP labels the ER lumen.

      For all the rest of our analysis, no automated colocalization package or plugin in Fiji was used for the analyses involving ANKER1-Halo with HDEL-GFP or MIC2/RON2 with Cb-EmeraldFP. Colocalization was assessed manually across all experiments by inspecting both full Z-stacks and maximum-intensity projections to ensure robust spatial overlap.

      For MIC2 and RON2, the presence of signal within the Cb-Emerald–positive filament was scored as either cytoplasmic, on the residual body or absence of colocalisation.

      In total, more than 300 and 500 vacuoles were analyzed for MIC2 and RON2–Cb-Emerald colocalization respectively (stable expression of both markers), and more than 150 vacuoles were analyzed for ANKER1-Halo and HDEL-GFP colocalization (transient expression of HDEL-GFP).

      This information has now been added to the Methods section.

      (1e) Statistical methods should be described on an experiment-by-experiment basis. The authors should justify why a one-tailed t-test was conducted. A two-tailed t-test seems more appropriate.

      We agree that statistical methods should be clearly justified on an experiment-by-experiment basis. All the statistical analysis have been performed using two tails and updated in the figures.

      (1f) In Figures 5E and 5F, using boxes to indicate the exact areas of the cell that were used in the fluorescence intensity measurements, rather than arrows, would make this data easier to interpret.

      The figure has been updated.

      Reviewer #3 (Recommendations for the authors):

      The current time-lapse images and videos do not clearly demonstrate microneme movement from the maternal parasite apical end to the residual body and back to the apical end of daughter parasites. As such, the route by which micronemes enter daughter parasites remains inconclusive. To strengthen their claims, the authors should employ higher temporal resolution imaging to definitively capture the movement of micronemes from the maternal apical region into the daughters. From the current data, it also seems plausible that the micronemes may be trafficked into the daughters through the conoid as well, as there is no evidence provided showing a movement of micronemes away from the apical end of the maternal parasite before being present in the daughter parasites.

      We agree with the reviewer that higher temporal resolution imaging would provide a more definitive view of microneme trafficking. However, long-term live imaging of replicating Toxoplasma gondii requires a compromise between temporal resolution and parasite viability. In our experiments, images were acquired every 15–30 min over periods of up to 16 h, as more frequent acquisition consistently induced phototoxicity and prevented completion of parasite replication.

      Despite this limitation, our imaging reliably tracked maternal micronemes over successive rounds of endodyogeny and consistently showed microneme signal associated with the residual body. These observations are in agreement with previous high-temporal-resolution studies, which demonstrated F-actin-dependent microneme trafficking within the residual body over shorter imaging periods (Periz et al., 2019).

      We cannot formally exclude the possibility that some micronemes are transferred directly to daughter parasites through the apical end. However, together with previous studies showing enrichment of F-actin at the basal region of developing daughter cells rather than at the apical tip (Periz et al., 2017), our observations support a model in which RB-mediated trafficking contributes to maternal microneme inheritance.

      The lack of a clear residual body marker needs to be addressed, as the distinction between the basal end of the maternal cell and a bona fide residual body must be explicitly defined to substantiate the major claim of the study. As it stands, it remains unclear whether micronemes and rhoptries as a whole travel through the residual body to be transported into the daughter parasites.

      We agree with the reviewer that a marker specific to the residual body would strengthen this study. Unfortunately, no such marker has been identified to date. The F-actin chromobody is currently the best available proxy, as previous studies have shown that the F-actin network is enriched within the residual body (Periz et al., 2017; Kellermeier et al., 2024). Moreover, high-resolution live-cell imaging has previously demonstrated F-actin-dependent microneme trafficking within this compartment (Periz et al., 2019). We have revised the manuscript to more clearly acknowledge this limitation.

      The conclusions drawn from the actin colocalization data in Figure 6C are based entirely on fixed samples, despite all experimental tools being compatible with live-cell imaging. Supplementing the fixed imaging with live cell data would increase its biological relevance. Published studies have shown that fixation of the actin chromobody results in the loss of resolution of an appreciable amount of the cytosolic F-actin network, and while the localizations analyzed here are primarily along the periphery, since the quantification and text make claims about the colocalization within the cytosol, this potential loss of cytosolic F-actin becomes an issue as there may be more actin available for analysis that is lost due to fixation within the cytosol of the parasites.

      We thank the reviewer for raising this important point and agree that conventional fixation can compromise preservation of the F-actin network. However, we used the same fixation protocol described by Periz et al. (2019), which allows reliable visualization of the RB-associated F-actin network. We have also corrected the description of the fixation protocol in the Materials and Methods.

      Fixation was necessary to image entire vacuoles with sufficient spatial resolution and signal-to-noise ratio for the volumetric analyses presented in Figure 6C. Although some loss of cytosolic F-actin cannot be excluded, this would be expected to reduce, rather than artificially increase, the detection of organelle–actin associations.

      Importantly, previous live-cell imaging studies demonstrated F-actin-dependent microneme trafficking (Periz et al., 2019), and our observations are consistent with these findings. Moreover, the defects observed following MyoF depletion provide independent functional evidence that the trafficking events described here rely on the actin–MyoF transport machinery.

      The statement of colocalization should be backed up by quantitative coefficients like Pearson's coefficient.

      We thank the reviewer for this helpful suggestion. Following this recommendation, we performed a Pearson correlation analysis of ANKER1-Halo and HDEL-GFP using the Coloc 2 plugin in Fiji. ANKER1-Halo showed a strong spatial correlation with HDEL-GFP (Pearson's R = 0.92 ± 0.04, n = 30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER. As ANKER1 is a transmembrane protein and HDEL-GFP labels the ER lumen, this analysis should be interpreted as evidence of co-distribution within the same organelle rather than direct molecular colocalization.

      In contrast, we do not consider Pearson's coefficient appropriate for evaluating the association of micronemes or rhoptries with F-actin. These organelles are predominantly concentrated at the apical pole and, when associated with F-actin, are typically positioned along rather than directly overlapping the filaments. Consequently, Pearson's coefficient would underestimate these biologically relevant associations. We therefore relied on morphological and spatial criteria, which we consider more appropriate for assessing organelle–cytoskeleton interactions.

      In addition, from the methods and presented figure images, specifically in Figure 6C for RON2, how the cytosolic and residual body actin is separated is difficult to discern, as there is a clear residual body actin signal overlapping a parasite. The methods for how this was separated and analyzed should be clearer to remove doubts about how this area was measured, as the current description raises concerns about the counting of the residual body actin within the cytosol.

      We agree with the reviewer that the distinction between cytosolic and residual body (RB)-associated F-actin required further clarification. All analyses were performed manually, as described in our response to Reviewer 2 (comment 1d). The RB was identified by the presence of thick, bundled F-actin filaments at the basal pole that formed a continuous structure connecting parasites within the vacuole, whereas cytosolic F-actin was defined as the thinner filamentous network within the parasite body.

      No automated or threshold-based segmentation was used because the marked differences in filament morphology and fluorescence intensity make reliable thresholding difficult and prone to misclassification. Manual annotation based on spatial localization and filament morphology was therefore considered the most appropriate approach. We have clarified these criteria in the Materials and Methods section.

      Line comments:

      (1) 113: round to rounds - "did not obtain maternal organelles after successive rounds of replication...".

      The text has been updated

      (2) 146: grammatical, "As consequence a progressive" -> "As a consequence", or "Consequently".

      The text has been updated

      (3) 164-166: "Autonomous duplication" implies the separation and duplication of the Golgi occurs on its own, i.e., without any outside intervention, when we know from Carmeille et al. 2021 and Figure 7C here that the Golgi becomes fragmented over rounds of division in the absence of MyoF. I think this is primarily a word choice error with "autonomous".

      We agree with the reviewer and the word autonomous has been removed

      (4) 184: The wording suggests that DC's refers to daughter IMC's, when DC has already been given as an abbreviation for daughter cells previously.

      The text has been updated to correct this error

      (5) 189: de novo is not italicized.

      The text has been updated

      (6) 192: The data shown so far do not show that the RB plays a selective role in organelle recycling.

      The text has been edited to fit better our results “Our data suggest that the organelles trafficking through the residual body (RB) have different fate”

      (7) 208: State that it depends on F-actin, but never show that it is dependent on F-actin through actin disruption, such as cytochalasin D treatment or a specific conditional disruption of F-actin.

      We agree with the reviewer that we did not repeat F-actin disruption experiments (e.g., cytochalasin D or jasplakinolide treatments) in this study. These experiments were performed in our previous work, where pharmacological disruption of F-actin was shown to impair microneme trafficking (Periz et al., 2019). We therefore chose not to repeat these assays.

      Instead, the present study provides complementary evidence by demonstrating that depletion of Myosin F (MyoF), a motor that uses F-actin as a transport track (Kellermeier et al., 2024), disrupts the trafficking of both maternal micronemes and rhoptries. Together, our previous F-actin perturbation experiments and the MyoF depletion data presented here support the interpretation that these trafficking events depend on the actin–MyoF transport machinery. 

      (8) 209: This suggests that the chromobody was transiently expressed in the RON2-Halo line, but the methods suggest MIC2-Halo and RON2-Halo were integrated into a parasite line stably expressing Cb-Emerald.

      The text has been edited.

      (9) 229: The section is confusing with the mention of (now maternal). If I understand correctly, the point being made is that the de novo synthesized MIC2 at stage 2 is now the maternal MIC2 for stage 4, but coloring-wise within the figure, the now maternal MIC2 at stage 4 from stage 2 would still be green. The methods suggest these images were all taken simultaneously, and not at specific timepoints of the same vacuole, so the now maternal line remains confusing.

      The text has been revised for clarity and now reads: “Because Toxoplasma gondii replicates asynchronously, vacuoles at different replication stages coexist within the same culture after 24 h. The second labeling step marks all proteins synthesized since the beginning of the experiment, allowing discrimination between proteins present in the original mother parasite and those synthesized during subsequent replication cycles. As daughter parasites form, they inherit material from their mother, such that proteins synthesized during one replication cycle become maternal proteins in the next. Under MyoF depletion, these newly synthesized protein pools accumulate within the residual body instead of being redistributed to daughter parasites during subsequent rounds of replication (Figure 7, stage 4).”

      (10) 238: The sentences here indicate that Golgi inheritance occurs without issue: "golgi inheritance remained unaffected by MyoF depletion". But, it is evident from the images shown that the Golgi is extremely fragmented, with many more Golgi fragments by stage 8 than there are parasites. This could be solved by rewording and including a line along the lines of "in accordance with the results found in Carmeille et al. 2021".

      The reference to this article is already stated later in the text now line 299-303 “This active role is further supported by the dependence of RB-mediated recycling on F-actin and the class XXII myosin MyoF, which we show to be essential for retrieval of maternal MIC2 and RON2 but dispensable for Golgi inheritance although we noticed a fragmentation of the Golgi, which has been described to depend on MyoF (Carmeille et al., 2021).”

      (11) 319: missing a comma, "Many of these, particularly...".

      The text has been edited.

      (12) 437: Images -> imaged.

      The text has been edited.

      (13) 439: a fresh media -> and fresh media.

      The text has been edited.

      (14) 439: Replication -> replicate.

      The text has been edited.

      (15) 440: images -> imaged.

      The text has been edited.

      Other grammatical issues within the methods:

      (1) Figure 5C: Within the y-axis label of Figure 5C, there is an asterisk with no asterisk explanation within the legend.

      We thank the reviewer for pointing out this oversight. The figure legend has been updated to include an explanation of the asterisk, providing a clearer understanding of the results.

      (2) Figure 5E: The addition of an IMC1 label to match the magenta color would be helpful to readers.

      The figure has been updated

      (3) Figure 6D: No statistics showing significance of colocalization.

      We thank the reviewer for this comment. In Figure 6D, we report the frequency of observed colocalization between organelles (MIC2 and RON2) and F-actin filaments, based on manual analysis across >300 vacuoles for MIC2 and >500 vacuoles for RON2. Specifically, we observed cytoplasmic F-actin colocalization in 94% of vacuoles for MIC2 and 79% for RON2, and residual body F-actin colocalization in 85% of vacuoles for MIC2 and 11% for RON2. This analysis is descriptive and is not intended to compare MIC2 versus RON2 quantitatively; rather, it illustrates the general association of these organelles with F-actin filaments. Standard statistical measures such as Pearson correlation are not meaningful in this context because the organelles are punctate and primarily localized along filaments rather than overlapping continuously. Importantly, the percentages reported represent the fraction of vacuoles in which colocalization can be observed, not the percentage of colocalization between Cb-Emerald and MIC2/RON2 within individual vacuoles.

      (4) Figure 7: The legend title of Figure 7 is at the end of the legend of Figure 6.

      The text has been edited.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, Pinto and colleagues set out to investigate whether the cow udder is a potential mixing site for the influenza virus. The authors have demonstrated that bovine mammary epithelial cells can be infected with both avian and human influenza A viruses, supporting the idea that the cow udder may be a potential site for reassortment. Furthermore, they demonstrate that the bovine-adapted IAV replicates to similar titers in avian epithelial cells when compared to an AIV precursor virus. Thus, suggesting there is no fitness trade-off, and confirms the potential for spill-back of the cattle B3.13 into poultry, which has already been observed. Overall, I believe the authors achieved their aims. However, there are instances in which the results do not entirely support the conclusions (noted in weaknesses). Given the ongoing questions surrounding highly pathogenic avian influenza A virus in dairy cows, this work provides valuable evidence for the potential of the cow udder as a site of reassortment. These findings highlight the need for surveillance of influenza A virus incursions into livestock species, particularly cows. Some specific strengths and questions regarding weaknesses have been outlined below.

      Strengths:

      (1) The authors use a diverse range of cell types and influenza A virus strains, as well as a wide range of techniques to address the questions at hand.

      (2) The use of cells from multiple bovine breeds for the MAC-T, bMEC and explants suggests the phenomenon is not unique to a single breed.

      (3) The results suggesting there is no fitness trade-off for Cattle Texas in an avian host are interesting, and confirm the potential for spill-back of the cattle B3.13 into poultry, which has been observed.

      Weaknesses:

      I have listed my complete questions/concerns below. However, there are two main weaknesses of the article in its current state. Firstly, there is no apples-to-apples comparison in terms of determining a preference for IAV to infect the cow udder over other organs (Q4). The mammary gland and respiratory tract are represented by epithelial cells, but for other organs, fibroblasts were chosen. I think the fairer comparison would be to compare epithelial cells from different organs to demonstrate a preference for the mammary gland. Secondly, the main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. Yet throughout the article, a 10x higher dose of IAV is used in the bMEC cells compared to everything else (Q6). This raises the question of how much of the results are due to a preference for the mammary epithelial cells, and how much is simply due to the increased dose.

      (Q4) When we set out to test if cow mammary gland cells were particularly susceptible to IAV infection compared to other bovine cell types, we used what was available in the Roslin Institute – a mix of primary and continuous cells from various anatomical sites: three epithelial cell types (two mammary, one respiratory tract) two immune cell types and four sets of fibroblasts from various organs. Given the representation of different anatomical sites, cell types and differentiation statuses, we considered this a suitably diverse panel with which to characterise infection dynamics of a broad range of IAVs, before more focussed investigations using the bMEC and explant tissues. Both mammary epithelial cell types grew our library of influenza challenge strains significantly better than the BAT-II respiratory epithelial cells, as well as the two immune cell types and all four fibroblast populations. Of the fibroblast cells, those derived from the brain grew IAV significantly better than the skin and turbinate fibroblasts, while blood-derived macrophages grew virus significantly better than the lymphocytes and non-brain fibroblasts. So there are “apple-to-apple” comparisons as well as apple-to-pear comparisons that give significant differences. We therefore think that our conclusions (in the abstract) that mammary cells are particularly replication competent for IAV, (at the end of the introduction) that “a wide range of cow-derived cells are susceptible” and that (in the results section) that “mammary cells showed the highest susceptibility” are justifiable. However, we agree that testing a wider variety of epithelial cells would be useful and have added text to the Discussion (lines 224-228) to acknowledge this.

      (Q6) We used a higher MOI for bMECs because test experiments with WT PR8 and the Cattle Texas 6:2 reassortant virus showed that MOI 0.01 infections gave more variable results than those run at MOI 0.1, perhaps because of the intrinsic variability of mixed primary cell populations. However, the end-point titres between the two conditions were not significantly different, so we therefore chose to go with the higher MOI. Accordingly, we do not think this choice is a confounding issue. This explanation (line numbers 340-345) and a new Supplementary Figure 11 showing the results of the two MOI tests have been added to the manuscript.

      Reviewer #2 (Public review):

      The authors use a library of influenza A viruses from different strains, classified in lab-adapted, human, avian, and swine according to the animal from which they were isolated. They propose that the cow mammary gland serves as a mixing vessel for influenza A viruses. As a first approach, the authors assess susceptibility to infection across different cell types, including continuous and primary cell lines, bovine mammary cells, and mammary explants. All these cells support polymerase activity. Then, they analyzed changes in the bovine virus's viral fitness relative to an avian precursor. The authors use single-gene replacement to study whether and which RNP segments improve viral transcription. As part of this section, they also test IFN-specific antagonism by NS1 to assess the input of segment 8. Quantitative glycomic analysis was performed on the continuous bovine mammary cell line to demonstrate the presence of both a2,3 and a2,6, which is consistent with their observation that these cells can be co-infected with human and avian IAVs simultaneously. The main question, however, is: what is the glycome in the explants, or directly from tissues?

      We report quantitative glycomics for the primary bovine mammary epithelial cells as well as the continuous line the referee highlights. However, we agree with R2 that a detailed glycomic analysis of primary bovine mammary tissue would allow a better understanding of the actual glycosylation status in vivo. This has been undertaken by the authors and is available as a bioRxiv preprint. This is now cited (ref 25) in the relevant part of the results (line 184-185)

      Overall, the manuscript is clearly written and provides new insights into the behaviour of the cattle isolate, now compared with a representative group of model or precursor HAs of different origins.

      It would be great if a consistent nomenclature for the IAV strains could be used in the study. There is a mix of origin (Texas), animal from which the virus was isolated (mallard), or abbreviations that do not follow guidelines (IAV07). Are the USSR and Udorn not lab-adapted?

      We chose the abbreviated names for a variety of reasons. Partly from common usage (e.g. PR8, Udorn), partly for consistency with other already published papers from the FluTrailMap consortia (e.g. Cattle Texas; Dholakia et al 2026), partly to make diversity obvious in certain figures (e.g. H3N1, H5N2 etc) and partly to avoid confusion between viruses that originate from the same geographic area (e.g. AIV07, AIV09, H5N8-20 etc which are all A/Ck/England/isolate numbers). Overall, we found it more confusing to use the expanded nomenclature. Re AIV07 which the referee criticises for not following naming guidelines – if this is a reference to the EURL nomenclature, AIV07 is the abbreviation for the specific virus A/Chicken/England/053052/2021, our representative virus for EURL genotype EA-2020-C, as we say in the text. This nomenclature has now been added to Table 1, to provide a fuller cross-reference for all the names.

      As to whether USSR and Udorn are lab-adapted – that depends on definitions. There is a continuum of adaptive changes and/or sequence drift starting from the very first growth cycle of an isolate in the laboratory. The viruses we define here as lab adapted are ones that have been deliberately adapted to other host species or which have very long passage histories in multiple laboratory systems resulting in known functionally significant changes; for example, one lineage of PR8 was passaged 77 times in mice, 717 times in cell culture, 30 times in chick embryos, 5 times in ferrets and a further 50 times in chick embryos (https://www.medscape.com/viewarticle/812621_3?form=fpf), rendering it unarguably lab-adapted. We admit that A/USSR/77 and A/Udorn/307/1972 are probably further along this adaptive pathway than more recent isolates such as A/Norway/3433/2018, but are unaware of any specific reason that would put them into our lab-adapted category.

      The experimental setup includes bovine mammary primary and continuous cells, as well as mammary explants. Some of the most significant differences, for example, in viral fitness studies and co-infection experiments, are observed in these explants. Perhaps there could be some additional focus on this observation. The implications in comparison to the results obtained in cultured cells could be described. How will the human and other HA subtype viruses fare in the explants?

      We agree that this is an important and interesting question, and had already tested the strains we used for co-infections: human seasonal pdm09 H1N1 “Norway” and low pathogenic avian influenza “H3N1”, in the mammary explants. Both replicate the avian virus to 20-fold higher titres. We have added this information to the revised manuscript as new Figure panels S6E-H, called out on line 200-201 of the results.

      Reviewer #3 (Public review):

      Summary:

      This excellent manuscript by Pinto, Sharp, and colleagues examines bovine tissue tropism for influenza viruses. They find that bovine flu, as well as other strains, has strong replication in mammary tissue. They also map the genetic changes to influenza that improve replication in bovine cells. Overall, the study is well designed and executed, and the results are very timely.

      Strengths:

      (1) The experiments are well-controlled.

      (2) The figures are well-constructed and easy to follow.

      (3) The Methods and legends are detailed, with sufficient information.

      Weaknesses:

      (1) A comparison to human cells would strengthen the overall impact of the results. Are human mammary cells also uniquely susceptible to influenza? Are bovine mammary cells special in some way?

      This is an interesting question, but we have not tested mammary gland cells from humans (or any other species of mammal). We have however reported elsewhere (Dholakia et al., Nat Commun. 2026 Jan 16;17(1):1603. doi: 10.1038/s41467-026-68306-6.) that Cattle Texas grows well in a variety of human respiratory cells. Here, we are considering the bovine mammary organ as a potential reassortment site for IAVs because of the ongoing viral mastitis epidemic in US dairy cattle; human mammary organs seem unlikely to create a similar opportunity.

      (2) For the virus infection studies with segment 8 swaps, it should at least be noted that some of the phenotypes could be driven by NEP.

      We agree; as Table S1 indicates, NEP has two changes (one shared with NS1) between AIV07 and our B3.13 isolate, so we should not have conflated segment and NS1. We have changed the text to acknowledge this throughout the results (lines 127, 137 and 149) and in the discussion (line 243-244).

      (3) The data demonstrating that bMEC can support co-infection are compelling and important, but would be strengthened with a comparison from a different cell type or species. Do mammary cells uniquely support higher co-infection?

      We have data showing that co-infection also occurs in the continuous MAC-T udder cell line and have now included these data in a revised Figure 4D (described/called out on lines 198-206). We have not tested bovine cells from other organs for co-infection potential as they do not seem to be significant sites of infection in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) How nasal turbinate and cardiac fibroblasts are acquired/cultures is missing from the methods.

      Apologies for the omissions and thank you, because rectifying this brought to light an error in cell naming. The cells originally called bovine cardiac fibroblasts were in fact a second independent preparation of skin fibroblasts. We have corrected the labelling in Figs 1, 2 and S2. The nasal turbinate cells were bought in from ATCC (code CRL-1390). This information has now been added to the methods (line 310).

      (2) Please specify what cell types make up the 2D enteroids, the mammary explants and the nasal turbinates.

      The composition of the 2D enteroids is described in detail in reference [60]. In precis, they are comprised predominantly of epithelial cells, including Paneth, goblet and enteroendocrine cells, as well as stem cells. This information has been added to the methods (lines 374-376) The mammary explants include duct epithelium, connective and muscle tissue, defined by H&E staining of cut sections (see new Figure S12 and text added to the methods on lines 384-386). We have also added the person who did the histology (Rebecca Ross) as an author and to the credit taxonomy (line 877). The nasal turbinate cells appear to be predominantly fibroblast morphology (Methods line 310-311).

      (3) Epithelial cells are the main target for IAV, so why were fibroblasts chosen as a comparison to mammary epithelial cells? This needs justification in text. Without justification, how much can be attributed to the results being mammary-specific, rather than epithelial-specific? The brain (choroid plexus epithelium), heart (epicardium) and skin all contain epithelial cells.

      We think this query is what the referee calls “Q4’ in the public part of their review. Please see our answer above.

      (4) Figure 1, S1, S2 and S3 seem to suggest that mammary cells are more susceptible to IAV infection than cells from other organs. But Figure 2 demonstrates that when it comes to Cattle Texas and AIV07, most of the cell types show high viral titers. If the question is about whether the cow udder is the primary mixing site, would it not be more relevant to investigate which cell type facilitates the best growth of the potential precursor viruses, similar to Figure 4A?

      Figs 1, S1, 2 and 3 all use “full” viruses whereas Fig 2 uses 6:2 reassortants between PR8 and the HPAIVs for biosafety reasons. WT PR8 replicates well in most of the bovine cells tested (Figs S1-3) so we do not see any contradiction. Figure 2 examines the contributions the internal genes make to replication in bovine cells. Re the question over the udder being a potential mixing site – this is where the virus is replicating in the real world; at least in part because of the transmission mechanism, but also because the mammary gland epithelium is highly susceptible to infection, as we show.

      (5) It is misleading to compare the bMEC infection of MOI 0.1, to all other infections of MOI 0.01. This is a consistent problem throughout the article - Figures 1, S1, 2, 4. Why was the bMEC infection at a 10x greater dose? The main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. If we compare the MOI 0.01 experiments alone, then the evidence relies on the immortalised MAC-T cells, compared to primary cell types. In this case, how much can be said about it being mammary specific, rather than immortalised vs primary? I do wonder, for example, how the epithelial nasal turbinates or type II pneumocytes would compare to the primary mammary epithelial cells if they were at the same MOI.

      Please see our answer to this query earlier in the rebuttal.

      (6) Why are the 2D enteroids excluded from Figure 1?

      We had limited supplies of a difficult-to-grow cell model, so we only used them to test the 6:2 viruses (Fig 2).

      (7) The colour scheme for Figure 3 is confusing. In Figure 3A, blue indicates European ancestry, and yellow represents North American ancestry. However, in Figure 3B, these colours now mean something different. To a reader, when there is a colour-coded schematic, it is instinctual to think that this then corresponds to the following panel(s). Since consistently throughout the article, yellow has been used for Cattle Texas, and blue has been used for AIV07, I would suggest choosing different colours to represent European and North American ancestry in Figure 3A

      We’ve changed the figure as the referee suggests and modified the Fig 3 legend accordingly (line 895)

      (8) I am unsure about the conclusions drawn from the results of Figure 3B. In the results, it is framed as trying to determine which segments contributed to the improved activity of Cattle Texas compared to AIV07. In lines 110-111, "... PB2 or PA from AIV07 significantly decreased Cattle Texas minireplicon activity". If PB2 is indeed significant, there is a missing yellow asterisk in Figure 3B.

      Apologies, there was indeed a missing asterisk on the figure; now added.

      Given the significance of PA, why was it not investigated in terms of growth kinetics similar to Figure 3C? Was it overlooked because it doesn't have a North American ancestry? The results of Figure 3B suggest that the 4 amino acid mutation in PA has significantly contributed to changes in polymerase activity.

      The PA changes do indeed matter for minireplicon activity – the key change is K497R, as detailed in our related publication in Nat Comms (citation 17). However, it is less important than changes in PB2, and the PA segment swap by itself has little effect on overall virus replication.

      (9) Similarly, in Figure 3C and lines 116-117, the error bars on the graph are overlapping at 48 hours, suggesting no difference in overall replication. The kinetics are slowed for AIV07 seg1-3, but not for AIV07 seg 1, indicating PB1 does not have an effect. This would then suggest that something in segment 2 or 3 contributes to the slowed kinetics in Figure 3C, which, from the Figure 3B results, is unlikely to be due to PB2. While it was not reassorted, the PA segment is potentially the driver, with its 4 amino acid mutations. I think it is worth performing growth kinetics with and without these 4 amino acid changes in PA.

      We agree that visually on a log<sub>10</sub> scale, the titres of the “WT” 6:2 Cattle Texas and 5:2:1 segment 1 reassortment appear close, but the average titres are 5 and 7-fold different at 24 and 48h respectively, while a 2-way ANOVA with Dunnet’s multiple comparison post-test gives statistical significance at 48h. We have added this information to the figure and its legend (lines 905-906).

      (10) In Figure 3C and E, why did you choose to perform the growth kinetics in the immortalised cell line, when you have access to primary cells? The primary cells would be a more accurate representation of what happens in situ.

      The primary cells were difficult to work with and only available intermittently, so we used what was available at the time.

      (11) In lines 145-147, "thus overall, the reassortment event that replaced segments 1, 2 and 8 alongside drift adaptations in segment 3 may have contributed to the ability of the B3.13 genotype virus to infect cattle". This is not clearly supported by the evidence presented. In terms of segment 1/PB2, the growth kinetics of Figure 3C have overlapping error bars at 48 hours. Where is any evidence presented for the role of segment 2/PB1? There is no change in Figure 3B.

      The referee is correct, calling out seg2 here was an error; we have revised the text (line 144).

      Segment 3 is overlooked in Figure 3 (as highlighted in Q9 and 10), and shows no difference in Figure S4.

      Please see response to Q8; we think segment 3 contributes via PA adaptation, not via PA-X.

      (12) In line 146 "... drift adaptation in segment 3". Make it clear here that you are talking about genetic drift. However, is this likely to be genetic drift? The 4 amino acid mutations are shown to have a significant impact on polymerase activity in Figure 3B, and in Figure 3C, PA potentially contributes to the reduced kinetics. When there are amino acid mutations that correspond to a beneficial phenotypic change, attributing this to drift alone rather than host adaptation is strange.

      Yes, wording clarified (line 145). “Drift” was used to distinguish it from reassortment but we agree this was an incorrect term in the context.

      (13) Figure 4A MAT-C cells: this is ostensibly the same experiment as Figure 3C in terms of the Cattle Texas and AIV07 viruses. If this is the case, how can you explain the difference in kinetics and overall titer? In Figure 3C, Cattle Texas reaches 10^6, and in Figure 4A it reaches almost 10^9. That's almost 3 log difference. Similarly, in Figure 3C Cattle Texas reaches 10^3, but in Figure 4A it reaches 10^6, a 3-log difference. At 24 hrs, they have roughly a 3-log difference between them in Figure 3C, but in Figure 4A this difference is much smaller. As far as I can tell, these are the same viruses, same dose and same cell model. The t0 titer is also vastly different between the two experiments.

      The experiments were done at different times (several months apart, so different cell passage numbers and/or serum batches) and by different people. We have no explanation other than biological variability. However, both groups of experiments include genuine biological replicates done over the course of 2-3 weeks, so in our view represent coherent tests within themselves.

      (14) In Figure 4, why weren't the a-2,6 a-2,3 proportions analysed for the explants and/or used for the co-infection experiment? It showed the greatest difference between the Cattle Texas and precursor viruses in Figure 4A.

      Our data on the proportion of 2,6 and 2,3 SA in bovine udder tissue are now available in a separate preprint (now cited as [25] in our MS). We did not use the explants for co-infection experiments because it would have been technically difficult to read out the outcome by flow cytometry.

      (15) Figure 4D requires a supplementary figure demonstrating the gating strategy, including one of the samples as an example.

      We have compiled a figure of this and added it as new Figure S13 (called out line 556).

      (16) In Figure 4D, why was an MOI of 5 chosen instead of the MOI of 0.01 used throughout the article for MAC-T cell infection? An MOI of 5 (so in a co-infection, a total of 10 virus particles per cell) is completely overloading the cells. At this dose, 10% of the cells were able to be co-infected, but how representative is this of a real co-infection scenario? While it demonstrates it is possible, it potentially remains highly unlikely, similar to the discussion around the swine respiratory tract in lines 235-240.

      We had also performed the co-infections at lower MOI (1) with very similar results – this is now included in Figure 4D. Furthermore, we redid the experiments at a lower MOI of 0.05 and still see co-infection; this now replaces the MOI 5 data in Figure 4D. The text has been revised accordingly (lines 202-205)

      (17) In Figure 4D, why were immortalised cells used when primary bMEC and mammary explants are available? Primary cells would provide more convincing evidence for the potential of the cow udder to be a mixing vessel. Considering that throughout the paper, a 10x higher viral dose is used in the bMEC culture, I wonder if you would need a significantly higher MOI than 5 to produce similar results in a co-infection experiment. The bMEC also has a more even a-2,3 to a-2,6 ratio compared to MAC-T in Figure 4C.

      The bMECs in the original figure are primary cells. In response to other queries, we now include data from the immortalised MAC-T cell line as well.

      Reviewer #2 (Recommendations for the authors):

      Figure 1A, the coloured underline to discriminate continuous and primary cells is lost upon printing... perhaps another way is better?

      We have changed the primary cells to italic text to make the distinction clearer.

      We have made some other minor changes to wording to correct grammatical errors or improve clarity as we went through.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper proposes a non-decision time (NDT)-informed approach to estimating timevarying decision thresholds in diffusion models of decision making. The manuscript motivates the method well, outlines the identifiability issues it is intended to address, and evaluates it using simulations and two empirical datasets. The aim is clear, the scope is deliberately focused, and the manuscript is well written. The core idea is interesting, technically grounded, and a meaningful contribution to ongoing work on collapsing thresholds.

      Strengths:

      The manuscript is logically structured and easy to follow. The emphasis on parameter recovery is appropriate and appreciated. The finding that the exponential NDT-informed function produces substantially better recovery than the hyperbolic form is useful, given the importance placed on identifiability earlier in the paper. The threshold visualisations are also helpful for interpreting what the models are doing. Overall, the work offers a well-defined, methodologically oriented contribution that will interest researchers working on time-varying thresholds.

      We appreciate the positive and constructive feedback. We have addressed your comments in the revised manuscript, as detailed in our responses to the specific comments below.

      Weaknesses / Areas for Clarification:

      A few points would benefit from clarification, additional analysis, or revised presentation:

      (1) It would help readers to see a concrete demonstration of the trade-off between NDT and collapsing thresholds, to give a sense of the scale of the identifiability problem motivating the work.

      Thank you for this constructive suggestion. We conducted a new simulation study in which we considered the non-decision time as a fixed parameter and estimated the remaining parameters of the collapsing threshold model. In this simulation study, we contaminated the non-decision time with noise at six levels (i.e., 0%, 2%, 4%, 6%, 8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters. The results of this simulation are presented in Appendix 1 in the new version of the manuscript.

      (2) Before moving to the empirical datasets, the manuscript really needs a simulation-based model recovery comparison, since all major conclusions of the empirical applications rely on model comparison. One approach might be to simulate from (a) an FT model with across-trial drift variability and (b) one of the CT models, then fit both models to each of the simulated data sets. This would address a longstanding issue: sometimes CT models are preferred even when the estimated collapse in the thresholds is close to zero. A recovery study would confirm that model selection behaves sensibly in the new framework.

      We are grateful for this constructive comment. The revised manuscript includes a model recovery study (see Simulation study 2, in particular Table 1 and the accompanying text). In this simulation, as the reviewer suggested, we generated data from FT-DDM with across-trial variability in drift rate and CT-DDM with hyperbolic and exponential collapsing thresholds. Then, for each generated dataset, we fitted two models and classified the datasets based on goodness-of-fit and estimated decay rate. The results show that incorporating the decay rate, which can be reliably estimated by the NDT-informed modeling framework, significantly improves the precision of model recovery.

      (3) An additional subtle point is that BIC is defined in terms of the maximised log-likelihood of the model for the data being modelled. In the joint model, the parameter estimates maximise the combined likelihood of behavioural and non-decision-time data. This means the behavioural log-likelihood evaluated at the joint MLEs is not the behavioural MLE. If BIC is being computed for the behavioural data only, this breaks the assumptions underlying BIC. The only valid BIC here would be one defined for the joint model using the joint likelihood.

      We thank the reviewer for raising this important methodological point. We agree that, strictly speaking, the behavioral log-likelihood evaluated at the joint maximum likelihood estimates is not guaranteed to be equal to the behavioral maximum likelihood estimate. Therefore, if one were to interpret BIC<sub>Behavior</sub> for joint (NDT-informed) models as a conventional BIC derived from a purely behavioral maximum likelihood fit, this would indeed violate the standard assumptions underlying BIC. Our intention, however, was not to claim that BIC<sub>Behavior</sub> represents a formally valid BIC in the strict information-theoretic sense. Rather, we used it as a diagnostic measure to assess how well the jointly estimated parameters account for the behavioral data relative to the uninformed models. Importantly, in our datasets, the behavioral log-likelihood evaluated at the joint estimates is higher than that obtained from the uninformed models. This suggests that the additional constraint introduced by the non-decision time information helps guide the optimization procedure toward parameter regions that provide a better account of the behavioral data. In other words, uninformed models appear to converge to suboptimal parameter estimates, but the joint modeling framework helps regularize the estimation process and yields better estimates of the optimal parameters. We fully acknowledge that this use of BIC<sub>Behavior</sub> for the joint models departs from standard practice in the cognitive modeling literature. For this reason, we rely primarily on the joint likelihood–based model comparison as the formally valid criterion. The behavioral BIC is reported only to provide additional intuition regarding the goodness of fit on behavioral data and is used exclusively to compare NDT-informed models with their uninformed counterparts under identical evaluation criteria.

      Also, to warn readers about this limitation, we included the following statement in the section where we defined BIC measures:

      “It is worth noting that comparing joint and behavioral models using BIC<sub>Behavior</sub> is uncommon and constitutes a limitation of the model comparison study.”

      (4) Table 1 sets up the Study 1 comparisons, but there’s no row for the FT model. Similarly, Figures 10 and 13 would be more informative if they included FT predictions. This matters because, in Study 1, the FT model appears to fit aggregate accuracy better than the BIC-preferred collapsing model, currently shown only in Appendix 5. Some discussion of why would strengthen the argument.

      In the revised manuscript, we have included the results for NDT-informed FT-DDM in the main text. However, we kept the FT-DDM with drift variability in the appendix, since none of the models in the main text include drift rate variability.

      (5) In Figure 7, the degree of decay underestimation is obscured by using a density plot rather than a scatterplot, consistent with the other panels of the same figure. Presenting it the same way would make the mis-recovery more transparent. The accompanying text may also need clarification: when data are generated from an FT model with across-trial drift variability, the NDT-informed model seems to infer FT boundaries essentially. If that’s correct, the model must be misfitting the simulated data. This is actually a useful result as it suggests across-trial drift variability in FT models is discriminable from collapsing-threshold models. It would be good to make this explicit.

      Regarding the visualization of decay rate estimation in Figure 7, we would like to clarify that, in the data-generating process for the FT model with across-trial drift variability, the decay rate is fixed at zero. That is, unlike the other parameters, there is only a single true value on the x-axis. If we were to present the decay rate recovery using a standard scatter plot (true vs. estimated values), all points would lie vertically above the single true value (zero), resulting in a vertical strip of points. While such a plot would technically be consistent with the other panels, it would not clearly convey the distributional properties of the estimated decay rates, specifically, which values are more likely under model misidentification. For this reason, we chose a density plot to more transparently illustrate the distribution of inferred decay rates when the true generating process is an FT model. We believe this representation more effectively communicates the extent and structure of mis-recovery. To avoid confusion, we have added a clarifying footnote in the revised manuscript explaining why a density plot was used in this specific panel.

      Regarding distinguishing fixed-threshold models from collapsing-threshold models, as suggested in your comment 2, we conducted an additional simulation, and the results showed that when using the NDT-informed diffusion model, we can distinguish between both models with high accuracy. Specifically, we showed that incorporating the estimated decay rate value provided by the NDT-informed modeling approach can significantly improve the model recovery accuracy.

      (6) Given the large recovery advantage of the exponential NDT-informed function over the hyperbolic one, the authors may want to consider whether the results favour adopting the former more generally. Given these findings, I would consider recommending the exponential NDT-informed model for future use.

      Consistent with the reviewer’s argument, we included the following text in the general discussion:

      “Importantly, the exponential collapsing threshold exhibited substantially better parameter recovery and superior model recovery performance, suggesting that this specification may be preferable in future cognitive modeling applications.”

      (7) In Study 2 (Figure 13), all models qualitatively miss an interesting empirical pattern: under speed emphasis, errors are faster than corrects, while under accuracy emphasis, errors become slower. The error RT distribution in the speed condition is especially poorly captured. It would be helpful for the authors to comment, as it suggests that something theoretically relevant is missing from all models tested.

      Thank you for mentioning this point. Because the models considered in the main text do not include across-trial variability in the starting point, the model cannot predict the fast error pattern observed in the speed condition of the second study. In the revised manuscript, we included a note on this point:

      “The NDT-informed models’ predictions are depicted in Figure 13. This figure shows that the FT-DDM overestimates the last RT quantiles for both correct and incorrect responses in both speed and accuracy conditions. However, the qualitative predictions of CT-DDMs align more closely with the empirical data. It is also worth noting that all the considered computational models misfit the incorrect responses in the speed condition. This misfit is to be linked to the presence of fast errors, specifically in the speed condition. Including the starting-point variability parameter in the model enables the model to predict fast errors. However, as the aim here was not merely to fit the data with the best possible model, but to test the NDT-informed modeling framework, we did not include starting-point variability in the model.”

      (8) The threshold visualisations extend to 3 seconds, yet both datasets show decisions mostly finishing by 1.5 seconds. Shortening the x-axis would better reflect the empirical RT distributions and avoid unintentionally overstating the timescale of the empirical decision processes.

      In the new version of the manuscript, we shortened the x-axis in these plots. See Figures 9 and 12 in the new version of the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript should explicitly state how critical the log-normal assumption for NDT is, and whether there are caveats if it doesn’t hold.

      We agree with this suggestion. Therefore, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and replicated the main results reported in simulation study 1 (see Appendix 8 in the new version of the manuscript). These simulation results reveal that independent of distributional assumptions on non-decision time measurements, constraining non-decision time can improve the estimation of the collapsing threshold.

      (2) On page 14, one paragraph refers to five models and another to six. It is unclear which is correct.

      We are sorry for the confusion. In the new version of the manuscript, we addressed this issue.

      (3) In the Discussion: Hawkins & Heathcote (2021) found that NDT estimates in the TRDM recover well, but the timer-offset parameter does not (and hence is set to a fixed value). It would be interesting to test whether the NDT-informed approach could extend to that parameter. The TRDM can also generate error RT distributions that are faster than correct RT distributions, which is the qualitative pattern that none of the present models capture in the speed condition of Study 2.

      We appreciate the reviewer’s suggestion, as it would be another important demonstration of the framework we built in the paper. Nevertheless, given the number of models already considered in the manuscript and the focus on collapsing threshold diffusion models, we believe that this addition would blur the focus of the present manuscript. Therefore, we have decided not to include TRDM in the manuscript. Nevertheless, we have addressed the reviewer’s concern regarding the non-decision time estimation in the TRDM in the revised manuscript:

      “Notably, this model can also predict faster error responses than the correct response, the pattern that is observed in the speed condition of Study 2. However, the authors reported poor parameter recovery for the onset of the timing process (Hawkins and Heathcote, 2021). Thus, the reliability issue here specifically concerns the estimation of the shift parameter of the timing accumulator. Informing the model with external estimates of non-decision time might, therefore, improve parameter recovery in this model as well.”

      Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model on estimates of single-trial non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and, in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      We appreciate your feedback and comments. Below, we provided a response for each comment.

      Weaknesses:

      (1) This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model? Unfortunately, it is not clear whether and how the model structure itself differs between the NDTinformed and non-NDT-informed cases. Ideally, they are the same actual model, but with one getting extra guidance on where to place the tau and/or sigma parameters from external measurements. The absence of sigma (non-decision time variance) estimates for the non-NDT-informed model, however, suggests it is different in structure, not just in its lack of constraints. If they were the same model, whether they do or do not possess non-decision time variability (which is not currently clear), the only possible reason that the NDT-informed model could achieve better BIC is because the non-NDT-informed model gets lost in the fitting procedure and fails to find the global optimum. If they are in fact different models - for example, if the NDT-informed model is endowed with NDT variability, while the non-NDT-informed model is not - then the fit superiority doesn’t necessarily say anything about an NDT-informed reliability boost, but rather just that a model with NDT variability fits better than one without.

      To respond to this comment, we would like to note that the structural difference between NDT-informed and uninformed models is the assumption about non-decision time. In principle, the behavioral parts (i.e., parameters related to the diffusion part) of both NDT-informed and uninformed models are identical. However, during the estimation, the non-decision time in the NDT-informed model is subject to an additional constraint imposed by the neural data. In other words, in the uninformed model, non-decision time is estimated using the behavioral data by maximizing the likelihood of a CT-DDM. However, the NDT-informed model incorporates an additional data type, resulting in a different likelihood function (see Equation (3)). Specifically, in the NDT-informed model, we make an additional assumption regarding the non-decision time: it is set to the mean of the trial-level non-decision time measurements distribution derived from neural data (Equation (3) specifies the joint model structure). This additional assumption constrains the search space of the non-decision time parameter in the NDT-informed model. As in the main text, we assumed that non-decision time measurements obtained from neural data are log-normally distributed, where sigma is the shape parameter of the log-normal distribution. Therefore, sigma can represent the variability of the non-decision time measurements, and it is not the trial-to-trial non-decision time variability parameter. Indeed, the non-decision time parameter is fixed across trials in both models. Also, it should be clear from Equation (3) that sigma only appears in the second term of the joint likelihood, which corresponds to non-decision time measurements and not the CT-DDM term. Therefore, the behavioral parts of the NDT-informed models and Uninformed models are identical, and the NDT-informed models include one additional parameter (i.e., sigma) corresponding to the variability in non-decision time measurements.

      Also, to explain how constraining non-decision time improves the BIC, we would like to clarify that constraining non-decision time using an additional data source constrains the search space for non-decision time, thereby leading to better parameter identification and, consequently, a better fit to behavioral data. We included the following text in the discussion section to make this explicit.

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) One reason this is unclear is that Footnote 4 says that this study did not allow trial-to-trial variability in nondecision time, but the entire premise of using variable external single-trial estimates of nondecision times (illustrated in Figure 2) assumes there is nondecision time variability and that we have access to its distribution.

      We are sorry for the confusion. To respond to this comment, we would like to highlight that this modelling approach does not include any mechanism for across-trial variability in the behavioral part, as the likelihood of the choice behavior (see Equation (3)) does not include any variability parameter. As mentioned before, sigma represents the variability in non-decision measurements. However, the model does not propagate the across-trial variability of the neural data on the behavior side (unlike the models in Ghaderi-Kangavari et al. 2023). In other words, in this approach, none of the diffusion model’s parameters include across-trial variability, and we have only considered and estimated neural data variability in the NDT-informed model. Also, to improve the manuscript’s coherence, we removed this footnote.

      (3) It is good that there is an Intro section to explain how the tradeoff between NDT and collapsing bound parameters renders them difficult to simultaneously identify, but I think it needs more work to make it clear. First of all, it is not impossible to identify both, in the same way as, say, pre- and postdecisional nondecision time components cannot be resolved from behaviour alone - the intro had already talked about how collapsing bounds impact RT distribution shapes in specific ways, and obviously mean (or invariant) NDT can’t do that - it can only translate the whole distribution earlier/later on the time axis. This is at odds with the phrasing “one CANNOT estimate these three parameters simultaneously.” So it should be first clarified that this tradeoff is not absolute. Second, many readers will wonder if it is simply a matter of characterising the bound collapse time course as beginning at accumulation onset, instead of stimulus offset - does that not sidestep the issue? Third, assuming the above can be explained, and there is a reason to keep the collapse function aligned to stimulus onset, could the tradeoff be illustrated by picking two distinct sets of parameter values for non-decision time, starting threshold, and decay rate, which produce almost identical bound dynamics as a function of RT? It is not going to work for most readers to simply give the formula on line 211 and say ”There is a tradeoff.” Most readers will need more hand-holding.

      We are grateful for this comment. In response to this comment, which was also partly mentioned by the first reviewer, we first highlight that, in the presence of a nonlinear collapsing threshold, the effect of non-decision time is no longer linear, as it forces the threshold to take a specific value at the final stopping point. To illustrate the tradeoff between imprecise non-decision time estimation and collapsing threshold estimation, we conducted a simulation study. The results for the simulation study are presented in Appendix 1. In this study, we contaminated the non-decision time with noise at six levels (i.e., 0%,2%,4%,6%,8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters.

      Regarding the second point, we would like to clarify that in this paper, consistent with other works on collapsing threshold, we assumed that the collapsing starts with evidence accumulation and not with stimulus onset, as the decision makers need a short amount of time for perceiving and encoding the stimulus (i.e., perceptual encoding time). Fixing the collapsing onset to the stimulus onset introduces an ad hoc assumption into the model (that the encoding time is zero), which is cognitively implausible. Therefore, although this assumption can mathematically resolve the issue, it is not cognitively plausible. However, one important point we did not consider in the previous version is the two-stage accumulation process models in which the collapsing onset occurs later than the evidence-accumulation onset. We have discussed the estimation of such models as a limitation in the general discussion:

      “The collapsing threshold dynamics considered in this work (e.g., exponential and hyperbolic) impose a monotonically decreasing threshold over time. Although these dynamics are theoretically well motivated (Fudenberg et al., 2018; Frazier and Yu, 2007) and have been employed in several previous studies (e.g., Olschewski et al., 2025; Milosavljevic et al., 2010; Voskuilen et al., 2016), some research has proposed delayed-collapsing threshold models (e.g., Diederich and Oswald, 2016), in which the onset of threshold collapse does not coincide with the onset of evidence accumulation. Estimating such delayed-collapsing models may require more than simply constraining non-decision time, as the onset of threshold collapse must also be identified. A promising approach for addressing this challenge is the HMP method, which may provide additional temporal information about distinct cognitive processing stages. In particular, HMP may allow the onset of threshold collapse to be estimated as a separate cognitive stage. Future research should therefore investigate the estimation of multi-stage evidence accumulation models (e.g., Diederich and Oswald, 2016; Diederich and Colonius, 2021) within the HMP framework.”

      (4) A lognormal distribution is used as line 231 says it “must” produce a right-skew. Why? It is unusual for non-decision time distribution to be asymmetric in diffusion modeling, so this “must” statement must be fully explained and justified. Would I be right in saying that if either fixed or symmetrically distributed nondecision times were assumed, as in the majority of diffusion models, then the non-identifiability problem goes away? If the issue is one faced only by a special class of DDMs with lognormal NDT, this should be stated upfront.

      We would like to clarify that this assumption is about the non-decision time measurements and not about the across-trial variability parameter of non-decision time. Although for computational convenience, non-decision time is often assumed to follow a normal distribution in diffusion modelling literature, some empirical studies have shown that the perceptual encoding time and total approximated non-decision time follow a right-skewed distribution. Moreover, as HMP estimations of non-decision time are usually right-skewed, we formalized the model with a log-normal distribution. However, we would like to note that the identifiability issue in collapsing-threshold diffusion models is not related to the distributional assumption over the non-decision time measurements, as poor parameter recovery of the collapsing threshold was also reported in Evans et al. (2020). To show that the distributional assumption about the non-decision time measurements does not affect the results, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and showed that constraining non-decision time improves parameter recovery for collapsing threshold parameters. The results for this simulation are presented in Appendix 8 in the new version of the manuscript. We also revised line 231 as follows (see line 236 in the new version):

      “Empirical studies on non-decision time measurement usually have reported a right-skewed distribution for their measurements (e.g., Weindel, 2021; Weindel et al., 2025). For instance, the measured perceptual encoding time and motor execution time reported by Weindel et al. (2025) are right-skewed. Therefore, we assume that the observed non-decision time measurements Z<sub>n</sub> follow an approximate log-normal distribution, which is right-skewed (we will discuss how this distributional assumption can affect the results later). Thus, we model the non-decision time measurements Z<sub>n</sub>, using a log-normal distribution with parameters µ and σ<sub>z</sub>. The available measurements (i.e., observed data) at trial n can be represented as

      follows:”

      (5) In the simulation study methods, is the only difference between NDT-informed and non-informed models that the non-NDT-informed must also estimate tau and sigma, whereas the NDT-informed model “knows” these two parameters and so only has the other three to estimate? And is it the exact same data that the two models are fit to, in each of the simulation runs? Why is sigma missing from the uninformed part of Figure 4? If it is nondecision time variability, shouldn’t the model at least be aware of the existence of sigma and try to estimate it, in order for this to be a meaningful comparison?

      As mentioned in the response to your first comment, the difference between NDT-informed and uninformed models is the access to an additional source of data related to non-decision time, and sigma represents the shape parameter of the non-decision time measurements distribution. Therefore, sigma belongs only to the NDT-informed model, and, as in the uninformed model, there is no additional data source, so the model does not include sigma.

      (6) I am curious to know whether a linear bound collapse suffers from the same identifiability issues with NDT, or was it not considered here because it is so suboptimal next to the hyperbolic/exponential?

      Thank you very much for this comment. The main reason we did not include linear collapsing threshold models was the assessment by Evans et al. (2020), which indicated that, with sufficient trials, these models can be estimated reasonably well. However, to investigate whether constraining non-decision time can also improve the estimation of linear collapsing threshold models, we conducted an additional simulation study, which is reported in Appendix 3 in the new version of the manuscript. The simulation results confirm those reported by Evans et al. (2020) and show that the parameters of the uninformed linear models can be identified using more than 500 trials. Constraining the non-decision time using the NDT-informed diffusion modeling framework still improves parameter estimation in the linear collapsing threshold model and reduces the required number of trials for reliable estimation to 250. Especially, the estimation of the starting threshold improves significantly.

      (7) The approach using HMP rests on the assumption that accumulation onset is marked by the peak of a certain neural event, but even if it is highly predictive of accumulation onset, depending on what it reflects, it could come systematically earlier or later than the actual accumulation onset. Could the authors comment on what implications this might have for the approach?

      Thank you for mentioning this point. We first would like to point out that Weindel et al. (2024) showed that HMP can predict the underlying generative distribution of cognitive states with high precision and without systematic bias. Second, it is worth highlighting the results reported in Appendix 5 (i.e., “Bias in non-decision time measurements”). In this appendix, we discussed how bias in non-decision time measurements (i.e., systematic underestimation or overestimation) affects the estimation of the collapsing threshold. Particularly, see Figure 3 in Appendix 5. To make these results clearer in the main text, we included the following paragraph at the end of the results section in simulation study 1:

      “Additionally, we examined the effect of systematic bias in non-decision time measurement on parameter recovery. Appendix 5 presents the parameter recovery simulation results in the presence of biased non-decision time measurement (i.e., systematically underestimated or overestimated). The results revealed that, even in the presence of biased non-decision time measurement, the actual generating parameters show a high correlation with the estimated parameters. Underestimation in non-decision time leads to overestimation in the starting threshold and decay rate. Conversely, overestimation in non-decision time leads to underestimation in both the starting threshold and non-decision time.”

      (8) Figure 7: for this simulation, it would be helpful to know the degree to which you can get away with not equipping the model to capture drift rate variability, when the degree of that d.r. variability actually produces appreciable slow error rates. The approach here is to sample uniformly from ranges of the parameters, but how many of these produce data that can be reasonably recognised as similar to human behaviour on typical perceptual decision tasks? The authors point out that only 5% of fits estimate an appreciable bound collapse but if there are only 10% of the parameter vectors that produce data in a typical RT range with typical error rates etc, and half of these produce an appreciable downturn in accuracy for slower RT, and all of the latter represent that 5%, then that’s quite a different story. An easy fix would be to plot estimated decay as a scatter plot against the rate of decline of accuracy from the median RT to the slowest RT, to visualise the degree to which slow errors can be absorbed by the no-dr-var model without falsely estimating steep bound collapse. In general, I’m not so sure of the value of this section, since, in principle, there is no getting around the fact that if what is in truth a drift-variability source of slow errors is fit with a model that can only capture it with a collapsing bound, it will estimate a collapsing bound, or just fail to capture those slow errors.

      Thank you for this comment. We would like to first note that the aim of this section is to illustrate that NDT-informed modeling enables us to distinguish between CT-DDM and FT-DDM with drift variability (i.e., the two competing models that are relatively hard to distinguish). Therefore, we changed the name of this section to “Simulation study 2: Model recovery”. In the new version of the manuscript, we included a cross-model fitting simulation and a model recovery simulation in this section. The estimated decay rates in the cross-model fitting study indicate that, when the underlying generating model is FT-DDM with drift variability, parameter estimation using the NDT-informed approach yields precise inference. This is due to the estimated decay rate, which is very close to zero. The model recovery results also confirmed that incorporating the decay rate value into the model inference improves precision.

      Moreover, to address the concern regarding the slow-error pattern in the simulated data, we examined the relationship between the slow–fast accuracy difference and the estimated decay rate. Specifically, we computed the difference between the accuracy of responses with response times below the median (ACC1) and those above the median (ACC2), and plotted this difference against the estimated decay rate. Author response image 1 presents the resulting scatter plot, with color indicating the drift rate. As shown in Author response image 1, incorrect inferences about the decay rate primarily occur at high drift rates. This pattern emerges because high drift rates produce very fast responses, leaving little time for the threshold to meaningfully collapse. Consequently, the behavioral signatures of collapsing-threshold and fixed-threshold models become increasingly similar under high drift conditions.

      Author response image 1.

      Illustration of the difference between the accuracies of responses with response time below the median and those with response time above the median against the estimated decay rate. Colour shows the drift rate value.

      Reviewer #2 (Recommendations for the authors):

      (1) Abstract: improves fit to behaviour in what way? Reliability or absolute quant fit to behavior? I.e., is it just helping constrain it so it finds the global opt?

      (2) Line 72 - Are these “neuroimaging” studies? Perhaps use the broader term “neuroscience.”

      (3) Line 95 - It’s important because it may confuse readers how something dynamic like a collapsing bound could be resolved with, say, fMRI.

      (4) Line 85 – “greater variability”... in what?

      (5) Line 87 -Revise to avoid misconstruing a time on task effect - e.g., “higher error rate for trials with longer RT” is more explicit.

      (6) Line 89-94 - It’s not clear what findings are being referenced here.

      (7) Line 103 - External biases such as priors and relative value?

      (8) Line 175 - Clarify this applies to any ddm, not just ctddm.

      (9) Line 268 - Explain what portion N200 accounts for.

      (10) Line 327 - Is this equation supposed to be for Delta-X, as opposed to X(t+delta-t)? If you want X(t+delta-t) on the LHS, then X(t) must be added to the RHS.

      We are very thankful for such a precise evaluation and constructive comments. We addressed all the comments raised by the reviewer in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The current paper addresses an important issue in evidence accumulation models: many modelers implement flat decision boundaries because the collapsing alternatives are hard to reliably estimate. Here, using simulations, the authors demonstrate that parameter recovery can be drastically improved by providing the model with additional data (specifically, an EEG-informed estimate of nondecision time). Moreover, in two empirical datasets, it is shown that those EEG-informed models provide a better fit to the data. The method seems sound and promising and might inform future work on the debate regarding flat vs collapsing choice boundaries. As an evidence-accumulation enthusiast, I am quite excited about this work, although for a broader audience, the immediate applicability of this approach seems limited because it does require EEG data (i.e., limiting widespread use of the method or e.g., answering questions about individual differences that require a very large N).

      We are very grateful for your positive evaluation and your comments.

      Reviewer #3 (Recommendations for the authors):

      This is a very decent study, very well written and properly executed. Most of my comments below are suggestions for the authors as to how to make the story more compelling.

      (1) I think the authors can do more to explain why the NDT-informed models fit better to empirical data. If the NDT-estimates are equal to the ground truth, then isn’t it possible that such models win simply because they have one parameter less? However, this is not what the authors are claiming, though, on l.567 it says that the fit is better because parameters are better estimated. I think it should be possible to dissociate these two accounts.

      Thank you for this important comment. For clarification, both NDT-informed and uninformed models estimate the non-decision time parameter, and neither treats it as equal to the ground truth. However, the difference between these two models is that, in the NDT-informed model, the non-decision time is subject to an additional constraint; therefore, the number of behavioral parameters (i.e., parameters of the diffusion part) is identical for both NDT-informed and uninformed collapsing threshold models. Although the number of behavioral parameters is identical in the NDT-informed model, one parameter is constrained, thereby limiting the model’s complexity and flexibility compared to uninformed models. Consequently, the improvement in the fit can only be attributed to better parameter estimation in the NDT-informed models, resulting from constraining the non-decision time using neural measurements. To clarify that in the manuscript, we included the following text in the revised version:

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) Given that so much of the writing focuses on the conclusion that collapsing boundary models ¿ flat models, it is very odd that there is no comparison to a flat boundary model reported in the text. Why is the model in Supplement 5 not just included in the main text (and in Table 1)? This would make it so much easier for the reader.

      To address this comment and also the similar point raised by Reviewer 1, we included the NDT-informed FT-DDM in the results section and compared the FT-DDM with CT-DDMs with respect to BIC<sub>Joint</sub>. However, we retained the other FT-DDM in the appendix because this model includes across-trial variability in drift rate, whereas the models in the main text do not.

      (3) I would have appreciated a bit more background about the importance of the number of trials per participant. Given that the proposed method requires collecting EEG data (which is time and labour-intensive) I wonder to what extent you get similar improvements in parameter recovery by collecting more data per participant (which is usually cheap). Put differently, I would appreciate an additional matrix in Figure 5 for vanilla models.

      The new version of Figure 5 in the revised manuscript now contains the goodness of parameter recovery for the uninformed CT-DDMs for different numbers of trials. As the simulation results suggest, even with 1000 trials, the parameters of the uninformed CT-DDMs are still not reliably identifiable. Therefore, these results suggest that although increasing the number of trials can slightly improve the parameter recovery of the CT-DDMs, it cannot fully resolve the reliability issue in their parameter recovery.

      In addition, it is worth clarifying that, although the paper focuses on extracting non-decision time from the EEG signal using the HMP method, as discussed in the general discussion, the NDT-informed approach can also be employed with purely behavioral methods.

      (4) L116-117: minor detail: I don’t think it’s fair to write that it’s an open question whether or not thresholds collapse. I think it’s fair to say that this is hard to show, and that the conditions under which it appears are unclear; but saying that it’s still unclear whether this occurs at all seems unfair with regard to previous work.

      Thank you for mentioning this point. We revised the mentioned sentence as follows in the new version of the manuscript:

      “This issue is particularly critical given that the conditions under which individuals adjust their decision thresholds during a single trial remain an open question.”

      (5) Figure 5: Minor detail: It would be useful to mention in the figure or caption how many trials underlie these simulations, which is now somewhat buried in the text.

      As Figure 5 shows sensitivity to the number of trials, and the number of trials is explicitly mentioned in the figure, we assume that the reviewer intended Figure 6. In the revised manuscript, we explicitly specify the number of trials for the results reported in Figure 6. We simulated 500 trials for each noise level in this graph.

      (6) Figure 10 and similar figures: I tried to figure out how corrects and incorrects differ, but couldn’t see the difference. Can the authors use something more colorblind friendly (hope I didn’t give up on being anonymous here)?

      We are so sorry for the inconvenience. In the revised manuscript, we used different symbols to distinguish correct and incorrect data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates the peptide-binding principles of promiscuous chicken MHC molecules. The data from crystallography, mass spectrometry, and modeling are convincing. However, the presentation would benefit from streamlining and clear links between data and conclusions. This paper will be of broad interest to immunologists and those interested in vaccine development.

      Overall, we are delighted and grateful to the eLIFE editors and the two reviewers for the careful and thoughtful assessments and reviews of our paper. We are glad that the strengths of the paper were apparent and appreciated. And of course, every paper has weaknesses, especially for a story as complex as this one.

      We made only minor changes to accommodate the reviewer comments, along with additions for which we only became aware upon this submission of a revised manuscript. In particular, we shortened the title and abstract to fit what is usual for an eLIFE paper, added Key Resources table with accompanying references, changed the numbering of the figures throughout the manuscript to ensure that each page represented a figure (rather than panels of a figure), moved the figure legends from the embedded figures to a list near the end of the manuscript, and split the supplemental spreadsheet into two renamed Data Source files.

      Before answering the comments and questions directly, perhaps a few points would help clarify why the paper is as it is.

      First, the experiments cover over three decades of work, with the first gas phase sequencing results done in 1992. Unlike some of the chicken class I alleles which immediately gave completely clear stringent motifs (B4, B12 and B15 in Wallny et al 2006 PNAS, B19 in Han et al 2023 J Immunol), we harvested nothing but confusion from the B21 class I results (Fig. 1). Initially, we thought that the lack of a clear motif for B21 was due to multiple well-expressed class I molecules but only one dominantly-expressed class I molecule was found (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol) and, to our surprise, bacterially-expressed BF2*21:01 heavy chain and b<sub>2</sub>-microglobulin refolded with two synthetic peptides without sequence in common, and the crystal structures showed that this molecule remodeled the binding site to accommodate two such disparate peptides (Koch et al 2008 Immunity). This was the beginning of our understanding of the spectrum of class I alleles from promiscuous generalists to fastidious specialists, which we have explored in a series of further papers (in particular, Chappell et al 2015 eLIFE, Tresgaskes et al 2016 PNAS, Kaufman 2018 Trends Immunol, Tregaskes and Kaufman 2022 Mol Immunol).

      Second, over these many years, we continued to explore the binding properties of BF2*21:01 in ever more detail, resulting in the current manuscript. We learned only slowly how to probe this unexpected promiscuity, unprecedented in the MHC literature, so that the experiments proceeded with our best understanding at the time, including taking advantage of new approaches as they become available. Each experiment built on the previous set of experiments and each brought us closer to an understanding.

      Third, having amassed a collection of data, we chose eLIFE exactly because it allows us to present the entire story from beginning to end without compromise, not just the highlights with the major points illustrated by a few main figures and with the supporting data in many supplementary figures. We include all the data, because it is all part of the story, and so interested researchers to look at the data from their own perspective. Although mostly we provide bar graphs, we include spreadsheets for the raw data (or close to them) for the final experiments (illustrated by Figs. 10 and 14-22) in the two source data files, so these can be assessed easily by others in the field, perhaps using approaches that we may not feel competent to perform.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

      Weaknesses:

      The manuscript is sometimes difficult to follow because the authors present a large amount of peptide-library, structural and immunopeptidomics data. without always clearly explaining how these datasets support the proposed simplifying principles.

      We are delighted and grateful to the reviewer 1 for the careful and thoughtful comments and questions concerning our manuscript. We are glad that the strengths of the paper were apparent and appreciated, and acknowledge the weaknesses that come with such a complex story with experiments performed over decades.

      Major Issues - Points Requiring Clarification or Additional Support:

      (1) Line 282-301, 537-545)

      The immunopeptidomics conclusions are mainly based on one B21 cell line with one biological replicate and at least two technical replicates. Given the complexity of the BF2*21:01 peptide repertoire, this is a major limitation. The authors should either provide additional biological replicates or clearly state this limitation in the Abstract, Results and Discussion.

      This limitation is clearly stated in lines 537-545, as part of a paragraph covering the various ways in which the data presented in this manuscript could be improved. In fact, we have performed immunopeptidomics of several different B21 cell types, with many replicates and found similar data as presented, giving us confidence in our interpretations. However, these other experiments belong in different stories, so it is not appropriate that the data be reported in this manuscript.

      (2) (Lines 290-313)

      The B21 cell preparations contain both BF2 and the lowly expressed BF1 molecule. Some peptides, especially 8-mers or peptides with atypical motifs, may derive from BF1*21:01. The authors should clarify how BF2*21:01-bound peptides were distinguished from possible BF1-derived peptides, or interpret the immunopeptidomics motif more cautiously. The authors should also provide or cite evidence confirming the B21 haplotype identity of the cell line and chicken materials used for immunopeptidomics.

      The concern about the contribution of BF1*21:01 to the immunopeptidomics is clearly stated in the manuscript, both lines 290-313 and as part of the paragraph describing the limitations of the experiments (lines 542-543). In fact, the expression of BF1 molecules has long been known to be less than 10% of BF2 molecules at the RNA level, and much less at the protein level (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol). The proportion of 8mers identified by immunopeptidomics is also low (Fig. 14), and it is not impossible that most 8mers are due to BF1*21:01. We have used assembly assays with peptide libraries, immunopeptidomics and a crystal structure to determine the peptide motif for typical BF1 molecules, of which BF1*21:01 is one and found it may contribute to 8mer peptides but very seldom to longer peptides. This work is unpublished but gives us confidence that the characteristics of BF2*21:01 are not misrepresented by the data in this manuscript.

      The sources of the chicken samples and the cell lines are described in detail under Materials and Methods (lines 577-590), citing relevant publications.

      (3) (Lines 217-221, 243-253)

      The authors acknowledge that MALDI-TOF cannot reliably distinguish peptide combinations with identical or similar masses, nor determine residue positions in some cases. Therefore, MALDI-TOF results should not be overinterpreted as precise evidence for residue preference. The authors should clearly indicate which conclusions are supported by LC-MS/MS.

      As described, the experiments follow each other in temporal sequence, so that we started with single peptides, then peptide libraries that varied in one position, then peptide libraries that varied in two positions first analysed by MALDI-TOF and later by LC-MS/MS. The final experiment (Fig. 10, with the original data in the supplementary spreadsheet) directly compares MALDI-TOF and LC-MS/MS results for six peptide libraries, so that the strength of the evidence for residue preference is clear. Throughout the manuscript, we do our best to not to overstate conclusions based on the data of any particular experiment.

      (4) (Lines 297-301, 316-330)

      The authors suggest that longer peptides may bulge in the middle or extend out of the groove at the C-terminal end. The rationale for the C-terminal extension is not clearly explained. Why is the C-terminal extension considered rather than the N-terminal extension? If the binding register is uncertain, long peptides should be analyzed separately from canonical-length peptides.

      When the first sequence of a chicken class I cDNA was determined, an immediate mystery was why one of the so-called invariant residues that coordinate the N- and C-termini of the bound peptide is not conserved (Kaufman et al 1992 J Immunol). In fact, this residue Tyr at position 86 in HLA-A2 and the equivalent position in all mammalian classical class I molecules is an Arg in the classical class I molecules of all non-mammalian vertebrates and is common with class II molecules (Kaufman et al 1995 Semin Immunol). Similar to class II molecules, this Arg in chicken class I molecules allows the peptide to extend out of the C-terminus, as shown by a crystal structure (Xiao et al 2018 J Immunol). The concern that we might be misidentifying the C-terminal amino acid was the basis for the analysis in Figs. 23 and 24, but in the absence of crystal structures, we are not able to provide a final answer this question. Perhaps relevant is the fact that a chicken class II molecule can bind exactly the same peptide in two conformations, one with a canonical 9mer core and the other with an unexpected 10mer core (Goryanin et al 2026 J Virol).

      By contrast, N-terminal extensions are only found for some class I alleles and thus far depend on the substitution of small amino acid sidechains for W166 (Li et al 2011 J Virol for bovine, Ma et al 2020 J Immunol for Xenopus, Wei et al 2022 J Immunol for ovine). Thus far, no chicken BF2 sequences have this substitution, consonant with the many crystal structures, including those for BF2*21:01 (Koch et al 2008 Immunity, Chappell et al 2015 eLIFE, this manuscript). However, in unpublished data, we find that most BF1 sequences have sequence differences that could allow N-terminal extensions, although we have no crystal structures to support this possibility.

      (5) (Lines 406-439)

      In vitro assembly assays show that several hydrophobic residues can be tolerated at Pc, whereas immunopeptidomics shows a strong Leu preference at this position. The authors should clarify whether this Leu preference reflects intrinsic BF2*21:01 binding specificity, TAP-mediated peptide transport, antigen processing, peptide loading, or a cell-line-specific effect. Additional experimental support, such as TAP transport analysis, would strengthen this conclusion.

      The preference for Leu at the final position of the peptide by immunopeptidomics of the B21 cell line is strong but not absolute and is certainly affected at the least by the length of the peptide (Figs. 23 and 24). Unpublished immunopeptidomics results (mentioned above) show that this is not a cell line-specific result. The evidence from assembly assays of various peptides is that several hydrophobic amino acids are tolerated with sufficient stability of BF2*21:01 that they are detected in the assay (Figs. 3, 5, 9 and 10). Thermostability assays (Fig. 6) show that peptides with these same hydrophobic amino acids are stable to at least body temperature of chickens. These experiments show that such stability is peptide-dependent (that is, whether a particular amino acid is tolerated depends on the stability conferred by the rest of the peptide). Finally, peptide translocation assays using B21 cells have been done (Tregaskes et al 2016 PNAS) and show that peptides with several hydrophobic amino acids can be pumped into the lumen of the endoplasmic reticulum. However, the assays are with single synthetic peptides, so the data are not extensive enough to separate the effects of the final amino acid from the rest of the peptide. Certainly, peptides with amino acids other than Leu at the C-terminus can be translocated. So, it is not yet clear at which point the preference for Leu at the C-terminus of the peptide arises.

      (6) (Lines 172-178, 243-279, 442-457)

      The structural analysis explains some residue combinations, such as Arg at P2 with Glu at Pc-2 or Trp at Pc. However, the structural interpretation is not fully integrated with the large-scale peptide library and immunopeptidomics results. Representative high- and low-frequency combinations should be discussed structurally.

      Six crystal structures show that BF2*21:02 remodels the binding to accommodate a variety of anchor residues (Koch et al 2008 Immunity, Chappel et al 2015 eLIFE). These crystal structures are representative of sequences found by the immunopeptidomics from very frequent (H-E at roughly 15% 8-12mers) to moderately frequent (E-L at roughly 6% 8-12mers) to infrequent (N-F, A-D and E-D at roughly 1.5%, 1.6% and 0.7% 8-12mers) based on Fig. 18. All but one of the structures has Leu at the C-terminus, with the last one having Val which is found but not frequently by immunopeptidomics.

      Similar numbers are found by LC-MS/MS of double-substitution libraries of the two original peptide sequences in Fig. 10 with H-E found frequently (8.1% in P390, 3.8% in P498) and the others infrequently (0.1, 0.9, 1.0, 0.3% in P390, 0, 1.4, 1.0, 0.3% in P498), as calculated from the numbers in the Supplementary data spreadsheet. As discussed in the manuscript, for single-substitution peptide libraries of the two original peptides, Ile/Leu at the C-terminus was very frequent but at the same or slightly less level as Phe, with Met less frequent and Val even less so (Fig. 7).

      In addition, there are two more structures along with models explicitly testing some substitutions (Fig. 5). Attempting more current modelling approaches, we found AlphaFold 3 was unable to correctly predict most of the conformations that are found in the crystal structures of BF2*21:01, so we don’t feel confident in using them to predict unknown structures of this kind.

      (7) The inference of co-variation between P2 and Pc-2, as well as the modulatory effects of P3 and Pc-3, should be better explained. At present, some conclusions appear to be based mainly on residue-frequency patterns, and the logical connection between these observations and the proposed binding principles is not always clear. Statistical analyses, such as mutual information, chi-square tests or permutation tests, and representative structural explanations would strengthen this conclusion.

      We endeavored to do our best to explain the data, our interpretations and our reasoning, so we apologise if we have not managed to be as clear as might be desired. We have included as close to raw data as possible for the LC-MS/MS and MALDI-TOF (Fig. 10) and for the immunopeptidomics (Fig. 14 and 18) in the Supplementary Data spreadsheet, exactly so that competent practitioners can carry out further analyses (including the sophisticated statistical tests mentioned).

      Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

      Weaknesses:

      The data overall support the analysis and conclusion well. The only caveat is linked to Figure 4, which does not describe the stability of the peptide-MHC complex, but instead shows refold yield, and the two are not always linked.

      We are grateful for the clear understanding of the strengths of the work. With regards to Fig. 4, we agree with the reviewer that there are differences in refold yield but that measure may not be correlated with stability of the peptide-MHC complex. However, we were basing our interpretation of stability on the position and quality of the monomer peak, as illustrated by the trace in Fig. 2, in which a sharp peak at the monomer position represents a stable complex (as seen for the 10 and 11mer peptides) and later peaks represent unstable complexes falling apart during the chromatography (as seen for the 7, 8 and 9mer peptides).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Issues: Editorial and Data Presentation Modifications

      (1) Lines 53-62, 155-170, 303-314: The terms Pc, Pc-2 and Pc-3 should be clearly defined early in the manuscript and figure legends.

      The Abstract introduces the abbreviations as “peptide positions P<sub>2</sub> and P<sub>c-2</sub>” followed by P<sub>C</sub>, which are standard usage and seem clear. The first usage in the text is “anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub>” which again seems clear, particularly in the context of the text and the figure. However, a parenthetical description has been added to read “…anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub> (position 2 and the position two before the C-terminus) …”. Given these usages, it seems unlikely that the reader will fail to understand P<sub>C</sub> and P<sub>C-3</sub>.

      (2) Lines 255-279: The term "peptide backbone" should be defined as the fixed sequence context outside the randomized positions, if this is what the authors mean. Suggest clarifying the meaning of "peptide backbone".

      The phrase reads “double-substitution libraries based on four peptide backbones”, which in context of the figures seems clear. However, a parenthetical description has been added: “The examination of double-substitution libraries analysed by MALDI-TOF was expanded to other backbones (that is, other sequences in which two positions were randomised): …”.

      (3) Several figures are complex. The authors should add brief take-home messages to figure legends.

      Every figure legend starts with a (sometimes quite long) take-home message. It is not clear what more should be added.

      (4) A concise summary table of the proposed binding rules, including preferred peptide length, P2, Pc-2 and Pc preferences, and the effects of P3/Pc-3, would be useful for readers.

      A concise summary table of binding rules would be very helpful, but the rules are complex, both qualitative and quantitative. For example, immunopeptidomics shows that 10mers are preferred, but that fails to capture the quantitation. The co-variation of P<sub>2</sub> and P<sub>c-2</sub>, which is the strongest and best characterized correlation, nevertheless is complex since the occupancy of P<sub>2</sub> by different amino acids (presumably independently of P<sub>c-2</sub>) varies considerably. At our current level of understanding, it is hard to imagine a table that is both concise and precise. With time and more data, perhaps code for quantitative prediction might be constructed (dare one suggest machine learning…), which is the hope of presenting all the available data in one place.

      (5) The Results section contains a large amount of peptide-library, structural and immunopeptidomics data. The authors should improve the logical flow and add clearer transition sentences to explain how each dataset supports the proposed simplifying principles.

      The reviewer is of course correct that any written text can be improved (although each critic may have a different idea about which part should be fixed), but we have done our best with the material and time available. We could respond productively to a more detailed critique.

      (6) Some statements in the Abstract and Discussion should be softened, especially those related to in vivo peptide preferences, BF2*21:01 promiscuity and peptide prediction, given the limited biological replication and uncertainty in peptide assignment.

      The statements in the both the Abstract and Discussion are very general, reflecting what we believe to be careful interpretations based on the data. We could respond productively to concerns about specific claims.

      Reviewer #2 (Recommendations for the authors):

      Overall, the data presented in this study are interesting; however, it is complex, and some results could be merged and simplified, as well as the figures. The data provide an in-depth analysis, using mass spectrometry, of the interplay between the different positions of the peptide and the residues favourable to bind within the antigen-binding cleft.

      (1) From the abstract, the concept of "promiscuous generalists and fastidious specialists" is not explored after or defined within the results.

      The Abstract introduces the concept of promiscuous generalists and fastidious specialists to provide the basis for exploring the peptide-binding specificity of the most promiscuous class I known, BF2*21:01. This overall concept is described in enormous detail in several publications cited in the current manuscript, but it is not particularly germane to the analyses.

      (2) From the abstract "These simplifying principles may eventually allow predictions of pathogen peptides", I'm not sure how "simplifying" the principles are with the data, if anything, it does show a rather complex interplay between the different residues of the peptide that enable the MHC to bind a large and diverse number of peptides.

      Compared to any combination of anchor residues being permitted at equal frequency, there are clear preferences which the experiments identify. Instead of an enormous range of possible peptide lengths, roughly 50% of peptides are 10mers. Of course, the structural reasons behind these results are certainly complex and likely must be understood in detail in order to attempt peptide predictions. 

      (3) Line 48. "less well-expressed". Less than what? Do you mean the level of expression was lower? And if yes, are there values of fold change for comparison?

      The sentence in the Abstract reads “Chicken BF2 alleles … are less well-expressed on the cell surface … while certain human HLA-B alleles … are well-expressed …”. Read as a complete sentence, the meaning is clear. This difference for chicken BF2 alleles has been quantified as reported in several publications cited in the current manuscript, ranging from 3-5 fold for peripheral blood lymphocytes to ten-fold for erythrocytes (Kaufman et al 1995 Immunol Rev, Chappel et al 2015 eLIFE), with similar numbers for a few HLA-B alleles on human peripheral blood lymphocytes and monocytes (Chappell et al 2015 eLIFE).

      (4) Line 149. "with individual peptides" which peptides are we referring to here?

      This introductory sentence to a paragraph outlines the general method used for the experiments in this section of the Results, “refolding in vitro … with individual peptides.” Which “individual peptides” are described in the following paragraphs, with the next section of the Results using “refolding in vitro … with peptide libraries”.

      (5) Line 170. If 9 mers and below are not stable, which is not really quantified or shown with the data on Figure 4, why is refolding material observed for peptides with different lengths of 9 aa and below on Figure 4? A lower yield of refolded material can have a different origin, and there is no association between stability and refold yield. The notion of stability here probably needs to be changed, as it does not apply to the data.

      As described in our response to a similar concern above, we are not basing our interpretation of stability on the quantity of refolded material, but on the position and quality of the monomer peak, as described clearly in the legend to Fig. 4: “The original 10mer (REVDEQLLSV) and 11mer (GHAEEYGAETL) peptides refold with BF2*21:01 to give stable monomers as do 11mer and 10mer derivative peptides, but 9mer, 8mer or 7mer peptides give heavy chain only peaks” and “The peptides 11mer GHAEEYAETL (top panel), 10mer REVDEQLLSV (middle panel), 11mer GHAEAAAAETL and 10mer GHAEAAAETL (bottom panel) gave sharp monomer peaks, while the 9mer GHAEAAETL, 8mer GHAEAETL and 7mer GHAEETL gave a delayed broad peak indicative of unstable binding or heavy chain.” This concept is illustrated by the trace in Fig. 2 (discussed in the text at the beginning of this section of the Results), in which a sharp peak at the monomer position represents a stable complex, and later peaks represent unstable complexes falling apart during the chromatography or free heavy chains.

      (6) Line 175. As the 3BEV structure had a P2-His, is a comparable structure expected?

      This sentence reads “Structures with amino acid substitutions in the 11mer peptide GHAEEYGAETL bind with Asp at P<sub>c-2</sub> and either His or Arg at P<sub>2</sub> (5AD0 and 5ADZ), comparable to the original structure (3BEV) (Fig. 5A).” Minor changes in positions and orientations of individual amino acid sidechains are expected and are clear from the crystal structures presented in Fig. 5A, but are overall comparable to the original peptide in 3BEV, which has a His at P<sub>2</sub> and a Glu at P<sub>c-2</sub>.

      (7) Line 175 "modelling the substituted". How was the modelling done?

      The sentence reads “Modelling the substituted peptide with Arg at P<sub>2</sub> and Glu at P<sub>c-2</sub> shows a steric clash that can explain why this peptide did not refold with BF2*21:01 (Fig. 5A).” The legend to Fig. 5A states that “modelling done as detailed in Materials and Methods”, but apparently that section was omitted. A section has been added now to the Materials and Methods which states “Beginning with known crystal structures, modelling was carried out using PyMol with the protein mutagenesis Wizard, the rotamer toggle, show bumps and show surfaces.”

      (8) Line 176 "shows a steric clash" with what? Figure 5 is too small to see, and there is no label on the residue to follow where the steric clash is coming from.

      As stated in the legend to Fig. 5, “Structures were determined for GHAEEYGAETL (3BEV), GHAEEYGADTL (5AD0) and GRAEEYGADTL (5ACZ), which all refolded successful to give stable monomers, while GRAEEYGAETL did not (see Fig. 3), all models of which showed steric clashes (one depicted, red arrow).” In the model shown, the clash is between R9 of the BF2*21:01 with the Glu at peptide position 9. Parenthetically, this depiction of the key MHC residues for the co-variation has been used repeatedly, starting with Chappell et al 2015 eLIFE. The picture can be zoomed to make it large enough to see.

      (9) Line 247 "many combinations". It is not clear here if combinations are referring to a set of double-substitutions or different peptides?

      Each peptide in the library has a different double-substitution, so the meaning is the same either way. The point of Fig. 10 is to compare the identification of peptides from LC-MS/MS (in which each peptide is identified exactly) with the identification of sets of peptides from MALDI-TOF (in which the order of the double substitution is not clear, as well as the exact amino acid in the cases of I/L and Q/K). As the figure shows, the numbers are generally very similar, but this sentence describes the percentage of cases for which only one method or the other identified a peptide.

      (10) Lines 252-253. The conclusion of the MS data that the LC-MS/NS is more accurate and sensitive than MALDI-TOF is not very surprising. What was the rationale for using both?

      Our examination of the binding properties of BF2*21:01 for peptides and peptide libraries developed over a long time-span, so in the beginning we looked at single peptides with size exclusion chromatography peaks as the measure, then single- and later double-substitution libraries, first by MALDI-TOF and later by LC-MS/MS. Each set of experiments built on the previous work, so that together they tell the story. For example, after we optimized the use of double-substitution libraries, we were worried about the effects of temperature, so we tested that by MALDI-TOF. While repeating the temperature experiment once we optimized the LC-MS/MS approach might have yielded some additional data, there were other questions to answer.

      (11) Line 272-273 "support the idea that P3 is an important position within the peptide (Figures 8-10) despite not contacting the MHC molecule" The structure of 3BEV clearly shows interaction between the P3-Ala and the Tyr156 of the MHC. Residues that are fully or partially buried in the MHC cleft almost all contact the MHC molecule. Maybe I've missed something, but I think this statement is inaccurate.

      We agree with the reviewer that nearly every peptide residue contacts the MHC molecule (but of course some much more than others). The statement is now changed in the text to read “support the idea that P<sub>3</sub> is an important position within the peptide (Figures 8-10) despite not being an anchor residue."

      (12) Line 294. Figure 14 clearly shows that the number of 8 and 9-mer peptides eluted is at the same level as the 11mer and above, so how does the data fit with the statement that 9mer and shorter peptides are not stable with the MHC?

      Fig. 4 shows that 7, 8 and 9mer derivatives of the original 11mer failed to refold to give a single peak of stable monomers, while Fig. 6 shows that the 10mer derivative of the original 11mer was more thermostable than both the 9mer derivative and the original 11mer. The reason why the 9mer yielded so little monomer in Fig. 4 while giving enough to test by thermostability in Fig. 6 is no longer remembered. A key point is that these results are peptide sequence-specific, so it is not impossible to imagine stable binding of an appropriate 8mer (or perhaps even a 7mer). Another uncertainty, mentioned in the Results and Discussion, is that BF1*21:01 molecules bind primarily 8mers, and the contribution of peptides from BF1*21:01 is not known with certainty.

      (13) Line 316. What was the rationale for choosing peptides > 12aa to see if there is some overhang or bulge? As even 11-12mer can exhibit such features.

      We were just looking for any obvious patterns, but we didn’t find any.

      (14) The section starting at line 366 would have benefited from some structure prediction or modelling to illustrate the findings.

      We would have been delighted to model peptides, but we have used AlphaFold3 to compare the models to our crystal structures for seven chicken class I alleles (including BF2*21:01) with one or more peptides. The models sometimes fit the experimental data but they often didn’t, often by a wide margin, and with BF2*21:01 the worst (presumably because the system is not trained on MHC molecules which utilise charge transfer). Therefore, we do not feel confident in using any such modelling approaches except in the most defined situations (such as illustrated in Fig. 5).

      (15) Typo - Alleles should be italic, and in vitro as well.

      Alleles of genes are in italics, but alleles of proteins are not. To write in vitro in italics is customary, which we have corrected.

      (16) Figures

      (a) Some of the figures could be merged together.

      Of course, any presentation can be improved, but which figures we should merge (some already being three pages in length) is not clear. We could productively respond to more detailed suggestions.

      (b) Figure 1. I can't see the different colours mentioned in the figure legend

      Our apologies if the colours are not clear enough, but they are present only as an aid (as elsewhere in the manuscript), with red D and E, blue H, K and R, green N, Q, S and T, and all other amino acids black.

      (c) Figure 12. Is this figure only with 11mer peptides?

      Figure 12 shows the percentage of peptides with particular amino acids at P<sub>2</sub> and P<sub>c-2</sub> for three 11-mer double-substitution peptide libraries, the sequences of which are written on the graphs and described in the figure legends.

      (d) Table 1. The name of the protein should be added to the table.

      OK.

    1. Author response:  

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystis research.

      We are grateful to Reviewer 1 for their careful and constructive assessment of our revised manuscript. We are pleased that the revisions have satisfactorily addressed their previous concerns, and we sincerely appreciate their recognition of the value of our work.

      Reviewer #3 (Public review):

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

      We are grateful to Reviewer 3 for their careful reading of the revised manuscript and for the additional constructive suggestions. We appreciate that these recommendations are aimed at improving clarity, accessibility, and presentation, and we have addressed them as detailed below.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      (1) Table 1

      (1a) The description of the column "Promoter Lengths Tested" is a bit confusing because it actually shows acronyms of the used promoters. Suggestion for clarity: either keep the name and mention just the lengths, or rename to "Promoter abbreviation" or "Promoter acronym".

      We thank the reviewer for pointing this out. To avoid confusion, we have renamed this column “Promoters Tested”, which better captures the information presented. The column contains both promoter identifiers and the corresponding promoter lengths, which are further clarified in the table notes.

      (1b) It is not entirely clear based on the formatting of the table, which columns are relevant for "Blastocystis ST4-WR1" and which for "Blastocystis ST7-B".

      We thank the reviewer for this helpful suggestion. We have revised the formatting of Table 1 to more clearly distinguish the data corresponding to Blastocystis ST7-B and Blastocystis ST4-WR1. Specifically, we have added a vertical divider between the relevant column groups to improve readability and make the species-specific information easier to follow.

      (2) Figure 2:

      If the authors wish to preserve the decorative colouring in the parts B and C, it would be better to at least keep the datapoint and boxplot hues consistent for individual conditions (or plot the datapoints in the same, neutral hue throughout, e.g., black or grey). Currently, this is only the case in the part C, but not in the part B.

      We thank the reviewer for this useful suggestion. We have revised Figure 2B to improve visual consistency between the data points and boxplots while preserving the overall colour scheme of the figure. This adjustment improves readability without changing the data or its interpretation.

      (3) Figure 3: With the modifications everything is clear. However, two issues are now apparent.

      (3a) After the part A was modified, the arrow colours changed (had been yellow and red, now are cyan and magenta), but the description in the legend remained (yellow and red), so the text should be corrected. [By the way, the original colours worked well.]

      We apologise for overlooking this inconsistency. The Figure 3 legend has now been corrected to match the revised arrow colours in panel A.

      (3b) The care taken by the authors to make the images more accessible is really greatly appreciated, but being a person on the colourblind spectrum, I can attest that the issue is not always just about red and green discrimination: the greyscale and blueish green used in the part A photo in particular are discernible only at huge magnification (to some people). If the signal in the part A were of the same hue as in the part B (yellowish green instead of blueish green), it would readily pop out. For that matter, what is the reason for the use of a different colour profile of the UnaG fluorescence in these two images?

      We sincerely thank the reviewer for this important accessibility-related comment. The difference in colour profiles between panels A and B was intentional because the images were acquired using different imaging modalities, as indicated in the figure legend. We wanted to avoid implying that the two panels were generated under identical imaging conditions. However, we appreciate the reviewer’s point that this distinction should not come at the expense of readability. We have therefore adjusted the display of the UnaG signal in panel A to improve contrast and visibility. 

      (4) Conclusion:

      The expression "bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements" feels a bit clunky. What the part "endogenous regulatory part discovery" alludes to is now clearer, but consider reformulating it to "discovery of endogenous regulatory elements". This would make it clear without the need to add the explanation "namely the identification of native promoter and terminator elements". [It is now clear that the accumulation of noun adjectives and the significance of the word "part" was what originally blurred the overall meaning.

      We thank the reviewer for this valuable suggestion. We agree that the original phrasing was unnecessarily clunky and have revised the sentence for clarity and flow. Lines 682–684 now read:

      “By integrating endogenous promoter and terminator discovery, DNA delivery, selection, clonal recovery, and reporter validation into a single pipeline, we provide a flexible foundation for routine transgene-based studies.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Several aspects of data interpretation, however, require clarification or reanalysis-most notably the single-channel analyses (event counts, Po metrics, and mixed parameters), the statistical treatment, and the interpretation of purported "OFF currents." Additional issues include PKD2L1-TRPP3 nomenclature consistency, kinetic comparison with ASICs, and the physiological relevance of the extreme acidification paradigm. Addressing these points will substantially improve reproducibility and mechanistic depth.

      Overall, this is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing, provided that the analytical and interpretative weaknesses are satisfactorily corrected.

      (1) The authors should re-analyze electrophysiological data, focusing on macroscopic currents rather than statistically unreliable Po calculations. Remove or revise the Po analysis, which currently conflates current amplitude and open probability.

      We agree with the reviewer that the Po analysis has strong limitations, particularly in experiments where the recording times are short, like when extracellular pH is changed either by photolysis (Figure 4D) or puff applications (Figure 3Aa). In order to circumvent that problem and not to rely only on Po estimations, we used alternative methods as well, including the analysis of the current membrane charge that we have used extensively during the manuscript (Figures 3A and 4D, for example) or the analysis of the event latencies (Figure 4G). Nevertheless, single-channel recordings clearly contain information that is not included in the macroscopic current analysis. We intend to stress in the revised version that the elementary current amplitude is conserved by manipulations such as pH changes, leaving the total number of channels (N) and the channel open probability (Po) as possible culprits for the current changes. Since these changes are rapid and reversible, it is likely that N stays constant and that Po changes. In order to address the reviewer’s concern, we propose the following changes/reanalysis: i) to report in each condition the minimum N (maximum observed openings; for example, in Figure 3Aa the minimum N goes from 4 in control conditions to 1 during the puff of the pH 6.4 solution). This method (to estimate N by counting the maximum number of open channels), while imperfect, provides a tentative estimate of Po; ii) following the previous point, we propose to reword the text (and images) and use the expression “apparent Po” instead of “Po”; iii) to report the fraction of time that the channels remain open. Also, we acknowledge that some traces are confusing (Figure 3Aa, top) as they seem to show macroscopic currents. We will modify those figures by plotting the amplitude histograms (as in Figure 1Bb) in order to show unambiguously that current recordings from CSFcNs only show single-channel activities.

      (2) PKD2L1-TRPP3 nomenclature should be clarified and all figure labels, legends, and text should use consistent terminology throughout.

      We agree with the reviewer that the nomenclature concerning polycystin members is confusing. In this manuscript we have followed the nomenclature that has been proposed in a recent, comprehensive review on polycystin channels by Palomero, Larmore and DeCaen (Palomero et al. 2023), where the authors refer to the channels by their gene name. In this review, the authors indicate that PKD2L1 channels correspond to TRPP2 (formerly TRPP3, their table 1). In another, recent review on TRP channels, however, the authors refer to the PKD2L1 channel as TRPP3 (Zhang et al. 2023). In order to avoid any confusion we will remove from the text any reference to the TRPP nomenclature and stick to the PKD2L1 name.

      (3) The authors should reinterpret the so-called OFF currents as pH-dependent recovery or relaxation phenomena, not as distinct current species. Remove the term "OFF response" from the manuscript.

      We concur with the reviewer that the term “OFF response” is not very helpful from the biophysical perspective and conveys the idea that it is another current. We will remove the term “OFF response” or “OFF current” in the revised manuscript and replace it by the term “photolysis-evoked PKD2L1 current”. Also, we will condense two sections (“The proton-induced current is an off-current” and “The off-current is mediated by the activation of PKD2L1 channels”) into a single new section (“The photolysis-induced current is mediated by PKD2L1 channels”), as separating the description of the photolysis-evoked PKD2L1 current compromises its description. Finally, we will rewrite the discussion to better describe this current.

      (4) Evidence for physiological relevance should be provided, including data from milder acidification (pH 6.5-6.8) and, where appropriate, comparisons with ASIC-mediated currents to place PKD2L1 activity in context.

      This is partly addressed in Figure 3. The data there indicates that PKD2L1 channels are very sensitive to pH variations around physiological pH. In order to make this conclusion stronger, we will add to the figure the EC50 values drawn from the fittings. In terms of the ASIC-mediated currents, one of our main conclusions is that ASIC channels are not present in the ApPr, as the effects of proton photolysis in the ApPr and are not blocked by ASIC channels blockers. Our results indicate that PKD2L1 channels are the exclusive pH sensitive channels in the ApPr, while ASIC channels are probably the acid-sensitive channels in the soma, although we have not studied the latter in detail. Following this and the editor’s comments, the subsection “the involvement of ASICs” in the Discussion has been modified.

      (5) Terminology and data presentation should be unified, adopting consistent use of "predominant" (instead of "exclusive") and "sustained" (instead of "tonic"), and all statistical formats and units should be standardized.

      Following the suggestions of the reviewer, an exhaustive work will be performed to unify terminology, data presentation and correct the text following the reviewer’s suggestions.

      (6) The Discussion should be expanded to address potential Ca<sup>2+</sup> -dependent signaling mechanisms downstream of PKD2L1 activation and their possible roles in CSF flow regulation and central chemoreception.

      This is indeed a very interesting and currently unresolved point in the physiology of CSFcNs. Published data indicate that calcium flowing into the cell through PKD2L1 channels is a key regulator of apical process physiology: on the one hand, PKD2L1 channels are calcium permeable and at the same time, they are inhibited by intracellular calcium (DeCaen et al. 2016). Also, ultrastructural data indicate that the ApPr is rich in mitochondria and tubulo-vesicular structures resembling the Golgi apparatus (Bjugn et al. 1988; Bruni et Reddy 1987), which are intracellular organelles that contribute to calcium homeostasis. Altogether, this evidence suggests that intraApPr calcium concentration needs to be finely regulated, both in space and in time, in order for the ApPr to fulfil its physiological roles. Based on what has been published in the literature, we can speculate that these calcium signals can be decoded by several systems: i) calcium can act as the second messenger linking the activation of the multimodal PKD2L1 channels to changes in CSFcNs excitability, which in turn regulate spinal neuronal networks controlling locomotor activity; ii) calcium could initiate neurosecretion of different molecules from the ApPr to the central canal (as has been proposed by the Wyart group in the zebrafish in the context of bacterial infections(Prendergast et al. 2023)); iii) calcium could activate the Hedgehog signaling pathways (as has been shown by (Delling et al. 2013)); iv) calcium could modulate CSF flow directly (by modulating ciliary activity) or indirectly (by modulating ependymal cells ciliary activity through paracrine interactions). Resolving these downstream pathways is essential to fully define the role of CSFcNs as integrators of CSF homeostasis. We will expand this issue in the Discussion of the revised ms.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments relate to their multimodal CSF-sensing function remains unclear.

      The authors confirm that in GATA3:GFP mice, where these cells are labeled, that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show that functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding of how these sensory cells can detect the changes of pH in the CSF so finely.

      Weaknesses:

      The following do not constitute weaknesses; rather, they are minor requests that this reviewer considers would complete this beautiful study.

      (1) It would be nice to quantify further the relation in spontaneous as well as in acidic or basic pH between the effects observed on channel opening and holding current: do they always vary together and in a linear way?

      Following the reviewer’s suggestion, we have performed a Spearman’s rank correlation test, which shows a significant correlation between the changes in the apparent open probability and holding current (paired experiments; ctrl vs pH 6.4 pressure applications; p < 0.05, Spearman r = 0.72 and critical value = 0.67). The Pearson correlation coefficient calculated on the same data set = 0.63 and the critical value is 0.632, which indicates that the correlation is not linear. We will add this analysis to the manuscript.

      (2) Since CSF-cNs also respond to changes in osmolarity (Orts Dell Immagine 2013) & mechanosensory stimulations in a PKD2L1 dependent manner (Sternberg NC 2018), it would be nice to test the same results whether the same results hold true on the role of PKD2L1 in AP for pressure application of changes in osmolarity.

      This is a very important point. As the reviewer mentions, previously published experimental evidence indicates that CSFcNs are also sensitive to osmolarity changes and mechanical stimulation in a PKD2L1-dependent manner. It is therefore reasonable to assume that, as for the pH sensitivity, osmotic and mechanical sensitivity depends on channels segregated to the ApPr. For the mechanosensitivity, the spatial segregation could be tested by “touching” either the ApPr or the soma with a piezo-controlled blunted pipette (see, for example, Hao et al. 2013). However, the sensitivity to osmotic changes is much more difficult to assess, as pressure application does not have enough spatial resolution to discriminate among compartments in such a compact cell such as the CSFcNs. In theory, a highly spatially localized osmotic jump could be reached with photolysis, but a caged compound releasing many osmotic particles simultaneously should be used. In typical photolysis experiments, a localized osmotic jump is produced, but it is very low (in the order of 1 to 2 mOsm).

      In mice, like in fish (Sternberg et al, NC 2018), we can observe throughout the figures that a large fraction of the channel activity occurs with partial and very fast openings of the PKD2L1 channel. I recommend the authors analyse the points below:

      (a) To what extent do these partial openings of the channel contribute to the changes in holding current and resting potential?

      As the reviewer indicates, these partial and very fast openings are a characteristic of PKD2L1 single-channel activity that seems to be present in different species. However, estimating what is the exact contribution of these events to the sustained current would require a detailed model of the channel that it is still lacking. Indeed, the exact mechanism by which CSFcNs show this prominent sustained current is unknown and should definitely being addressed in future works.

      (b) In the trace from the outside out AP, it looks like the partial transient openings are gone. Can the authors verify whether these partial openings are only present in somatic recordings?

      The outside-out recordings from the ApPr also show some partial openings (please look at the upper trace in Figure 4Db). We will specifically mention this important point in the revised version of the ms.

      (3) Previous studies have observed expression of metabotropic Glutamate receptors in CSF-cNs (transcriptome from Prendergast et al CB 2023). The authors only used blockers for ionotropic glutamate receptors in their recordings: could it be that these metabotropic receptors influence the response to uncaging of MNI-Glu when glutamate is co-released with a proton?

      We thank the reviewer for pointing out the presence of metabotropic glutamate receptors in CSFcNs. However, our evidence indicates that there is no contribution of metabotropic receptors when uncaging MNIglutamate because: i) the response obtained when uncaging MNI-gLGG (where there is no glutamate release; Figure 5Ab) and ii) the response obtained when uncaging protons from DPNIGABA (a GABA cage that has similar photochemistry than MNI cages which also release a proton when photolysed; data not shown), are the same. Indeed, in both experiments (MNI-gLGG or DPNI-GABA uncaging) a clear photolysis-evoked PKD2L1 current can be observed.

      (4) In the outside out patch of the AP, PKD2L1 unitary currents appear rare. Could it be that the disruption in the cilium or underlying actin/myosin cytoskeleton drastically alter the open probability of the channel?

      Although we have not quantified it, the reviewer is right that the opening frequency of PKD2L1 channels in the outside-out patches is lower than in the whole-ApPr recordings. We interpreted this difference as a difference in channel number. However, another plausible interpretation is that, as the reviewer suggests, the biophysics of the channels are affected because the protein is taken out from its normal ionic environment and/or loses important interactions with regulatory proteins.

      (5) Could the authors use drugs against ASIC to specify which ASIC channels contribute to the pH response in the soma?

      As described in the manuscript, we did perform experiments with ASIC channel blockers, although we did not attempt to characterize the specific ASIC channel involved in the somatic response. Based on what has been published in the literature, we used both psalmotoxin-1 (which blocks ASIC1 channels) and APETx2 (which blocks ASIC3 channels). The presence of ASIC1 channels in mice CSFcNs has been shown by (Orts-Del’Immagine et al. 2012; Orts-Del’Immagine et al. 2016), while the presence of ASIC3 in the lamprey CSFcNs has been shown by (Jalalvand et al. 2016). When we puff an acidic solution aiming at the soma, we can record an inward current that is blocked by psalmotoxin-1, although there is always a small component remaining (as originally shown by Orts-Del’Immagine in the aforementioned articles); however, we have not attempted to block this small component that remains after psalmotoxin-1 bath application.

      (6) This is out of the scope of this study, but we did observe in fish a very rarely-opening channel in the PKD2L1KO mutant. I wonder if the authors have similar observations in the conditions where PKD2L1 is mainly in the closed state.

      We have never seen such kind of openings in our recordings (when the channel is closed or in the presence of dibucaine).

      Bjugn, R, H K Haugland, et P R Flood. 1988. “Ultrastructure of the mouse spinal cord ependyma.” Journal of Anatomy 160 (octobre): 117‑25.

      Bruni, J. E., et K. Reddy. 1987. “Ependyma of the Central Canal of the Rat Spinal Cord: A Light and Transmission Electron Microscopic Study”. Journal of Anatomy 152 (juin): 55‑70.

      DeCaen, Paul G., Xiaowen Liu, Sunday Abiria, et David E. Clapham. 2016. “Atypical Calcium Regulation of the PKD2-L1 Polycystin Ion Channel”. eLife 5 (juin): e13413. https://doi.org/10.7554/eLife.13413.

      Delling, Markus, Paul G. DeCaen, Julia F. Doerner, Sebastien Febvay, et David E. Clapham. 2013. “Primary cilia are specialized calcium signalling organelles”. Nature 504 (7479): 311‑14. https://doi.org/10.1038/nature12833.

      Hao, Jizhe, Jérôme Ruel, Bertrand Coste, Yann Roudaut, Marcel Crest, et Patrick Delmas. 2013. “Piezo-Electrically Driven Mechanical Stimulation of Sensory Neurons”. In Ion Channels, édité par Nikita Gamper, vol. 998. Methods in Molecular Biology. Humana Press. https://doi.org/10.1007/978-1-62703-351-0_12.

      Jalalvand, Elham, Brita Robertson, Peter Wallén, et Sten Grillner. 2016. “Ciliated Neurons Lining the Central Canal Sense Both Fluid Movement and pH through ASIC3”. Nature Communications 7 (janvier): 10002. https://doi.org/10.1038/ncomms10002.

      Orts-Del’Immagine, Adeline, Riad Seddik, Fabien Tell, et al. 2016. “A Single Polycystic Kidney Disease 2-like 1 Channel Opening Acts as a Spike Generator in Cerebrospinal Fluid Contacting Neurons of Adult Mouse Brainstem”. Neuropharmacology 101 (février): 549‑65. https://doi.org/10.1016/j.neuropharm.2015.07.030.

      Orts-Del’immagine, Adeline, Nicolas Wanaverbecq, Catherine Tardivel, Vanessa Tillement, Michel Dallaporta, et Jérôme Trouslard. 2012. “Properties of Subependymal Cerebrospinal Fluid Contacting Neurones in the Dorsal Vagal Complex of the Mouse Brainstem”. The Journal of Physiology 590 (16): 3719‑41. https://doi.org/10.1113/jphysiol.2012.227959.

      Palomero, Orhi Esarte, Megan Larmore, et Paul G. DeCaen. 2023. “Polycystin Channel Complexes”. Annual Review of Physiology 85 (Volume 85, 2023): 425‑48. https://doi.org/10.1146/annurev-physiol-031522-084334.

      Prendergast, Andrew E., Kin Ki Jim, Hugo Marnas, et al. 2023. “CSF-Contacting Neurons Respond to Streptococcus Pneumoniae and Promote Host Survival during Central Nervous System Infection”. Current Biology 33 (5): 940-956.e10. https://doi.org/10.1016/j.cub.2023.01.039.

      Zhang, Miao, Yueming Ma, Xianglu Ye, Ning Zhang, Lei Pan, et Bing Wang. 2023. “TRP (Transient Receptor Potential) Ion Channel Family: Structures, Biological Functions and Therapeutic Interventions for Diseases”. Signal Transduction and Targeted Therapy 8 (1): 261. https://doi.org/10.1038/s41392-023-01464-x.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers were very impressed with your work and definitely feel it is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing. They, however, request some re-analysis of the data and discussion with minimum new experiments, if any. I think, if feasible, this will improve the quality of the study and would look forward to receiving a revised version.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1 - Molecular identity and localization of PKD2L1

      Major

      Nomenclature clarity.

      Please clarify the distinction between PKD2L1 and TRPP3. Several parts of the text and figure labels appear to conflate these names. PKD2L1 corresponds to the TRPP3 subfamily member and should not be interchanged with PKD2/TRPP2. Please confirm and update the nomenclature consistently throughout the manuscript (text, figure labels, and captions).

      We have addressed this issue in response to reviewer 1. There seems to be some confusion in the literature concerning the nomenclature of PKD2L1 channels as in some recent publications the PKD2L1 channels are still named as TRPP3. However, the nomenclature of PKD2L1 channels or TRPP2, was updated in 2016 (Wu, Sweet and Clapham, Pharmacological Reviews, 2010). As indicated in response to reviewer 1, we have removed from the text any reference to the TRPP nomenclature and stuck to the PKD2L1 name.

      Physiological meaning of apical restriction.

      Expand the discussion of why apical-restricted localization matters. Specifically, address how segregation to the ApPr could support directional sensing of CSF flow and/or detection of localized pH gradients.

      Following the reviewers and editor’s comments, we have revised the discussion in order to take this and other comments into account.

      Open-state annotation (O3).

      Please include the O3 state in panels Ba and Bb; the figure clearly shows an additional open level consistent with O3.

      The editor is right in that there is another state that presumably corresponds to O3. Following the editor’s recommendation, we now indicate this 3rd level and add a short sentence explaining this in figure 1 legend.

      Minor

      Indicate the ROI definition and background-subtraction method used for fluorescence quantification (ApPr vs soma).

      Not applicable for this figure.

      Ensure the same intensity scale (lookup table and range) is used across panels to enable direct comparison.

      Done.

      (2) Figure 2 - Electrophysiological characterization of ApPr and somatic recordings Major

      Definition of "PKD2L1-dependent current."

      Define this term precisely at its first appearance. Specify whether it denotes currents inhibited by dibucaine, abolished in PKD2L1-knockout preparations, or both.

      Done

      Statistical power of single-channel analysis.

      The number of observed openings (< 1000 events) is too low to estimate open probability (Po) reliably. Please re-analyze the data using macroscopic current traces rather than Po-based kinetics.

      Confounded Po analysis.

      The current Po analysis mixes current amplitude and Po in the same calculation, conflating independent variables. Re-evaluate or remove this analysis.

      Unknown channel count.

      Because the number of channels in each patch is unknown, Po and "closed probability" values cannot be interpreted meaningfully. Focus instead on the averaged macroscopic current density.

      General analytical validity.

      The single-channel analyses in Figure 2 are not interpretable under these experimental conditions. Closed-time distributions and Po-based metrics (e.g., "Po1," "P2") depend critically on channel number and event sampling. Moreover, the manuscript applies essentially the same Po methodology across conditions (Po1 vs P2), which adds no mechanistic resolution and risks circular interpretation.

      Actionable recommendation:

      Remove Po- and closed-time-based analyses from Figure 2 and from the manuscript as a whole. Reanalyze the data using metrics that remain valid when the channel number is unknown:

      Macroscopic current analysis (leak-subtracted current density, I-V relationships, activation time constants).

      Single-channel conductance only (amplitude histograms and unitary slope conductance), without attempting Po or dwell-time inference.

      Report filtering bandwidth and sampling rate, and restrict statistical treatment to these robust parameters.

      Following the reviewers (see above) and editors’ recommendations, we have reanalyzed the data in order to avoid the analysis based on Po and Pc. We have instead calculated from the recordings other 2 parameters, n<sub>max</sub> (the maximum number of channels that open simultaneously during a 500 ms time window) and the total open time of a single channel during the same 500 ms time window. The main text, figures and corresponding figure legends, and the Materials and Methods section have been changed accordingly. Notably, the Po and Pc analysis were removed from Figures 3, 4 and Supplementary Figure 3, and replaced by the above-mentioned parameters. Also, the fact that the recordings are not long enough to calculate Po is now specifically mentioned in the Materials and Methods section, lines 785 to 790. In addition, the analysis in Figure 3Ce has been redone so that the activity of the channel as a function of pH is now plotted as the normalized apparent Po (relative to the apparent Po value at pH 7.4).

      Minor

      State whether input-resistance values (1.8-4.4 GΩ) were leak-subtracted and series-resistance-compensated.

      As already mentioned in the methodology section (line 693), series resistance was not compensated for during the experiments. We have now added a sentence in the methodology section indicating that in the voltage range that was chosen for the analysis of the input resistance, no voltage-dependent conductance was activated (lines 809 to 812).

      Ensure unit consistency: use Po or normalized Po rather than frequency (Hz) throughout.

      (3) Figure 3 - pH-evoked currents and kinetics

      Major

      Invalid Po analysis (Fig. 3Ca-Ce).

      The Po- and Pc-based single-channel analyses in panels 3Ca-3Ce should be deleted. As noted earlier, the event count is insufficient, and the number of active channels in each patch is unknown. Under these conditions, Po and Pc values have no quantitative meaning and could mislead readers. These panels do not contribute additional mechanistic insight beyond the macroscopic current data and therefore, should be removed. If retained for illustrative purposes, they must be explicitly labeled as representative traces without any statistical quantification.

      As we mentioned above, we removed the Po and Pc analysis from the manuscript.

      Minor

      Present regression equations and r<sup>2</sup> values for the linear fits shown in Fig. 3D.

      Done (lines 947951).

      Confirm that all axes include units and identical scaling between conditions for direct comparison.

      Done.

      (4) Figure 4 - Laser photolysis and local stimulation experiments

      Major

      Laser timing annotation.

      Clearly mark laser-pulse timing (e.g., arrow or shaded region) on all current traces to facilitate interpretation.

      We thank the editor for pointing out the inconsistencies in terms of the laser pulse timing. To indicate the laser pulses, we have now added an arrowhead in cases where a single sweep is shown (for example, Figure 3D), and an arrowhead and a dotted magenta vertical line in cases where multiple sweeps are shown (for example, Figure 3C).

      pH calibration within the laser spot.

      Provide quantitative calibration of pH changes induced by laser photolysis, including information on spot size, local diffusion, and estimated pH recovery kinetics.

      This is an important point and we thank the editor for mentioning it. We have now completed the subsection untitled “photolysis” where we provide information on the lateral and axial dimensions of the photolysis laser spot used in this work (lines 734 to 737). We have also rewritten part of Figure 5A legend to highlight the fact that the experiments presented there (photolysis on top of the ApPr and next to it) are compatible with a high spatial resolution of proton release (lines 1023 to 1026).

      On the other hand, we have attempted to perform pH calibrations in the setup using the pH-sensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2, and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, when trying to perform a calibration under the conditions of a real experiment,

      where photolysis occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml), we encountered the following problem, which made it impossible to obtain any useful data: the 405 nm uncaging pulse bleaches the dye, and any useful information is lost. Also, our imaging system is not fast enough to follow the pH change. As it is discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 814), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond time range.

      Repeated stimulation effects.

      Discuss whether repeated photolysis induces adaptation or desensitization of PKD2L1 currents, and indicate whether current amplitude decreases across successive trials.

      This issue is now specifically mentioned in the Materials and Methods section, lines 739 to 741.

      Invalid interpretation of the "OFF response."

      The interpretation of the so-called "OFF response" in Figure 4C is not supported by the presented data. There is no evidence for a bona fide OFF current, and the literature cited does not demonstrate such a phenomenon for PKD2L1 alone. Rather, previous studies implicate PKD1L3-dependent mechanisms in similar biphasic responses. Please reconsider the cited references and remove claims of an OFF current attributed to PKD2L1.

      Done.

      Actionable recommendations:

      Do not use the term "OFF response" throughout the manuscript. Recast these transients as pH dependent recovery or relaxation of current following cessation of acidification.

      Done. We have performed extensive rewriting and reorganization of the Results and Discussion in order to take into account both the reviewer’s and editor’s comments. Please also take a look at comment #3 of Reviewer 1 and point 11 below.

      Include continuous-illumination controls (sustained local acidification) to test whether a steady state current is maintained. This will clarify whether the post-stimulus transient reflects recovery kinetics rather than a distinct current species.

      We thank the editor for suggesting this experiment. However, continuous laser illumination is a difficult manipulation and does not necessarily lead to an acidification of the illuminated volume. Indeed, with continuous illumination the cage is lost from the illumination spot and needs to be replaced by diffusion from the non-illuminated volume, leading to non-homogeneous concentrations. Also, the chances of inducing photo damage are higher. We thus designed a similar experiment where instead of performing continuous illumination we photolysed with short and high frequency trains in order to produce a long-lasting acidification. The results of these experiments have been added to the manuscript as part of the results section and in Figure 5H. Similarly to what is seen with single illuminations, the photolysis trains induce a current that appears almost exclusively at the end of the train, implying that the current is indeed a PKD2L1-dependent recovery current.

      Align the current time course with measured or estimated local pH (or calibrated proxy) to demonstrate causal coupling and avoid implying a separate conductance.

      We have added the calculated pH change to the inset of Figure 4C as an example.

      Revise the schematic/model figure and textual description accordingly, restricting the framework to phasic vs sustained activation modes without invoking a separate OFF current for PKD2L1.

      Done.

      Minor

      Include scale bars, sample numbers (n), and laser parameters (duration, power) in all panels.

      In order not to make the figure and the panels very heavy in the original version, we tried to limit the number of scale bars. We have now performed some modifications, added the missing scale bars, and changed the figure legend in order to take into account the editor’s comments. We have also corrected a few values that were wrongly reported.

      Standardize p-value formatting (e.g., p = 6 × 10 ⁶) throughout the figure and legend.

      Done.

      (5) Figure 5 - Single-channel recordings

      Major

      Mixed parameters (current amplitude and Po).

      The current analysis improperly mixes single-channel current amplitude and Po within the same figure, conflating distinct parameters. These quantities must be analyzed and presented separately, or the Po data should be removed entirely if not independently supported.

      Insufficient event count.

      Given the very limited number of observed openings, Po-based statistics are not meaningful. Please report only representative single-channel traces and corresponding amplitude histograms without attempting quantitative Po estimation.

      Minor

      Convert frequency (Hz) values to Po for consistency with earlier analyses, or remove frequency metrics altogether if Po analysis is omitted.

      Figure 5 does not include Po or event frequency analysis, so we think there must be a misquotation of the figure. However, the Po issue has already been addressed before and alternative analysis have been proposed.

      (6) Introduction

      The introductory paragraph mentions the "five senses" as a framing concept. However, this statement lacks scientific grounding in the context of CSF-contacting neurons and chemosensory physiology. The traditional "five senses" classification is not an evidence-based neurophysiological framework and may be misleading to readers. I recommend removing or rephrasing this part, focusing instead on molecular and cellular mechanisms of sensory transduction (e.g., chemical, mechanical, and pH sensing) rather than on classical sensory categories.

      Following the editor’s recommendation, we have removed this part.

      The manuscript refers to PKD2L1 using the term TRPP2 in some parts of the introduction. This is incorrect, as PKD2L1 corresponds to TRPP3, not TRPP2. Please correct this nomenclature and ensure consistent use of "PKD2L1 (TRPP3)" throughout the entire manuscript to avoid confusion with the distinct PKD2/TRPP2 protein, which belongs to a different subfamily with separate physiological roles.

      We thank the editor for pointing this out. As we mentioned in the responses to the “public reviews”, the literature is confusing, so we decided to remove from the manuscript any mention to TRPP channels.

      (7) Discussion

      The current Discussion reads largely as a descriptive summary of results and lacks conceptual depth. It does not effectively integrate the biophysical properties of PKD2L1 with its physiological role as a neuronal pH sensor, nor does it develop a broader interpretation relevant to CSF homeostasis or chemoreception.

      Following the reviewers and editor’s recommendations, we have now added a new section in the Discussion untitled “PKD2L1 downstream signaling mechanisms”.

      (8) Insufficient biophysical analysis

      The discussion of channel gating and pH dependence is superficial and does not explore the energetic or structural mechanisms underlying proton sensitivity. The authors should analyze their data in the context of known PKD/TRPP family biophysics-for example, protonation sites, subunit composition, or gating kinetics-and explain how these confer bidirectional (acidic vs alkaline) sensitivity within physiological ranges.

      In this work, we studied the pH sensitivity of PKD2L1 channels in the context of CSFcN sensory physiology. From a pure biophysical perspective, the pH sensitivity of PKD2L1 channels has been studied by multiple groups; however, it is still unknown how the gating of the channel responds to pH changes, although it can be proposed that some polar residues in the protein regulate the state of the pore. Likewise, the mechanism of the “off-response” is also unknown. To the best of our knowledge, there is only one article in which the authors have attempted to relate pH, PKD2L1 channel structure, and function. In this work (Su et al., Nature Communications 2018), the authors compare PKD2L1 channels with another pH-sensitive member of the TRP family, TRPML3, whose structures at pH 7.4 and 4.8 are known (Zhou et al., Nature Structural and Molecular Biology, 2017). We have rewritten some sentences of the Discussion in order to be more specific about the pH dependence of PKD2L1 channels and its proposed mechanisms.

      (9) Weak physiological context

      The manuscript does not adequately address how PKD2L1 functions as a true physiological pH sensor. The discussion should connect channel activity to realistic CSF pH fluctuations (6.8-7.6) and to relevant physiological or pathophysiological conditions (e.g., respiratory acidosis, neurogenic regulation of CSF composition). Without this, the relevance of large, artificial acidification (pH 3-3.5) remains unclear.

      We have added a new section in the Discussion where we speculate on how PKD2L1 channels may be activated in physiological and pathophysiological conditions. However, we would like to insist here that the main goal of the photolysis experiments (which induce short and large acidifications) was to assess the spatial segregation of PKD2L1 channels. We now mention this point specifically and also speculate on the conditions that could eventually give rise to the “recovery” current.

      (10) Over-interpretation of unsupported points

      The paragraph describing voltage propagation from the ApPr to the soma/axon is speculative and unsupported by any data in the manuscript. Please delete this section entirely, including the citation to Orts-Del'immagine et al., unless new electrophysiological evidence is added.

      We think this point (the propagation of signals originating from the ApPr to the soma) is important in the context of our work, so we have decided to make new experiments in order measure directly the degree of coupling between the 2 compartments. To do that we made simultaneous, current-clamp and voltage-clamp recordings from the ApPr and the soma. In these conditions we were able to measure experimentally and for the first time both the coupling coefficient and coupling conductance, which confirm that the propagation of voltage signals from the ApPr to the soma is extremely efficient. These new results are now described in a new subsection and in a new Figure 6.

      (11) Clarify the role of "OFF currents."

      The Discussion repeatedly refers to an "OFF response," but this phenomenon is not experimentally demonstrated for PKD2L1 alone. It likely represents pH-dependent recovery rather than an independent current. All discussion of "OFF currents" should be removed or reformulated accordingly.

      Following the editor and reviewer’s comments, we have deleted the term “off response” and “off currents” from the ms and have replaced them with the term “recovery current”. We have also changed the discussion accordingly.

      (12) Integration with ASICs and compartmental sensing

      While the manuscript briefly mentions ASIC involvement, it does not articulate how PKD2L1- and ASIC-mediated signals might complement each other in different compartments (ApPr vs soma). The authors should discuss the potential division of labor between these sensors and how such compartmentalization enhances pH detection in CSFcNs.

      Following the editor’s comments, we have rewritten the part of the subsection ‘the involvement of ASICs’ in the Discussion.

      (12) Broadened physiological perspective

      The Discussion should close by considering Ca<sup>2+</sup> -dependent downstream pathways activated by ⁺ PKD2L1 and their implications for CSF flow regulation, neurosecretion, and central chemoreception. These translational aspects would substantially improve the impact and readability of the manuscript.

      Done

      Overall, the Discussion must evolve from a descriptive narrative to a mechanistically and physiologically integrative synthesis, highlighting why PKD2L1 is not merely present in the ApPr but is a key molecular transducer linking ionic microenvironment to neuronal excitability.

      As it has been detailed above, we have performed several changes in the Discussion that follow the reviewer’s and editor’s recommendations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We believe that the manuscript has been substantially strengthened through the revision process. The main changes are summarized below:

      We substantially revised the Introduction and Discussion sections to better position our work relative to previous studies on starvation-dependent thermotaxis plasticity, neuropeptidergic modulation, and AWC function.

      We clarified throughout the manuscript the distinction between negative thermotaxis in innocuous thermal ranges and thermonociceptive responses to noxious heat. We now discuss more explicitly that these behaviors involve at least partly distinct molecular, cellular, and circuit-level mechanisms.

      We performed new experiments in ins-1 mutants. Unlike what was previously reported for thermotaxis plasticity, ins-1 does not appear required for starvation-dependent thermonociceptive plasticity in our paradigm (new Figure 6—figure supplement 1).

      We revised the analysis and terminology used for AWC calcium imaging data. We no longer use the “deterministic/stochastic” terminology and instead describe a starvation-induced shift from predominantly excitatory responses to a mixed distribution of excitatory and inhibitory responses. We also added new quantitative analyses and histogram representations of response distributions, as directly suggested by reviewers, to better illustrate this point.

      We performed new genetic interaction experiments using eat-4; flp-6 double mutants. These analyses revealed that glutamatergic and FLP-6 signaling act largely in parallel to mediate heat-evoked reversals after early food deprivation, while prolonged starvation reveals a hierarchical interaction between these pathways.

      We revised and clarified the mechanistic model figures accordingly, particularly regarding the proposed ASI → AWC signaling pathway and the role of ASI-derived neuropeptides.

      We improved the presentation and statistical rigor throughout the manuscript, including:

      - Replacement of heating power values by corresponding temperature increases,

      - Clarification of the rationale for using the 1-hour off-food condition as reference,

      - Expanded statistical reporting and multiple-comparison procedures,

      - Additional methodological details for calcium imaging, rescue validation, and cell ablation approaches,

      - Clarification of replotted datasets in figure legends.

      We also simplified the manuscript by removing experiments whose interpretation remained ambiguous (notably the nsy-1 and nsy-7 analyses).

      Below, we provide a detailed point-by-point response to all reviewer comments.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Thapliyal and Glauser investigates the neural mechanisms that contribute to the progressive suppression of thermonociceptive behavior that is induced under conditions of starvation. Several previous studies have demonstrated that when starved, C. elegans alters its preferences for a variety of sensory cues, including CO2, temperature, and odors, in order to prioritize food seeking over other behavioral drives. The varied mechanisms that underlie the ability of internal states to alter behavioral responses are not fully understood; however, there is growing evidence for a role of neuropeptidergic signaling as well as the capacity for functionally distinct microcircuits, formed by distinct internal states, to trigger similar behavior outcomes.

      Within the physiological range of C. elegans (~15-25{degree sign}C), starvation triggers a profound reduction in temperature-driven thermotaxis behaviors. This reduction involves the recruitment of the amphid sensory neuron pair AWC. The AWC neurons primarily act to sense appetitive chemosensory cues; however, under starvation conditions begin to display temperature responses that previous studies have linked to the reduction in thermotaxis navigation. Here, Thapliyal and Glauser investigate the impact of starvation on thermonociceptive responses, innate escape behaviors that are triggered by exposure to noxious temperatures above 26{degree sign}C or rapid thermal stimuli below 26{degree sign}C. They compare the strength of thermonociceptive behaviors, specifically heat-triggered reversals, in worms experiencing either early food deprivation (1 hour off food) or prolonged starvation (6 hours off food). Their experiments demonstrate a progressive loss of heattriggered reversals that is mediated by AWC and ASI neurons, as well as both glutamatergic and neuropeptidergic signaling.

      At the level of neural activity, this study reports that the transition from early food deprivation to prolonged starvation reconfigures the temperature-driven activity of AWC neurons from largely deterministic to stochastic. This finding is interesting in light of previous work that reported the opposite transition (from stochastic to deterministic) in temperature-driven AWC responses when comparing well-fed worms to those kept from food for 3 hours. This study also identifies neural and genetic mechanisms that contribute to differences in thermonociceptive responses at +1 versus +6 hours of starvation; confusingly, these mechanisms are partially distinct from those that contribute to differences in negative thermotaxis behaviors in well-fed and +3 hours of starvation worms (Takeishi et al, 2020). A limitation of this manuscript is that these differences are not particularly acknowledged or addressed, other than the hypothesis that independent mechanisms underlie negative thermotaxis versus thermonociceptive stimuli. However, this suggestion is not experimentally verified.

      We thank this reviewer for pointing to the interest of our work. The difference between previous work focusing on negative thermotaxis in the range of innocuous temperatures and our work focusing on thermo-nociceptive response is important and indeed deserves further clarification and a deeper discussion in the manuscript.

      Two major empirical evidence for a distinction between negative thermotaxis (as assessed in previous studies) and thermonociceptive plasticity (as assessed in our paradigm) were already included in the initial article version. First, we reported a decrease in average response in AWC neurons due to a shift in the distribution of response polarities from mostly up-response to a mix of ‘up-response’ and ‘downresponse’ after starvation, while previous results showed an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to starvation-evoked thermotaxis adaptation, ASI neurons are required to orchestrate starvation-evoked plasticity in thermonociception. These observations already indicate differences at the circuit and cellular level. For the revision, we conducted further experiments to address the molecular level. We tested ins-1 mutants (see also specific point 3 by reviewer 2, below) and deepened this aspect in the discussion section of the revised manuscript. Previous study found that INS-1 signaling from the intestine is a major mediator of negative thermotaxis plasticity. In contrast, our new data show that INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (see new Figure 6-supplement 1). Taken together these three lines of empirical evidence support the notion that negative thermotaxis and thermonociceptive starvation-evoked plasticity involves at least partially distinct mechanisms and it seems therefore inappropriate to qualify this notion as purely hypothetical.

      Modification in the revised manuscript include extended introduction about the known thermotaxis regulation mechanisms (Introduction section), new Figure 6-supplement 1 about ins-1 and accompanying text in the result section as well as extended discussion about these differences (Discussion section)

      Multiple additional aspects of this study make the results difficult to synthesize with existing knowledge, including

      (1) Differences in - and insufficient discussion of - the magnitude and kinetics of thermal stimuli;

      We have included a better description of the stimuli characteristics in the revised methods section. The discussion section was deepened to better emphasize that different types of thermal stimuli have been used in different studies.

      (2) This study's use of "heating power" rather than temperature values when presenting behavioral results;

      Thanks for noting this point, which was indeed an unnecessary complication in the result display of the initial manuscript. We have changed ‘power values’ to corresponding ‘temperature increase’ in the revised figures.

      (3) The use of +1 hours starvation as a baseline instead of well-fed worms. Indeed, this last point reflects a noticeable experimental result that differs from previous studies, namely that at room temperature, the basal movements of well-fed and starved worms are not different. Such a surprising result warrants further quantification of worm mobility in general and could have prompted a set of experiments directly testing previously published thermal conditions to demonstrate that the new effects reported arise specifically from the use of thermonociceptive stimuli, as hypothesized.

      The consideration of on-food and off-food behavioral state is an important point indeed. We found that, at room temp, worms shift from dwelling on food (a state with high spontaneous reversal rate) to global search off-food (a state with low spontaneous reversal, after 1hr starvation) (see Figure 1B). Therefore, unlike the reviewer’s statement, the reported data highlighted key differences in the basal locomotion of worms in fed and 1hr starved conditions. Furthermore, these behavioral states have been characterized very deeply using high-content worm behavioural tracking in our recent publication: Thapliyal et al. 2023 (PMID: 37236963). Our choice of using 1hr as a baseline is primarily driven by the fact that an elevated baseline of spontaneous reversals on food decreased the dynamic range to monitor changes in heat-evoked reversals. 1-hour early food deprivation reduced spontaneous reversals and led to a mild attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. Additionally, we observed a clear progressive decrease in heat-evoked reversals with increased duration of food-deprivation, which we further used to dissect the mechanism underlying this plasticity. The choice of the 1hr food deprivation timepoint as a reference is further justified below.

      Finally, a previous report (Yeon et al, 2021) demonstrated differences in the impact of chronic versus acute neural silencing on starvation-dependent plasticity in the context of negative thermotaxis. We therefore wonder whether similar developmental compensation impacts the neural circuits that contribute to starvation-dependent plasticity in the thermonociceptive responses.

      Indeed, this is an interesting question. Our conclusions are so far based on ablation (with chronic effects). In order to gain insight on this question, future studies could address the impact of chronic vs acute silencing approach in starvation-dependent thermonociceptive responses. We have added an opening on this question in the discussion, as follows:

      “Another open question is whether ASI action takes place during development (prior to starvation), or more acutely with active signaling after starvation.”

      A weakness of this manuscript is that the introduction is insufficiently scholarly in terms of citations and the description of current knowledge surrounding the impact of internal state on sensory behavior, particularly given previous work on the impact of feeding state on thermosensory behavioral plasticity (Takeshi et al 2020, Yeon et al 2021) and chemosensory valence (Banerjee et al 2023, Rengarajan et al 2019, etc).

      To address this weakness, we have revised the introduction section of the manuscript and cited previous relevant research on the impact of internal states on animal behavior, including the papers suggested by the reviewer. We note that 2 out of 4 suggested citations were already present in the initial manuscript (though in the discussion section).

      Similarly, the authors' commanding knowledge of the distinction between thermotaxis navigation (especially negative thermotaxis) and thermonociceptive behaviors could be communicated in more depth and clarity to the readers, in order to contextualize this study's new findings within the previous literature.

      As mentioned above, we have deepened this aspect in the discussion section of the manuscript (with a dedicated paragraph). It is quite clear that starvationinduced plasticity in negative thermotaxis and thermonociceptive behaviors engage distinct mechanisms (at least in part). These differences include distinct alterations in AWC calcium activity, role of ASI neurons and INS-1 neuropeptide.

      Nevertheless, this study represents a solid addition to the growing evidence that C. elegans sensory behaviors are strongly impacted by internal states, and that neuropeptidergic signaling plays a key role in mediating behavioral plasticity. To that end, the authors have provided solid evidence of their claims.

      We thank this reviewer for the efforts in evaluating our manuscript, for the positive assessment of our work, and for highlighting some weaknesses which, we believe, have been addressed through the revision.

      Reviewer #2 (Public review):

      In this work, Thapliyal and Glauser tried to provide a mechanistic understanding by which animals modulate their neural circuit responses to control nociceptive behavior on the basis of the dynamic internal feeding state. It is an important study that adds to the growing body of evidence coming from multiple model systems. They have used elegant genetics, behavioral, and Ca-imaging experiments to demonstrate how the auxiliary thermosensory neuron pair, AWC, and one of the internal state-sensing interneuron pairs, ASI, respond to dynamic internal starvation state to modulate behavioral response to noxious heat. Interestingly, these neuron pairs use distinct molecular mechanisms along with some other unidentified neurons to suppress heat-induced reversal response under short-term and prolonged starvation. The experiments are well performed, supporting most of the claims and providing an important framework for future studies.

      I have some queries that, if answered, will certainly enhance the study.

      (1) The results suggest that ASI is one of the primary drivers for the starvation-evoked behavioral plasticity, which regulates AWC activity under prolonged starvation. It raises many important questions, including: (a) how starvation modulates ASI response to heat?, and (b) under prolonged starvation, whether ASI also promotes other, non-AWC, glutamatergic inhibitory neurons to suppress heat-induced reversal, and how?

      We agree with this reviewer that the mechanisms by which ASI detects and mediates starvation-evoked changes in our model is a very interesting (unsolved) question. However, addressing these questions empirically represents a substantial body of work that would go beyond the scope of the present report. E.g., is temperature-dependent activity in ASI even relevant? At present, we envision that ASI could either work acutely (during heat stimuli) or be modulated over much longer time frames (hours of starvation) as an internal state sensor. Therefore, there will be quite some exploration needed before we figure out the ASI-level regulation more fully (including the critical temporal aspect regarding cell activity, as well as quantitative and qualitative transmission aspects). It will be very interesting in future work to address these questions.

      (2) How does ASI regulate AWC activity? In the proposed model (Figure 8) authors suggested an independent, unknown signal, other than INS-32 and NLP-18, from ASI to regulate AWC activity. However, from the results, the existence of another signal is not very clear.

      Thanks for raising this point, which reveals a weakness in our graphical representation (in Fig. 8) that was not properly conveying our point. Our current work shows INS-32 and NLP-18 to be important in modulating heat-evoked reversals upon starvation. However, at the moment, we don't know if INS-32, NLP-18, both, and/or other neuropeptides from ASI modulate AWC activity patterns. The calcium imaging experiments in single, double and potentially triple mutants would answer these questions but are not within our current reach, given the time needed to carry out these experiments. However, we acknowledge this point and have changed the figure and its legend to state that the arrow connecting ASI to AWC activity pattern could potentially reflect the action of these neuropeptides.

      (3) Previously, Takeishi et. al. showed that ins-1 dynamically modulates AWC-AIAmediated thermotaxis behavior based on the feeding state of the animal. It raises questions whether ins-1 also contributes to noxious heat-induced reversal behavior.

      We thank the reviewer for this question. We have now quantified the phenotype of ins-1 mutant in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity, unlike plasticity in thermotaxis behavior (See Figure 6- Supplement 1). Together with the differential activity patterns in AWC and the differential need for ASI neurons, these new data further consolidate the notion that starvation-evoked thermotaxis adaptation and noxious-heat avoidance engage separable molecular, cellular and circuit-level modulatory mechanisms. A specific discussion paragraph was added too.

      (4) Experiments with AWC fate conversion mutants (nsy-1 and nsy-7) were very good ideas; however, the results obtained were confusing. flp-6 mutant data suggest AWCoff would be essential for heat-induced reversal, especially at the low intensity stimulus level. However, the nsy-1 mutant-forming two AWCon neurons showed complete rescue at the low heat level, which is quite opposite. Similarly, although less prominent, eat-4 rescue experiments suggested both nsy-1 and nsy-7 should behave normally at high heat conditions, which was not the result observed.

      We appreciate this comment and the legit attempt to infer what we should expect from a worm with two AWCon or two AWCoff, respectively. From previous studies so far, it's not quite clear if cellular properties of newly formed AWCs in nsy-1 and nsy-7 mutants, including response to sensory cues, formed synapses and their partners, expression of neuromodulator and gap junctions, synaptic output are similar or different. We think further studies are required to first establish if FLP-6 and glutamate signaling (expression, release and action) from altered AWCs in nsy-1 and nsy-7 mutants are the same or different. Therefore, direct comparison between cell fate conversion mutants with flp-6 and glutamate would rely on too many assumptions at this stage. Considering this comment, the limited additional value of the data with nsy1 and nsy-7 mutants (in the absence of additional analyses) and the confusion it could trigger, we have decided to remove these non-essential data of the manuscript.

      Reviewer #3 (Public review):

      Summary:

      Thapliyal and Glauser show that hunger alters how C. elegans responds to noxious thermal stimuli. Using targeted neural ablation, mutant analysis, and live-cell functional imaging, the authors demonstrate that hunger changes the properties of AWC sensory neurons, which sense noxious heat. The authors further show that the effects of hunger on nociception require ASI neurons, which are known to respond to hunger and mediate the effects of food deprivation on behavior. Finally, the study uses mutant analysis to implicate glutamate and specific neuropeptides in thermal nociception and in the modulation of nociceptors by hungerresponsive neurons.

      Strengths:

      The study clearly shows a strong effect of hunger on nociception and documents a striking effect of hunger on the intrinsic properties of AWC sensory neurons, which respond to noxious heat. The study also clearly and compellingly demonstrates that ablation of hunger-responsive ASI neurons blocks the effects of hunger on nociceptive AWCs. These data, which constitute the kernel of the manuscript, are striking and exciting.

      Weaknesses:

      The study has some weaknesses that the authors should address.

      (1) Ablation of AWC neurons alters the basal sensitivity to noxious heat stimuli. This should be clearly noted in the description of the result and warrants some discussion.

      We thank this reviewer for raising this legitimate point. We have clarified this aspect in the results section of the revised manuscript, reading as follows:

      “Removal of AWC nearly abolished heat-evoked reversal behavior across all stimulus intensities and timepoints (Figure 2B and E). While one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines, this observation suggests that AWC plays an essential role in mediating the thermonociceptive response under both early food deprivation and prolonged starvation. Notably, in AWC-ablated animals, the residual response level was unaffected by starvation, suggesting that AWC might also be required for the expression of starvation-dependent plasticity.”

      The contrast with known function of the best-characterized sensory neurons mediating thermal nociception (AFD and FLP) is discussed as follows:

      “...Therefore, noxious heat-evoked activity in AWC varies widely according to context, which is in line with previous literature [18, 20, 41]. Interestingly, the role of AWC is distinct from that of AFD and FLP neurons, which are canonically linked to thermosensation and nociception [5, 14, 42, 43], but contribute only modestly to heat-evoked behavior in our assay conditions with between 1 and 6 hrs of food deprivation.”

      (2) Throughout the study, it seems that data are replotted in multiple figure panels. The authors should clearly indicate in the figure legends when this occurs. Also, the authors should ensure that statistical tests requiring multiple comparisons are correctly implemented and reflect the number of times experimental data are compared to a single set of control data.

      Thanks for raising this important point. We have clarified this aspect in the revised figure legends of the manuscript, and in the method section. In some instances, we reconducted some analyses to be perfectly rigorous in multiple comparison accounting. This did not lead to significantly different conclusions. The one exception was that the small effect of eat-4 mutation on spontaneous reversal went below significance threshold. We therefore removed this aspect of the result reporting and of the corresponding interpretation scheme, which became slightly simpler (Figure 3). Globally, this makes the story more focused.

      (3) How ASIs modulate AWCs remains unclear. The authors find that loss of INS-6, an insulin-like peptide provided by ASIs, partially recapitulates the effect of ASI ablation. This observation is not further developed, and instead, the authors characterize other secreted factors that seem to mediate sensitization of animals to noxious heat stimuli. While it is interesting that there are multiple opposing inputs into the nociceptor circuit, the essential connection between ASIs and AWCs that underlies the foundational observations in Figures 1 and 2 is not sufficiently characterized.

      Whereas we agree that how ASI modulates AWCs is only partially solved by our study, we should emphasize that our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal upregulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI ablated line, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint.

      Following this reviewer’s comment, we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data (with a negative result) remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS-6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required go deeper in this question, which could be performed in future work.

      “...Additional neuropeptides (such as INS-6) may also be involved, but in the absence of direct evidence for their origin from ASI, they were not included in this scheme.”

      (4) The assertion that 'starvation reshapes AWC responses from deterministic to stochastic' is not clearly supported by the data. AWC neurons seem capable of showing different responses to thermal stimuli, and the probabilities associated with these responses change after fasting. The different kinds of responses are seen under basal and fasted conditions.

      We thank this reviewer for the comment. There is an activity response shift that is quite solidly described, including with new quantitative analyses of distributions (histograms in new Fig. 4CD and new Fig. 5C-D, accompanied by Kruskal-Wallis tests). Yet, we totally agree that the wording choice was inappropriate. We have furthermore changed our terminology to avoid using the terms “stochastic” or “deterministic” that were indeed a cause of confusion. We now use the terms “stimulus-locked responses” and describe the shift as “shift from mostly excitatory responses to a mix of both excitatory and inhibitory responses”. We have also included detailed methodology for characterization of traces and statistical analysis in the revised method section of the manuscript, together with the new analyses on peak polarity distribution.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers agree that the study is clearly presented and makes good use of behavioral, genetic, and imaging approaches to link starvation state with changes in AWC and ASI function. To strengthen the manuscript and ensure clarity for readers, we ask you to address the following points in revision:

      (1) Positioning and citations.

      Clarify how your findings relate to Takeishi 2020, where the opposite trend in AWC activity was reported, and make a clear distinction between thermonociception and thermotaxis. The introduction should also include additional citations in two specific areas: prior work on AWC and noxious thermal stimuli, and studies demonstrating starvation-dependent behavioral changes via altered neuropeptide release (e.g., Banerjee 2023; Rengarajan 2019).

      We have clarified this aspect with extension of the work cited in the introduction and extensive rewriting of the discussion sections.

      Our data shows that mechanisms underlying starvation dependent changes in thermonociception and thermotaxis show differences at the molecular, cellular and circuit levels. First, we see a decrease in average response in AWC neurons due to shift from mostly excitatory to a mix of excitatory and inhibitory responses in response to noxious heat after starvation, while previous study found an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to thermotaxis behavior ASI neurons are required to orchestrate starvation evoked plasticity in thermonociception. And, finally, previous study found that INS-1 signaling from the intestine regulates thermotaxis behavioral plasticity while INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (new data in Figure 6 supplement 1).

      We have revised the introduction section of the manuscript and cited previous relevant research on AWC and noxious thermal stimuli and studies demonstrating starvationdependent behavioral changes via altered neuropeptide release including the papers suggested by reviewers. The extended discussion section reads as follows:

      “Starvation regulates thermonociceptive and negative thermotaxis plasticity via at least partly different mechanisms

      Previous studies showed that AWC plays an important role in starvationdependent plasticity in the negative thermotaxis behavior in an innocuous thermal range between 15 and 25°C [26, 33]. Negative thermotaxis involves the detection of thermal changes created by animal movement in spatial thermogradient (0.5°C/cm), the magnitude of the expected thermal changes approximating 0.01°C/s [26]. The starvation impact on negative thermotaxis was shown to (i) involve an up-regulation of AWC cell activity, (ii) rely on INS1 neuropeptide produced in the intestine and (iii) to occur independently of ASI neurons. In contrast, our study used thermo-nociceptive stimuli, with faster raising thermal slopes (~0.5-2°C/s, hence 50-200 times faster than those occurring for thermotaxis) and covering noxious temperatures (up to 28°C). Our results indicate that the regulation of thermo-nociceptive response by starvation (i) is linked to a shift in the distribution of AWC activity response polarities from mostly excitatory to a mix of excitatory and inhibitory response, (ii) relies on ASI and specific neuropeptide produced in ASI, and (iii) works independently of INS-1 neuropeptide. Therefore, our study complements our understanding of the modulation of temperature-dependent behavior in C. elegans with previously undocumented mechanisms at the circuit, cellular and molecular levels.”

      (2) ASI → AWC mechanism.

      Because ASI is central to your conclusions, please expand on how ASI is thought to act on AWC and/or other neurons. If an additional ASI signal is proposed beyond INS-32/NLP-18, mark this as speculative unless further rationale can be provided, and adjust the model figure accordingly.

      We have revised Figure 8 and its legend to clarify what is still hypothetical in the way ASI could affect AWC activity patterns and reversals. Note that the figure was also modified to integrate the conclusions made from epistasis analysis of eat-4 and flp-6.

      (3) AWC response description.

      The data support a shift in response distributions rather than a categorical switch from "deterministic to stochastic." Please adjust the language accordingly and provide a clear description of how traces were classified as "up, variable, or down," ideally with a quantification of the distributional shift.

      We agree that the term “stochastic” can convey different things, and because it was used for something different for AWC in the past, we should have avoided it. We have revised the nomenclature. What we observe can indeed be better described as a shift in the response polarity distribution. The article was revised accordingly. We also included the quantitative analysis and histogram representation, suggested in one of the specific comments, and added detailed methodology on the categorization of traces.

      When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (4) Presentation and statistics.

      In figure legends, indicate where datasets are replotted across panels and confirm that multiple-comparison corrections take account of repeated comparisons to the same controls.

      We have included these details in the revised figure legends, and a statement in the method section.

      (5) Methods clarity.

      Provide justification for using 1-h off-food as the baseline, with quantification of baseline mobility/reversal rates. Expand the calcium-imaging methods to describe the processing pipeline (ΔR calculation, baseline period, drift correction), and add a brief rationale if the approach deviates from common normalization procedures. Clarify how cell ablations were performed and verified for specificity, and how cell-specific rescues were confirmed. Please also acknowledge the potential for developmental compensation with chronic ablation.

      The justification of using 1hr off-food as baseline was made more prominent in the revised manuscript.

      Revised result section:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with much-reduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvation-dependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      In addition, the method section was modified as follows:

      - Calcium imaging details were added regarding ΔR calculation, baseline period, drift correction.

      - We now explicitly refer to the original respective articles describing the neuroablation lines.

      - We clarify that cell-specific transgene expression for rescue was confirmed using SL2::mCherry co-marker

      In the result section, we now explicitly address potential developmental compensation in genetic ablation backgrounds in the result section as follows: “…one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines”.

      Reviewer #1 (Recommendations for the authors):

      (1) The data availability statement is missing from the reviewed manuscript and should be included.

      Thanks, we have included the data availability statement in the revised manuscript.

      (2) We request additional information on how n's were determined for individual experiments, as well as the inclusion of post-hoc power measurements for all quantification.

      n were determined in agreement with previous studies using similar measures. No a priori power analyses were performed. A posteriori power analyses are not informative beyond the reported effect sizes and p-values (now reported in File S2). We clarified this in the statistical subsection of the method section.

      (3) In many cases, the specific statistical tests used are not clear or justified; more details should be provided, including the non-post-hoc test used. Are all tests one-way ANOVAs? For comparisons across genotype and starvation duration, two-way ANOVAs would likely be more appropriate. Also, the authors switch between Bonferroni post-hoc tests and Holm-Bonferroni post-hoc tests. What determined the use of one versus another?

      We have now clarified the statistical analyses used and provided full details in File S2. We have now more systematically applied two-way ANOVAs across all relevant analyses (with detailed parameters reported in File S2). When particularly relevant (e.g epistasis analysis between eat-4 and flp-6 mutations) the results of the two-way ANOVAs, is also explicitly stated in the result section.

      We also note that all multiple-comparison corrections were performed using the Bonferroni method. Previous mentions of Holm-Bonferroni correction were inaccuracies, and we apologize for this confusion; these mentions have now been corrected throughout the manuscript.

      (4) The use of heating power instead of the temperature experienced by the worms is an unwelcome abstraction. We strongly recommend revisiting that choice.

      We do agree. We have revised the figures to label the axis with temperature increase.

      (5) For calcium imaging, how are the traces categorized into "calcium up", "calcium down", or "no change"? Were those determined blindly - i.e., by individuals unaware of the experimental condition? Did the response direction need to be consistent across different temperatures? Did the change from baseline need to hit a specific threshold, consistent with previous studies in the field (i.e., +/- 3xSD for a minimum amount of time)? We encourage the authors to include these details in their methods section.

      We have complemented the method section to clarify the criteria for the qualitative classification of traces. More importantly, new quantitative peak polarities comparisons were added (see specific points below and above, about histograms).

      (6) For the experiments showing that exposure to food odor does not prevent response reduction, we suggest that feeding worms heat-killed bacteria would be a helpful control for the importance of bacterial nutritional status. In addition, showing that the impact of starvation was reversible with re-feeding would have been a useful experiment in line with standard experimental design in the starvation field.

      Thanks, indeed with our current work we cannot pinpoint the role of additional sensory cues (except food odor) to be mediating starvation-evoked plasticity. Together with refeeding, these are all extremely interesting questions that we aim to answer and potentially link with ASI and AWC activity in our future work.

      (7) For the various AWC rescue experiments, we found it curious that there wasn't an AWCon+off rescue, only each neuron individually.

      Previous studies have identified similar or opposite responses of both AWCs for distinct sensory cues. Though our calcium imaging experiments point to both AWC on and off having similar response patterns to heat, we cannot rule out the possibility that their output (ability of control reversals) is distinct possibly due to recruited neuromodulators. Therefore, in the present work, we examined where these cell types act via the same or distinct combinations of neuromodulators to control reversals.

      Reviewer #2 (Recommendations for the authors):

      Experiments suggested:

      (1) The authors should look into the Ca-dynamics in ASI.

      How does the spontaneous and heat-evoked activity of ASI differ in fed, early food-deprivation and prolong starvation and its link to releases of neuromodulators, modified AWC activity to alter output of thermal nociception are very interesting questions. However, these questions are extremely exploratory (see argumentation above in response to the public review) and addressing them goes beyond the scope of our current manuscript.

      (2) The authors should check AWC activity in ins-32 and nlp-18 mutant animals.

      In this study, we focused on the roles of INS-32 and NLP-18 released from ASI in modulating heat-evoked reversals, as these mutants exhibit relatively strong behavioral phenotypes. However, these effects remain less pronounced than those observed following ASI ablation. In addition, we cannot exclude the contribution of additional signaling molecules, including INS-6 and other neuropeptides.

      A comprehensive analysis of AWC activity in this context would require calcium imaging across multiple genetic backgrounds, including single, double, and potentially higher order peptide and receptor mutants, combined with cell-specific rescue experiments. While we appreciate the suggestion, such an approach would represent a substantial extension of the present work and will be important to pursue in future studies to further elucidate the underlying mechanisms.

      (3) Short-term food deprivation completely eliminated heat heat-induced reversal response to 100W stimulus, while the response to 400W stimulus remained unaffected. This suggests fed, short-term starvation, and prolonged starvation are three distinct states, and authors should also test the response of AWC and ASI ablated animals in the fed conditions.

      We agree that analyzing thermal nociception in fed states, in addition to short-term and prolonged food deprivation states is an important and interesting question, as these 3 conditions likely represent 3 distinct internal states that may recruit different neural pathways.

      Several reasons led us to set the fed condition aside for this study, and we realize we insufficiently explain them in the initial manuscript. There are 2 main reasons.

      (1) It is of paramount importance to consider the ‘baseline’ reversal rate (spontaneous reversals not triggered by heat, but visible in our dataset as the first point in the ‘dose-response’ curve). In Fed animals spontaneous reversal rate is very high (~38%) compared to the 1hr and 6hr food-deprivation conditions (>10%). This has two consequences: first a decreased dynamic range for quantify heat-evoked reversal, and, second, the difficulty in judging quantitative differences in heat-evoked reversals with such major differences in baseline reversals.

      (2) Experimentally, assessing calcium responses in truly fed animals presents technical challenges. With our current setup, animals must be removed from food for at least ~5 minutes prior to recording (followed by ~5 minutes of imaging), which effectively corresponds to a “freshly starved” condition rather than a fully fed state. While previous studies have used serotonin to mimic aspects of the fed state, such manipulations can be difficult to interpret in this context.

      A systematic comparison including fully fed animals, as well as AWC- and ASI-ablated conditions across these states, would be a valuable direction for future work, in particular once the methodological barriers associated with point 2, have been overcome.

      We have clarified these choices in the result section as follows:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with muchreduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvationdependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      (4) The authors should test the effect of ins-1 in noxious heat-mediated dynamic reversal behavior.

      We thank the reviewer for this valuable suggestion. We have now tested the phenotype of ins-1 mutants in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity unlike plasticity in thermotaxis behavior (New Figure 6 Sup1). This molecular aspect adds to our initially presented evidence at the cell activity and circuit levels, that noxiousevoked reversal and thermotaxis behaviors are regulated in a clearly separable manner.

      (5) Whether Glutamate and flp-6 work in parallel or in the same pathway to regulate reversals?

      We thank the reviewer for this question. We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat-evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4; flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup></sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig.

      8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food-deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (6) The authors should perform flp-6 and eat-4 mutant/rescue experiments in the nsy-1 and nsy-7 background to clarify the results.

      Our results indicate that glutamate release via EAT-4 from both AWC<sup>ON</sup> and AWC<sup>OFF</sup>, as well as FLP-6 from AWC<sup>OFF</sup>, contributes to heat-evoked reversals regulation. The experiments suggested by the reviewer would, in principle, provide further insight into the interaction between AWC subtype identity and the respective roles of glutamatergic and peptidergic signaling.

      However, as discussed in more details above in the public review, the extent and nature of AWC<sup>ON/OFF</sup> remodeling in nsy-1 and nsy-7 mutant backgrounds remain incompletely understood. This introduces significant uncertainty in interpreting results obtained from combining these mutations with eat-4 and flp-6 manipulations. As a result, such experiments would be difficult to interpret in a definitive manner at this stage.

      We therefore consider this an important direction for future work, once the roles of nsy1 and nsy-7 in AWC subtype specification and function are more clearly established. As our preliminary results with nsy-1 and nsy-7 mutants added more confusion than clarity, we have chosen to set them aside (former Fig. 3-figure supplement 2 has been removed).

      Minor comments:

      (1) What is food odor? The experiment should be clearly mentioned.

      Thanks for spotting this unintended omission. Food odor experiments were performed by adding OP50 bacteria on the inward side of the petri dish lid instead of the NGM surface. We have added this description in the method section of the revised manuscript.

      (2) Panel 3C is coming before 3B. This should be rearranged.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 3, with the addition of eat-4 x flp-6 genetic interaction analysis.

      (3) In Figure 6, if the panels are arranged horizontally, it would be easier to follow.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 6, with the addition of eat-4 flp-6 genetic interaction analysis.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors should consider moving measurements of AFD-ablated animals into the main Figure 1. AFD is a well-known thermosensor, and it is worth showing that responses to noxious thermal stimuli persist in animals lacking AFD.

      Thank you for this suggestion. We have moved measurements of AFDablated animals to the main figure (revised Fig. 2).

      (2) AFD ablation does affect responses to noxious heat. The authors could consider ablating/silencing AFDs and AWCs simultaneously to determine whether these two neurontypes account for the behavior.

      We agree that investigating the combinatorial contributions of thermosensory neurons, including AFD and AWC, to thermal nociception is an important and interesting question. In principle, simultaneous ablation or silencing of these neuron types could indeed reveal unexpected interactions.

      In our experimental paradigm, however, we observe only a minimal contribution of AFD neurons to heat-evoked reversals, whereas ablation of AWC nearly abolishes the response. Based on these observations, we chose to focus the present study on AWC, which appears to play a more prominent role in this behavior.

      A more detailed dissection of the potential interactions between AFD and AWC, including combinatorial manipulations, would be a valuable direction for future work. In particular, the possibility that AFD exerts a modulatory influence remains an interesting hypothesis to explore.

      (3) Given that EAT-4/VGLUT and FLP-6 neuropeptides each contribute to nociception, the authors should consider testing an eat-4; flp-6 double mutant to determine whether this combination of neurochemical signals accounts for AWC signaling to downstream circuits.

      We thank the reviewer for this suggestion, which is similar to point 5 of Reviewer 2 (above).

      We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4;flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup>2</sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig. 8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (4) The authors should consider representing AWC responses to thermal stimuli as histograms to illustrate how fasting increases the probability of some responses and decreases the probability of others.

      We thank this reviewer for the suggestion. The proposed histograms nicely convey the concept of “shift in response polarity distribution” that we observed (using the new terminology we now use, instead of using the term “stochastic”). To create such histograms, we computed the magnitude of the peaks on a trial-by-trial basis. When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (5) It seems important to better understand the ins-6 mutant phenotype and determine whether ASI-to-AWC signaling involves this insulin-like peptide (ILP). The authors should consider using some of the tools available for disrupting ILP signaling to more clearly demonstrate that a specific neurochemical signal mediates modulation of AWCs by ASIs.

      Our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal up-regulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint (as follow up data indicated at least for nlp-18).

      Following this reviewer’s comment (and the similar one in the public review), we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required to go deeper in this question, which could be performed in future work.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The authors investigated the response of worms to the odorant 1-octanol (1-oct) using a combination of microfluidics-based behavioral analysis and whole-network calcium imaging. They hypothesized that 1-oct may be encoded through two simultaneous, opposing afferent pathways: a repulsive pathway driven by ASH, and an attractive pathway driven by AWC. And the ultimate chemotactic outcome is likely determined by the balance between these two pathways.

      It is not surprising that 1-octanol is encoded as attractive at low concentrations and repulsive at higher concentrations. However, the novel aspect of this study is the discovery of the combinatorial coding of 1-oct in the periphery, where it serves as both an attractant and a repellent. Furthermore, the study uses this dual encoding as a model to explore the neural basis of sensory-driven behaviors at a whole-network scale in this organism. The basic conclusions of this study are well supported by the behavioral and imaging experiments, though there are certain aspects of the manuscript that would benefit from further clarification.

      A key issue is that several previous studies have demonstrated a combinatorial and concentration-dependent coding of odorant sensing in the nematode peripheral nervous system. Specifically, ASH and AWC are the primary receptors for repellent and attractive responses, respectively. However, other neurons such as AWB, AWA, and ADL are also involved in the coding process. These neurons likely communicate with different interneurons to contribute to 1-oct-induced outputs. The authors' conclusion that loss of tax-4 reduces attractive responses and that osm-9 mutants reduce repulsive responses is not entirely convincing. TAX-4 is required for both AWC (an attractive neuron) and AWB (a repulsive neuron), and osm-9 is essential for ASH, ADL, and AWA (attraction-associated). Therefore, the observed effects on the attractive and repulsive responses could be more complex. Additionally, the interpretation of results involving the use of IAA to reduce the contribution of AWC at lower concentrations lacks clarity.

      The authors did not observe any increased correlation between motor command interneurons and sensory neurons, which is consistent with the absence of a consistent relationship between state transitions and 1-oct application. Furthermore, they did not observe significant entrainment of AIB activity with the 2.2 mM 1-oct application. This might be due to the animals being anesthetized with 1 mM tetramisole hydrochloride, which could affect neural activity and/or feedback from locomotion.

      Comments on revisions:

      The authors have addressed all my previously raised concerns.

      Reviewer #2 (Public review):

      Summary:

      The authors used whole-network imaging to identify sensory neurons that responded to the repellant 1-octanol. While several olfactory neurons responded to the initial onset of odor pulses, two neurons consistently responded to all the pulses, ASH and AWC. ASH typically activates in response to repellants, and AWC typically activates in response to the removal of attractants. However in this case, AWC activated in response to the removal of 1-octanol, which was unexpected because 1-octanol is a harmful repellant to the worm. The authors further investigated this phenomenon by testing different concentrations of 1-octanol in a chemotaxis assay, and found that at lower (less harmful) concentrations the odor is actually an attractant, but becomes repulsive at higher concentrations. The amplitude of the ASH response appeared to be modulated by concentration, but this was not true for AWC. The authors propose a model where the behavioral response of the worm is the result of integrating these two opposing drives, where repulsion is a result of the increased ASH activity over-riding the positive drive from AWC. The authors further tested this theory by testing mutants that ablated the AWC response (tax-4 or AWC::HisCl) or ASH response (osm-9 or ASH::HisCl). The chemo-silencing (HisCl) and tax-4 experiments were consistent with their hypothesis, while the osm-9 mutation had a limited impact on chemotaxis behavior, highlighting the potential role of osm-9-independent signaling in ASH in response to 1-octanol. While the interneuron(s) that integrate these signals to influence behavior were not identified, the authors did find that increasing concentrations of 1-octanol did increase the likelihood of AVA activity, a neuron which drives reversals (and hence, behavioral repulsion).

      Strengths:

      This was simple and elegant work that identified specific neurons of interest which generated a hypothesis, which was further tested with mutants that altered neuronal activity. The authors performed both neuronal imaging and behavioral experiments to verify their claims.

      Weaknesses:

      The authors note that other sensory neurons likely contribute to 1-octanol chemotaxis. Given the NeuroPAL data, it would have been nice to identify these other neurons as well. However, the reviewer is aware that this is tangential to the primary focus of this study.

      Reviewer #3 (Public review):

      Summary:

      This work describes how two chemosensory neurons in C. elegans drive opposite behaviors in response to a volatile cue. Because they have different concentration dependencies, this leads to different behavioral responses (attraction at low concentration and repulsion at high concentration). It has been known that many odorants that are attractive at low concentrations are aversive at high concentrations, and the implicated neurons (at least AWC for attraction and ASH for repulsion) have been well established. None the less, by studying behavior and neural responses in a common context (odor pulses, as opposed to gradients) this provides a clear picture of how these sensory neurons may guide the dose dependent response by separately modulating odor entry and odor exit behaviors.

      Strengths:

      (1) This work provides good evidence that worms are attracted to low concentrations and repelled by high concentrations of 1-oct. Calcium imaging also makes it clear that dose-dependence of this response is stronger for ASH than AWC.

      (2) This work presents calcium imaging and behavior with the same stimulus (sudden pulses in volatile odor concentration), while previous studies often focus on using neuronal responses to pulses to understand navigation of gentle gradients.

      Weaknesses:

      (1) As a whole it is not clear precisely how important AWC is (compared to other cells) for the attractive response (as the authors correctly acknowledge).

      (2) The evidence that AIB minus AVA contains relevant information is weak. It appears the entrainment index in Fig. 6H for AIB-AVA could easily be explained by the negative entrainment between AVA and the stimulus (along with no effect or role for AIB). This is suggested by the similar p-values and similar distribution of random EIs (stretched and mirrored) between the first and last rows of this figure.

      (3) The model in Figure 7 would be strengthened if it was demonstrated that IAA is attractive when worms are saturated in a 1/10^4 concentration. Panel 7G (and ref. 39) indicate that 10^-4 IAA activates ASH, which would suggest a different explanation for the change from attraction to repulsion in 7C.

      It was previously published that 1x10^-4 IAA is attractive in a similar microfluidics context (Albrecht and Bargmann), and we have confirmed its attractiveness in our experimental setup (not shown). Specifically, worms accumulate in zones of the arena where 1x10^-4 IAA is present: they are attracted upon first encounter and maintain attraction after prolonged exposure. A sentence to this effect has been added to Results (line 402).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      This looks great! My only small recommendation is to include the 1-octanol concentration used (2.2 mM) in Fig 4 F,G so that the reader can easily compare these data to Fig 4 D,E. You can do this in the figure or legend, though in the figure would be preferable (as you did for Fig 5 and Fig 6).

      [1-oct ] added to Fig. 4D, E

      Reviewer #3 (Recommendations for the authors):

      (1) The AWC traces in Fig. 3C do not match, despite them being from the same animal. There appear to be too many ASH traces. NeuroPAL data/traces from other animals were not available on Zenodo when this review was prepared. This should be fixed.

      Single-animal traces added to figure so reader can see the relationship in the single animal shown and across the full dataset. The averaged ASH/AWC traces are retained in the updated figure because this observation is central to the main hypothesis, so it is important to demonstrate it is a consistent phenomenon. Legend modified (lines 1016-17), ‘consistent’ added to Results (line 147). Data will be uploaded to Zenodo upon finalization of the Version of Record

      (2) The normalization of traces in Fig. 4 is not clear (panel A is not min-max, as panel B appears to be). This normalization may be critical to the interpretation that ASH responds dose-dependently. The code linked on Zenodo was not available when this review was prepared. This should be fixed.

      The reviewer correctly points out a slight error in normalization of the individual worm traces for the ascending concentration series (4A, left panel; the maximal value shown was slightly under 1.0). This is now corrected, interpretation unchanged.

      (3) Chemogenetic data (4F,G and 5D) should specify the promoters used, as they are not ASH and AWC-specific.

      Figures 4F, G and 5D altered to specify the promoters used, as requested. As before, the full range of neurons in which these promoters are active is stated in the Discussion